Skip to content

Added agents image generation docs - #559

Open
dannyjameswilliams wants to merge 12 commits into
mainfrom
agents/image_generation
Open

dannyjameswilliams wants to merge 12 commits into
mainfrom
agents/image_generation

Conversation

@dannyjameswilliams

Copy link
Copy Markdown
Contributor

What's being changed:

Support for new query agent feature - image generation via structured outputs.

Type of change:

  • Documentation content updates (non-breaking change to fix/update documentation )
  • Bug fix (non-breaking change to fixes an issue with the site)
  • Feature or enhancements (non-breaking change to add functionality)

How has this been tested?

  • Local build - the site works as expected when running yarn start

@orca-security-eu orca-security-eu Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Orca Security Scan Summary

Status Check Issues by priority
Passed Passed Infrastructure as Code high 0   medium 0   low 0   info 0 View in Orca
Passed Passed SAST high 0   medium 0   low 0   info 0 View in Orca
Passed Passed Secrets high 0   medium 0   low 0   info 0 View in Orca
Passed Passed Vulnerabilities high 0   medium 0   low 0   info 0 View in Orca

image_field: list[Annotated[QAImage, Field(description="<image guidance/style description here>")]] = Field(
max_length = 4 # max_length must be specified for lists of images
)
# END AnnotateListImageExample

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, I wasn't aware of this? Does this let you give extra "system_prompt" style influence over the image generation in description ?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah precisely, so the model will generate an image prompt, but the description field attached to any image fields will always be passed down to the image model

Comment thread docs/query-agent/guides/ask_mode.md Outdated

## Image Generation

Ask mode allows you to also request images to be generated, which will be based on any retrieved data from the search. [See the page on image generation for more details.](../reference/image_generation.md)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe remove "from the search". For example, I think it could calculate the average price and then display that right?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good point, I will adjust :)

</TabItem>
</Tabs>

and display it with [PIL](https://pypi.org/project/pillow/) in Python, or save it as a PNG file in JavaScript/TypeScript:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe an idea to also have, save it back in Weaviate? And a code block on how to do that?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Implemented! I added two tabs, one to save/display the image and another to save/search the image in weaviate. Both are dropdowns to keep it less cluttered, let me know what you think

@CShorten CShorten left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Super cool, awesome work with this Danny! 🔥

Love the examples! model <> t-shirt, slide deck, series of adverts, temperature charts, ... haha, plenty to choose from! 🚀

* `"landscape"` (default): 1536×1024
* `"portrait"`: 1024×1536

## Non-client usage

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We don't really support non-client usage and don't have any REST APIs documented.
We' re using clients to exactly encapsulate such implementation details, so maybe we're better off to remove this section not to confuse users?

@danmichaeljones danmichaeljones Oct 5, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The existing output_format does allow plain dict[str, Any] in addition to pydantic BaseModels, so I agree we should probably document how to use image generation with raw JSON schemas. But that's not non-client usage as such (you're still using our client, just we don't constrain where you get your JSON schemas from) so maybe we want a different heading and to have a fuller example that calls our client with a dict schema?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated this section now, so it's still using the client but raw json schemas instead of pydantic/zod


**Images are generated independently**, meaning that if you want a consistent theme amongst your requested images, you should add a consistent description to your image field. Try specifying specific layout instructions, hex color codes and stylistic choices.

**Timeouts**: image generation adds latency. In Python, when your output format contains images, the client's default timeout rises to 60 seconds. If you set your own `timeout`, make sure it's long enough.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

still wording says rises even 60s is the default I think

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants