Repository navigation
fix: update Gemini models - #1133
Conversation
cabljac
left a comment
There was a problem hiding this comment.
Thanks Corie, the core of this is right. Dropping the model allowlist and letting ids pass through to the API is the correct long-term shape, and the createModelReference cleanup checks out (the withVersion branch was dead code in plugin 1.31.0). With vector-search dropped (tracked in #1140), the remaining comments are all on chatbot/multimodal. The one I'd dig into before release is the CANDIDATE_COUNT>1 path, the rest are smaller.
bd3e158 to
2e664a0
Compare
Dropping the model allowlist widened which configs reach the Genkit client and which reach the legacy SDK clients, and the new `global` Vertex AI default reaches a client that cannot use it. Review: PR #1133. - The legacy Vertex AI client, used when CANDIDATE_COUNT is above one, passed `global` to `@google-cloud/vertexai`, which builds `<location>-aiplatform.googleapis.com` and so cannot resolve it. Pass `apiEndpoint` for that case so the request reaches the global host with a `locations/global` resource path. - Both legacy chatbot clients read `parts[0].text`, which is wrong for thinking models: they can lead with a thought part or answer in a later one. Add a shared `answerText` helper that skips thought parts, and fail with a clear error rather than dereferencing an empty candidate list. - Multimodal dropped per-call `safetySettings` from `generateOnCall`: the spread put them at the top level while Genkit only reads `config.safetySettings`. Merge them into `config` instead. - Multimodal answered from the prompt alone when IMAGE_FIELD was configured but the document had no image, writing a COMPLETED status for a malformed doc. Throw instead, matching the legacy Google AI client. - The multi-candidate rule lived in two places that could drift. Extract `wantsMultipleCandidates` and use it for both client selection and the candidates write. - `MODEL` had no validation once the allowlist went, so a typo installed cleanly and failed on every write. Add a loose shape regex, not an allowlist. - Drop the `pro-vision` routing special case: the model is retired, so sending it to a legacy SDK cannot help, and any future id containing that substring would divert silently. - Use the multimodal wording for the missing-model error, which no longer says "Model not found."
Parity with the same review on GoogleCloudPlatform/firebase-extensions#1133, limited to the issues that exist here. The `global` endpoint problem does not apply: the kits legacy Vertex client uses `@google/genai`, which handles `global` natively. - Both legacy clients read `parts[0].text`, which is wrong for thinking models: they can lead with a thought part or answer in a later one. Add a shared `answerText` helper that skips thought parts. - The multi-candidate rule was genuinely divergent here: the client gate read the deploy-time `candidateCount` while the write path read the per-discussion override. A discussion asking for two candidates therefore selected the single-candidate Genkit client and then failed the candidates write. Extract `wantsMultipleCandidates`, thread the effective count through `getGenerativeClient`, and use the one predicate in both places. - `MODEL` had no validation once the allowlist went. Add a shape check on the resolved param and a matching `validationRegex` on the prompt, not an allowlist. - Use a missing-config error message that does not say "Model not found.", which now means something else.
|
@CorieW With the Gemini 2.5 model retirement date approaching, can we get an update on when this will be available please? Thank you for making the change! :) |
The chatbot default and the Genkit allowlist still pin retiring Gemini 2.5 ids. Parity with firebase/extensions#2943.
Mirror the firestore-genai-chatbot change: default the MODEL param to gemini-3.6-flash and let createModelReference fall through to googleAI.model() / vertexAI.model() for ids that are not in the known list, instead of returning null. Unknown ids no longer bypass Genkit and land on the legacy generative-ai / vertex-ai clients, so current Gemini releases work without an extension update. Gemini 2.5 models retire in October 2026. The known list carries Gemini 3.6 / 3.5 / 3.1 plus 2.5 so version aliases still resolve. createModelReference is now non-null, so the dead null checks in createGenerateOptions are gone.
googleAI.model()/vertexAI.model() already resolve any id.
…arams Gemini 3.x is served from the Vertex AI `global`, `us` and `eu` endpoints only, so the previous "same as Cloud Functions location" default would 404 for the new `gemini-3.6-flash` default. VERTEX_AI_MODEL_LOCATION (chatbot) and VERTEX_AI_PROVIDER_LOCATION (multimodal) now default to `global`. Gemini 3.x also deprecates temperature/topP/topK - the Vertex AI model card for gemini-3.6-flash states custom values are ignored. The params are kept for Gemini 2.5 configurations and the limitation is documented in the installer text and generated READMEs.
Dropping the model allowlist widened which configs reach the Genkit client and which reach the legacy SDK clients, and the new `global` Vertex AI default reaches a client that cannot use it. Review: PR #1133. - The legacy Vertex AI client, used when CANDIDATE_COUNT is above one, passed `global` to `@google-cloud/vertexai`, which builds `<location>-aiplatform.googleapis.com` and so cannot resolve it. Pass `apiEndpoint` for that case so the request reaches the global host with a `locations/global` resource path. - Both legacy chatbot clients read `parts[0].text`, which is wrong for thinking models: they can lead with a thought part or answer in a later one. Add a shared `answerText` helper that skips thought parts, and fail with a clear error rather than dereferencing an empty candidate list. - Multimodal dropped per-call `safetySettings` from `generateOnCall`: the spread put them at the top level while Genkit only reads `config.safetySettings`. Merge them into `config` instead. - Multimodal answered from the prompt alone when IMAGE_FIELD was configured but the document had no image, writing a COMPLETED status for a malformed doc. Throw instead, matching the legacy Google AI client. - The multi-candidate rule lived in two places that could drift. Extract `wantsMultipleCandidates` and use it for both client selection and the candidates write. - `MODEL` had no validation once the allowlist went, so a typo installed cleanly and failed on every write. Add a loose shape regex, not an allowlist. - Drop the `pro-vision` routing special case: the model is retired, so sending it to a legacy SDK cannot help, and any future id containing that substring would divert silently. - Use the multimodal wording for the missing-model error, which no longer says "Model not found."
46dab5e to
7bbff78
Compare
|
@neilpatrickadams Hi there, sorry for the delay, priorities have been on some other extensions work. I'm going to try and get it across the line ASAP |
The legacy client now reads parts from candidates instead of response.text(), so the mock has to expose them.
…esponse generateContentStream's aggregate concatenates every part into parts[0].text and drops the thought flag, so answerText() could never skip thought parts on the Vertex path. Nothing consumed the stream; the client awaited the aggregate. generateContent() returns the parts as sent.
Reading parts directly bypassed the SDK text() accessor, which used to throw with the block or finish reason. A blocked prompt or a SAFETY/RECITATION candidate surfaced as a bare "No text returned candidate". Both legacy clients now name the prompt block reason or the first candidate's finish reason in the thrown error.
- Legacy Google AI client refuses a SAFETY or RECITATION candidate even when it carries partial text, matching what the SDK text() accessor did. - answerText() joins every non-thought text part instead of taking the first; gemini-2.5-flash on Google AI returns a candidate split across parts. - Offer the us and eu multi-region Vertex locations the descriptions already name. - Changelogs say existing installs keep their stored model and location, and multimodal discloses that ids outside the old allowlist now run through Genkit with sampling and safety config applied. - Multimodal docs drop the pro-vision text-only note; a missing image is now an error on both providers. - Bump chatbot to 0.1.0 and multimodal to 1.1.0.
Hey @cabljac, any update on this one? |
|
@neilpatrickadams sorry for the lack of update, i was waiting on a fix to genkit, which is one of the core dependencies of the extension. The fix is in on genkit, so i can now proceed with this. I'll keep you in the loop |
The new us and eu location options built us-aiplatform.googleapis.com, which does not serve Vertex AI. Genkit's Vertex plugin only routes multi-region locations to aiplatform.<loc>.rep.googleapis.com from 1.40.0, so bump genkit, @genkit-ai/google-genai and @genkit-ai/firebase to ^1.42.0 in both extensions. The chatbot's legacy Vertex client (used for CANDIDATE_COUNT > 1) builds its own host, so map global, us and eu to the right apiEndpoint there. 1.42 narrows VertexPluginOptions.apiVersion, which broke the multimodal plugin setup that passed a union of both option types; build each plugin's options in its own branch instead.
…nt has no image Throwing when IMAGE_FIELD is configured but the document has no image turned every image-optional document from COMPLETED into ERRORED on update. Now that Genkit serves every config, keep the pre-1.1.0 Genkit behaviour of generating from the prompt alone. Also add the us/eu endpoint fix and Genkit 1.42 bump to both CHANGELOGs.
…ertex client The Google AI client already treats a SAFETY or RECITATION candidate as having no answer, but the Vertex client stored its partial text as the response. Filter blocked candidates first, so the response comes from the next usable candidate, or the message ends in an ERROR status that names the finish reason. Tests pin both cases and the apiEndpoint the client passes for the default global location. mockReset in beforeEach stops a queued mock response leaking into the next test. The CHANGELOGs now list us and eu as added locations rather than a fix, since they are new in this release, and name all three bumped Genkit packages.
Apply the multimodal initializePlugin(config) shape to the chatbot, so each provider's options keep their own type without casts, and cut the new JSDoc to one line each to match the surrounding code.
|
@neilpatrickadams the release is out! Thanks so much for your patience. |
Resolves #1134
Summary
firestore-genai-chatbotandfirestore-multimodal-genaitogemini-3.6-flash. Genkit no longer throwsModel not found.for unknown ids — falls through togoogleAI.model()/vertexAI.model().global(VERTEX_AI_MODEL_LOCATIONon the chatbot,VERTEX_AI_PROVIDER_LOCATIONon multimodal). Gemini 3.x is only served fromglobal,usandeu, so the previous "same as Cloud Functions location" default would have failed for the new default model.TEMPERATURE,TOP_PandTOP_K— the params stay for Gemini 2.5 configurations.Vertex AI location
gemini-3.6-flashand the other Gemini 3.x models are only served from the Vertex AIglobal,usandeuendpoints (locations). A single region such asus-central1returns a 404. Both extensions resolved the Vertex location to the Cloud Functions region when the location param was left at itsnulldefault, so a Vertex install taking the new default model would have broken.globalis now the default option in both selects and is listed first; a single region is still selectable for anyone pinning a model that is served there (for example a Gemini 2.5 model). Existing installs keep their stored value — the change only affects new installs and anyone who reconfigures.Sampling controls
Gemini 3.x deprecates
temperature,topPandtopK, and the Vertex AI model card forgemini-3.6-flashstates that custom values are ignored. Rather than remove the params (breaking for existing installs that set them, and still meaningful for Gemini 2.5 until October 2026), the installer text and generated READMEs now say the controls are ignored by Gemini 3.x.Parity with firebase/extensions#2943. Translate and resize live in that repo (#2942), not here.
Test plan
gemini-3.6-flashand locationglobalgemini-3.6-flashand with an unlisted current model idus-central1fails ongemini-3.6-flash, and thatglobalsucceedsnpm run generate-readme(diff is limited to the changed param text)Follow-up commits after review
candidates, which the legacy client now reads.generateContent(); the streaming aggregate flattens parts and drops the thought flag, so thought parts could not be skipped there.text()accessor did.answerText()joins all non-thought text parts.gemini-2.5-flashon Google AI returns a candidate split across parts.usandeumulti-region options added to both Vertex location selects.IMAGE_FIELDis set, on both providers.E2E against
dev-extensions-testingwith the pinned legacy SDKs andgemini-3.6-flash: both legacy clients work, including Vertex onglobal.candidateCount: 2works on Vertex and is rejected by Google AI, but the extension never forwardscandidateCountto the legacy clients, so that is tracked with #1154.Pre-existing bugs surfaced by review, tracked separately: #1153 (multimodal never writes the candidates field), #1154 (chatbot per-discussion overrides ignored on the Genkit path).