Skip to content

fix: update Gemini models - #1133

Merged
cabljac merged 13 commits into
mainfrom
fix/gemini-model-lifecycle
Sep 30, 2026
Merged

cabljac merged 13 commits into
mainfrom
fix/gemini-model-lifecycle

Conversation

@CorieW

@CorieW CorieW commented Aug 13, 2026 •

Copy link
Copy Markdown
Collaborator

Resolves #1134

Summary

  • Default firestore-genai-chatbot and firestore-multimodal-genai to gemini-3.6-flash. Genkit no longer throws Model not found. for unknown ids — falls through to googleAI.model() / vertexAI.model().
  • Default the Vertex AI location params to global (VERTEX_AI_MODEL_LOCATION on the chatbot, VERTEX_AI_PROVIDER_LOCATION on multimodal). Gemini 3.x is only served from global, us and eu, so the previous "same as Cloud Functions location" default would have failed for the new default model.
  • Document that Gemini 3.x deprecates TEMPERATURE, TOP_P and TOP_K — the params stay for Gemini 2.5 configurations.

Vertex AI location

gemini-3.6-flash and the other Gemini 3.x models are only served from the Vertex AI global, us and eu endpoints (locations). A single region such as us-central1 returns a 404. Both extensions resolved the Vertex location to the Cloud Functions region when the location param was left at its null default, so a Vertex install taking the new default model would have broken.

global is now the default option in both selects and is listed first; a single region is still selectable for anyone pinning a model that is served there (for example a Gemini 2.5 model). Existing installs keep their stored value — the change only affects new installs and anyone who reconfigures.

Sampling controls

Gemini 3.x deprecates temperature, topP and topK, and the Vertex AI model card for gemini-3.6-flash states that custom values are ignored. Rather than remove the params (breaking for existing installs that set them, and still meaningful for Gemini 2.5 until October 2026), the installer text and generated READMEs now say the controls are ignored by Gemini 3.x.

Parity with firebase/extensions#2943. Translate and resize live in that repo (#2942), not here.

Test plan

  • Confirm chatbot and multimodal installer defaults are gemini-3.6-flash and location global
  • Smoke chatbot with gemini-3.6-flash and with an unlisted current model id
  • Confirm a Vertex install pinned to us-central1 fails on gemini-3.6-flash, and that global succeeds
  • Run multimodal Genkit client tests
  • Regenerate both READMEs via npm run generate-readme (diff is limited to the changed param text)

Follow-up commits after review

  • Google AI test mock returns candidates, which the legacy client now reads.
  • Legacy Vertex client uses generateContent(); the streaming aggregate flattens parts and drops the thought flag, so thought parts could not be skipped there.
  • Legacy clients name the prompt block reason or candidate finish reason when no answer comes back, and the Google AI client refuses SAFETY/RECITATION candidates as the SDK text() accessor did.
  • answerText() joins all non-thought text parts. gemini-2.5-flash on Google AI returns a candidate split across parts.
  • us and eu multi-region options added to both Vertex location selects.
  • Versions bumped to chatbot 0.1.0 and multimodal 1.1.0. Both extensions change default model and location, and multimodal now errors on a document without an image when IMAGE_FIELD is set, on both providers.

E2E against dev-extensions-testing with the pinned legacy SDKs and gemini-3.6-flash: both legacy clients work, including Vertex on global. candidateCount: 2 works on Vertex and is rejected by Google AI, but the extension never forwards candidateCount to the legacy clients, so that is tracked with #1154.

Pre-existing bugs surfaced by review, tracked separately: #1153 (multimodal never writes the candidates field), #1154 (chatbot per-discussion overrides ignored on the Genkit path).

@cabljac cabljac left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks Corie, the core of this is right. Dropping the model allowlist and letting ids pass through to the API is the correct long-term shape, and the createModelReference cleanup checks out (the withVersion branch was dead code in plugin 1.31.0). With vector-search dropped (tracked in #1140), the remaining comments are all on chatbot/multimodal. The one I'd dig into before release is the CANDIDATE_COUNT>1 path, the rest are smaller.

Comment thread firestore-genai-chatbot/functions/src/generative-client/genkit.ts Outdated
Comment thread firestore-genai-chatbot/extension.yaml
Comment thread firestore-genai-chatbot/functions/src/generative-client/genkit.ts
Comment thread firestore-genai-chatbot/functions/src/generative-client/genkit.ts
Comment thread firestore-multimodal-genai/functions/src/generative-client/genkit.ts Outdated
@CorieW
CorieW force-pushed the fix/gemini-model-lifecycle branch from bd3e158 to 2e664a0 Compare August 19, 2026 17:31
CorieW added a commit that referenced this pull request Aug 20, 2026
Dropping the model allowlist widened which configs reach the Genkit client and
which reach the legacy SDK clients, and the new `global` Vertex AI default
reaches a client that cannot use it. Review: PR #1133.

- The legacy Vertex AI client, used when CANDIDATE_COUNT is above one, passed
  `global` to `@google-cloud/vertexai`, which builds
  `<location>-aiplatform.googleapis.com` and so cannot resolve it. Pass
  `apiEndpoint` for that case so the request reaches the global host with a
  `locations/global` resource path.
- Both legacy chatbot clients read `parts[0].text`, which is wrong for thinking
  models: they can lead with a thought part or answer in a later one. Add a
  shared `answerText` helper that skips thought parts, and fail with a clear
  error rather than dereferencing an empty candidate list.
- Multimodal dropped per-call `safetySettings` from `generateOnCall`: the spread
  put them at the top level while Genkit only reads `config.safetySettings`.
  Merge them into `config` instead.
- Multimodal answered from the prompt alone when IMAGE_FIELD was configured but
  the document had no image, writing a COMPLETED status for a malformed doc.
  Throw instead, matching the legacy Google AI client.
- The multi-candidate rule lived in two places that could drift. Extract
  `wantsMultipleCandidates` and use it for both client selection and the
  candidates write.
- `MODEL` had no validation once the allowlist went, so a typo installed cleanly
  and failed on every write. Add a loose shape regex, not an allowlist.
- Drop the `pro-vision` routing special case: the model is retired, so sending
  it to a legacy SDK cannot help, and any future id containing that substring
  would divert silently.
- Use the multimodal wording for the missing-model error, which no longer says
  "Model not found."
CorieW added a commit to firebase/extensions that referenced this pull request Aug 20, 2026
Parity with the same review on GoogleCloudPlatform/firebase-extensions#1133,
limited to the issues that exist here. The `global` endpoint problem does not
apply: the kits legacy Vertex client uses `@google/genai`, which handles
`global` natively.

- Both legacy clients read `parts[0].text`, which is wrong for thinking models:
  they can lead with a thought part or answer in a later one. Add a shared
  `answerText` helper that skips thought parts.
- The multi-candidate rule was genuinely divergent here: the client gate read
  the deploy-time `candidateCount` while the write path read the per-discussion
  override. A discussion asking for two candidates therefore selected the
  single-candidate Genkit client and then failed the candidates write. Extract
  `wantsMultipleCandidates`, thread the effective count through
  `getGenerativeClient`, and use the one predicate in both places.
- `MODEL` had no validation once the allowlist went. Add a shape check on the
  resolved param and a matching `validationRegex` on the prompt, not an
  allowlist.
- Use a missing-config error message that does not say "Model not found.",
  which now means something else.
@neilpatrickadams

Copy link
Copy Markdown

@CorieW With the Gemini 2.5 model retirement date approaching, can we get an update on when this will be available please? Thank you for making the change! :)

The chatbot default and the Genkit allowlist still pin retiring Gemini 2.5 ids.

Parity with firebase/extensions#2943.
Mirror the firestore-genai-chatbot change: default the MODEL param to
gemini-3.6-flash and let createModelReference fall through to
googleAI.model() / vertexAI.model() for ids that are not in the known
list, instead of returning null. Unknown ids no longer bypass Genkit and
land on the legacy generative-ai / vertex-ai clients, so current Gemini
releases work without an extension update. Gemini 2.5 models retire in
October 2026.

The known list carries Gemini 3.6 / 3.5 / 3.1 plus 2.5 so version
aliases still resolve. createModelReference is now non-null, so the dead
null checks in createGenerateOptions are gone.
googleAI.model()/vertexAI.model() already resolve any id.
…arams

Gemini 3.x is served from the Vertex AI `global`, `us` and `eu` endpoints only,
so the previous "same as Cloud Functions location" default would 404 for the
new `gemini-3.6-flash` default. VERTEX_AI_MODEL_LOCATION (chatbot) and
VERTEX_AI_PROVIDER_LOCATION (multimodal) now default to `global`.

Gemini 3.x also deprecates temperature/topP/topK - the Vertex AI model card for
gemini-3.6-flash states custom values are ignored. The params are kept for
Gemini 2.5 configurations and the limitation is documented in the installer
text and generated READMEs.
Dropping the model allowlist widened which configs reach the Genkit client and
which reach the legacy SDK clients, and the new `global` Vertex AI default
reaches a client that cannot use it. Review: PR #1133.

- The legacy Vertex AI client, used when CANDIDATE_COUNT is above one, passed
  `global` to `@google-cloud/vertexai`, which builds
  `<location>-aiplatform.googleapis.com` and so cannot resolve it. Pass
  `apiEndpoint` for that case so the request reaches the global host with a
  `locations/global` resource path.
- Both legacy chatbot clients read `parts[0].text`, which is wrong for thinking
  models: they can lead with a thought part or answer in a later one. Add a
  shared `answerText` helper that skips thought parts, and fail with a clear
  error rather than dereferencing an empty candidate list.
- Multimodal dropped per-call `safetySettings` from `generateOnCall`: the spread
  put them at the top level while Genkit only reads `config.safetySettings`.
  Merge them into `config` instead.
- Multimodal answered from the prompt alone when IMAGE_FIELD was configured but
  the document had no image, writing a COMPLETED status for a malformed doc.
  Throw instead, matching the legacy Google AI client.
- The multi-candidate rule lived in two places that could drift. Extract
  `wantsMultipleCandidates` and use it for both client selection and the
  candidates write.
- `MODEL` had no validation once the allowlist went, so a typo installed cleanly
  and failed on every write. Add a loose shape regex, not an allowlist.
- Drop the `pro-vision` routing special case: the model is retired, so sending
  it to a legacy SDK cannot help, and any future id containing that substring
  would divert silently.
- Use the multimodal wording for the missing-model error, which no longer says
  "Model not found."
@cabljac
cabljac force-pushed the fix/gemini-model-lifecycle branch from 46dab5e to 7bbff78 Compare September 18, 2026 14:30
@cabljac

cabljac commented Sep 18, 2026

Copy link
Copy Markdown
Collaborator

@neilpatrickadams Hi there, sorry for the delay, priorities have been on some other extensions work. I'm going to try and get it across the line ASAP

The legacy client now reads parts from candidates instead of response.text(),
so the mock has to expose them.
…esponse

generateContentStream's aggregate concatenates every part into parts[0].text
and drops the thought flag, so answerText() could never skip thought parts on
the Vertex path. Nothing consumed the stream; the client awaited the
aggregate. generateContent() returns the parts as sent.
Reading parts directly bypassed the SDK text() accessor, which used to throw
with the block or finish reason. A blocked prompt or a SAFETY/RECITATION
candidate surfaced as a bare "No text returned candidate". Both legacy
clients now name the prompt block reason or the first candidate's finish
reason in the thrown error.
- Legacy Google AI client refuses a SAFETY or RECITATION candidate even when
  it carries partial text, matching what the SDK text() accessor did.
- answerText() joins every non-thought text part instead of taking the first;
  gemini-2.5-flash on Google AI returns a candidate split across parts.
- Offer the us and eu multi-region Vertex locations the descriptions already
  name.
- Changelogs say existing installs keep their stored model and location, and
  multimodal discloses that ids outside the old allowlist now run through
  Genkit with sampling and safety config applied.
- Multimodal docs drop the pro-vision text-only note; a missing image is now
  an error on both providers.
- Bump chatbot to 0.1.0 and multimodal to 1.1.0.
@neilpatrickadams

Copy link
Copy Markdown

@neilpatrickadams Hi there, sorry for the delay, priorities have been on some other extensions work. I'm going to try and get it across the line ASAP

Hey @cabljac, any update on this one?

@cabljac
cabljac self-requested a review September 30, 2026 14:49
@cabljac

cabljac commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator

@neilpatrickadams sorry for the lack of update, i was waiting on a fix to genkit, which is one of the core dependencies of the extension. The fix is in on genkit, so i can now proceed with this. I'll keep you in the loop

The new us and eu location options built us-aiplatform.googleapis.com,
which does not serve Vertex AI. Genkit's Vertex plugin only routes
multi-region locations to aiplatform.<loc>.rep.googleapis.com from
1.40.0, so bump genkit, @genkit-ai/google-genai and @genkit-ai/firebase
to ^1.42.0 in both extensions.

The chatbot's legacy Vertex client (used for CANDIDATE_COUNT > 1) builds
its own host, so map global, us and eu to the right apiEndpoint there.

1.42 narrows VertexPluginOptions.apiVersion, which broke the multimodal
plugin setup that passed a union of both option types; build each
plugin's options in its own branch instead.
…nt has no image

Throwing when IMAGE_FIELD is configured but the document has no image
turned every image-optional document from COMPLETED into ERRORED on
update. Now that Genkit serves every config, keep the pre-1.1.0 Genkit
behaviour of generating from the prompt alone.

Also add the us/eu endpoint fix and Genkit 1.42 bump to both
CHANGELOGs.
…ertex client

The Google AI client already treats a SAFETY or RECITATION candidate as
having no answer, but the Vertex client stored its partial text as the
response. Filter blocked candidates first, so the response comes from
the next usable candidate, or the message ends in an ERROR status that
names the finish reason.

Tests pin both cases and the apiEndpoint the client passes for the
default global location. mockReset in beforeEach stops a queued mock
response leaking into the next test.

The CHANGELOGs now list us and eu as added locations rather than a fix,
since they are new in this release, and name all three bumped Genkit
packages.
Apply the multimodal initializePlugin(config) shape to the chatbot, so
each provider's options keep their own type without casts, and cut the
new JSDoc to one line each to match the surrounding code.
@cabljac
cabljac merged commit 6b5ea3d into main Sep 30, 2026
16 checks passed
@cabljac
cabljac deleted the fix/gemini-model-lifecycle branch September 30, 2026 15:36
@cabljac

cabljac commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator

@neilpatrickadams the release is out! Thanks so much for your patience.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

gemini-3.5 models not supported on firestore-multimodal-genai extension

3 participants