Skip to main content

Systems and Workflows

Gemini 3.8 TTS: Choose Flash or Flash-Lite, Then Migrate

Compare Flash and Flash-Lite on one approved workload, then inspect transcript metadata, speaker turns, voice selection and WAV handling before moving a TTS application.

Two orange model paths with arrows lead to an audio file, with a small Gemini mark.
On this page
  1. Start with the pipeline you already have
  2. Treat the launch date as context, not an access check
  3. Choose a model against the application’s workload
  4. Separate transcript text from delivery instructions
  5. Rebuild dialogue as explicit turns
  6. Give vocal events a separate rule
  7. Decide where voice selection belongs
  8. Audit the output bytes before changing the handler
  9. Run a migration trial that can fail clearly
  10. Decide whether to migrate, narrow or hold

Choosing between Flash and Flash-Lite is only the first production decision. A text-to-speech migration can fail even when the chosen model generates a convincing voice. The request may send stage directions as words to be spoken. A response handler may add a WAV header to audio that already has one. A dialogue request may omit a speaker assignment that the new schema expects. Each failure sits at a different point in the pipeline, so a successful audio sample cannot clear the whole integration.

Google announced Gemini 3.8 Flash TTS and Flash-Lite TTS on September 23, 2026. Its current English Flash-Lite guide recommends Flash-Lite as the replacement for gemini-3.1-flash-tts-preview. The English Flash guide describes the shared schema and the creative-tier choice. Both document request and audio-output changes that matter to an existing integration. This article builds a migration review from those current guides; it does not report a completed implementation or test.

The engineering decision has two parts: which model fits the application’s volume, listening standard and approval cost, and whether its pipeline can preserve the intended words, directions, speakers and file handling after migration. A team should compare candidates on the same workload, map the existing request and file contracts, then validate the chosen path through its intended account and downstream tools.

Start with the pipeline you already have

Consider a hypothetical application that turns approved support articles into spoken answers. It sends text to Gemini 3.1 Flash TTS Preview, receives audio, wraps the returned bytes as a WAV file and passes that file to a player and an archive. A second feature produces a short exchange between an assistant and a caller. The team wants to move both features to Gemini 3.8.

That description is deliberately incomplete. Before migrating, the engineers need to inspect the requests and handlers that actually run. Does the application put phrases such as “say this reassuringly” inside the text? Does it identify speakers in prose, structured fields or both? Does the output handler assume headerless PCM? Does a downstream player rely on a particular sample format? An answer based on the model name alone would miss those dependencies.

Make one inventory for each production path: the request builder, the configured model, the voice selection, the response format, the byte-handling code, the stored file and every consumer of that file. Include the fallback path and any batch job that uses a different request builder. The purpose is to locate assumptions that a short playground demonstration would not exercise.

For an actual application, the first review should produce two sanitized examples from its existing code: one single-speaker request and one two-speaker request. Trace each through the output handler to a file the product can play. Those examples become migration fixtures. They are evidence of what the current application asks for; they are not evidence that Gemini 3.8 will respond as intended.

Choose Actual size to read the graphic closely.

Treat the launch date as context, not an access check

Google’s API release notes list both Gemini 3.8 TTS models and the Voices endpoint as generally available on September 22, 2026. Its September 23 announcement separately described rollout across the Gemini API, AI Studio and other product surfaces, with Gemini Enterprise API access then described as coming soon. API model status and a particular production account’s feature, region or surface access are separate checks.

Before scheduling the change, check the intended API surface with that account. Confirm the model identifier it accepts, the voices the application needs, the response formats it can request and any applicable usage limits. If an enterprise route is required, confirm its status there rather than carrying forward the launch page’s “coming soon” wording as a current fact.

The current English Flash-Lite and Flash guides name gemini-3.8-flash-lite-tts and gemini-3.8-flash-tts. Confirm those identifiers, the current SDK, response formats and active project limits on the intended route. The public rate-limit guide sends developers to AI Studio for project and model limits; it does not establish this account’s quota.

Choose a model against the application’s workload

The current English Flash-Lite guide presents that model as the high-volume, lower-latency, cost-conscious replacement for gemini-3.1-flash-tts-preview. The Flash guide positions Flash for fidelity, acting nuance and broader dialect coverage. These are Google’s workload descriptions, not measured performance for this application. The guides say the two Gemini 3.8 TTS models share an API schema and structured prompting format, allowing a model change through one parameter once the request has been migrated.

Shared schema is helpful, but it does not make the models interchangeable for every product promise. The hypothetical application has routine spoken answers and occasional dialogue. It could begin its evaluation with Flash-Lite for the routine path because that matches Google’s stated positioning. It should include Flash when a line requires more demanding direction or pronunciation. Neither choice is a measured winner for this application.

Compare both candidates on the same approved transcript and review rule. Record whether the words are correct, speakers are assigned correctly, the file works downstream, and the output meets the product’s listening standard. Then measure actual turnaround, retries, output tokens and review effort. Google’s pricing page, captured September 30, lists standard paid API rates through December 31, 2026 of $0.50 per million text input tokens for either model, plus $9 per million Flash audio output tokens or $6 per million Flash-Lite audio output tokens. It lists higher rates beginning January 1, 2027.

These are dated list rates, not a measured cost per approved answer. Rise’s cost-per-accepted-task method helps frame the calculation; its other model prices are not Gemini TTS rates. Use Google’s rate-limit guide to locate the account’s active model and project limits before forecasting volume.

Hume’s three-rater VoiceEQ evaluation of preview versions offers bounded voice-quality context and discloses a non-exclusive Google licensing agreement. It reports weaker speaker similarity and some control limits alongside stronger overall voice results. Those preview-version observations cannot select a migration model for this application. TechTarget’s report frames the release as a refinement of existing voice AI capabilities, not an independently proven production result.

A mixed result is possible. The application might use one model for routine answers and another for selected dialogue. Such a design also creates routing, monitoring and review work. It becomes worthwhile only if tested output and operating cost justify the extra path. The migration decision should concern approved audio delivered to users, rather than which model sounds best in an isolated sample.

Choose Actual size to read the graphic closely.

Separate transcript text from delivery instructions

The most consequential request change may be a distinction the old code does not make. The current English Flash-Lite migration guide says Gemini 3.8 TTS treats input text strictly as a verbatim transcript. It warns that embedded instructions such as “Say cheerfully: Hello!” or a textual “Speaker 1:” label may be spoken aloud. The guide directs turn-level speaker and style information into structured speech_metadata instead.

Suppose the hypothetical support application currently sends this text to produce a friendly answer:

Say reassuringly: Your replacement is on its way.

Under the documented Gemini 3.8 approach, the words to be spoken are the sentence about the replacement. The reassuring delivery belongs in the request’s style metadata. This example illustrates the mapping; it is not a tested request or a promise that a particular style phrase will work. The team must use the current schema for its chosen API and listen to the result.

A useful audit searches the request builder for instructional prefixes, speaker labels, bracketed production notes and concatenated prompt templates. The risky text may come from more than one place. A content editor might enter the transcript, while application code prepends a direction that no longer belongs there. A migration that changes only the model identifier leaves the prefix in place.

Classify each piece of input before rewriting it. Spoken words remain transcript. Delivery direction becomes style metadata. Speaker identity becomes speaker metadata. A genuine vocal event may remain in the transcript using the documented event syntax. If an instruction has no clear destination, record it for review instead of guessing. That classification creates a traceable reason for every request change.

The acceptance check is literal: compare the approved transcript with what the file says. Listen for the expected sentence, missing words and accidentally spoken directions. Transcript fidelity deserves its own result even when the voice sounds natural. A fluent performance of the wrong words is still a migration failure.

Choose Actual size to read the graphic closely.

Rebuild dialogue as explicit turns

A two-speaker feature adds a second way for prose labels to leak into audio. The current English Flash-Lite guide says each turn in a multi-speaker request should explicitly assign a speaker in speech_metadata, and that assignment should match a configured speaker. It describes attaching speaker and style metadata to each text block or part in the relevant API structure.

The migration review should therefore follow the dialogue one turn at a time. For every line, identify the words to be spoken, the intended speaker and any line-specific delivery direction. Then check that the application builds a separate, correctly assigned turn. A global instruction saying that there are two characters does not replace the turn-level mapping described in the guide.

For the hypothetical support exchange, imagine that the assistant asks whether the caller has an order number and the caller says they need to find it. The request builder must keep those two roles distinct even if the application later inserts a new line between them. A string-based approach that depends on labels such as “Assistant:” and “Caller:” may be easy to assemble, but the English guide warns against relying on such labels inside the transcript.

Validate more than whether two voices are audible. Check turn order, word assignment and what happens when a turn is inserted, removed or regenerated. If the wrong speaker delivers a correct sentence, the file is unsuitable for the conversation even though every word appears. Record the defect at the request level so the team can fix speaker mapping rather than repeatedly adjusting a voice description.

This review is also a chance to simplify templates. When speaker identity and style have explicit fields, an old block of prose instructions may no longer be necessary. Remove it only after confirming that each instruction has either moved to an appropriate field or been retired intentionally. A shorter request is useful when its meaning is clearer, not merely because it contains fewer characters.

Choose Actual size to read the graphic closely.

Give vocal events a separate rule

The English migration guide makes a narrower exception for short vocal events. It says momentary non-speech events and pauses can be represented with inline tags such as <laugh>, <sigh>, <cough>, <breath> or <short pause>. It directs delivery styles such as whispering into speech_metadata.style.

That gives the hypothetical application a practical editing rule. A laugh at a particular beat is an event in the performed transcript. A calm tone across a line is a delivery instruction. Putting both into the same free-form text field makes it harder to predict what should be spoken and what should shape performance.

Do not add event tags merely because the new model accepts them. Each one changes what a listener may hear and what an editor must approve. Test only events the product needs, and keep their placement visible in the approved script. If an event is decorative, removing it may produce a clearer spoken answer. If it conveys meaning, verify that the output contains it at the right moment and that it does not obscure the words around it.

Confirm exact syntax and behavior in the intended account before encoding tags in a production template. The English guide supplies a classification for review; it does not substitute for testing the actual request.

Decide where voice selection belongs

A migration may also expose a request template that repeatedly describes the same voice in a long prose block. The current English Flash-Lite guide advises designing custom voice personas upstream and then passing the resulting voice_... identifier to synthesis requests with minimal or empty style strings. It describes preset voices, an extended voice library, custom personas and replication. The Voice design guide says stored voices use store=true, have a 200-per-project limit and a one-year lifetime. These are documented capabilities and limits, not account results for this application.

For engineers, the decision is where the approved voice is represented. If the application uses a preset, record the preset selection and the reason it meets the product’s needs. If it uses a custom persona, record the approved identifier, its creation and expiry dates, and how the request builder retrieves it. Plan a new design and approval check before that stored voice expires; do not assume a replacement sounds identical. Avoid silently mixing a stored voice identifier with an old paragraph that tries to redesign the voice on every request. When both are present, reviewers may struggle to tell which input explains a change in output.

Voice replication adds a separate rights and consent dependency. Google’s announcement says a replica can start from a 30-second reference sample of the user’s own voice or one they have rights to use, and requires a matching verbal consent recording from the voice owner. Its footnote says replication through AI Studio is unavailable in Illinois, Texas, the EEA, the UK, Switzerland and India. That is an AI Studio restriction as stated in the announcement; the captured text does not establish rules for every other surface.

If the hypothetical application does not need a replica, its migration plan should not acquire one merely to use the new model. If it does need one, treat rights, consent, surface access and identifier handling as prerequisites. A code path that can reference a voice does not establish permission to generate with it.

Choose Actual size to read the graphic closely.

Audit the output bytes before changing the handler

The other documented change sits after generation. The current English Flash-Lite migration guide says Gemini 3.8 TTS defaults to WAV (audio/wav) for unary requests, including a standard RIFF header. It contrasts that with gemini-3.1-flash-tts-preview and earlier TTS models, which it says defaulted to headerless raw PCM (audio/l16). The guide tells developers who previously wrapped raw bytes in a WAV header to remove that manual wrapping and write returned WAV bytes directly to a .wav file. It also describes explicitly requesting raw audio/l16, audio/mulaw or audio/alaw when a pipeline needs those formats.

That distinction matters to the hypothetical application because its old handler adds a WAV header. If the new response already contains one, repeating that step could produce an invalid or misleading file. The precise failure must be established by inspecting the actual response and trying it in the application’s consumers; no file was generated for this article.

Review the handler in terms of bytes and declared format. What does the request ask for? What type does the response report? Does the first file produced by the handler open in a normal player? Does the archived file contain the same playable audio? Does the product’s player accept it? Those questions follow the audio through the pipeline instead of stopping at a successful API response.

If the application needs raw PCM for a downstream component, an explicit raw format may preserve part of its existing handling. That still requires a deliberate choice of format and an end-to-end test. If it accepts WAV, removing the old wrapper may be simpler. Neither path should be selected solely from the file extension: a .wav name cannot repair bytes assembled under the wrong assumption.

Keep an unmodified response-derived file during the trial. When an editor or player rejects the processed version, compare it with that source before changing prompts or voices. A format defect and a performance defect require different fixes. Separating them reduces the chance that the team spends a day tuning delivery for audio its handler has corrupted.

Choose Actual size to read the graphic closely.

Run a migration trial that can fail clearly

An effective trial uses the intended account, current documentation and the application’s real request builder. It starts with a small set of approved, rights-cleared transcripts. The hypothetical support application could choose a normal answer, an answer that previously used an embedded direction, a short two-speaker exchange and a file that exercises its archive and player. These are proposed cases, not tests performed for this article.

For each case, save the input transcript, intended style and speaker mapping, selected voice, model identifier, requested format and result. Then compare the generated words with the approved transcript. Listen for spoken instructions and labels. For dialogue, check which voice delivers each turn. Inspect the returned file before and after the output handler, and play the final file in the actual product path. Record retries and manual repair instead of retaining only a successful take.

The trial should include a deliberate negative check. Put an old request template beside its migrated version and identify which text moved into metadata. This is a review of request construction, not a recommendation to send a known-bad template to users. It gives a code reviewer a concrete way to see whether the migration removed the source of accidentally spoken directions.

Also test the fallback and any batch path. A migration may work in the interactive feature while a background job still uses the old wrapper or old model name. If the application has several audio consumers, play the same approved file through each relevant one. Passing a local player alone would not establish that the archive, editor and client player handle it.

Use a short written acceptance rule. For example, the team might require exact transcript content, correct speaker turns, playable files in every intended consumer and a documented outcome for the chosen voice. Quality judgment can be part of that rule, but it should identify the listening standard and reviewer. A vague “sounds fine” cannot tell the next engineer which defect is acceptable.

A completed trial should leave artifacts that another engineer can inspect: sanitized requests, response-format observations, sample files, review notes and the code change. If a case fails, the record should indicate whether the problem began in transcript construction, speaker or style metadata, model behavior, byte handling or a downstream consumer. That classification turns a failed sample into a useful engineering decision.

Choose Actual size to read the graphic closely.

Decide whether to migrate, narrow or hold

After comparing Flash and Flash-Lite on the same approved workload, the hypothetical application can migrate a path when its production account accepts the intended model, current documentation supports the request, the output says the approved words with the correct speakers, and the final file works throughout its delivery chain. The team should also understand its current usage terms and the review effort required to produce acceptable audio.

It can narrow the migration when routine single-speaker answers pass but dialogue or a particular consumer does not. Moving the proven path first is a defensible engineering choice if routing is explicit and the remaining path keeps working. The decision record should name which feature moved and which did not.

It should hold a path when account access, request syntax, voice permission, output format or final playback remains unresolved. Changing a model identifier before those checks pass would make the product depend on assumptions the current documentation cannot verify for that application.

A model choice is supported only by the application’s own matched outputs, dated cost inputs and active limits. The migration is complete only when the chosen path delivers approved audio through its intended route. Google’s current English guides, release notes and dated pricing identify changes and planning inputs; they do not provide a production-account result, a tested request, a measured model comparison or an inspected file. The next practical step is to map one current request and its output handler, confirm the schema on the intended route, and run the small end-to-end trial against the account that would serve users.

Choose Actual size to read the graphic closely.

Checked for this article

Sources

  1. Google, “Gemini 3.8 text-to-speech says hello”
  2. Google AI for Developers, “Gemini 3.8 Flash-Lite TTS”
  3. Google AI for Developers, “Gemini 3.8 Flash TTS”
  4. Google AI for Developers, “Voice design”
  5. Google AI for Developers, “Release notes”
  6. Google AI for Developers, “Gemini Developer API pricing”
  7. Google AI for Developers, “Rate limits”
  8. Hume AI, “Gemini 3.8 Flash TTS VoiceEQ evaluation”
  9. TechTarget, “Gemini 3.8 text-to-speech refines voice AI capabilities”

Keep going

All articles