generate(text, **options) returns a complete SpeechResult: audio and timeline. generate_dialogue(segments, **options) accepts SpeechSegment objects or dictionaries with text, synthesis options, and optional pause_after_ms. Blank text is rejected. pause_between_segments_ms supplies the default gap; segment options override operation options.
Options include voice_id, speed, emotion, intensity, reference_audio, temperature, seed, and inference_steps. Model pages specify support, voices, and defaults. Reference audio accepts the transcription input formats; raw arrays need sample_rate. A resolved segment cannot select both a reference and a voice.
result.audio is an Audio object with owned mono float32 NumPy samples and sample_rate. It supports save(path) and wav_bytes(). Timeline entries include segment_index, text, text_start, text_end, start_ms, and end_ms. Timing precision is model-dependent, not necessarily word-level.
generate_stream() and generate_dialogue_stream()
# With the model still loaded:
withmodel.generate_stream("First sentence. Second sentence.") asstream:
# Send chunk.audio.samples to your application's audio/file consumer.
Streaming retains no complete recording: consume or store chunks yourself. Delivered chunks remain valid after iteration advances or the model unloads. Each chunk has audio, start_ms, and timeline; there is no stream result() that collects everything.
Pass cancel_event=threading.Event() to an operation and set it from another thread to cancel. TTS cancellation raises OperationCancelledError; already delivered chunks remain usable. Breaking iteration and exiting the context cleans up without requiring an error handler.
Python does not provide speak() or microphone/playback management. Use application-owned playback and unload the model with unload() or a with block.
generate(text, options?) retains generated audio. generateDialogue(segments, options?) accepts { text, voiceId?, pauseAfterMs?, ... } segments. Blank text is rejected. pauseBetweenSegmentsMs sets the dialogue default; segment options override operation options.
Options include voiceId, speed, emotion, intensity, referenceAudio, temperature, seed, and inferenceSteps. Availability, voices, and defaults depend on the model; check its model page. Reference audio and voice selection are alternatives, not simultaneous inputs.
The generation exposes finished, result(), audio (an async iterable), speak(), and dispose(). These read the same retained generation; reading chunks does not discard its audio. dispose() stops unfinished work and releases SDK references, including dependent playback.
Each chunk has { audio, startMs, timeline }. Audio is mono { samples: Float32Array, sampleRate }. Timeline entries contain segmentIndex, text, text offsets textStart/textEnd, and audio times startMs/endMs. Alignment precision depends on the model; word-level timing is not guaranteed. Treat returned samples as read-only; copy before modifying or transferring them.
speak() and speakDialogue()
// From a user action, with the model already loaded:
Playback starts when enough audio is buffered; it does not wait for the whole text. Its handle has pause(), resume(), and cancel(). Completion is reported through onPlayback, not a playback finished promise.
States: buffering, playing, paused, finished, cancelled, failed. highlight contains the current text range or is null during silence/terminal states; failed also includes error.
Use generation.speak({ onPlayback }) to play retained audio from the beginning, including while generation continues. Each call creates a separate playback handle. Direct model.speak() discards played audio; generate() retains it until disposal. New speech on the same instance pauses its previous playback; different instances can overlap.
Native playback
Playback options include backgroundBehavior ("pauseUntilResumed" by default, "pauseAndAutoResume", or "continue") and audioFocus ("interruptOthers" by default, "duckOthers", or "mixWithOthers"). Resume paused speech through its handle.
Continued background audio requires iOS UIBackgroundModes: audio and Android foreground-service setup; see the microphone/background setup. Playback-only Android apps need the mediaPlayback service type and permission, not microphone permission. Background playback does not promise arbitrary background inference jobs.
generate(text, options?) retains generated audio. generateDialogue(segments, options?) accepts { text, voiceId?, pauseAfterMs?, ... } segments. Blank text is rejected. pauseBetweenSegmentsMs sets the dialogue default; segment options override operation options.
Options include voiceId, speed, emotion, intensity, referenceAudio, temperature, seed, and inferenceSteps. Availability, voices, and defaults depend on the model; check its model page. Reference audio and voice selection are alternatives, not simultaneous inputs.
The generation exposes finished, result(), audio (an async iterable), speak(), and dispose(). These read the same retained generation; reading chunks does not discard its audio. dispose() stops unfinished work and releases SDK references, including dependent playback.
Each chunk has { audio, startMs, timeline }. Audio is mono { samples: Float32Array, sampleRate }. Timeline entries contain segmentIndex, text, text offsets textStart/textEnd, and audio times startMs/endMs. Alignment precision depends on the model; word-level timing is not guaranteed. Treat returned samples as read-only; copy before modifying or transferring them.
speak() and speakDialogue()
// From a user action, with the model already loaded:
Playback starts when enough audio is buffered; it does not wait for the whole text. Its handle has pause(), resume(), and cancel(). Completion is reported through onPlayback, not a playback finished promise.
States: buffering, playing, paused, finished, cancelled, failed. highlight contains the current text range or is null during silence/terminal states; failed also includes error.
Use generation.speak({ onPlayback }) to play retained audio from the beginning, including while generation continues. Each call creates a separate playback handle. Direct model.speak() discards played audio; generate() retains it until disposal. New speech on the same instance pauses its previous playback; different instances can overlap.
Browser playback may require a click or tap. Call speak() from that interaction after loading the model.