Blocks and returns a TranscriptionResult. Inputs: a PCM WAV file path, a Wfloat Audio object (including TTS result.audio), or a NumPy array with sample_rate. Arrays use (frames,) or (frames, channels); Wfloat downmixes/resamples without changing your input. File decoding supports uncompressed 8/16/24/32-bit PCM WAV; decode other formats yourself and pass samples.
Options: language, task="transcribe" | "translate", timestamps="segment" | "word", hotwords, on_transcript, and cancel_event. Unsupported model options reject. Current adapters expose segment timestamps for Whisper but no word timestamps.
on_transcript receives a current whole-recording text preview, not a delta. Results include text, optional segments / words, and stop_reason ("complete" or "cancelled"). Optional timing.start_ms / end_ms are relative to the input. Cancellation preserves partial work and may include provisional. Failures raise TranscriptionError with partial_result.
load_streaming_speech_to_text() and create_session()
Live events replace provisional text by utterance id (starting at "0"); is_final commits it. An empty update may clear a previous provisional transcript. Final segments retain utterance IDs.
push() processes available audio synchronously; your application owns microphone capture and its input queue. Avoid inference inside time-sensitive device callbacks. finish() ends input and flushes pending speech; cancel() preserves partial results. result() retrieves the terminal result after finishing/cancelling, or re-raises failure. One live session owns the model; conflicting operations reject.
Session options also accept on_error and cancel_event. on_error receives failures with preserved partial_result. Whisper-like models use overlapping windows; native streaming models process incrementally. Real-time performance depends on the model and device.
Inputs: mono { samples: Float32Array, sampleRate } or { uri } pointing to an accessible file:// / Android content:// audio file. Wfloat downmixes/resamples as needed. TTS result.audio is also valid input.
transcribe(audio, options?)
Returns a handle with result() and cancel(). onTranscript({ text }) contains the current whole-recording preview, not a text delta. Some models only provide a final update.
Results contain text, optional segments / words, and stopReason ("complete" or "cancelled"). Timing, when available, uses { startMs, endMs } relative to the input. Cancellation preserves partial text and may include provisional; failures reject with TranscriptionError.partialResult and cause.
Options on transcribe() or createSession(): language, task: "transcribe" | "translate", timestamps: "segment" | "word", and hotwords. Support is model-specific; unsupported requests reject. Current adapters expose segment timestamps for Whisper but no word timestamps.
awaitsession.startMicrophone();// Start from a user action.
// Later, when the user stops recording:
constresult=awaitsession.finish();
awaitlive.unload();
Live updates replace provisional text by utterance id (starting at "0"); isFinal commits it. Empty updates can clear a previous provisional transcript. The final result includes segments with those IDs.
Alternatively, await session.push({ samples, sampleRate }) supplies mono PCM. It acknowledges acceptance, not completed recognition or backpressure. finish() ends input and drains pending work; result() waits without ending input; cancel() stops and preserves partial results. One live session owns an instance; conflicting work rejects.
There is no default backlog limit. Set maxBufferedAudioMs on createSession() to bound pending audio; exceeding it fails with preserved transcript. Wfloat warns about growing backlog and never silently drops audio. Whisper-like models use overlapping windows; native streaming models process incrementally. Neither guarantees real-time speed on every device.
Shared microphone
createMicrophoneCapture() creates a reusable source. Attach it to STT/VAD sessions with attachMicrophone(source) before pushing or starting capture, then call source.start(). Call source.stop() to stop shared capture; finish each session separately. A session cannot mix attached-microphone and manually pushed input.
Microphone and background options
startMicrophone(options?) and createMicrophoneCapture(options?) accept backgroundBehavior, voiceProcessing, and onCaptureState. Default "pauseUntilResumed" pauses capture when the app backgrounds; call startMicrophone() again (or shared source.start()) to resume. "pauseAndAutoResume" resumes on foregrounding; "continue" requests background capture. voiceProcessing: true requests native echo/noise processing for conversational audio; effectiveness depends on the device and route.
iOS requires NSMicrophoneUsageDescription; background audio additionally needs UIBackgroundModes containing audio. On Android, recording needs RECORD_AUDIO. For background recording, also declare:
For background playback, use mediaPlayback and FOREGROUND_SERVICE_MEDIA_PLAYBACK; apps using both declare both service types and permissions. Android shows a foreground-service notification while active. Start capture while the app is foregrounded.
Inputs: mono { samples: Float32Array, sampleRate }, browser File/Blob, or AudioBuffer. File formats depend on browser decoding support. Wfloat downmixes/resamples as needed. TTS result.audio is also valid input.
transcribe(audio, options?)
Returns a handle with result() and cancel(). onTranscript({ text }) contains the current whole-recording preview, not a text delta. Some models only provide a final update.
Results contain text, optional segments / words, and stopReason ("complete" or "cancelled"). Timing, when available, uses { startMs, endMs } relative to the input. Cancellation preserves partial text and may include provisional; failures reject with TranscriptionError.partialResult and cause.
Options on transcribe() or createSession(): language, task: "transcribe" | "translate", timestamps: "segment" | "word", and hotwords. Support is model-specific; unsupported requests reject. Current adapters expose segment timestamps for Whisper but no word timestamps.
awaitsession.startMicrophone();// Start from a user action.
// Later, when the user stops recording:
constresult=awaitsession.finish();
awaitlive.unload();
Live updates replace provisional text by utterance id (starting at "0"); isFinal commits it. Empty updates can clear a previous provisional transcript. The final result includes segments with those IDs.
Alternatively, await session.push({ samples, sampleRate }) supplies mono PCM. It acknowledges acceptance, not completed recognition or backpressure. finish() ends input and drains pending work; result() waits without ending input; cancel() stops and preserves partial results. One live session owns an instance; conflicting work rejects.
There is no default backlog limit. Set maxBufferedAudioMs on createSession() to bound pending audio; exceeding it fails with preserved transcript. Wfloat warns about growing backlog and never silently drops audio. Whisper-like models use overlapping windows; native streaming models process incrementally. Neither guarantees real-time speed on every device.
Shared microphone
createMicrophoneCapture() creates a reusable source. Attach it to STT/VAD sessions with attachMicrophone(source) before pushing or starting capture, then call source.start(). Call source.stop() to stop shared capture; finish each session separately. A session cannot mix attached-microphone and manually pushed input.
Microphone input requires HTTPS or localhost and user permission.