Accepts the same audio inputs as speech to text; arrays need sample_rate. Returns a DetectionResult with segments and stop_reason ("complete" or "cancelled"). Each segment has id, start_ms, and end_ms; return_audio=True adds an Audio clip.
Option
Meaning
speech_threshold, silence_threshold
Speech entry/exit thresholds.
min_speech_duration_ms
Minimum accepted speech duration.
min_silence_duration_ms
Silence needed to end a speech segment.
speech_padding_ms
Extra audio around detected speech.
return_audio
Include clips; default false.
on_probability
Receive probability, start_ms, and end_ms.
cancel_event
Cooperative cancellation using a threading.Event.
Current defaults: speech threshold 0.5, silence threshold 0.15 below it, minimum speech 250 ms, minimum silence 500 ms, padding 30 ms. Thresholds/scores are not interchangeable across models. Padding is clamped to available audio and previous segment boundaries to avoid overlapping clips. Times are relative to supplied audio, not wall-clock time.
Sessions process application-supplied audio through synchronous push() and support on_speech_start, on_speech_end, and on_error. One session owns the model. Your application manages capture and buffering.
Live clips are delivered through on_speech_end; the final result contains timing ranges, not every clip. Retain callback audio yourself if needed. finish() flushes the last segment, cancel() preserves completed ranges, and result() retrieves the terminal result. Silence produces no speech segments. Failures raise VadError with partial_result and notify session on_error.
Accepts the same audio inputs as speech to text. Returns a handle with result() and cancel(). Results contain segments and stopReason: "complete" | "cancelled". A segment has id, startMs, and endMs; returnAudio: true adds its audio.
Option
Meaning
speechThreshold, silenceThreshold
Speech entry/exit thresholds.
minSpeechDurationMs
Minimum accepted speech duration.
minSilenceDurationMs
Silence needed to end a speech segment.
speechPaddingMs
Extra audio around detected speech.
returnAudio
Include completed speech clips; default false.
onProbability
Receive { probability, startMs, endMs } model scores.
Current defaults: speech threshold 0.5, silence threshold 0.15 below it, minimum speech 250 ms, minimum silence 500 ms, padding 30 ms. Scores and threshold effectiveness are not interchangeable across models. Padding is clamped to available audio and previous segment boundaries to avoid overlapping clips.
awaitsession.startMicrophone();// From a user action.
// Later:
constsummary=awaitsession.finish();
Sessions also support push({ samples, sampleRate }), attachMicrophone(source), result(), and cancel(). finish() ends input and flushes the last segment; result() only waits. One session owns the model at a time.
Live onSpeechEnd delivers clips when returnAudio is enabled. The final session result retains timing ranges, not all audio clips; keep callback clips yourself if needed. Times are relative to the supplied audio/session. Silence produces no speech segments. Failures use VadError.partialResult and session onError.
See speech to text for shared microphone setup and platform-specific capture behavior.
Accepts the same audio inputs as speech to text. Returns a handle with result() and cancel(). Results contain segments and stopReason: "complete" | "cancelled". A segment has id, startMs, and endMs; returnAudio: true adds its audio.
Option
Meaning
speechThreshold, silenceThreshold
Speech entry/exit thresholds.
minSpeechDurationMs
Minimum accepted speech duration.
minSilenceDurationMs
Silence needed to end a speech segment.
speechPaddingMs
Extra audio around detected speech.
returnAudio
Include completed speech clips; default false.
onProbability
Receive { probability, startMs, endMs } model scores.
Current defaults: speech threshold 0.5, silence threshold 0.15 below it, minimum speech 250 ms, minimum silence 500 ms, padding 30 ms. Scores and threshold effectiveness are not interchangeable across models. Padding is clamped to available audio and previous segment boundaries to avoid overlapping clips.
awaitsession.startMicrophone();// From a user action.
// Later:
constsummary=awaitsession.finish();
Sessions also support push({ samples, sampleRate }), attachMicrophone(source), result(), and cancel(). finish() ends input and flushes the last segment; result() only waits. One session owns the model at a time.
Live onSpeechEnd delivers clips when returnAudio is enabled. The final session result retains timing ranges, not all audio clips; keep callback clips yourself if needed. Times are relative to the supplied audio/session. Silence produces no speech segments. Failures use VadError.partialResult and session onError.
See speech to text for shared microphone setup and platform-specific capture behavior.