Smart Turn v3.x

Smart Turn is an open‑source semantic Voice Activity Detection (VAD) model that tells you whether a speaker has finished their turn by analysing the raw waveform, not the transcript.

Model architecture

Backbone: Whisper Tiny encoder
Head: shallow linear classifier
Params: 8M
Checkpoint: 8 MB ONNX (int8 quantized), 32MB ONNX (unquantized)

How to use

Please see the blog post and GitHub repo for more information on using the model, either standalone or with Pipecat.

Thanks

Thank you to the following organisations for contributing audio datasets:

Downloads last month: -; Downloads are not tracked for this model. How to track

Inference Providers NEW

Voice Activity Detection

This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

pipecat-ai
/

smart-turn-v3

Smart Turn v3.x

Links

Model architecture

How to use

Thanks

Datasets used to train pipecat-ai/smart-turn-v3

Smart Turn v3.x

Links

Model architecture

How to use

Thanks

Datasets used to train pipecat-ai/smart-turn-v3

Smart Turn v3.x