Smart Turn v3.x
Smart Turn is an open‑source semantic Voice Activity Detection (VAD) model that tells you whether a speaker has finished their turn by analysing the raw waveform, not the transcript.
Links
- Blog post: Smart Turn v3
- GitHub repo with training and inference code, and more information
- Datasets
Model architecture
- Backbone: Whisper Tiny encoder
- Head: shallow linear classifier
- Params: 8M
- Checkpoint: 8 MB ONNX (int8 quantized), 32MB ONNX (unquantized)
How to use
Please see the blog post and GitHub repo for more information on using the model, either standalone or with Pipecat.
Thanks
Thank you to the following organisations for contributing audio datasets:
Inference Providers
NEW
This model isn't deployed by any Inference Provider.
🙋
Ask for provider support