NVIDIA Parakeet
Fast, locally runnable speech-to-text models for transcribing interviews, meetings and recordings with punctuation and timestamps
huggingface.coParakeet is NVIDIA’s family of automatic speech-recognition models for turning audio into text. For researchers, the practical starting point is Parakeet TDT 0.6B v3: a downloadable 600-million-parameter model that adds punctuation and capitalisation, returns word- and segment-level timestamps, and can process long recordings. That makes it useful for transcribing interviews, focus groups, meetings, lectures and audiovisual archives without sending sensitive recordings to a transcription service. The checkpoint is released under CC BY 4.0 and works through Hugging Face Transformers or NVIDIA NeMo. Its central limitation for Indian research is language coverage: v3 supports English and 24 other European languages, but no Indic languages, so it is most useful for English-language material in India. It is optimized for Linux systems with NVIDIA GPUs; test accuracy on local accents, noisy field recordings and specialist vocabulary before relying on a transcript.
Details
- Kind
- Tool
- Run by
- Commercial
- Topics
- Where
- Global, India
- Licence
- Open licence
- Cost
- Free
- Licence terms
- CC BY 4.0