NVIDIA Canary

STT.ai does not run this model — this page is reference information about it. Skryf aan met die modelle wat ons wel hardloop →
3.5%
WER
4
Tale
45.0x
Spoed
CC-BY-4.0
Lisensie

Aangaande NVIDIA Canary

NAVIDIA - Kanariese is'n 1B parametermodel wat by Engels, Duits, Frans en Spaanse transkripsie uitmunt.'n Vinnige Verkragkerkodeerder met'n transformator dekodeerder en ondersteun outomatiese taalopsporing en vertaling.

Tale wat ondersteun word deur NVIDIA Canary

Model InligtingName
  • VerskafferNVIDIA
  • Argitektuur-
  • LisensieCC-BY-4.0
  • OpgedateerCancel: Meeting NameMeetingMar 2026

Vrae wat dikwels gevra word

NVIDIA Canary is a speech-to-text model by NVIDIA. STT.ai does not currently run NVIDIA Canary — this page is reference information about the model. For transcription on STT.ai, the model picker offers the Whisper family (Turbo, Large V3, Medium, Small).

On standard benchmarks, NVIDIA Canary achieves around 3.5% Word Error Rate. Real-world accuracy depends on audio quality, accent, and language; for noisy or accented recordings, expect a few percentage points higher WER.

NVIDIA Canary is not available on STT.ai. Its own licence terms govern running it yourself. STT.ai's free tier covers the Whisper models we do run — 600 minutes to start, then a monthly free top-up.

NVIDIA Canary is released under CC-BY-4.0, a permissive open-source license. You can self-host NVIDIA Canary on your own hardware or use our hosted version — both are commercially usable.

NVIDIA Canary supports 4 languages. Auto-detection picks the right language for most audio; you can also specify it manually for a small accuracy lift.

NVIDIA Canary processes audio at about 45.0x real-time on our GPUs. A 1-hour audio file finishes in under 1 minutes; longer files queue and notify by email when done.

NVIDIA Canary has 1B parameters. Larger models tend to be more accurate but slower; STT.ai hosts NVIDIA Canary on GPU so the parameter count doesn't affect your client-side performance.

NVIDIA Canary accepts every format STT.ai supports — MP3, WAV, M4A, FLAC, OGG, MP4, MKV, MOV, WebM, AVI, and others. Output as TXT, SRT, VTT, DOCX, JSON, or PDF.

Yes. Speaker diarization runs alongside NVIDIA Canary for every transcription — each speaker is labeled and you can rename them in the editor afterwards.

Yes. NVIDIA Canary runs in our managed environment — audio is processed and deleted by default and never used for training without explicit opt-in. Pro plans add client-side encryption for transcripts at rest.

Use the compare-stt tool to run NVIDIA Canary against any other supported model on the same audio — you'll see WER, segment count, speaker labels, and confidence scores side-by-side. The NVIDIA Canary vs Whisper Large V3 comparison is the most commonly run.

Ja. Spesifiseer "nvidia-canary" as die model parameter op die /v1/aanteken eindepunt. Python en Node.js SDKs sluit 8800 voorbeelde in. Vry 'nPI-vlak sluit 100 minute/month in.

Yes. Because NVIDIA Canary is CC-BY-4.0-licensed, you can self-host it. STT.ai's open-source page lists the project repo and weights. Most production teams use our hosted version to skip GPU procurement, model swaps, and ops.