How Together AI built the world’s fastest speech-to-text stack
AI Classified by Officially
NVIDIA TensorRT multi-profile engines, conditional NVIDIA CUDA graphs, evented I/O, shared memory, and the Python GC fix behind Together’s ASR latency results.
40+ Models Chosen for Production...40+ Models Chosen for Production...40+ Models Chosen for Production...
A 1M-token text prompt can fit the entire Harry Potter series and still only weigh around 5 MB. That scale sounds enormous, but the input itself is compact. Text also arrives almost ready for inference: tokenize it, batch it, and move it through the model.
Audio changes the shape of the problem. The same Harry Potter corpus as audiobooks is 5 to 10 GB, roughly three orders of magnitude larger than the text. Before any of it reaches the GPU, the server has to decode the container, resample, filter noise, run VAD, segment speech, and compute audio features.
The model side flips too. LLMs these days have hundreds of billions or trillions of parameters, so serving work naturally concentrates inside the GPU: quantization, KV cache, attention kernels, batching, and parallelism. Speech-to-text models are much smaller, often in the hundreds of millions to low billions of parameters, so the surrounding data path matters much more.
That makes ASR serving a full-path systems problem spanning GPU execution, CPU preprocessing, memory movement, transport, connection scheduling, and runtime behavior. The same stack also has to serve two different regimes: offline transcription, where throughput matters most, and streaming transcription, where latency and jitter dominate.
Together’s ASR stack serves the two lowest-latency speech-to-text models ranked by Artificial Analysis: NVIDIA’s Parakeet-TDT 0.6B v3 and OpenAI’s Whisper Large v3. The faster of the two, NVIDIA Parakeet-TDT 0.6B v3, can transcribe roughly 20 hours of speech, about the runtime of the Harry Potter film franchise, in under 10 seconds.
This is an extract. The publication continues at the source.
Read the original at the source: https://www.together.ai/blog/how-together-ai-built-the-worlds-fastest-speech-to-text-stack
Officially imported this from Together AI’s own source and shows an extract. If you work there, claiming the profile and verifying the domain lets you choose to show the full text here.
Provenance
- Organization
- Together AI — imported from official source
- Official source
- https://www.together.ai/blog/rss.xml RSS
- Imported
- September 20, 2026 19:52
- Versions
- 1 recorded
- Identity
https://www.together.ai/blog/how-together-ai-built-the-worlds-fastest-speech-to-text-stack