What does 99.9% uptime mean for inference?
AI Classified by Officially
40+ Models Chosen for Production...40+ Models Chosen for Production...40+ Models Chosen for Production...
Reliability numbers are easy to publish. What’s hard is explaining what they mean: which failure domains the architecture actually covers, whether the provider controls the infrastructure at those layers, and what happens when something breaks at 3 a.m.
Together runs inference for teams like Cursor, Decagon, Cartesia, and Yutori. We’ve been paged for most of what follows; here’s what we’ve learned.
When inference goes down, someone’s product goes down with it. GPU inference fails differently from conventional services. The hardware has failure modes; CPU infrastructure doesn't, and the systems are tuned hard for performance. Hitting 1M tokens per minute per GPU at 200 TPS, or sub-50ms TTFT on voice models with custom kernels, leaves limited slack. Adding reliability to a system like that is exponentially harder with each nine.
The useful mental model is layers, and the failure modes in each one are distinct:
This is an extract. The publication continues at the source.
Read the original at the source: https://www.together.ai/blog/99-9-uptime-for-inference
Officially imported this from Together AI’s own source and shows an extract. If you work there, claiming the profile and verifying the domain lets you choose to show the full text here.
Provenance
- Organization
- Together AI — imported from official source
- Official source
- https://www.together.ai/blog/rss.xml RSS
- Imported
- September 20, 2026 19:52
- Versions
- 1 recorded
- Identity
https://www.together.ai/blog/99-9-uptime-for-inference