What does 99.9% uptime mean for inference?

Imported from official source

AI Classified by Officially

40+ Models Chosen for Production...40+ Models Chosen for Production...40+ Models Chosen for Production...

  • The short version: Each reliability tier maps to a specific failure domain, and each one requires its own architecture to survive it. Roughly speaking: 
  • 99% means your architecture can survive node-level failures: GPU hardware faults, driver crashes, thermal events. Getting there generally takes automated health checking, node draining, and fast replica replacement within a single DC.
  •  99.9% means your architecture can survive a full data center failure. That usually means model weights deployed across two facilities, enough capacity on each side to absorb the full load, and live traffic routing to both, not a cold standby.
  • 99.99% means your architecture can survive a regional outage. That typically calls for multi-region deployment with AZ redundancy and reserved failover capacity.
  • Reliability numbers are easy to publish. What’s hard is explaining what they mean: which failure domains the architecture actually covers, whether the provider controls the infrastructure at those layers, and what happens when something breaks at 3 a.m.

    Together runs inference for teams like Cursor, Decagon, Cartesia, and Yutori. We’ve been paged for most of what follows; here’s what we’ve learned.

    When inference goes down, someone’s product goes down with it. GPU inference fails differently from conventional services. The hardware has failure modes; CPU infrastructure doesn't, and the systems are tuned hard for performance. Hitting 1M tokens per minute per GPU at 200 TPS, or sub-50ms TTFT on voice models with custom kernels, leaves limited slack. Adding reliability to a system like that is exponentially harder with each nine.

    The useful mental model is layers, and the failure modes in each one are distinct:

    This is an extract. The publication continues at the source.

    Read the original at the source: https://www.together.ai/blog/99-9-uptime-for-inference

    Officially imported this from Together AI’s own source and shows an extract. If you work there, claiming the profile and verifying the domain lets you choose to show the full text here.

    Provenance

    Organization
    Together AI — imported from official source
    Official source
    https://www.together.ai/blog/rss.xml RSS
    Imported
    September 20, 2026 19:52
    Versions
    1 recorded
    Identity
    https://www.together.ai/blog/99-9-uptime-for-inference

    Officially records where a publication came from, not whether it is true. Imported records are reproduced from an organization's own official source.