Introducing preemptible compute: the same compute, half the price
AI Classified by Officially
Now in public preview for Together GPU Clusters: interruptible GPU compute, with a five-minute drain window and automatic refill toward your target.
40+ Models Chosen for Production...40+ Models Chosen for Production...40+ Models Chosen for Production...
Today we're announcing the public preview of preemptible compute for Together GPU Clusters, available on Kubernetes clusters in all regions. Preemptible nodes give teams a lower-cost way to run interruption-tolerant work — short experiments, inference bursts, batch jobs — on the same GPU infrastructure they already use, billed sub-hourly at a flat 50% of the on-demand rate. You can add preemptible capacity to a new or existing cluster starting today.
Preemptible compute adds a second compute type to Together GPU Clusters. Standard nodes are fulfilled synchronously and are never preempted. Preemptible nodes use the same NVIDIA accelerated compute, draw from un-used capacity, and can be reclaimed when that capacity is needed elsewhere.
Preemptible nodes are billed at a flat 50% of the on-demand rate. The rate remains fixed rather than moving with a spot market.
When a node is reclaimed, the cluster follows a five-minute drain (maximum 5 minutes) sequence:
TogetherPreempted Kubernetes event fires, and your pods receive SIGTERM.terminationGracePeriodSeconds) to checkpoint and exit.The cluster retains its preemptible target after the node is removed and automatically refills toward that target as capacity becomes available. You do not need to request replacement capacity.
Billing is sub-hourly, with usage metered every one to two minutes. A node that runs for 12 minutes is billed for approximately 12 minutes.
This is an extract. The publication continues at the source.
Read the original at the source: https://www.together.ai/blog/introducing-preemptible-compute-the-same-compute-half-the-price
Officially imported this from Together AI’s own source and shows an extract. If you work there, claiming the profile and verifying the domain lets you choose to show the full text here.
Provenance
- Organization
- Together AI — imported from official source
- Official source
- https://www.together.ai/blog/rss.xml RSS
- Imported
- September 20, 2026 19:52
- Versions
- 1 recorded
- Identity
https://www.together.ai/blog/introducing-preemptible-compute-the-same-compute-half-the-...