Together AI
together.ai
Imported from official source
Together AI — publications from its own official source.
- Type
- Company
- Scope
- US · national
- Website
- together.ai
- Feed
- Atom
Publications 27
-
How to train your own Jev for $17
We just launched our own Jev-like classifier, together/Tev1-4B-experimental, on top of Qwen3.5 4B on Together’s serverless platform. In this blog post we’ll show you how to fine-tune your own version!
-
Canary rollouts: upgrade models in production without downtime
A hard model swap exposes every user at once, and rolling back means cold-starting the old deployment under pressure. Here's how staged traffic ramps, metric gates, and automatic rollback work on d...
-
How a global fintech scaled coding agent traffic with Dedicated Model Inference
Inside a global bank's shift to self-serve dedicated inference: how Together's DMI gave engineering teams direct control over scaling, models, and testing.
-
Migrating from closed to open source models, Together
Moving from closed to open source models can take weeks, not years. A five-stage playbook: discover, evaluate, adapt, decide, and production.
-
Together AI expands fine-tuning service with more models, live metrics, and finer controls
Together Fine-Tuning adds the latest open-weight models, live experiment tracking, Expert LoRA, early stopping, tokenized dataset previews, pre-flight validation, and lower training prices on selec...
-
To Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72!
We ported ThunderKittens to NVIDIA's Vera Rubin NVL72 and rebuilt our NVFP4 GEMM around the new hardware, taking it from 42% of roofline to over 22 PFLOPS — competitive with cuBLAS and CuTe DSL. He...
-
Introducing preemptible compute: the same compute, half the price
Together GPU Clusters now supports preemptible compute: the same GPU capacity at a flat 50% of the on-demand rate, with a five-minute drain window.
-
The Open Source AI Stack
A deep dive into the open model AI stack — model, inference, gateways and routers, harness, and tools — and how keeping each layer independent lets you swap in a new open model in minutes instead o...
-
GLM-5.3 vs. Claude Fable 5 on DeepSWE: Cost, Coding, and Routing
We ran 904 DeepSWE rollouts on GLM-5.3 and Claude Fable 5. A tie on pass@1, but GLM-5.3 wins pass@4 and costs 5.4x less: \$3.99 per rollout vs. \$21.63.
-
GLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing
We ran 904 DeepSWE rollouts on GLM-5.3 and GPT-5.6 Sol. Sol leads pass@1 by 3.7 points; GLM-5.3 wins pass@4 at half the cost, and a GLM-first cascade hits 85.9%.
-
A/B test models in production
Shadow traffic proves a candidate is operationally sound. It can't tell you if users like it better. Run the split at the endpoint instead of in your app code.
-
DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing
We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and Claude Fable 5. Fable leads pass@1 at 90x the cost; Pro wins pass@4, and a Pro-first cascade hits 82.7%.
-
Kimi K3: the complete developer guide
Kimi K3 is the first open 3T-class model. See how it benchmarks, what it costs, and how to call it on the Together AI API, with copy-paste code examples.
-
ThunderAgent: 2x Faster Agentic Inference for Synthetic Data Generation at Scale
ThunderAgent is a program-aware scheduler for agentic inference. By treating each agent workflow as a schedulable program, it eliminates KV cache thrashing to deliver more than 2x single-node throu...
-
Together AI announces strategic partnership with Moonshot AI to natively serve Kimi models
Together AI partners with Moonshot AI to natively serve Kimi models, starting with the 2.8T parameter Kimi K3, with day zero access and post-training.
-
The production platform for open-weight AI inference
Run open models in production with full control over performance, cost, and quality. Deploy in minutes, roll out safely, and scale to your SLOs.
-
Together AI and Y Combinator partner to launch the first dedicated GPU cluster for the YC community
No more two-year compute contracts. Together AI and YC just gave YC startups a faster way to get GPUs.
-
What does 99.9% uptime mean for inference?
Reliability numbers are easy to publish. We break down what 99%, 99.9%, and 99.99% uptime actually require, the failure domains each tier has to survive, and the questions to ask any inference prov...
-
New in Together GPU Clusters: Reliability and control for production GPU clusters
See how Together AI is improving production GPU clusters with passive health checks, node repair, stronger Slurm reliability, OIDC, and startup scripts.
-
Together AI brings Thinking Machines Lab’s new model Inkling on day 0
Together AI offers day zero access to Inkling, Thinking Machines Lab's multimodal mixture-of-experts model for text, image, and audio reasoning.
-
Announcing our $800M Series C to accelerate the shift to open-source AI
We raised $800M to accelerate the shift to open-source AI. Here's why the economics of closed models don't scale, and what we're building next.
-
Together AI at ICML 2026: frontier research across the full stack
Nine papers at ICML 2026 across the full stack. The research that becomes the Together platform. Find us at booth B714 in Seoul.
-
ParallelKernelBench: Frontier LLMs can't write fast multi-GPU kernels (yet)
ParallelKernelBench tests whether LLMs can write fast multi-GPU CUDA kernels across 87 real workloads. The best model solves under a third, but a few generated kernels beat any public implementation.
-
Kimi K2.7 Code vs Claude Fable 5: Landing pages that cost 94% less
We generated 12 landing pages with Kimi K2.7 Code and Claude Fable 5. Kimi cost 94% less and scored within a few points on every page. Here's what actually moved the needle.
-
Building trust in enterprise AI: Together AI earns ISO 27001:2022 certification
Together AI has earned ISO 27001:2022 certification, validating our commitment to enterprise-grade security for production AI workloads.
-
Serving MiniMax-M3 for efficient inference: Unlocking 1M-Token Context and Multimodality Without Regrets
How Together served MiniMax-M3 efficiently with KV-block-major sparse attention, paged MSA decode, optimized index scoring, and a Rust-based multimodal gateway.
-
How Together AI built the world’s fastest speech-to-text stack
Together AI built the fastest speech-to-text stack on Artificial Analysis by treating ASR as a full-path systems problem, not just a GPU inference problem.