How a global fintech scaled coding agent traffic with Dedicated Model Inference

Imported from official source

AI Classified by Officially

40+ Models Chosen for Production...40+ Models Chosen for Production...40+ Models Chosen for Production...

A global fintech runs its coding assistant on GLM 5.2 through Together's Dedicated Model Inference, handling spiky, engineering-hours traffic that static capacity planning couldn't keep up with.

With DMI, the customer's engineers scale endpoints, roll out models, and test changes themselves, no tickets, no waiting on Together. The result: infrastructure that moves as fast as the teams adopting it.

This customer ships financial products to millions of users across dozens of markets, and growth shows no sign of slowing. Sustaining that pace is an engineering problem before anything else, and the company's engineers lean on AI coding agents to do it.

That puts inference on the critical path of how fast the company ships, rather than inside any single customer-facing feature. The workload runs on GLM-5.2, the mixture-of-experts model built for long-horizon coding and agentic work, served on Together. Traffic follows the working day: spiky, concentrated in engineering hours, and it climbs every time another team adopts agents into its workflow.

The constraint: capacity planning couldn't keep up with adoption

The customer came to Together after running coding workloads with other inference providers, and first consolidated onto our earlier dedicated offering. That offering worked, but wasn't built for how this workload actually behaves. The coding-assistant traffic isn't steady; it's peak-load and relatively low-TPS, concentrated in engineering hours, with sharp bursts in concurrency and prompt size as more teams put agents into their daily workflow. That shape is precisely why concurrency, not raw throughput, was the design priority when the workload moved to GLM-5.2.

This is an extract. The publication continues at the source.

Read the original at the source: https://www.together.ai/blog/global-fintech-scales-coding-agent-traffic-with-dedicated-model-inference

Officially imported this from Together AI’s own source and shows an extract. If you work there, claiming the profile and verifying the domain lets you choose to show the full text here.

Provenance

Organization
Together AI — imported from official source
Official source
https://www.together.ai/blog/rss.xml RSS
Imported
September 20, 2026 19:52
Versions
1 recorded
Identity
https://www.together.ai/blog/global-fintech-scales-coding-agent-traffic-with-dedicated-...

Officially records where a publication came from, not whether it is true. Imported records are reproduced from an organization's own official source.