GLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

Imported from official source

AI Classified by Officially

Sol wins the first try, GLM-5.3 wins the rest at half the price, and the cascade beats both: 85.9% at \$6.61 a task.

40+ Models Chosen for Production...40+ Models Chosen for Production...40+ Models Chosen for Production...

Don't pick one. Run GLM-5.3 first, escalate to GPT-5.6 Sol when the tests fail. That cascade solves 85.9% of DeepSWE tasks at \$6.61 each. Sol alone solves 72.7% at \$8.37. Thirteen points better, 21% cheaper.

  • Sol wins the single shot, narrowly. 72.7% pass@1 against GLM-5.3's 69.0%, a 3.7 point gap that sits inside a couple of standard deviations.
  • GLM-5.3 wins every retry after that. It ties Sol at pass@2 (81.1 vs. 81.0) and leads pass@4 (87.6% vs. 85.8%).
  • The price gap is 2.1x. \$3.99 per rollout against \$8.37. Per \$100 spent, GLM-5.3 solves 17 tasks and Sol solves 9.
  • Sol is faster and steadier: 19 minutes and 61 steps per rollout against GLM's 35 and 124, with 61 tasks solved four for four against 48.
  • GLM-5.3's failures are cleaner. It breaks tests that already passed in 11% of its failures, against 20% for Sol. Gate Sol's diffs on regressions.
  • The two diverge (0.43 per-task correlation) and cover 106 of 113 tasks between them, which is what makes the cascade work.
  • We ran GLM-5.3 (max) against GPT-5.6 Sol (max) on all 113 DeepSWE tasks, four trials each, from the published per-trial records: 904 rollouts in total, 452 per side. Sol is the precision flagship. GLM-5.3 is the open-weight challenger that closed the gap. Every figure below comes from this run, so it can differ from other public GLM-5.3 vs. GPT-5.6 Sol scorecards.

    GPT-5.6 Sol still holds the single-shot crown on DeepSWE, a benchmark that tests a model's software engineering ability across many task types and programming languages. GLM-5.3 arrives less than four points behind it at half the price and pulls ahead the moment you allow more than one attempt. This is the closest the open tier has come to the frontier, and the question worth answering is what Sol's remaining premium actually buys.

    This is an extract. The publication continues at the source.

    Read the original at the source: https://www.together.ai/blog/glm-5-3-vs-gpt-5-6-sol-on-deepswe-cost-coding-and-routing

    Officially imported this from Together AI’s own source and shows an extract. If you work there, claiming the profile and verifying the domain lets you choose to show the full text here.

    Provenance

    Organization
    Together AI — imported from official source
    Official source
    https://www.together.ai/blog/rss.xml RSS
    Imported
    September 20, 2026 19:52
    Versions
    1 recorded
    Identity
    https://www.together.ai/blog/glm-5-3-vs-gpt-5-6-sol-on-deepswe-cost-coding-and-routing

    Officially records where a publication came from, not whether it is true. Imported records are reproduced from an organization's own official source.