GLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

Together AI Version 1 original current

Imported from official source

We ran 904 DeepSWE rollouts on GLM-5.3 and GPT-5.6 Sol. Sol leads pass@1 by 3.7 points; GLM-5.3 wins pass@4 at half the cost, and a GLM-first cascade hits 85.9%.

This version

Version
1 of 1
Recorded
September 20, 2026 19:52
Change
Initial
Content hash
b1c78ae2e8452bb623257b099c8ad8c0
All versions
Revision history

Officially records where a publication came from, not whether it is true. Imported records are reproduced from an organization's own official source.