Open weights just hit the frontier. Read the token bill.
An open-weight model just matched the frontier on long-horizon agentic coding and MCP tool use. If you run air-gapped, the build-vs-buy math quietly flipped this week.
GLM-5.2 shipped MIT-licensed, full weights, 1M context. The numbers (vendor and third-party, so benchmark your own):
- SWE-bench Pro 62.1, up from 58.4 on 5.1.
- FrontierSWE 74.4, within a point of Opus 4.8 at 75.4.
- Terminal-Bench 2.1 climbed from 63.5 to 81.
- MCP-Atlas tool-use sits near Opus.
For on-prem that flips the calculus. Frontier-class long-horizon coding now runs inside the air gap: no egress, no per-call metering, weights you own outright. The model-ownership red line got cheaper to hold.
The trap: on-prem you don't pay per token, you pay in decode tok/s and KV cache. GLM-5.2 is token-hungry, by early reviews one of the least efficient in its class. A model that wins the leaderboard while burning 2-3x the tokens can still lose your GPU budget.
Own the weights. Then benchmark the tok/s, because the leaderboard doesn't pay your GPU bill.