Last month we showed that the harness, not the model, decided whether a frontier AI got the physics right. That raised a follow-up question: if the harness does so much of the work, do you still need an expensive frontier model? We ran a set of sealed DyadBench physics problems to find out. The grader doesn't read code; it simulates whatever model the agent commits and checks it against a sealed reference.
In a physics harness, the open models are close

With every model in the same harness, the Dyad Agent, the field is tight. Opus 5.5 still leads, but three open-weight models land within a dozen points of it. The frontier is no longer far ahead once the harness is doing its job.
Model | Effort | Passed | Weighted score | Cost per pass | Model calls per trial | Failed tool calls |
|---|---|---|---|---|---|---|
Opus 5.5 | max | 26/27 | 0.944 | $3.85 | 40.0 | 45 |
GLM 5.3 Flash | max | 24/27 | 0.796 | $0.27 | 48.8 | 203 |
DeepSeek V4.1 Flash | max | 23/27 | 0.759 | $0.40 | 71.8 | 239 |
Kimi K3 | max | 23/27 | 0.778 | $3.04 | 35.7 | 100 |
Qwen 3.8 Max | xhigh | 21/27 | 0.537 | $2.90 | 45.0 | 238 |
Grok 4.7 | xhigh | 20/27 | 0.537 | $3.11 | 45.9 | 168 |
And they are an order of magnitude cheaper

This is the chart that matters for a budget. GLM 5.3 Flash gets most of the way to Opus for a small fraction of the cost per trial, which makes it the value pick by a wide margin. The other half of the picture is just as useful: open weights do not automatically mean cheap. Two of the open models cost nearly as much as Opus while scoring lower.
Take away the harness and the cheap models fall apart

To check whether the harness was really the reason, we moved models into a general-purpose coding agent, Claude Code, with the same Dyad tools and docs available. The frontier model barely notices: Opus 5.5 does as well or better. The cheaper models do notice. GLM 5.3 Flash and Grok 4.7 both lose about half their passes, and they get slower and, for Grok, more expensive too.
That is the result in one picture. A frontier model brings enough discipline to succeed in a generic loop. A cheaper model needs the harness to supply it.
The remaining gap is narrow and subtle

Inside the Dyad Agent, most problems are a near-sweep for everyone. What separates the frontier model from the rest is concentrated in one hard problem, a relativistic electron in a laser pulse. Notably, across every trial on that problem, no model weakened its own checks to get there, the failure mode that hurt the generic harness last month.

The misses look like this: a model that is right at the end but wrong along the way. The open models' errors were small choices their checks couldn't see, like a sign or a phase. The frontier model's edge is noticing when a check is blind and building a better one. The cheaper models can do that too, just less consistently.
What it means
The harness is still the multiplier, and the multiplier is largest on the models you can afford. Put an inexpensive open-weight model in a harness built for physics and it gets close to the frontier at a fraction of the price. Open weights also let teams run the model where their data lives. Physical AI no longer requires a frontier-priced model. It requires the right harness.
Method in brief: nine sealed DyadBench problems (seven shared ones for the harness-swap chart), three trials each, graded by simulating the committed model against a sealed reference. Opus 5.5 ran on the Anthropic API, the open models on Fireworks, Grok 4.7 on Amazon Bedrock at effort xhigh; costs are list prices as of 2026-10-02. Three trials per problem is a small sample, so single-trial differences are noise. GLM's Claude Code cost is unknown because cached input was not reported. The full write-up, with every table and caveat, is in the companion post.









