/

/

Open-Weight Models Can Do Physical AI, but Only With Specialized Harnesses

Open-Weight Models Can Do Physical AI, but Only With Specialized Harnesses

Open-Weight Models Can Do Physical AI, but Only With Specialized Harnesses

Date Published

Contributors

Share

Date Published

Contributors

Share

Last month we showed that the harness, not the model, decided whether a frontier AI got the physics right. That raised a follow-up question: if the harness does so much of the work, do you still need an expensive frontier model? We ran a set of sealed DyadBench physics problems to find out. The grader doesn't read code; it simulates whatever model the agent commits and checks it against a sealed reference.


In a physics harness, the open models are close

Scoreboard of six models in the Dyad Agent: pass rate, cost per trial and minutes per trial

With every model in the same harness, the Dyad Agent, the field is tight. Opus 5.5 still leads, but three open-weight models land within a dozen points of it. The frontier is no longer far ahead once the harness is doing its job.

Model

Effort

Passed

Weighted score

Cost per pass

Model calls per trial

Failed tool calls

Opus 5.5

max

26/27

0.944

$3.85

40.0

45

GLM 5.3 Flash

max

24/27

0.796

$0.27

48.8

203

DeepSeek V4.1 Flash

max

23/27

0.759

$0.40

71.8

239

Kimi K3

max

23/27

0.778

$3.04

35.7

100

Qwen 3.8 Max

xhigh

21/27

0.537

$2.90

45.0

238

Grok 4.7

xhigh

20/27

0.537

$3.11

45.9

168

And they are an order of magnitude cheaper

Cost per trial against pass rate

This is the chart that matters for a budget. GLM 5.3 Flash gets most of the way to Opus for a small fraction of the cost per trial, which makes it the value pick by a wide margin. The other half of the picture is just as useful: open weights do not automatically mean cheap. Two of the open models cost nearly as much as Opus while scoring lower.

Take away the harness and the cheap models fall apart

Pass rate in the Dyad Agent against Claude Code for three models

To check whether the harness was really the reason, we moved models into a general-purpose coding agent, Claude Code, with the same Dyad tools and docs available. The frontier model barely notices: Opus 5.5 does as well or better. The cheaper models do notice. GLM 5.3 Flash and Grok 4.7 both lose about half their passes, and they get slower and, for Grok, more expensive too.

That is the result in one picture. A frontier model brings enough discipline to succeed in a generic loop. A cheaper model needs the harness to supply it.

The remaining gap is narrow and subtle

Trials passed out of 3, per model and problem, hardest problem first

Inside the Dyad Agent, most problems are a near-sweep for everyone. What separates the frontier model from the rest is concentrated in one hard problem, a relativistic electron in a laser pulse. Notably, across every trial on that problem, no model weakened its own checks to get there, the failure mode that hurt the generic harness last month.


Graded trajectory for a passing and a failing model

The misses look like this: a model that is right at the end but wrong along the way. The open models' errors were small choices their checks couldn't see, like a sign or a phase. The frontier model's edge is noticing when a check is blind and building a better one. The cheaper models can do that too, just less consistently.

What it means

The harness is still the multiplier, and the multiplier is largest on the models you can afford. Put an inexpensive open-weight model in a harness built for physics and it gets close to the frontier at a fraction of the price. Open weights also let teams run the model where their data lives. Physical AI no longer requires a frontier-priced model. It requires the right harness.

Method in brief: nine sealed DyadBench problems (seven shared ones for the harness-swap chart), three trials each, graded by simulating the committed model against a sealed reference. Opus 5.5 ran on the Anthropic API, the open models on Fireworks, Grok 4.7 on Amazon Bedrock at effort xhigh; costs are list prices as of 2026-10-02. Three trials per problem is a small sample, so single-trial differences are noise. GLM's Claude Code cost is unknown because cached input was not reported. The full write-up, with every table and caveat, is in the companion post.




Authors

Dr. Chris Rackauckas is the VP of Modeling and Simulation at JuliaHub, the Director of Scientific Research at Pumas-AI, Co-PI of the Julia Lab at MIT, and the lead developer of the SciML Open Source Software Organization. He is the lead developer of the Pumas project and has received a top presentation award at every ACoP in the last 3 years for improving methods for uncertainty quantification, automated GPU acceleration of nonlinear mixed effects modeling (NLME), and machine learning assisted construction of NLME models with DeepNLME. For these achievements, Chris received the Emerging Scientist award from ISoP.

Venkateshprasad Bhat is a Software Engineer at JuliaHub, where he works on DyadAgent, an agentic harness that brings AI agents to modeling and simulation. A contributor to SciML, he's interested in the many ways AI can accelerate scientific discovery.

Authors

Dr. Chris Rackauckas is the VP of Modeling and Simulation at JuliaHub, the Director of Scientific Research at Pumas-AI, Co-PI of the Julia Lab at MIT, and the lead developer of the SciML Open Source Software Organization. He is the lead developer of the Pumas project and has received a top presentation award at every ACoP in the last 3 years for improving methods for uncertainty quantification, automated GPU acceleration of nonlinear mixed effects modeling (NLME), and machine learning assisted construction of NLME models with DeepNLME. For these achievements, Chris received the Emerging Scientist award from ISoP.

Venkateshprasad Bhat is a Software Engineer at JuliaHub, where he works on DyadAgent, an agentic harness that brings AI agents to modeling and simulation. A contributor to SciML, he's interested in the many ways AI can accelerate scientific discovery.

Authors

Dr. Chris Rackauckas is the VP of Modeling and Simulation at JuliaHub, the Director of Scientific Research at Pumas-AI, Co-PI of the Julia Lab at MIT, and the lead developer of the SciML Open Source Software Organization. He is the lead developer of the Pumas project and has received a top presentation award at every ACoP in the last 3 years for improving methods for uncertainty quantification, automated GPU acceleration of nonlinear mixed effects modeling (NLME), and machine learning assisted construction of NLME models with DeepNLME. For these achievements, Chris received the Emerging Scientist award from ISoP.

Venkateshprasad Bhat is a Software Engineer at JuliaHub, where he works on DyadAgent, an agentic harness that brings AI agents to modeling and simulation. A contributor to SciML, he's interested in the many ways AI can accelerate scientific discovery.

Learn about Dyad

Get Dyad Studio – Download and install the IDE to start building hardware like software.

Read the Dyad Documentation – Dive into the language, tools, and workflow.

Join the Dyad Community – Connect with fellow engineers, ask questions, and share ideas.

Contact Us

Want to get enterprise support, schedule a demo, or learn about how we can help build a custom solution? We are here to help.