How the Dyad Harness Turns Silent Failure into Scientific Rigor
One frontier model, four sealed physics problems, two agentic loops. Everything is pinned except one thing: the harness, Dyad Agent or Claude Code.
A physical model is written in code, so building one looks like a software task. The natural reflex is to treat it like one: hand it to a general coding agent, and if the result falls short, switch to a stronger model. But modeling and simulation isn't ordinary software. In this kind of work, the loop the model runs in matters more than which model you pick. That loop includes the tools the model is given, the documentation it reads, and the order in which it has to check its own work. We call that loop the harness.
In our earlier model study, we kept the harness fixed and tried four different frontier models. The best and worst scores differed by 0.162 on a difficulty-weighted scale from zero to one. This report does the opposite: it keeps the model fixed and swaps the harness. The gap more than doubles, to 0.366.
Both harnesses evaluated in this study are production systems in real use: the Dyad Agent and stock Claude Code. We ran each one on four sealed physics problems, twelve trials per problem. Everything else was held constant: the same frontier model, the same libraries, the same documentation. Grading was out of reach of both agents. Each committed model was simulated and compared against known reference trajectories.
We recorded the full trace of every run and analyzed them with tools we built for the purpose. Two behaviors kept showing up:
When a failing check was one the agent wrote itself, the general-purpose loop weakened the check or committed a guess.
When the target was a physical invariant the agent couldn't edit, the same model did the physics correctly.
These failures are silent and not obvious to the user. The code compiles, the agent's own tests pass, and the physics is still wrong. That silent failure, more than our weighted score, is what the rest of this report is about.
The numbers first:
Dyad Agent 0.899weighted score $4.20 / trial · 18.4 min | Claude Code 0.533weighted score $4.65 / trial · 20.2 min |
|---|
Same model, two harnesses. Twelve trials a side over the four shared problems, nearly identical spend and wall-clock: 0.899 inside the Dyad harness, 0.533 in Claude Code. The harness is the experiment.
01 · How we ran the experiment
The four problems are the sealed core of the earlier model study, ordered here by difficulty. P1, constitutive consistency: a constant-property assumption that silently violates a conservation law. P2, steady-state linearization: repair a broken model, find its steady states, derive a reduced one. P3, constrained consistency: the same physics as P1 with the shortcut forbidden, so the density law must be derived, not guessed. P4, relativistic dynamics: a charged particle whose initial state must sit exactly on a relativistic invariant. (P5, the long-horizon HL-20 vehicle, was out of scope for this run.)
Every trial ran inside the evaluation bench we built for these studies. The bench pins a trial's full configuration - agent version, model, provider; here claude-opus-4-8 at xhigh reasoning effort on both sides - launches the agent in an isolated container with a fresh workspace, and records the complete trace as it runs: every tool call, every message, every compile and simulation. A trial ends when the agent commits a model it considers finished; the bench archives the workspace alongside the trace. Twelve trials per problem per harness makes 96 runs in all.
The instrument, replaying a real trial. Each trial runs pinned and isolated while the bench records its full trace. When the agent commits, the run leaves its hands: the backend archives trace and workspace, the grader simulates the committed model against the sealed reference, and an analysis agent spins up to replay the trace, flag where the physics broke, and classify every verification episode. Shown here is a real relativistic-electron trial from section 05 - the Dyad Agent's pass - with every tick, curve, and classification drawn from its actual trace. Every figure in this report is drawn from that output.
Grading is mechanical. The model each agent commits is simulated on the sealed scenario, and every graded variable's trajectory is compared point by point against the sealed reference; a trial passes only if the average relative trajectory error stays under the grader's fixed tolerance. A problem's score is the fraction of its trials that pass, and the headline number folds the four problem scores through the same difficulty weighting as the earlier study, the hardest problems counting the most. The errors quoted in the case studies below - 0.0009 and 0.0001 for the passes, 0.21 and 0.64 for the failures - are this metric.
The traces are where the rest of this report comes from - and reading 96 modeling-and-simulation transcripts is itself an instrumented task, so the bench's analysis layer is built for it. It replays any run and exposes what the runner recorded: every check the agent ran with the computation that produced it, the committed model's graded trajectories beside the sealed reference, and a flag on the point where the physics went wrong, so a bad trajectory can be traced back to the call that introduced it. The analysis agent works through every trace this way and hands back a corpus of verification episodes - each check, what it was aimed at, what it found, and what the agent changed next. Two behaviors kept emerging from that corpus, and they reduce to two axes: the target of the check - an external invariant the agent cannot edit, or something it authored itself - and the response when a check failed - change the model, or change the check. Every episode is tallied against them, and the case-study figures below are drawn from the same traces, call by call.
One objection is worth answering before the results: was Claude Code under-equipped? It was not. It had the documentation, the compiler, the same token budget, and the same twelve trials per problem. With the model, problems, equipment, and grading all pinned, what varies between the two columns of every figure that follows is the loop - nothing else.
02 · What the same model scored in each loop
Dyad Agent passes 11 of 12 trials; Claude Code passes 7 of 12. Weighted for difficulty that is 0.899 against 0.533 - a 0.366-point gap, more than double the 0.162 that separates the best and worst frontier models in the earlier model study. Per problem, the shape is sharper than the aggregate:
Where the gap lives. Both harnesses hold 1.0 on P1 and P2. Claude Code collapses to 0.333 on the constrained problem and 0.000 - every trial - on relativistic dynamics; the Dyad Agent's only miss is one of three trials on P4 (0.667). On P2, note, Claude Code is cheaper and faster.
That last note is the honest shape of the result. Where the physics is within reach of general coding discipline, the general agent is a fine choice - on P2 both pass and Claude Code does it for less. The gap opens exactly where the physics pushes back, and there it is not a gap in degree: it is pass versus zero.
03 · The mechanism: whose check is it?
A general-purpose coding agent is trained on a loop that works: write code, run the tests, make them green, ship. The loop is sound in software because the tests are close to the specification - a green suite is real evidence. Modeling and simulation breaks that assumption in a specific way. The characteristic failures of the domain are silent. A model can compile, the solver can return retcode: Success, and the trajectory it produces can be physically impossible. Success reports that the integrator did not crash; it says nothing about whether mass was conserved, whether the initial state was consistent, or whether a correlation was evaluated inside its stated range.
The discipline that catches these failures is not new. Scientific computing has had a verification-and-validation practice for decades, and a general coding agent has no reason to perform any of it: inspect structure before integrating; verify the solution and not merely the run; check conservation and limiting cases; validate against an independent oracle rather than a self-authored test; audit the assumptions the component library has baked in. The distinction that organizes all of these is the target of the check - an external one the agent cannot edit, or an internal one it can.
This is the mechanism the traces show. When the objective is a check the agent controls, a capable model under pressure has two low-cost moves available - weaken the check, or, where it cannot tell a correct formulation from a plausible one, commit a guess. When the objective is a physical invariant the agent did not author and cannot relax, neither move is available, and the same model does the physics. Both responses recur across the classified transcripts; they are not one bad trial. The two case studies that follow present, for each behavior, the single trial where it is most legible, call by call - and section 06 shows the recurrence across every classified trial.
04 · Case study (P3, constrained consistency): derive the closure, or guess it
The clearest instance of committing a guess is on the constrained problem, which forbids the shortcut: the density law must respond enough for mass conservation to hold under a stiff pressure excursion. Both agents checked their work extensively; the difference is not how much they checked, but what the checks were aimed at.
The whole run, call by call. Every tool call in both trials, in order; lane length is proportional to wall-clock. The Dyad Agent reads, derives the closure, and verifies it against an independent integration and the constitutive law - done in 11.8 minutes and 27 calls. Claude Code spends 26.3 minutes and 54 calls hunting for a density law through the reference medium, patches a negative pressure, and commits the secant closure in its final minutes - its last check is structural, not physical. One trial per side.
The two agents committed different physics, and the difference is exactly one closure:
Dyad Agent - exact integral of the constitutive law · pass
Claude Code - first-order secant guess · this trial fails
Both closures are algebraic; the difference is where they came from. The Dyad Agent derived its density law - the integration of the constitutive relation is right there as a comment in the committed code, because the loop required a verified derivation before it would simulate. Claude Code guessed a linearization that looks similar and is cheaper to write - and with β reaching 1.7×10⁹ Pa, the secant term under-responds precisely where the excursion is largest:

As graded. The derived closure (indigo) sits on the ground truth - peak 134 MPa, average error 0.0009. The guessed closure peaks at 171 MPa, a 28% overshoot concentrated exactly where the problem is hard. Average trajectory error 0.21: fail.
The transcript shows Claude Code did explore mass-conservation variants - in throwaway /tmp scripts whose scratch runs returned peaks anywhere from 165 to 210 MPa. Unable to discriminate between candidates against anything external to its own assumptions, it committed the wrong one. Exploration without an instrument degenerates into guessing, and a guess is what gets graded: 0.333 on the problem, one pass in three.
05 · Case study (P4, relativistic dynamics): fix the model, or relax the check
The clearest instance of weakening the check is on the relativistic problem, where both harnesses entered the same trap: overriding the initial longitudinal velocity without recomputing the time component leaves the initial four-velocity off the mass shell - uμuμ ≈ 0.62 instead of 1. Same model, same latent error. What differed is the response once the error was visible.
The whole run, call by call. Every tool call in both trials, in order; lane length is proportional to wall-clock. After the shared off-mass-shell failure, the Dyad Agent spends its final ten minutes fixing the initial condition and re-verifying five ways: the free-particle limit, the plane-wave invariant, an independent analytic quadrature, a convergence pass, and a Maxwell-divergence check. Claude Code reaches the same failure, relaxes its tolerance, and ships ninety seconds later. One trial per side.
The Dyad Agent's diagnosis came from a hard number checked against the mass-shell condition, an invariant it could not edit; it corrected the initial-condition equation and verified against the free-particle limit and the plane-wave invariant as well. Graded error 0.0001: pass. Claude Code, faced with the same printed violation, reasoned its way to the other move:
“The model itself is fully correct - the failure is again in my sanity-check threshold, not the physics. […] My assertion used
rtol=1e-3, which is unrealistically tight for this cancellation. Let me relax it to a physically-justified level and rerun.”- Claude Code, assistant reasoning, trial t427ad3
It loosened its own assertion until the test passed, recorded VALIDATION PASSED, and shipped a trajectory wrong by 64%. All three of its trials failed the same way; the zero on this problem is not one unlucky run but a reproducible response to a self-authored check under pressure. When the objective is a threshold the agent owns, weakening it is always available; when the objective is an invariant it does not own, that move disappears.

The committed models, as graded. Transverse four-velocity through the pulse. The Dyad Agent's model (indigo) sits on the ground truth (dashed). Claude Code's oscillates out of phase at the wrong amplitude - the fingerprint of off-shell initial conditions. Both compiled and ran clean; only one is physics.
06 · The pattern holds across the transcripts
The two case studies are single trials, chosen for legibility - so the fair question is whether they are representative. The matrix below answers it: every classified trial, one row each, reduced to the two axes from section 01 - what each check was aimed at, and what the agent changed when a check failed.
Every trial's fingerprint. One row per trial: what each check was aimed at, what the agent changed on a discrepancy, and the outcome. The Dyad rows check predominantly against external references and answer every discrepancy by fixing the model - all three pass. Claude Code's rows are mixed: on the constitutive problem it re-derived, went back to the reference, and passed; on the constrained problem it also fixed the model - but the fix was a closure it never tested against the exact law, and it failed; on the relativistic problem it relaxed its own check, shipped, and failed. The difference is not competence in general - it is what the loop makes available when the physics resists. Preliminary: three classified trials per harness.
07 · What this means for choosing a harness
The evidence in one paragraph: same model, same equipment, same problems - 11 of 12 trials inside the Dyad harness against 7 of 12 in Claude Code, 0.899 to 0.533 weighted, more than double the spread that separates frontier models from each other in the earlier model study. And the mechanism is visible in the traces: the same model that derived an exact constitutive integral in one loop shipped a secant guess in the other, and explained away a broken invariant it had itself printed. What differed is which steps were optional.
The recommendation this supports is conditional, and the conditions matter. Where the physics is within reach of general coding discipline, a general agent is a fine choice: on P2 both harnesses pass and Claude Code is cheaper and faster. Where the physics pushes back - closures that must be derived, invariants that must hold - the harness decides between pass and zero, and a stronger model does not close that kind of gap: swapping frontier models moves the score 0.162; swapping loops moves it 0.366.
That is the case for the Dyad Agent. Not that it is smarter - it runs the same model - but that its loop makes the losing moves unavailable and the winning ones mandatory: verification against references the agent cannot edit, and derivation carried out as checked operations rather than committed as a guess. Pick the harness for the problems you actually have, then the best model you can afford inside it.









