How Dyad Outperforms Claude Code in Physical Modeling Tests
An AI-generated model can compile, run, and still violate physics. Before using it to guide design decisions, engineers need evidence that its equations, numerical behavior, and physical assumptions hold up.
This white paper examines how the agent harness—the tools, instructions, and verification checks surrounding an AI model—affects the reliability of physical modeling results. Learn how Dyad combines focused engineering tools with a workflow that requires the agent to derive, simulate, and verify its models.
Inside the white paper
A controlled comparison: A full 2×2 evaluation tests two frontier models in both Dyad and Claude Code, separating the effects of the model and the harness.
Measurable results: With the same frontier model, Dyad achieved a difficulty-weighted score of 0.90 versus 0.53 for Claude Code at nearly identical cost in JuliaHub’s internal evaluation.
Physical verification in practice: See why successful execution and a close fit to calibration data can miss violations of conservation laws, numerical convergence, and equilibrium behavior.
Criteria for evaluating engineering AI: Identify the checks that help determine whether a generated model can support engineering decisions.
The paper presents JuliaHub’s internal evaluation alongside complementary findings from an independent, peer-reviewed crystallization study, with the scope and limitations of each explained.








