Anthropic's Claude Fable 5 scored 0.889 in JuliaHub's physical AI benchmark, outperforming OpenAI's GPT-5.6 family across five sealed modeling and simulation problems.
JuliaHub tested four frontier models on problems ranging from constitutive consistency to NASA's HL-20 flight vehicle dynamics. Claude Fable 5 led with a 0.889 weighted score at $9.60 per trial, while GPT-5.6-sol scored 0.814 at $1.74 per trial.
Performance gaps widen on complex physics
The evaluation used JuliaHub's Dyad AI agent harness, which reads engineering specifications, derives physics equations, compiles models, and simulates trajectories. Each model received identical configurations: 1M context window, 128k token budget, and "xhigh" reasoning effort.
GPT-5.6-terra scored 0.786 at $1.25 per trial, while GPT-5.6-luna managed 0.727 at $3.26 per trial. The benchmark specifically targets physical AI's core challenge: models that compile and run cleanly while encoding impossible physics.
JuliaHub's evaluation pipeline compares simulated trajectories against sealed ground truth rather than scoring code quality. The company noted that agents often pass self-written tests while relying on simplifications that fail in real-world deployment.
The hardest problem required reading aerodynamic data, deriving six-degree-of-freedom motion equations, and iterating through simulation verification. JuliaHub ran three trials per model on four core problems, plus one long-horizon trial each on the flight vehicle challenge.
JuliaHub ships Dyad with multiple agent backends across vendors, positioning the company to recommend whichever model delivers the best user experience. The evaluation held all variables constant except the underlying language model.
Physical AI applications face unique verification challenges compared to web development or compiler tasks, where correct behavior is more easily contained and checkable. The gap between passing tests and matching real-world physics creates trust issues with existing benchmarks.
JuliaHub plans to expand the evaluation framework as frontier models advance in scientific computing capabilities.
💬 Discussion
Sign in to join the discussion.
Sign in →No comments yet — be the first.