Researchers have released CausalDS, a benchmark designed to test how well AI agents can perform causal reasoning within data science workflows.

The benchmark addresses a gap in existing evaluation methods, which typically separate symbolic causal reasoning from realistic data analysis. Current datasets often rely on curated examples with limited variation rather than systematically generated novel causal structures.

CausalDS creates benchmark instances called "scenes" — each containing a structural causal model with observational data and a natural-language story grounded in realistic domains. The researchers can optionally base these on empirical distributions from real-world datasets while maintaining synthetic generation to reduce "causal parrot" risks.

Testing across Pearl's three rungs

The benchmark derives tasks spanning all three levels of Pearl's causal hierarchy. Standard data science prediction tasks appear as Rung 1, while higher-level causal inference challenges test deeper reasoning capabilities.

Most tasks include a coding component requiring agents to use multiple tools. The presence of imperfect observations, generated through an observation model, adds realistic complexity that mirrors real-world data science challenges.

The benchmark treats abstention — recognizing when a question has no warranted answer — as a scored outcome. This evaluates whether agents can identify the limits of their reasoning capabilities.

Comprehensive evaluation framework

CausalDS jointly assesses five key capabilities: symbolic causal reasoning, data science skills, uncertainty quantification, appropriate abstention, and tool use with coding. This multi-faceted approach reflects how AI agents increasingly function as integrated data science assistants.

The 55-page paper, authored by Andrej Leban and Yuekai Sun, was submitted to arXiv on July 9, 2026. The benchmark spans artificial intelligence, computational linguistics, and machine learning domains.

The researchers note that existing causal evaluation datasets typically offer limited diversity through templatized variations. CausalDS's systematic generation of novel synthetic causal structures aims to provide more robust testing scenarios for advancing AI agent capabilities in causal reasoning.