CompMat-Bench
A benchmark of 94 tasks from 15 recently published computational materials studies. Each task is one research step. The expensive simulations are run in advance, and answers are graded by fixed rules with no LLM judge.
Department of Materials Science and NanoEngineering, Rice University
* Corresponding author.
01 / Benchmark design
Research steps grounded in published studies.
Each study is split into research steps. The expensive simulations are run once, in advance. Agents prepare inputs or analyze precomputed outputs, and fixed rules grade their answers against the reference results of our reproduction.
View full-size figure ↗
02 / Agent performance
Guidance and workflow length change agent performance.
With full guidance, agents built on three LLMs pass 66.0 to 90.4% of the 94 single tasks. Preparing simulation inputs is the hardest kind of task for every agent.
The strongest agent falls only when long workflows meet reduced guidance. Either factor alone lowers the pass rates of the weaker agents. With both, all three agents do worse than on single tasks with full guidance, and pass rates fall to 50 to 75%.
View full-size figure ↗
03 / Failure analysis
Agents miss what an experienced researcher would notice.
Most failures are scientific errors. Agents do not recognize anomalies in their own results, and agents built on different LLMs lack the same implicit knowledge, so they fail in the same way. In one workflow, the warning sign was printed in the agent's own output in nine failing runs. The agent continued anyway.
Research paper
CompMat-Bench: Benchmarking AI Agents for Computational Materials Science
Chenmu Zhang, Levi Felix, Jun-Jie Zhang, Xingfu Li, Xuelian Jiang, Tao Jiang, Subhendu Mishra, Xixi Qin, and Boris I. Yakobson*
Department of Materials Science and NanoEngineering, Rice University
* Corresponding author.