AI News HubLIVE
站内改写6 分钟阅读

待翻译:Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Terminal-Bench-Science 0.1 Terminal-Bench-Science evaluates AI agents on workflows from researchers' own work. Scientists, not model developers or data vendors, set the bar for scientific capability in AI. Terminal-Benc…

来源Hacker News AI作者: matt_d

AI 服务暂时不可用,以下为来源正文,待恢复后补全翻译。

Terminal-Bench-Science 0.1 Terminal-Bench-Science evaluates AI agents on workflows from researchers' own work. Scientists, not model developers or data vendors, set the bar for scientific capability in AI. Terminal-Bench-Science is a benchmark led by researchers at Stanford University and built by the team behind Terminal-Bench in collaboration with domain experts from a range of scientific disciplines and research institutions around the world. It measures the AI agent capabilities through a diverse set of challenging, expert-curated workflows drawn from scientific research. Terminal-Bench-Science is a continuous benchmark that evolves alongside frontier AI, creating a feedback loop between scientific needs and AI development. Our first release includes 70 tasks from the life, physical, Earth, mathematical, and engineering sciences. The strongest model evaluated, Claude Opus 5, achieves a 30% resolution rate on Terminal-Bench-Science 0.1. Terminal-Bench-Science 0.1 Leaderboard Resolution rates across 70 scientific workflow tasks on Terminal-Bench-Science 0.1 Overview While Terminal-Bench has driven progress in AI agents for software engineering, Terminal-Bench-Science brings the same ambition to science. Our goal is to drive the development of agents with scientific capabilities that make them useful research assistants. These agents should execute technically demanding and time-consuming workflows, freeing scientists to focus more of their time on the parts of science where human judgment matters most: defining research questions, forming hypotheses, interpreting and validating results, and communicating findings. In this role, AI agents can extend what researchers accomplish and help accelerate scientific discovery. Achieving this requires benchmarks that reflect real scientific practice, provide verifiable evidence of capability, and evolve alongside the AI frontier. We need benchmarks drawn from real scientific workflows. Scientific capability should be evaluated on real research practice, not textbook questions or standardized exercises, contributed by practicing scientists themselves. Terminal-Bench-Science gives scientists across domains a direct voice and a shared platform to set the bar for AI progress on the problems they care about. The stakes in science are too high, and its benchmarks must reflect the scientific community's priorities rather than outside interests. We need verifiable evidence of scientific capability. Without reliable evaluation, we cannot tell whether agent capabilities are improving or where their limitations remain. Terminal-Bench-Science evaluates agents in realistic environments and grades concrete artifacts such as analyses, simulations, proofs, code, and data products with reproducible, task-specific tests. We need a benchmark that keeps pace with the frontier. Too often, scientific benchmarks are treated as papers to publish rather than mechanisms for driving progress. They are released once and then abandoned as models advance and known limitations persist. Terminal-Bench-Science is a continuous benchmark that evolves alongside the AI frontier. Through regular releases, scientists can contribute new workflows, improve existing tasks, and create a feedback loop between scientific needs and AI development. Terminal-Bench-Science Feedback Loop Terminal-Bench-Science Feedback Loop Tasks Terminal-Bench-Science 0.1 includes 70 tasks across the life, physical, Earth, mathematical, and engineering sciences. Tasks span scientific data analysis, statistical inference, simulation, optimization, theorem proving, image reconstruction, signal processing, inverse problems, sensor calibration, model fitting, classification, and scientific machine learning. Terminal-Bench-Science 0.1 Task Coverage 70 expert-curated tasks across five scientific domains Tasks are contributed by researchers through an open process on GitHub, with discussion and feedback in the #tb-science channel on Discord. Contributions begin as proposals, where reviewers discuss each idea, leave feedback, and approve those that look like a strong fit: scientifically grounded workflows worth measuring in the benchmark. Approved proposals are implemented as pull requests, where reviewers confirm that each task is objectively verifiable, genuinely challenging for AI agents, and not something today's frontier systems already solve easily. To merge, domain reviewers assess scientific validity and realism, technical reviewers inspect task construction and verification, and a bar raiser performs a final quality check. Of 920 proposals, 464 were approved for implementation and 386 pull requests were opened, but only 70 tasks made it into Terminal-Bench-Science 0.1. That selectivity reflects how difficult it is to create tasks that are scientifically interesting, challenging for frontier agents, and sufficiently well specified for rigorous evaluation. Progress across proposals, pull requests, and reviews is tracked on the public task dashboard. Terminal-Bench-Science 0.1 Contributions From 920 task proposals to 70 landing in Terminal-Bench-Science 0.1376 contributors across 22 countries from proposals, reviews, or pull requests Results Terminal-Bench-Science 0.1 leaves substantial room for progress on AI agents for scientific research. Each evaluated model ran three independent trials per task across all 70 tasks. Claude Opus 5 with Claude Code achieves the highest resolution rate at 30.0%, followed by GPT-5.6 Sol with Codex at 22.4% and Claude Fable 5 with Claude Code at 21.4%. Claude Opus 4.8 sits in the middle at 10.5%. GPT-5.6 Terra, Kimi K3, and Grok 4.6 all resolve less than 10% of tasks. GLM 5.3 is the strongest open model at 8.1%, and GPT-5.6 Luna is last at 3.3%. Terminal-Bench-Science distinguishes between systems about as well as Terminal-Bench 3.0 while pushing resolution rates down by more than 10 percentage points for every model evaluated on both. That gap is deliberate: during review, tasks were calibrated to challenge the newest frontier models. Terminal-Bench-Science vs. Terminal-Bench Resolution Rates on Terminal-Bench 2.1, Terminal-Bench 3.0, and Terminal-Bench-Science 0.1 Performance is only one dimension of progress. The cost-resolution plot shows total evaluation cost across all 70 tasks against resolution rate. GPT-5.6 Luna, Kimi K3, and GPT-5.6 Terra occupy the low-cost end of the frontier. GPT-5.6 Sol and Claude Opus 5 reach the highest resolution rates at greater cost, with Opus 5 at $7.0k. GPT-5.6 Sol matches Claude Fable 5's performance at less than a third of the cost ($4.2k vs $14.2k). Token usage shows a different frontier. Claude Fable 5 matches GPT-5.6 Sol's performance while using about a quarter fewer tokens (6.4B vs 8.4B). Kimi K3 anchors the low-token end and Claude Opus 5 the high-resolution end. Only Kimi K3 and Claude Opus 5 appear on both Pareto frontiers. Terminal-Bench-Science 0.1 Cost vs. Resolution Rate Pareto frontier of cost and resolution rate across evaluated systems Terminal-Bench-Science 0.1 Tokens vs. Resolution Rate Pareto frontier of token usage and resolution rate across evaluated systems Resolution rates also vary by scientific domain. Anthropic and OpenAI models take the top two spots in every domain except the engineering sciences, where Grok 4.6 ties GPT-5.6 Sol for second place (14.8%) at lower cost and token usage. Claude Opus 5 leads both GPT-5.6 Sol and Claude Fable 5 in every domain except the mathematical sciences, where Claude Fable 5 (33.3%) and GPT-5.6 Sol (31.4%) take the top two spots. The full breakdown by domain is available on the leaderboard. Terminal-Bench-Science 0.1 Resolution Rates by Domain Opus 5 GPT-5.6 Sol Grok 4.6 Domain resolution rates: Opus 5, GPT-5.6 Sol, and Grok 4.6 Conclusion and Roadmap Terminal-Bench-Science 0.1 is a community effort by researchers across the life, physical, Earth, mathematical, and engineering sciences, together with the Terminal-Bench and Harbor team. It is the most rigorous benchmark of scientific agent capabilities we could build in the open, and we are only getting started: Terminal-Bench-Science 0.1 is the first release of a continuous benchmark. Regular releases will add tasks, broaden coverage across the five scientific domains, retire tasks that agents saturate or that review reveals to be underspecified, and keep the leaderboard current as new frontier models are released. Each release is calibrated against the frontier at the time. For Terminal-Bench-Science 0.1, this was Claude Opus 5 and GPT-5.6 Sol, and as stronger models emerge we will use them to evaluate and calibrate new and improved tasks. Tasks are versioned so that trials can be re-used, re-graded, or re-run with a single Harbor command, which keeps the cost of updating results low. Progress is tracked in the open on the task dashboard, and every release is tagged on GitHub and Harbor Hub. Work on Terminal-Bench-Science 0.2 is already underway, with a pull request deadline of October 5, 2026. If you are a researcher with a workflow that frontier agents should be able to do but cannot yet, we want it in the benchmark. The contribution flow is Propose → Build → Review: propose your task through the task proposal form, build it following the contributing guide, and it will go through automated checks, parallel domain and technical review, and final bar-raiser approval before merge. Join the effort in #tb-science on Discord and on GitHub, and drop into our weekly meetings and office hours via the project calendar. Let's let scientists define what scientific capability in AI looks like, and measure it rigorously, together. Citation If you find this work useful, please cite it. You can use the "Cite this repository" button on GitHub (generated from CITATION.cff) or cite manually using the information below. @software{Terminal-Bench-Science_Team_Terminal-Bench-Science_Evaluating_AI_2026, author = {{Terminal-Bench-Science Team}}, doi = {10.5281/zenodo.22110254}, license = {Apache-2.0}, month = aug, title = {{Terminal-Bench-Science: Evaluating AI agents on research workflows across scientific domains}}, url = {https://github.com/harbor-framework/terminal-bench-science}, version = {v0.1.0}, year = {2026} } Acknowledgements Thank you to all of the task contributors, reviewers, and advisors behind Terminal-Bench-Science. Special thanks to our project lead advisors Ludwig Schmidt and Sanmi Koyejo; our senior reviewers Allen Hart, Ivan Bercovich, Joseph Janssen, Jiaming Hu, Steffen Bollmann, and Sergey Aganezov; our AI research advisors Ryan Marten, Alex Shaw, Lin Shi, Benjamin Feuer, Mike A. Merrill, Alex Dimakis, Jenia Jitsev, Bodhisattwa Majumder, Peter Clark, Thomas Wolf, Braden Hancock, and Andy Konwinski; and our scientific advisors Sara Beery, Jo Dunkley, J. Nathan Kutz, Ching-Yao Lai, Scott Linderman, Emma Lundberg, Russ Poldrack, Aviv Regev, and Risa Wechsler. Terminal-Bench-Science is an open academic collaboration hosted by Stanford University and the Laude Institute, in partnership with the Stanford AI Lab (SAIL), the Stanford Institute for Human-Centered Artificial Intelligence (HAI), Stanford AI Measurement Science (AIMS), the NSF AI Institute for Foundations of Machine Learning (IFML), the Allen Institute, and the Allen Institute for AI (Ai2). As part of the Terminal-Bench franchise, it is built by the Terminal-Bench and Harbor team together with a community of scientific contributors. We thank the Laude Institute for support through the Slingshots program, Snorkel AI for support through the Open Benchmarks Grants program, the 2077AI Open Source Foundation for PP API credits supporting task review and curation, and UniPat AI and Modal for their support of Terminal-Bench-Science. We thank Bespoke Labs, Anthropic, Google, Moonshot AI, SpaceXAI, and Z.ai for API credits supporting leaderboard evaluations. Written by: Steven Dillmann