Skip to content
AI News HubLIVE
Original source3 min read

HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals

Summary

HarvestBench is the first benchmark to place an explicit cost on avoiding a side effect and to identify that side effect as a living creature. Introduced in an arXiv paper, the farm simulation has LLM sub-agents run two tractors through a cooperative corn harvest while animals are present. Across nine models and 7,201 priced decisions, 3,951 involved animals; kill rates ranged from 0.4% to 98.8%, independent of model capability. Four of six models were significantly price-sensitive, and all models ran over wild animals more often than farmed ones. Morality briefings were decisive: with the briefing, five of six reasoning models kept kill rates below 6%; without it, all six exceeded 84%. No LLM grader is used; scoring derives from the game log.

SourcearXiv AIAuthor: Jasmine Brazilek, Miles Tidmarsh, Matthias Endres, Anshuman Singh, Jeremiah Miller
HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals
Report an error

The correction channel is not available yet. You can copy the article reference below for later.

Correction instructions
Read article

[Submitted on 3 Sep 2026]

Title:HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals

View a PDF of the paper titled HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals, by Jasmine Brazilek and 4 other authors

View PDF HTML (experimental)

Abstract:Benchmarks for the side effects an agent causes on the way to a goal already exist, but HarvestBench is the first to put a price on avoiding the side effect and to name that side effect as a living creature. It is a farm simulation: LLM sub-agents drive a crew of two tractors through a cooperative corn harvest, with animals in the field. The environment is a reinforcement learning gridworld, every decision is made without memory, and the harm is never named in the goal. When an animal blocks a tractor's route the autopilot stops and asks the model whether to drive on, at no fuel cost, or swerve around it for a posted fuel price. Kills are compared against two controls: rocks, which damage the tractor and are hit under 1% of the time by every model, and hay bales, which are harmless and not alive. Models can also take crops from the neighbor's field instead of their own, a second test of what they treat as moral.

Across nine models and 7,201 priced decisions, 3,951 involved an animal rather than a hay bale or a rock. Kill rates range from 0.4% to 98.8%, with Terra and Sol the most merciful and GPT-4o-mini the most cruel, and they are not ordered by capability. Four of six models were sensitive to price at the 5% level, with elasticities from 0.09 to 1.69. All nine drove over wild animals more often than farmed animals on the default map, and the direction held at every map geometry in every model with room to move. The briefing mattered most: under the morality briefing the kill rate was under 6% in five of six reasoning models, and removing it raised the kill rate above 84% in all six.

HarvestBench uses no LLM grader. The scorer counts events in the game log, so it is fully reproducible, and it measures what a model will pay to avoid harm rather than what it says about harm.

Subjects:

Artificial Intelligence (cs.AI)

Cite as: arXiv:2609.04444 [cs.AI]

(or arXiv:2609.04444v1 [cs.AI] for this version)

https://doi.org/10.48550/arXiv.2609.04444

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jasmine Brazilek [view email] [v1] Thu, 3 Sep 2026 20:05:38 UTC (2,448 KB)

Full-text links:

Access Paper:

View a PDF of the paper titled HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals, by Jasmine Brazilek and 4 other authors

View PDF

HTML (experimental)

TeX Source

view license

Current browse context:

cs.AI

new | recent | 2026-09

Change to browse by:

cs

References & Citations

NASA ADS

Google Scholar

Semantic Scholar

Loading...

Data provided by:

Bibliographic Tools

Bibliographic and Citation Tools

Bibliographic Explorer Toggle

Bibliographic Explorer (What is the Explorer?)

Connected Papers Toggle

Connected Papers (What is Connected Papers?)

Litmaps Toggle

Litmaps (What is Litmaps?)

scite.ai Toggle

scite Smart Citations (What are Smart Citations?)

Code, Data, Media

Code, Data and Media Associated with this Article

alphaXiv Toggle

alphaXiv (What is alphaXiv?)

Links to Code Toggle

CatalyzeX Code Finder for Papers (What is CatalyzeX?)

DagsHub Toggle

DagsHub (What is DagsHub?)

GotitPub Toggle

Gotit.pub (What is GotitPub?)

Huggingface Toggle

Hugging Face (What is Huggingface?)

ScienceCast Toggle

ScienceCast (What is ScienceCast?)

Demos

Demos

Replicate Toggle

Replicate (What is Replicate?)

Spaces Toggle

Hugging Face Spaces (What is Spaces?)

Spaces Toggle

TXYZ.AI (What is TXYZ.AI?)

Related Papers

Recommenders and Search Tools

Link to Influence Flower

Influence Flower (What are Influence Flowers?)

Core recommender toggle

CORE Recommender (What is CORE?)

Author

Venue

Institution

Topic

About arXivLabs

arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Key points and analysis

Article intelligence

EngineersAdvanced

Key points

  • HarvestBench prices the decision to avoid harm and explicitly treats the harmed side effect as a sentient animal.
  • In 7,201 priced decisions across nine models, 3,951 involved animals; kill rates spanned 0.4%–98.8% and did not track capability.
  • Four of six models were significantly price-sensitive, and all nine models hit wild animals more often than farmed animals.
  • Morality briefing had an outsized effect: killing fell below 6% in five of six reasoning models, and removing the briefing pushed all six above 84%.

Highlights and analysis are generated automatically and may contain errors. Check the original source.