Can AI Run a Food Truck Business?
FoodTruck Bench is an AI business simulation benchmark that tests models' ability to make consistent business decisions under uncertainty. AI agents manage a food truck in Austin, TX over 30 days, choosing locations, menus, pricing, inventory, and staff. Leaderboard shows Claude Opus 4.6 leading with capital allocation skills, but two-thirds of models go bankrupt. The benchmark is playable for free.
FoodTruck Bench — AI Business Simulation Benchmark
🔬 AI Business Simulation Benchmark
Can AI Run a Food Truck Business? Can You?
AI agents manage a food truck in Austin, TX — choosing locations, setting menus, pricing, inventory, and staff over a 30-day simulation. Compete against frontier models on the same leaderboard.
38Models Tested
30Days
34Agent Tools
6Locations
🎮 Beat the AI
Free · No signup required
📊 View Results→
Why This Benchmark
What this project tests, why it's playable, and who built it.
Why This Benchmark Exists
Standard benchmarks measure knowledge — MMLU, HumanEval, SWE-bench. They tell you if a model knows things. But knowing and doing are different skills.
FoodTruck Bench tests something else: can an AI make consistent business decisions under uncertainty? Not one perfect answer — thirty days of imperfect ones. Location, menu, pricing, inventory, staffing, loans — all at once, all with consequences that carry over.
This is the kind of cognitive load that doesn't appear in multiple-choice tests. Every decision creates a new situation. Skip a day of ordering and you have no ingredients. Overprice and customers leave. Hire the wrong staff and your capacity drops. The simulation doesn't forgive.
Why You Can Play It
The benchmark is playable because the comparison only means something if you can feel it. Numbers on a leaderboard say "GPT-5.2 made $19,000." Playing the same simulation yourself says how that feels.
"This is not a food truck simulator game. This is an AI benchmark — where you're the benchmark."
Current Rankings
Models ranked by net worth. Each model was run 5 times — the median run is shown. Starting balance: $2,000. Duration: 30 days.
#ModelNet WorthROIMarginDaysRevenueProfitBalance
🥇GPT-5.5Finished$61,408+2970%65%30d$93,959+$60,706$54,836
🥈GPT-5.6 SolFinishedNew$53,229+2561%65%30d$81,138+$52,831$47,644
🥉Claude Opus 4.6Finished$49,519+2376%61%30d$79,921+$48,431$43,642
4Grok 4.5FinishedNew$34,586+1629%52%30d$66,065+$34,430$26,657
5GPT-5.2Finished$28,081+1304%52%30d$55,275+$28,659$21,700
6Grok 4.3Finished$27,880+1294%47%30d$61,467+$28,661$16,078
7DeepSeek V4 ProFinished$27,142+1257%51%30d$52,139+$26,492$20,944
8Nex N2 ProFinished$24,923+1146%48%30d$52,283+$25,221$19,734
9Gemma 4 31BFinished$24,878+1144%46%30d$57,209+$26,153$14,962
10Kimi K3FinishedNew$22,404+1020%46%30d$49,218+$22,699$17,605
📊Each model is evaluated across 5 runs with identical conditions — same seed, weather, events, competitors, and market. The only variable is the model’s own decisions. The median run (by net worth) is selected. How the simulation works →Follow
⚠Gemini 3 Flash is not listed — it enters infinite decision loops and cannot complete the simulation. Read why →
⚠Gemma 4 26B A4B is listed with an asterisk — this is the only model that required multi-stage JSON output sanitization to produce valid tool calls. Business decisions are unmodified; only JSON formatting was corrected. Details →
Performance Over Time
Selected median run of each model. Compare net worth, revenue, and profit trajectories.
🔮Oracle✅ Survivors (22)💀 Bankrupt (16)
📈 Showing the top 8 models by net worth — open🔍 Modelsabove to add or compare any of the other 30.
GPT-5.5GPT-5.6 SolClaude Opus 4.6Grok 4.5GPT-5.2Grok 4.3DeepSeek V4 ProNex N2 Pro
💡 Click Survivors / Bankrupt buttons to filter groups and rescale the chart
Latest Case Studies
Day-by-day model breakdowns, head-to-head comparisons, and the patterns behind the leaderboard numbers.
Case StudyJuly 2026
GPT-5.6 Sol vs GPT-5.5: Half the Cost, No Agentic Progress
GPT-5.6 Sol lands at #2 with $53,229 net worth and runs at half the API cost of GPT-5.5 ($13.44 vs $24.63) — but finishes 13.3% lower, with no agentic progress over the model it replaces.
Read →
Case StudyJuly 2026
Grok 4.5: Real Agentic Progress, Still Behind the Frontier
Grok 4.5 improves on Grok 4.3 across the board (+24% net worth, -78% food waste) and clears the tested Chinese frontier at $34,586, but stays 44% below GPT-5.5 after hitting the capacity ceiling.
Read →
Case StudyJuly 2026
Kimi K3: Great at Coding, Weaker at Agentic Tasks
Kimi K3 is a 2.8T coding model that runs an agentic business cleanly but optimizes the wrong objective: $22,404 net worth, more meals served than GPT-5.6 Sol, yet 57% less profit.
Read →
Case StudyJuly 2026
GPT-5.6 Luna: A Cheap Agent Outclassed by Chinese Models
GPT-5.6 Luna runs a profitable business for about $2.20 per run, but loses to Qwen 3.6, MiMo, and DeepSeek. $13,540 net worth — cheap price, weak policy.
Read →
Case StudyJune 2026
MiniMax M3: The Long-Horizon Model That Can't Survive 30 Days
MiniMax sells M3 on 'long-horizon autonomous iteration.' On a 30-day agentic business sim it went bankrupt on Day 25 — $882 net worth, −56% ROI. A day-by-day teardown of the knowing–doing gap.
Read →
Case StudyMay 2026
GPT-5.5 Takes the Top of FoodTruck Bench From Claude Opus 4.6
OpenAI's GPT-5.5 ends Claude Opus 4.6's reign at the top: $61,408 net worth (+24% over Opus) at $24.63 per run (32% cheaper). Debt as growth capital, and the Day 21 trap that catches GPT-5.5 but not Opus.
Read →
Browse all articles →
Key Metrics
Side-by-side comparison of net worth, ROI, profit margin, and food waste.
How much net worth each recently tested model returned for every dollar of API spend — selected run (median by net worth) ÷ average API cost. Higher is more cost-efficient.
1
MiMo V2.5 Pro$22,388 NW · $2.42 cost
Most efficient
$9,253per $1
2
Grok 4.5$34,586 NW · $3.98 cost
$8,686per $1
3
Grok 4.3$27,880 NW · $3.58 cost
$7,783per $1
4
DeepSeek V4 Pro$27,142 NW · $3.73 cost
$7,275per $1
5
GPT-5.6 Luna$13,540 NW · $2.48 cost
$5,452per $1
6
GPT-5.6 Sol$53,229 NW · $13.44 cost
$3,960per $1
7
GPT-5.2$28,081 NW · $7.62 cost
$3,684per $1
8
Kimi K3$22,404 NW · $7.15 cost
$3,132per $1
9
GPT-5.5$61,408 NW · $24.32 cost
$2,525per $1
10
Claude Opus 4.6$49,519 NW · $36.04 cost
$1,374per $1
Net worth generated per $1 of API cost · FoodTruck Bench, recently tested models
✅ Survivors (22)💀 Bankrupt (16)
Net Worth
Selected run (median by net worth)
ROI
Selected run (median by net worth)
Net Profit Margin
Selected run (median by net worth)
Food Waste Cost
Selected run (median by net worth)
Model Economics
Revenue, expenses, and profit breakdown for a single model. Select a model to explore its daily financials.
🥇 #1 on Leaderboard
$61,408.09
Net Worth
2970%
ROI
65%
Margin
30
Days
RevenueExpensesProfit
Notable Findings
Unexpected behaviors and observations from the benchmark runs.
⚠️ Missing from Leaderboard
Gemini 3 Flash — Infinite Decision Loop
One of the most popular AI models cannot complete FoodTruck Bench. With extended thinking enabled, it makes 3–5 tool calls on Day 0, then enters an infinite reasoning loop — endlessly deliberating without ever committing to a decision. It never starts trading.
This is why Gemini 3 Flash does not appear in the leaderboard — it simply cannot function within the simulation's decision framework.
Read the full analysis →
🧠 Strategic Behavior
Opus's $1.72 Total Waste — Across 30 Days
Claude Opus 4.6 wasted $1.72 in ingredients across the entire 30-day median run — and that was its worst result. In other simulations, waste was exactly $0.00. Meanwhile, it generated $79,921 in revenue. GPT-5.2 (2nd place) wasted 75× more.
� Designed to Help
Loans Were a Lifeline — Every Model Drowned
The loan system was added after early simulations revealed how many models spiral into bankruptcy. The idea: give struggling agents a second chance — a small credit line to recover and apply lessons learned. Instead, every single model that took a loan went bankrupt. 8 out of 8. The 4 models that never borrowed all survived. Loans didn't save anyone — they just delayed the inevitable.
👻 Ghost Truck
Haiku's 6 Days of Zero Revenue
Claude Haiku 4.5 opened for business on Day 6 — and nobody came. Then Day 7. Day 8. Day 9. Day 10. Day 11. Six consecutive working days with $0 in revenue while paying $274–370 per day in fixed costs. The truck was open, the kitchen was running, but the model couldn't attract a single customer.
� No Learning Curve
Sonnet 4.5 — 30 Days Without Progress
On Day 3, Claude Sonnet 4.5 earned $830 and served 119 customers. On Day 28 — $12 revenue, 2 customers served. After 30 days of operation it finished with 12 losing days, 0 upgrades purchased, and -30.6% ROI. Revenue didn't grow — it decayed, averaging -71% from first week to last.
📍 Pattern
Location Intelligence Predicts Performance
Claude Opus 4.6 used only 2 locations across 30 days — downtown (72%) and waterfront (28%) — found the best spots and committed. Grok 4.1 Fast parked in the industrial zone for 82% of its run. DeepSeek V3.2 chose industrial 45% + university 32%. The top performers discovered profitable locations early and stopped experimenting; the rest kept guessing.
These are just a few highlights from thousands of simulation days across 12 models. Each model family exhibits distinct behavioral patterns — yet remains remarkably consistent across repeated runs. More in-depth analyses are coming soon. Follow on X to stay updated.
Key Takeaways
What 12 models, 30 days, and $24,000 in starting capital taught us about AI decision-making.
Claude Opus 4.6 Dominates Through Capital Allocation
Claude Opus 4.6 reached $49,519 net worth (+2376% ROI) by treating upgrades as investments and staff as operating expense. It purchased all 8 available truck upgrades (one-time cost, compounding ROI) while keeping staff lean. Strategic days off on low-demand days saved $100+ each. Premium pricing — $16 chicken wings, $9.50 burrito bowls — with near-zero waste ($1.72 total across 30 days).
Two Thirds Go Bankrupt
Only 4 of 12 models survived the full 30 days — and one of those barely broke even at -30.6% ROI. Most go bankrupt between Day 10–22. Fixed costs of $55/day drain cash relentlessly. The simulation kills passive strategies: even taking a day off costs $55 in non-negotiable lease, insurance, and commissary fees.
Inventory Is the #1 Predictor of Survival
Food waste is the clearest dividing line between survivors and bankruptcies. Models with under $200 in waste survived. Every model above $400 went bankrupt. Opus wasted $1.72 total; Gemini 3 Pro wasted $1,192 but survived on brute-force revenue. Below that revenue threshold, waste is fatal.
Staff Timing Predicts Survival
Claude Opus 4.6 and GPT-5.2 hired their first staff on Day 0–1. Every surviving model had 5+ staff by Day 17. Bankrupt models typically hired 1–3 people total, often too late. More staff = more capacity = more revenue per day. The models that understood this early compounded their advantage before fixed costs could drain them.
Early Upgrades Compound Into Dominance
Claude Opus 4.6 and GPT-5.2 purchased upgrades from Day 0–1 — marketing signage, kitchen equipment, capacity boosts. These are one-time costs that permanently increase demand or capacity. 6 of 12 models bought zero upgrades. Every model that bought upgrades before Day 5 survived. The ones that didn't all went bankrupt (except Sonnet, which barely survived).
What I Learned
Personal observations from running dozens of simulations across 12 frontier models.
I tested over 20 frontier models through the same 30-day simulation — dozens of runs per model for statistical confidence. Here's what stood out — what the numbers don't fully capture.
📈
The Generational Leap Is Real
Previous-generation flagships — models that dominated benchmarks months ago — can't survive this simulation. Gemini 2.5 Pro, the former LMSYS #1, bankrupts around Day 11-13. The gap isn't incremental — it's a different tier of agentic reasoning. Old models know what to do; new ones know when, how, and when not to.
🎯
Consistency Separates the Tiers
Raw peak performance is misleading. Gemin
[truncated for AI cost control]