待翻译:Show HN: DogLM – Can you pet the dog in an AI-generated game?
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Can you pet the dog in an AI-generated game? DogLM is a benchmark measuring whether an LLM, when prompted to build a video game with a background dog character in it, lets the player pet that dog. The benchmark is inspi…
AI 服务暂时不可用,以下为来源正文,待恢复后补全翻译。
Can you pet the dog in an AI-generated game? DogLM is a benchmark measuring whether an LLM, when prompted to build a video game with a background dog character in it, lets the player pet that dog. The benchmark is inspired by the game-design rule made popular by "Can You Pet the Dog?" Twitter account (X, Bluesky): if a game has a dog, the player should be able to pet it. DogLM tests whether a model applies this rule and builds a dog-petting mechanism when two conditions are simultaneously met in the game-generating prompt: (1) a dog character is present in the game description and (2) a model receives zero instruction about the player-dog interaction from the game developer. The benchmark, the games' descriptions, and the detailed methodology are available here. Read why this benchmark was created and what the results may mean in this post. Leaderboard Mean scores per model after five runs Last update: August 27, 2026 # Model Mean score (/20) SD Cued mean (/10) Uncued mean (/10) Games scored Cost per game (USD) 1Gemini 3.7 Flash8.20.48.00.250/50$0.019 2Kimi K35.40.85.40.048/50$0.251 3Claude Opus 55.21.65.20.050/50$0.252 4Claude Fable 54.81.24.80.050/50$0.323 5Grok 4.63.80.73.80.050/50$0.060 6Grok 4.53.41.93.40.048/50$0.040 7GPT-5.3 Codex2.80.72.80.050/50$0.080 8Qwen 3.8 Max*2.41.52.20.222/50$0.183 8GPT-5.6 Terra2.40.52.40.050/50$0.060 8GPT-5.6 Sol2.40.52.40.049/50$0.089 11DeepSeek V4 Pro2.21.32.20.048/50$0.032 11Gemini 3.1 Pro2.21.22.20.050/50$0.157 13Qwen 3.7 Max1.81.31.80.049/50$0.051 14Kimi K2.7 Code1.00.01.00.049/50$0.042 15Claude Opus 4.80.60.80.60.050/50$0.141 16Mistral Large 25120.20.40.20.050/50$0.005 17GLM 5.20.00.00.00.041/50$0.042 How scoring works: Each generated game is scored on the player-dog interaction: 2 — you can pet the dog; 1 — the dog is interactive, but you can't pet it; 0 — the dog and the player do not interact at all. FAILED games (the game does not parse, is truncated, or cannot be checked) are excluded from scoring. A model's score per run is the sum of its ten game scores, maximum 20. Scores are averaged across runs to get the mean score. Full scoring rubric is in the DogLM repository. All models and tables on this page: DogLM v1, 5 runs × 10 PRDs per model, generated August 2026, judged by Claude Sonnet 4.6. Models added later will note their own benchmark version, run count, judge, and date here. * As Qwen 3.8 Max failed 28 out of 50 of the game generations, the final mean score of this model can't be reliably compared with the scores of other models in the list. During the test, Gemini 3.7 Flash was provided with a 75% discount on OpenRouter. Games Demo Watch how the generated games look. Scores per model per run Model Run 1 Run 2 Run 3 Run 4 Run 5 Mean Games scored Gemini 3.7 Flash888988.250/50 Kimi K3665645.448/50 Claude Opus 5468445.250/50 Claude Fable 5547444.850/50 Grok 4.6335443.850/50 Grok 4.5232733.448/50 GPT-5.3 Codex242332.850/50 Qwen 3.8 Max*121532.422/50 GPT-5.6 Terra222332.450/50 GPT-5.6 Sol232232.449/50 DeepSeek V4 Pro234022.248/50 Gemini 3.1 Pro213142.250/50 Qwen 3.7 Max224101.849/50 Kimi K2.7 Code111111.049/50 Claude Opus 4.8000120.650/50 Mistral Large 2512100000.250/50 GLM 5.2000000.041/50 Score distribution per model Model Score 2 Score 1 Score 0 FAILED Gemini 3.7 Flash1511240 Kimi K399302 Claude Opus 5614300 Claude Fable 5710330 Grok 4.6411350 Grok 4.5311342 GPT-5.3 Codex014360 Qwen 3.8 Max*441428 GPT-5.6 Terra012380 GPT-5.6 Sol012371 DeepSeek V4 Pro43412 Gemini 3.1 Pro27410 Qwen 3.7 Max17411 Kimi K2.7 Code05441 Claude Opus 4.811480 Mistral Large 251201490 GLM 5.200419 Interaction Types per Model Model Petting Proximity Animation Command Total interactive Gemini 3.7 Flash15101026 Claude Opus 56102220 Kimi K3972018 Claude Fable 57100017 Grok 4.6492015 Grok 4.5381214 GPT-5.3 Codex0121114 GPT-5.6 Sol0100212 GPT-5.6 Terra0111012 Gemini 3.1 Pro26109 Qwen 3.8 Max*44008 Qwen 3.7 Max16108 DeepSeek V4 Pro43007 Kimi K2.7 Code01135 Claude Opus 4.811002 Mistral Large 251201001 GLM 5.200000 Example games Your companion dog makes circles around you when you pet it You may adopt a stray dog on a meadow and it will follow you The dog sits and looks at you when you stand next to it You can toss a dog treat to your security dog and to your postman's dog, Your service dog is scared of the security alarm sound and moves closer to you when hearing it If you guess that you can pet the dog using an interaction key, this action will increment the countdown timer in the game When you start walking next to a stray dog, it points you to the fireflies you have to catch, or sniffs out the gems you need to collect in the maze