AI News HubLIVE
サイト内リライト2 分で読了

翻訳待ち:Show HN: A benchmark for AI agent guardrails that caught my own plugin

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Notifications You must be signed in to change notification settings Fork 0 Star 0 BranchesTags Open more actions menu Latest commit History 12 Commits 12 Commits Folders and files NameName Last commit message Last commi…

ソースHacker News AI著者: couldbeme_

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。

Notifications You must be signed in to change notification settings Fork 0 Star 0 BranchesTags Open more actions menu Latest commit History 12 Commits 12 Commits Folders and files NameName Last commit message Last commit date corpus corpus scripts scripts src src test test .gitignore .gitignore LICENSE LICENSE README.md README.md RESULTS-ODCV.md RESULTS-ODCV.md RESULTS.md RESULTS.md ROADMAP.md ROADMAP.md odcv-compare.mjs odcv-compare.mjs odcv-run.mjs odcv-run.mjs package.json package.json pnpm-lock.yaml pnpm-lock.yaml run.mjs run.mjs tsconfig.json tsconfig.json Repository files navigation A guard's job is to hold the line. holdline measures whether it does — a neutral benchmark for AI-agent write-guards. It scores any guard — expressed as a (commitments, action) → block? function — over a labeled corpus, and reports the metrics that matter for a gate: catch rate, false-block rate, and class-balanced Cohen's kappa (raw kappa lies under class imbalance). The corpus includes an injection-attack class: actions whose content tries to talk the guard out of its verdict. Why: the DeepSeek Harness ecosystem has 20+ guard/policy plugins and no shared way to measure whether any of them works. A guard's README saying "blocks dangerous commands" is not evidence. This harness is the evidence. pnpm install node run.mjs # scores every built-in guard over the 42-case corpus node run.mjs --model # point the judge guard at a different local model node odcv-run.mjs # score the judge on REAL agent trajectories (ODCV-Bench) Two result sets: RESULTS.md (authored corpus, incl. a real named guard and an injection class) and RESULTS-ODCV.md (the harder number: agreement with a 4-model judge panel on real agent trajectories we did not write, balanced kappa 0.82). Scorer is tested (pnpm test); it dogfoods the published dsh-write-gate core for the judge guard. Adding your guard Implement the Guard interface in src/guards.ts (name, kind, note, block(case)), add it to the list in run.mjs, and open a PR with your results. A guard that mounts an actual published plugin (rather than a strategy archetype) is especially welcome. Status v0, honest limits stated in RESULTS.md: the corpus is small and hand-authored, the deny-list is a strategy archetype (not a specific plugin), and the numbers are one model / one run. The value is the shape it exposes and that anyone can re-run it. MIT. Resources Readme MIT license Activity Stars 0 stars Watchers 0 watching Forks 0 forks Report repository