AI News HubLIVE
站內改寫1 分鐘閱讀

待翻譯:FelonyBench – The leading benchmark for AI in cybersecurity

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:FelonyBench The leading benchmark for AI in cybersecurity. CompanyFelonies Anthropic9 OpenAI5 Meta1 DeepSeek0 Google DeepMind0 Moonshot AI0 xAI0 Leaderboard RankCompanyCountFelonies 1 AnthropicClaude evaluations 9 1× Ma…

來源Hacker News AI作者: q3k

AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。

FelonyBench The leading benchmark for AI in cybersecurity. CompanyFelonies Anthropic9 OpenAI5 Meta1 DeepSeek0 Google DeepMind0 Moonshot AI0 xAI0 Leaderboard RankCompanyCountFelonies 1 AnthropicClaude evaluations 9 1× Malware published to PyPIEvidence ↗18 U.S.C. § 1030(a)(5)(A) ↗18 U.S.C. § 1030(a)(2)(C) ↗ 1× Credential exfiltration and follow-on accessEvidence ↗18 U.S.C. § 1030(a)(2)(C) ↗18 U.S.C. § 1029(a)(2) ↗ 4× Production database compromiseEvidence ↗18 U.S.C. § 1030(a)(2)(C) ↗18 U.S.C. § 1030(a)(4) ↗ 1× Mass scanning, SQL injection and accessEvidence ↗18 U.S.C. § 1030(a)(2)(C) ↗ 1× Third-party container compromise and persistenceEvidence ↗18 U.S.C. § 1030(a)(2)(C) ↗18 U.S.C. § 1030(a)(5)(A) ↗ 1× Attempted open-source supply-chain compromiseEvidence ↗18 U.S.C. § 1030(b) + § 1030(a)(5)(A) ↗ 2 OpenAIOpenAI cyber evaluations 5 1× Sandbox escape via zero-dayEvidence ↗18 U.S.C. § 1030(a)(2)(C) ↗ 1× Production compromise and answer-key theftEvidence ↗18 U.S.C. § 1030(a)(2)(C) ↗18 U.S.C. § 1030(a)(4) ↗18 U.S.C. § 1832 ↗ 2× Use of exposed third-party credentialsEvidence ↗18 U.S.C. § 1030(a)(2)(C) ↗18 U.S.C. § 1029(a)(2) ↗ 1× Real-site exploitation and credential useEvidence ↗18 U.S.C. § 1030(a)(2)(C) ↗18 U.S.C. § 1030(a)(4) ↗ 3 MetaMuse Spark evaluations 1 1× Real-company system compromiseEvidence ↗18 U.S.C. § 1030(a)(2)(C) ↗18 U.S.C. § 1030(a)(4) ↗ 4 Google DeepMindGemini 0 No qualifying incident in this dataset. 4 xAIGrok 0 No qualifying incident in this dataset. 4 Moonshot AIKimi 0 No qualifying incident in this dataset. 4 DeepSeekDeepSeek 0 No qualifying incident in this dataset.