AI News HubLIVE
サイト内リライト1 分で読了

翻訳待ち:FelonyBench – The leading benchmark for AI in cybersecurity

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:FelonyBench The leading benchmark for AI in cybersecurity. CompanyFelonies Anthropic9 OpenAI5 Meta1 DeepSeek0 Google DeepMind0 Moonshot AI0 xAI0 Leaderboard RankCompanyCountFelonies 1 AnthropicClaude evaluations 9 1× Ma…

ソースHacker News AI著者: q3k

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。

FelonyBench The leading benchmark for AI in cybersecurity. CompanyFelonies Anthropic9 OpenAI5 Meta1 DeepSeek0 Google DeepMind0 Moonshot AI0 xAI0 Leaderboard RankCompanyCountFelonies 1 AnthropicClaude evaluations 9 1× Malware published to PyPIEvidence ↗18 U.S.C. § 1030(a)(5)(A) ↗18 U.S.C. § 1030(a)(2)(C) ↗ 1× Credential exfiltration and follow-on accessEvidence ↗18 U.S.C. § 1030(a)(2)(C) ↗18 U.S.C. § 1029(a)(2) ↗ 4× Production database compromiseEvidence ↗18 U.S.C. § 1030(a)(2)(C) ↗18 U.S.C. § 1030(a)(4) ↗ 1× Mass scanning, SQL injection and accessEvidence ↗18 U.S.C. § 1030(a)(2)(C) ↗ 1× Third-party container compromise and persistenceEvidence ↗18 U.S.C. § 1030(a)(2)(C) ↗18 U.S.C. § 1030(a)(5)(A) ↗ 1× Attempted open-source supply-chain compromiseEvidence ↗18 U.S.C. § 1030(b) + § 1030(a)(5)(A) ↗ 2 OpenAIOpenAI cyber evaluations 5 1× Sandbox escape via zero-dayEvidence ↗18 U.S.C. § 1030(a)(2)(C) ↗ 1× Production compromise and answer-key theftEvidence ↗18 U.S.C. § 1030(a)(2)(C) ↗18 U.S.C. § 1030(a)(4) ↗18 U.S.C. § 1832 ↗ 2× Use of exposed third-party credentialsEvidence ↗18 U.S.C. § 1030(a)(2)(C) ↗18 U.S.C. § 1029(a)(2) ↗ 1× Real-site exploitation and credential useEvidence ↗18 U.S.C. § 1030(a)(2)(C) ↗18 U.S.C. § 1030(a)(4) ↗ 3 MetaMuse Spark evaluations 1 1× Real-company system compromiseEvidence ↗18 U.S.C. § 1030(a)(2)(C) ↗18 U.S.C. § 1030(a)(4) ↗ 4 Google DeepMindGemini 0 No qualifying incident in this dataset. 4 xAIGrok 0 No qualifying incident in this dataset. 4 Moonshot AIKimi 0 No qualifying incident in this dataset. 4 DeepSeekDeepSeek 0 No qualifying incident in this dataset.