本文にスキップ
AI News HubLIVE
原典の内容 · 翻訳・分析待ち2 分で読了

翻訳待ち:Can an Open Model Do Security Research? Cantina’s apex-flash-1 Solves 40 of 60 Held-Out Bug Tasks

記事の要約

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Cantina Security, with Yeta Labs, has released apex-flash-1, an open-weights model trained specifically for vulnerability research. It is a reinforcement learning fine-tune of Z.ai’s GLM-5.3-Flash, released on Hugging Face under the MIT license. Is it deployable? Yes, the MIT weights serve on vLLM, SGLang or Transformers, but BF16 needs roughly 640 GB of GPU memory. […] The post Can an Open Model Do Security Research? Cantina’s apex-flash-1 Solves 40 of 60 Held-Out Bug Tasks appeared first on MarkTechPost.

ソースMarkTechPost著者: Michal Sutter
翻訳待ち:Can an Open Model Do Security Research? Cantina’s apex-flash-1 Solves 40 of 60 Held-Out Bug Tasks
誤りを報告

訂正窓口はまだ利用できません。記事情報をコピーして保存できます。

訂正案内
本文へ

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。

Cantina Security, with Yeta Labs, has released apex-flash-1, an open-weights model trained specifically for vulnerability research. It is a reinforcement learning fine-tune of Z.ai’s GLM-5.3-Flash, released on Hugging Face under the MIT license. Is it deployable? Yes, the MIT weights serve on vLLM, SGLang or Transformers, but BF16 needs roughly 640 GB of GPU memory. What Cantina Built apex-flash-1 has 321.3B total parameters, per its Hugging Face safetensors metadata. The GLM-5.3-Flash base is a Mixture-of-Experts model with 18B active parameters. Cantina trained it with GRPO using a rank-256 LoRA plus selective full-parameter training. The data covers 150 tasks built from 50 real vulnerability cases. Each case appears in 3 variants: guided whitebox, focused whitebox and focused blackbox. Authorization, identity and scope flaws make up 72% of cases. Accounting and numerical precision bugs add 18%. Time validation, business rules and SSRF cover the rest. As per the model card on HF, RL rollouts ran inside the Codex agent harness on production-like software and protocol environments. Benchmark Results Cantina evaluated 60 tasks from 20 held-out vulnerability cases. Each model ran the set once, with costs estimated from provider pricing. apex-flash-1: 40/60 solved (66.7% pass@1), about $2.38 GLM-5.3-Flash (base): 36/60 solved (60.0%), about $4.56 Claude Opus 5 High: 43/60 solved (71.7%), about $74.68 Opus solved 3 more tasks but cost about 31x more per run. That is roughly $0.06 per solved task for apex-flash-1 versus $1.74 for Opus. These are company-reported numbers on an internal benchmark. A Worker Model, Not an Orchestrator Cantina positions apex-flash-1 as a worker orchestrated by a larger model. The card lists code reading, tool use, exploit development and verification as target skills. An experimental apex-flash-1-abliterated variant ships with modified refusal behavior. It was not separately evaluated. Cantina’s rationale is that defenders need capable models they can run and control locally. Interactive Explainer How It Compares Featureapex-flash-1Aikido Altar-1Cisco Foundation-Sec-8B-ReasoningGLM-5.3-Flash DeveloperCantina Security + Yeta LabsAikido SecurityCisco Foundation AIZ.ai Base modelGLM-5.3-FlashGLM-5.3 (pruned)Llama 3.1 8BOwn pretraining Size321.3B total, BF16328 GB, INT4 (W4A16)8B320B total, 18B active LicenseMITInherits GLM-5.3 licenseCustom (see NOTICE.md)MIT Security methodGRPO RL on 50 real vulnerability casesExpert pruning (REAP) + quantizationInstruction tuning + RLHF on security QAGeneral-purpose base Primary useAgentic vuln research workerAir-gapped autonomous pentestingSOC triage and threat defenseGeneral coding and agents HardwareMulti-GPU node (~642 GB BF16 weights)4x H200 with vLLMSingle GPUMulti-GPU node Published security result66.7% pass@1, 60 tasks60.4% recall, 32-CVE internal setCisco-reported security benchmarks60.0% on Cantina’s set Sources: Cantina, Aikido, Cisco, Hugging Face model cards. © Marktechpost Key Takeaways apex-flash-1 is a 321.3B open-weights security model under MIT. GRPO training on 50 real vulnerability cases produced 150 tasks. It scored 66.7% pass@1 versus 71.7% for Claude Opus 5 High. Its 60-task run cost about $2.38 versus $74.68 for Opus. BF16 needs a multi-GPU node; community 4-bit ports exist. Check out the Technical details. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post Can an Open Model Do Security Research? Cantina’s apex-flash-1 Solves 40 of 60 Held-Out Bug Tasks appeared first on MarkTechPost.

要点と分析を開く

記事インテリジェンス

エンジニア上級

要点

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Cantina Security, with Yeta Labs, has released apex-flash-1, an open-weights model trained specifically for vulnerability research. It is a reinforcement learning fine-tune of Z.a…

要点と分析は自動生成され、誤りを含む場合があります。原典をご確認ください。