AI news ·
Cantina Security releases apex-flash-1, a 321.3B open model for vulnerability research
apex-flash-1 solves security tasks at $0.06 each versus $1.74 for Claude Opus 5 High-a 31x cost gap on Cantina's benchmark.

Cantina Security and Yeta Labs released apex-flash-1 on October 5, 2026, a 321.3 billion-parameter open model built for vulnerability research and released under an MIT license. The model costs roughly 31 times less per solved security task than Anthropic's Claude Opus 5 High in Cantina's internal testing, a gap that changes the economics of running AI-powered bug hunting at scale.
What apex-flash-1 is and how it was built
apex-flash-1 is an open-weights model fine-tuned from Z.ai's GLM-5.3-Flash. Cantina Security and Yeta Labs used GRPO (Group Relative Policy Optimization) with a rank-256 LoRA adapter plus selective full-parameter training. The training data drew from 150 tasks derived from 50 real vulnerability cases. The model is available on Hugging Face under an MIT license, meaning organizations can download, modify, and deploy it without restriction.
Inference requires substantial hardware. Running the model at BF16 precision demands roughly 640 GB of GPU memory. It works with vLLM, SGLang, or Hugging Face Transformers. Cantina positions apex-flash-1 as a "worker model" that larger orchestration systems can direct - a component for security workflows rather than a standalone product.
Performance and cost compared to Claude Opus 5 High
On Cantina's internal benchmark of 60 held-out vulnerability tasks, apex-flash-1 solved 40 tasks, achieving a 66.7% pass@1 rate at an estimated cost of $2.38 per run. Claude Opus 5 High solved 43 tasks for a 71.7% pass@1 rate, but at $74.68 per run. Per solved task, that works out to roughly $0.06 for apex-flash-1 versus $1.74 for Opus.
The 5-percentage-point gap in accuracy is real, but the cost differential reshapes what's practical. A security team running hundreds of vulnerability scans per week faces dramatically different budgets depending on which model sits in the pipeline. For defenders who need a locally controllable model they can audit and fine-tune, the open license adds a dimension that closed APIs cannot match.
Why this matters for IT and development teams
For security engineers and DevSecOps leads, apex-flash-1 represents a concrete option for running vulnerability research workloads on infrastructure they control. The MIT license removes legal friction around commercial use and modification. The 640 GB memory requirement means this is not a drop-in replacement for smaller models - it targets teams with access to multi-GPU nodes or high-memory instances - but the per-task cost makes a clear case for organizations already running that class of hardware.
For executives evaluating security tooling budgets, the numbers are blunt: $0.06 per solved task versus $1.74. Even accounting for infrastructure costs, the order-of-magnitude gap persists. AI Security Analytics Courses can help teams build the skills to integrate models like apex-flash-1 into existing vulnerability management pipelines. The open-weight approach also means security teams can fine-tune on their own codebases and threat models - something closed API models do not permit.