Are you a CISO or Engineering Leader worried about the Security and Cost of LLM code generation?
.png)
Review the SCW AI Trust Index, our proprietary LLM benchmarking data, before going all-in on an AI model
The cyber industry is in a state of overwhelm, with tools like Anthropic’s Mythos (and the swiftly caged-then-released Fable 5) proving to be a catalyst for many security leaders to start rapidly augmenting their security programs to keep up with the vast changes in the speed of exploitation, not to mention the much higher volume of code being produced, brought about by AI technology.
In these circumstances, it can be compelling to dive into preventative (and, let’s be honest, reactive) security measures to combat a perceived onslaught of AI-powered breach risk. However, there is a far more pervasive issue today that likely affects every CISO: the choice of LLM models in software production environments. While the consequences of shadow AI in general are well documented, this is only part of the concerning risk picture. We shouldn’t be paying tokens to generate code, paying more tokens to use AI to find secure coding issues, and then paying tokens again to generate and apply security fixes. Every AI model is susceptible to producing weaknesses and vulnerabilities at varying rates, especially when comparing performance across frameworks or programming languages. In this aspect, a determination needs to be made on the right model choice to support the existing tech stack, languages, and frameworks in use, and this information has been difficult to source, until now.
Secure Code Warrior partnered with key researchers at RMIT University, and we’ve spent months asking a deceptively simple question: when you hand your codebase over to an AI model, exactly how much security risk are you inheriting? The answer, based on our empirical study of 1700+ complete application codebases generated across leading frontier and budget models and eleven language frameworks, is: more than you think, and in ways that will surprise you.
These are the five things you must take away from our research paper as you bring your security program into the AI era:
1. Cost is not a proxy for security.
It is tempting to assume that paying more buys safer generated code. Our data shows that the truth is more nuanced. GPT 5.1 achieved the highest overall normalized security score in our original study — 79.6 out of 100 — at an average cost of $5.60 per run. Claude Sonnet 4.5 scored 71.2, meaningfully lower, at $43.99 per run. That's nearly 8x the cost for a worse security outcome.
The cost gap isn't driven by per-token pricing alone. The Anthropic models executed significantly more tool-calling turns within the agentic coding loop, generating 176 million output tokens across Sonnet 4.5's runs compared to 42 million for GPT 5.1, as an example. More activity, more tokens, and more cost did not provide more security. Budget is not quality, so don’t let it be a substitute for measurement.
2. Every model has a vulnerability fingerprint.
Across more than 1700 codebases and 16 models, we identified 83 unique CWE categories with validated true-positive findings. These aren't random; each model exhibits a systematic pattern that persists regardless of framework or language. In our original 6-model study, 39 of the 87 CWE categories identified were produced by every single model tested, showing that a large core of vulnerability types is model-agnostic, not vendor-specific. Anthropic’s Claude models, for example, showed a recurring tendency toward hard-coded credentials (CWE-798). Gemini 2.5 Flash consistently left debug code in production builds (CWE-489).
These fingerprints suggest training-data- or alignment-driven biases rather than random noise. That predictability is actually useful: once you know your model's blind spots, you can configure targeted scanning rules and post-generation checks around them.
3. Some vulnerabilities are universal… and no tested model avoids them.
For example, one of the most pervasive findings in our current living index is CWE-532 (Insertion of Sensitive Information into Log File), and it has appeared as a true positive 6,765 times.
Every model worked from the same ordinary application spec, the kind of plain-language brief a product team writes every day, not a security checklist and not a stripped-down prompt designed to induce failure. Rather than indicating an exotic edge case, most of the rest of the vulnerabilities show failures of omission: credentials left in source, overly permissive defaults, missing authorization checks; the kind of gap that closes the moment you explicitly ask for it, because the model won't infer it from an ordinary spec on its own. The exception is the single biggest finding itself: logging sensitive data isn't something models forgot to add, it's something they actively wrote, unprompted, the same way an unskilled human developer might. This is entirely volunteered by the model, without explicitly being asked by a human to adopt that behavior. The model faithfully implements what is described in the prompt, including everything not mentioned from a security-hardening perspective.
4. Framework context shapes model performance.
This finding should reshape how organizations think about AI-assisted development. The study makes it clear that model strengths vary across technology stacks. A model that performs well in one framework can underperform in another, so teams should evaluate models where they’ll actually be used.
In our current SCW AI Trust Index, the gap between frameworks dwarfs that between models. JavaScript Basic — aggregated across all 13 models we've tested — carries a total risk score of 6,770. C#/.NET Basic, the safest framework in the study, sits at just 111. That's a 61x difference in security risk, driven solely by framework choice.
Our original study found the likely reason why: frameworks with opinionated security defaults, e.g., C# with ASP.NET's built-in anti-forgery tokens, Java Spring's declarative security model, produced dramatically safer code regardless of which model generated it. Raw environments like JavaScript Basic amplify model differences catastrophically. Choose your framework with security intent first, if possible, in your organization.
5. No model wins everywhere, and model selection must be framework-aware.
In our original study, GPT 5.1 led overall, but ranked last in Swift iOS. Sonnet 4.5 produced the safest Java Spring code in the entire study, then ranked last in Java EE API. Gemini 2.5 Pro dominated Python and Swift, yet was middling in Java. There is no universally safe model.
Organizations with polyglot codebases need to either accept model-specific blind spots across their stack, or adopt a routing strategy that matches each framework to the model with the strongest security profile for that context. And regardless of which model you choose, security scanning is not optional. Even the best-performing model produced validated vulnerabilities across every single framework we tested.
That cost pressure isn't going away. As agentic workflows scale, so does the bill: more tool-calling turns, more regeneration cycles, more tokens per task. It's a fair bet that rising per-task and per-token costs will push some organizations toward older or cheaper models just to keep spending under control. Our data says that this decision needs to be made with evidence, not instinct: GPT 5 Mini was both the cheapest model in our original study and the least secure model we tested. Cost pressure is real, and you must let it drive you to measure rather than guess.
The question has never been whether to scan; rather, security leaders must prioritize gathering the right data, risk signals, and AI observability to know what to scan for in the first place. This is where AI governance plays a major role, and we have a comprehensive solution to deliver this vital information right now.
>>> LEARN MORE: SCW Trust Agent: AI
>>> DOWNLOAD: eBook
>>> DISCOVER: Our benchmarking microsite (full research paper will be linked here)
Govern AI-driven development before it ships
Measure AI-assisted risk, enforce secure coding policy at commit, and accelerate secure delivery across your SDLC.
Explore more blogs
우리는 이 방법을 잘 알고 있습니다. 우리는 이 두 가지 축복을 골고루 살기 위해 노력하고 있습니다.
%252520%252520(3).avif)
Supercharged Security Awareness: How Tournaments are Inspiring Developers at Erste Group
Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique. Duis cursus, mi quis viverra ornare, eros dolor interdum nulla, ut commodo diam libero vitae erat. Aenean faucibus nibh et justo cursus id rutrum lorem imperdiet. Nunc ut sem vitae risus tristique posuere.

Security as culture: How Blue Prism cultivates world-class secure developers
Learn how Blue Prism, the global leader in intelligent automation for the enterprise, used Secure Code Warrior's agile learning platform to create a security-first culture with their developers, achieve their business goals, and ship secure code at speed

One Culture of Security: How Sage built their security champions program with agile secure code learning
Discover how Sage enhanced security with a flexible, relationship-focused approach, creating 200+ security champions and achieving measurable risk reduction.
Secure AI-driven development before it ships
See developer risk, enforce policy, and prevent vulnerabilities across your software development lifecycle.

