AI Trust Index

AI-generated code carries 15 vulnerabilities per codebase,
on average

Tested across 16 leading AI models and 1,760 codebases, we identified that the risk is predictable by model and framework, giving security leaders the data to close the gap.

Download visual report
Explore the data
AI models tested
16
Frameworks evaluated
11
Codebases generated
1,760
Unique CWEs found
86
At a glance

What is the SCW AI Trust Index?

The SCW AI Trust Index is a living benchmark that measures how often leading AI models generate code with security vulnerabilities. It's updated as new models are released — not published once and left to age.

Download visual report
This is a dynamic heading with tag and style options

Vulnerabilities follow predictable patterns

Vulnerabilities follow predictable patterns. Across 1,760 codebases, the same weakness types recur — logging failures, injection flaws, and broken access control lead every time. AI-generated risk is repeatable, not random.

This is a dynamic heading with tag and style options

Every model has a distinct security fingerprint

Each of the 16 models leaves its own repeatable mix of OWASP categories, so teams can anticipate — not just react to — where a given model is likely to introduce risk.

This is a dynamic heading with tag and style options

No single model wins everywhere

Security leadership shifts by framework — GPT-5.1 leads in Java Enterprise API, Claude Opus 4.8 leads in Python Django. The right model depends on where a team is building.

This is a dynamic heading with tag and style options

Cost doesn't predict security

Across all 16 models, the priciest model tested (Claude Fable 5, $174.94 per run) ranks 2nd overall — while the cheapest (Gemini 2.5 Flash, $0.58 per run) scores below average. Price doesn't predict security either way.

Vulnerabilities

The most common flaw is far from exotic, and mimics pervasive human erro

86 unique CWEs were identified across the dataset. Three account for a disproportionate share of all confirmed findings.

Download visual report
CWE IDNameTrue PositivesAverage Risk
CWE-532Insertion of Sensitive Information into Log File8,54334
CWE-79Improper Neutralization of Input During Web Page Generation ('Cross-site Scripting')2,94940.5
CWE-798Use of Hard-coded Credentials1,34840.6
-Critical-or-high-severity issues per codebase, on average-4+
All 16 models

Overall Trust Index

Normalized score, 0–100. Higher means fewer vulnerabilities relative to the rest of the field — a score of 100 means fewest in this dataset, not vulnerability-free.

Download visual report

Scores for the original 6 models differ from the published RMIT study because the normalization pool expanded from 6 to 16 models. Scores are always relative to the full cohort.

Claude Sonnet 5
80.4
Claude Fable 5
76.4
GPT 5.3 Codex
75.3
Claude Opus 4.8
74.5
GPT 5.1
71.6
Gemini 3.1 Pro
66.6
GPT 5.5
66.5
Gemini 2.5 Pro
65.6
Claude Sonnet 4.5
65.2
Gemini 3.5 Flash
59.1
Devstral 2
53.1
Claude Haiku 4.5
50.7
Claude Sonnet 4.6
45.5
Gemini 2.5 Flash
39.4
Qwen3 Coder
32.7
GPT 5 Mini
21.6
Frameworks

The leader changes by framework

A model's overall rank doesn't guarantee it's the safest choice for a specific stack.

Download visual report
FrameworkLeading model
Java Enterprise APIGPT-5.1
Java SpringClaude Sonnet 4.5
Python DjangoClaude Opus 4.8
C# (.NET)GPT-5.5
CClaude Fable 5
Research

How we did this

Sixteen AI models each generated the same complex expense-management application — authentication, file upload, payments, database access — across 11 frameworks, 10 independent times per combination. Every codebase was scanned with the same open-source static analysis tools (Semgrep, Bearer, Bandit) and every finding was verified for false positives using the same agentic process.

Download visual report
Distinct verified vulnerability findings
1,554
Confirmed vulnerabilities (true positives)
27,000+
Independent SAST tools per codebase
3
Runs per model/framework pair
10
Methodology

Developed in partnership with RMIT

The original study and methodology were developed in partnership with RMIT University and covered the first six models tested. Secure Code Warrior independently extended the research to 16 models using that same methodology.

Read full methodology

Frequently asked questions

What is the SCW AI Trust Index?

A living benchmark that measures how often leading AI models generate code with security vulnerabilities. Its methodology was originally developed with RMIT University and has since been independently extended by Secure Code Warrior to 16 models. It's updated as new models are released rather than published as a one-time study.

How many vulnerabilities does AI-generated code have, on average?

An average of 15 confirmed vulnerabilities per codebase across the 1,760 codebases tested, with 4 or more typically rated critical or high severity.

Which AI model produces the most secure code?

No single model wins in every context. Claude Sonnet 5 scored highest overall at 80.4 out of 100, but the leading model changes by framework.

Does a model's API cost predict how secure its code is?

No. Across all 16 models tested, the most expensive model (Claude Fable 5, $174.94 per run) ranks 2nd overall, while the cheapest (Gemini 2.5 Flash, $0.58 per run) scores below average.

What is the most common security flaw in AI-generated code?

CWE-532 — writing sensitive information like passwords or tokens into application log files — appeared 8,543 times, more than any other flaw in the dataset.

How was the research conducted?

Sixteen AI models each generated the same application across 11 frameworks, 10 times each, scanned with the same tools and verification process every time.

Still have questions?

Support details to capture customers that might be on the fence.

Contact

Get the data your team needs to govern AI-generated code with confidence

Download the visual summary for a fast, shareable overview of the findings, or explore the full interactive dataset — by model, by framework, by vulnerability type.

Download the EBOOK
Explore the data