Skip to content

Research · Open for participation

The vibe-coded SaaS security study

We are assessing 100 SaaS applications built primarily by AI coding assistants, using one fixed test plan, and publishing the frequency of what we find. The methodology is below. The results are not, because we have not collected them yet.

Why this page has no findings on it

Plenty of security companies publish research with impressive numbers and no method. We are doing it the other way round: protocol first, results when they exist. If you are reading this before the study closes, the honest answer to “what did you find” is that we are still finding it.

Protocol

How the study runs.

Published in advance so the method cannot be adjusted after the data starts arriving.

  1. 01

    Sample definition

    Production or pre-production SaaS applications where the majority of application code was generated by an AI coding assistant. Self-reported at intake, verified during scoping by reviewing commit patterns and the team's own account of how the product was built.

  2. 02

    Target sample size

    100 applications. Results will be published at 100, with an interim report at 25 if the pattern distribution is already stable. We will publish the achieved sample size, not a rounded one.

  3. 03

    Assessment method

    A fixed test plan applied identically to every participant: data-layer policy review, two-tenant authorization sweep, built-bundle secret extraction, endpoint enumeration including routes the interface never calls, and rate limit testing on authentication and paid-API endpoints.

  4. 04

    Classification

    Every finding classified against CWE and, where applicable, the OWASP Top 10 for LLM Applications. Severity assigned by a consultant using exploitability and business impact, with the rationale recorded so the distribution can be audited.

  5. 05

    Anonymisation

    No participant is named, and no finding is published in a form that could identify an application. Published data is aggregate: frequency by category, severity distribution, and anonymised code patterns rewritten to remove identifying detail.

  6. 06

    Disclosure

    Findings go to the participant first, in full, before anything is aggregated. Participants receive their individual report regardless of whether the study proceeds, and may withdraw their data from the aggregate at any point before publication.

  7. 07

    Publication

    Full methodology, the achieved sample size, the frequency table and the limitations section published openly on this site under a permissive licence. No gate, no email capture, no summary-only version.

Limitations

What this study will not be able to tell you.

Stated up front, because a limitations section written after the results is a limitations section written to protect them.

  • Self-selection. Companies who volunteer for a free security assessment are not a random sample of AI-built SaaS, and are plausibly more security-conscious than average. This biases results toward fewer findings, not more.
  • Our test plan is fixed, which makes results comparable across participants but means we will systematically miss classes it does not cover.
  • Severity is assigned by human judgement. We will publish the rationale so the distribution can be challenged.
  • The tooling landscape moves quickly. A result about code generated in 2026 may not hold for code generated in 2027, and we will date everything accordingly.

Prior work

What has already been published.

This study is not the first look at AI-generated code security. These are the results it builds on, attributed to the researchers who produced them.

Pearce et al., 'Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions', IEEE S&P 2022

Analysed Copilot completions across scenarios drawn from the CWE Top 25 and found a substantial proportion of generated code contained security-relevant weaknesses.

Why it matters hereThe first large systematic look at assistant-generated code. Establishes that the question is empirical rather than theoretical.

Perry et al., 'Do Users Write More Insecure Code with AI Assistants?', ACM CCS 2023

Participants with access to an AI assistant produced less secure solutions on several tasks, and were more likely to believe their code was secure than participants without one.

Why it matters hereThe confidence gap matters more than the defect rate. A developer who believes the code is fine does not review it.

OWASP Top 10 for Large Language Model Applications, 2025

Consensus categorisation of LLM application risk, including prompt injection, improper output handling, excessive agency and unbounded consumption.

Why it matters hereThe reference taxonomy we classify against, so results are comparable with other work rather than using a private scheme.

MITRE CWE Top 25 Most Dangerous Software Weaknesses

Ranked list of the weakness types most prevalent and impactful in reported vulnerabilities.

Why it matters hereClassification baseline. Using an established scheme is what makes a frequency table meaningful to anyone else.

Summaries above are our characterisation of other researchers’ work, not quotations. Read the original papers before citing them; we have linked the titles and venues so you can find them.

Questions

About taking part.

Or read what we test for in AI-built applications.

Because publishing a protocol before collecting data is what separates research from marketing. It means we cannot quietly adjust the method once we see which findings would make a better headline, and it lets anyone reading the eventual results judge whether the design supports them.

Participation is free and you receive a full individual report. The exchange is that we may use your anonymised results in the aggregate, and you can withdraw that permission at any point before publication. We will also, obviously, know your product well enough to offer you further work. You are under no obligation to take it.

No, not even with permission. Naming any participant makes the others identifiable by elimination in a sample this size, and a study where being included implies you had findings is a study nobody sensible joins.

You get told immediately, not at the end of the study, and you get the reproduction steps and the fix. Aggregate publication happens long after, and only in a form that cannot identify you. If you would rather withdraw entirely at that point, you can.

Then we publish that. A finding that AI-built applications are no worse than hand-written ones would be a genuinely useful result and we would report it as readily as the alternative. Committing to publish before you know the answer is the point.

Put your application in the study.

You get a full security assessment at no cost and your results before anyone else sees an aggregate. We get one more data point toward a study worth citing.

Volunteer your applicationsecurity@innsecs.com

No sales sequence. A scoping call and a written proposal cost nothing.