Auditing code with AI agents

Published:  05/10/2026 11:00

Introduction

Did you know AI can write insecure code?

I've had frontier models adding unsanitized user provided command line arguments to a runtime or even a database query, not doing much in terms of securing said user input.

However, the same model should be able to find the vulnerability as well, the problem being that you have to ask it to because as much as it looks like the model thinks for you, it does not.

The agentic workflow

Modern AI assisted software development uses AI coding agents, providing a harness around a model with tool-use capabilities to read, write, execute, browse the web, do whatever your computer can do really (which is another security issue on its own but we'll talk about it another time).

The agentic software will also provide context compaction and most importantly the logic around having the model loop until it achieves a certain goal or task.

Agentic workflow schematic with subagents

In truth there's way more than that, including:

  • Coordinating multiple agents or subagents;
  • Follow a plan, step by step;
  • Ask specific design questions to operator;
  • Connect to MCP servers, LSPs, ...

The field of possibilities when it comes to different models and context/prompts combined with the existence of multiple agentic software can make this hard to do effectively.

Sticking to one of the flagship TUIs is probably a good idea at this time.

The Codex agent from OpenAI in a terminal

For this article, we'll use TUI agents and especially Opencode.

The most popular of these is probably Claude Code.

Sample project

The sample project is a simple Go API about creating "tasks" (it's basically a variant on the TODO list).

It was originally created by having an agent follow the docs/overview.md file.

A few obvious vulnerabilities were added to it for testing purposes as we'll see later.

Link to the repository: https://github.com/net7be/vulnerable-demo-app

AGENTS.md

Before we start helping agents audit our code, we have to mention the AGENTS.md file.

It became a standard file that AI agent software look at when started in a project directory.

It's a markdown file that basically serves as README.md but for AI agents.

They're often kept short, laying out a high level overview of the project and what it does, some of the project structure as well as commands to build, run and test the project.

Coding style and conventions are often found in that file, as well as some things that the authors do not want the AI to do.

The file can link to other, more fleshed out markdown file with more in-depth documentation or outside (web based) documentation sources.

How the file will be interpreted exactly depends on the model, the agentic software and what you're asking it to do.

Our sample project has a simple AGENTS.md:

# AGENTS.md

## Style guidelines
- Use `gofmt` to format all written code

## Building the project
```
cd cmd/tasks-server && go build
```
The binary `tasks-server` or `tasks-server.exe` is created.

## Project structure
- Keep cmd/tasks-server/main.go minimal
- Most code goes in packages in `pkg`
- Can save reports or artefacts in `docs`

Summary of project structure can be found in `docs/structure.md` and the initial blueprint is in `docs/overview.md`.

Skills

Skills are just more markdown files to be conditionally added to the model context.

They have a frontmatter with the fields name and description being mandatory.

Other usually accepted fields are described in the official specification.

Agentic software will look for skills either in globally known places, such as ~/.claude/skills/<skill_name>/SKILL.md or in specific places local to the current directory, for instance in .agents/skills/<skill_name>/SKILL.md.

Using .agents instead of .claude is supposed to create a better standard for the different agents TUI available.

The agent generates a special initial prompt with the available skills using all of the discovered descriptions and names to let the model know about them.

The model will then automatically attempt to load the skill when relevant to its current prompt.

In essence, skills are akin to just copy/pasting the same prompt over and over again when you want the agent to perform a certain task.

They can include directives, provide guidance through a project or documentation, provide links and more importantly, commands to use and code examples required to perform certain tasks.

They can be like a manual explaining how to do certain things or use certain tools.

As with AGENTS.md, they may also describe actions that the model should not do.

Other files from the skill directory can be referenced by SKILL.md as well.

Lots of skills exist and they're very free-form due to the nature of LLMs and available tooling.

Some skills repositories exist but one should always carefully read the files before using them as skills have been a source of prompt since their inception.

Proposed security audit skill

A few of these obviously already exist.

Our belief is you should create your own, maybe with tools that are specific to your projects and places that you know require extra attention.

Would you want to go further, we recommend having a look at the Cloudflare security-audit-skill which is under MIT license so you're free to use that one as a starting point.

However, that skill alone is a lot of tokens already, on top of possibly having the review the whole codebase.

Depending on your own token-abundance situation you might want to start with something lighter.

Our version is only 37 lines long, feel free to use it a starting point as well or to make it even smaller.

To run the audit on the entire project, prompting "security audit" or "run security audit" should work.

Example showing the skill being loaded:

opencode showing the message "loading skill security-audit"

You can refine the prompt to scope it to parts of the project or the current changes in a git repository, and that should use less of the context window and can produce different results sometimes, especially if it can happen with no compaction.

For what we're usually doing we don't like global skills being loaded when we don't want them to be so we tend to just copy (or link) the needed skills to the project directory when needed.

That approach also allows for easy local modifications to skills with the downside that you have to update the skills in all of your projects when a global change is to be made.

Expected findings

Linking the project again for those interested: https://github.com/net7be/vulnerable-demo-app

Most obvious expected findinds:

  • SQL injection — in store.go — Allows extracting password hashes or anything else in the database really;
  • Path traversal — in handlers.go — Makes it possible to browse the filesystem (but not read or write anything).

Whether the path traversal is considered critical or not is a matter of opinion but it's not good in any cases.

We also got a few less important issues:

  • No throttling for login endpoint — Makes brute-force attacks easier;
  • Possibly too much information in error messages — In store.go some databases error also print the Go error description;
  • Dockerfile copies the whole project directory with no .dockerignore — Increases the risk of Docker images having local secrets on them or otherwise unwanted data (not (yet) the case in the project itself).

The last one is actually pretty bad because a local database file will be copied to the Docker image.

We also planted logic flaws or questionable design choices or whatever you want to call these:

  • Tasks can be assigned to another user by anyone — There's an "admin" level user so that is a bit strange especially since task deletion is only available to admin users — A comment also says the endpoint is admin-only but it's not;
  • Getting the tasks list doesn't require being logged in — A bit strange but could be a design choice;
  • A task marked as done can be reassigned — Not a security issue but a logic bug.

Results for a few models

We ran a full audit of the example project with a few models including some local ones, albeit with very limited maximum context.

GPT6 Luna (on Codex)

We ran GPT 6 Luna on Codex and it produced one of the least verbose reports from all the models.

Keep in mind it's the "cheapest" of the newer OpenAI models.

For some reason, it also didn't format the report as markdown even though it's required by the skill. But you can always ask it to do it in a followup request.

Findings summary for Luna:

Vulnerability found Priority
SQL InjectionHigh
Task assignment allowed to non-adminHigh
Tasks list is publicOther
Path traversalOther
Login endpoint is brute-force vulnerableOther

It also noted that HTTP request bodies are capped at 1MB, which is true, and that password hashes are omitted from JSON responses due to the "-" decorator in the User struct.

A one-liner fix was proposed for each point.

The full report can be found here.

Big Pickle

Big Pickle (the ridiculous name is intentional) is one of the free models currently offered through Opencode, it's just that they keep the underlying model secret (it might be changing too). Keep in mind they can use your data for training when using these models.

The model produced a detailed report while still keeping things short, including the proposed fix for each issue.

Vulnerability found Priority
SQL InjectionHigh
Path traversalHigh
Task assignment allowed to non-adminHigh
Tasks list is publicHigh
Any user can mark any task as doneMedium
Auth tokens never expireMedium
Login endpoint is brute-force vulnerableMedium
Server doesn't use HTTPSMedium
No .dockerignore and "COPY . ." in DockerfileMedium
Secret written to stdout on first runLow
Plaintext password returned when creating usersLow
Store details in error messages - Though not sent through APILow
Discloses DB location at startup on serverLow
Generated password have low entropy (debatable)Low
handleCreateTask decodes the request body before checking for authenticationLow

I hadn't noticed that last one when creating the app but the model is right, that order of operations is strange.

The initial admin secret being printed to console is part of the spec and is fine but it's nice to point it out, so are the plaintext passwords from the create users endpoint.

Big Pickle basically found everything I wanted in what I'd still call a concise report.

GLM 5

GLM is currently very strong in hacking and coding benchmarks. In essence, it should be better than Big Pickle but since the previous model did find everything we wanted already, maybe it'll be similar.

We tried GLM through https://build.nvidia.com/ as it's one of the free model. Keep in mind they can use your data for training or review.

We got the same base findings as the previous model except this one doesn't care about HTTPS (which is fine) or the plain text password returned when creating users.

It also doesn't care about the error messages or the database location being logged at start, which is also fine.

Interestingly it didn't find the logic error in the order of operations in handleCreateTask. But it found extra stuff:

Vulnerability found Priority
/reports requires no authenticationOther
No HTTP server timeoutsOther
Brittle unique-violation detectionOther
Logic error in .gitignoreOther

The last 3 are quite interesting.

We had no idea the Go HTTP server doesn't have any timeout by default. But it doesn't. So, when publicly exposed, a denial of service is possible by opening a very high amount of connections until some OS limit is reached (e.g. the max amount of opened files).

We always put Go APIs behind a proxy but that is still important to know. Cloudflare has a write-up on the subject.

Next, the unique violation issue is related to how unique constraint violations are detected in store.go. Specific error text is looked in from the error description whereas we should check the error code instead.

That issue is more code quality than security but it's interesting to know.

The logic error in .gitignore is related to the presence of reports/*.txt when txt files exist in subdirectories such as reports/past/report1.txt.

That was a mistake from our part when creating these files and it went under the radar, GLM found it, underlining how one can miss questionable elements in the initial code review when the review scope is an entire new project.

There's a price to pay when generating a lot of code fast. Hopefully the productivity gain makes up for it.

Creating a separate skill dedicated to code quality review would also be a good idea, and many of these do exist.

The full report for GLM 5 is on the repository.

Writing proof-of-concept exploits

GLM 5 is the only model that actually wrote exploits to test the vulnerabilities, successfully extracted a password hash and assigned a task to someone else, before cleaning up all of these changes afterwards and removing all the temporary files.

It's pretty clear these models were trained on hacking and do not have extreme guardrails like some other mainstream ones (for instance the Microsoft Phi models just refused to run the audit). We have to be extra careful with the permissions given to this agent.

Here's the working SQL injection it wrote in Golang:

SQL injection routing written by the AI including the creation of a test user

It made a test user to avoid using real data.

NB: If all of that makes you nervous, agent software often have a special "explore" or "plan" agent that isn't allowed to edit anything and (normally) won't ask you the permission to do so.

The audit can run as that agent with no issue, it just won't be able to write the report but you can switch to the default agent and ask it to write it at the end.

MiMo 2.6 Flash

Another Chinese open weight model, MiMo Flash is fast and basically found everything Big Pickle did, just missing a few of the extras GLM found.

It did mention we may have a theoretical user enumeration by timing stating that the bcrypt routine only runs when the username exist. So, by measuring the response time it could be possible to test whether a username exists or not.

That "attack" is very interesting. They suggest always running the bcrypt function even when the user doesn't exist but against a dummy hash. That makes sense.

Link to the full report.

Local model: Qwen 3.5 9B params

While we're at it, let's test the full audit on a modest local model.

Depending on the project size you will need to extend the context window, 32k tokens being the minimum.

On most projects it'll be necessary to scope the audit to avoid having too much compaction.

Local audits are interesting for sensitive projects to avoid sending all of your data to some remote service. As a reminder, Anthropic always uses your data for training by default, even on consumer paid plans. You have to manually opt out, and trust them.

The model got these findings very quickly:

Vulnerability found Priority
SQL InjectionHigh
Store details in error messagesHigh
/reports requires no authentication & Path traversalHigh
Task assignment allowed to non-adminHigh
Plaintext password returned when creating usersHigh
Discloses DB location at startup on serverHigh
Login endpoint is brute-force vulnerableHigh
No input sanitization for JSON (non-issue)Medium
Missing security headersMedium
No HTTP server timeoutsLow
Auth tokens never expireLow

The report is a bit too verbose and priorities are a bit strange but otherwise it's pretty good.

For some reason it diluted the path traversal in the fact that /reports doesn't require authentication. But the path traversal itself is way worse of a vulnerability.

It's the first time headers like X-Frame-Options are mentioned though these are better handled by a reverse proxy.

Qwen didn't mention the absence of a .dockerignore file but did find the missing timeout settings for the HTTP server that only GLM found otherwise.

Scoped audit with GPT OSS 20B

As an exercise I tasked another local model to just look at handlers.go which has all of the public entry points.

The workflow was extremely fast as it basically ran the grep commands I would've thought of:

The model is grepping various patterns like looking for "sql."

It found the directory traversal immediately as well as the hint that handleLogin returns the entire "user" object. It ignored the fact that the password hash is excluded from serialization but that's still a good point to double check.

The full report is on the repository.

Conclusion

The experiment has taught us a few things about Golang and its HTTP server which is always nice.

It looks like you need to run the audit with several models to get the best possible overview, and sometimes run the audit several times with the same model (might be fixed by temperature settings?).

For instance, only MiMo Fast told us about possible user enumeration because bcrypt is only invoked when the user exists, and bcrypt is artificially made to take some time to compute.

As a side note, none of the models found it strange that a task marked as done could be reassigned to another user. That kind of semantic weirdness is lost to LLMs unless they're seriously guided into finding it so let's not pretend the technology can think for us because it can't.

On the other hand, I believe there will be elements we also miss as humans when reviewing code, and hopefuly they're not the same as the ones the AI missed.

Scoping the audit is also a good idea when you know where to start exploring. The audit can also be planned in multiple stages with some explorer agent generating a map of the most interesting (and susceptible to have security issues) places to explore in the code.

Using less of the context window is always a boon when running this kind of operation, especially with very memory-constraint local models.

As is often the case, finding the right prompts, tools, skills and model combinations will produce better results than just throwing the most expensive model out there at the problem and there isn't any standard (yet) as how to make AI audit code.

That means it's the best time to go out and experiment!

Comments

Loading...