Figma’s security team reviews every pull request with an AI agent at a median cost of about 50 cents, and Google Cloud runs agents that scan their repositories, write their own fuzzing harnesses and draft their own patches. Both published how they did it this year. The machinery in both cases is cheap enough for a company of ten people, and the method is written down in public.
What makes both systems work is not the model they chose. It is a document describing what “insecure” means in their specific codebase, and that document is the part you have to write yourself.
What Figma built, and what it cost them to get right
Figma’s write-up describes agents at three points in their development process: reviewing every pull request before merge, guiding developers towards safe patterns while they write, and auditing a decade-old monorepo for things nobody has looked at in years. All three run from one shared threat model, expressed as 68 rules drawn from vulnerabilities they had actually seen. Their summary of the lesson is that the policy is the threat model.
Their numbers are worth reading carefully, because they are more honest than most vendor material. Precision in the first week was 15%, meaning roughly five findings in six were noise, and they tuned it to above 70% before it was useful. Measured against 66 real vulnerabilities drawn from their bug bounty and past incidents, the system catches 64.2% when weighted by payout. The first full repository audit surfaced more than 100 latent vulnerabilities, two of them critical, that their existing static analysis had not found.
The sequencing advice in that post is the part most teams get backwards. Figma argue you should improve precision before recall, because historical bugs can only tell you what you are missing, while live developers are the only source of information about whether your findings are worth anyone’s time. A system that is thorough and noisy gets ignored within a fortnight, and once it is ignored the recall number stops mattering.
What Google Cloud built
Google’s write-up covers the same idea with considerably more money behind it. Their code scanning framework, Mantis, splits the work across specialised agents: a strategist that decides which parts of a codebase deserve attention, research agents that examine the files it selects, and deduplicator, reviewer and critic agents that discard noise before a human sees anything. It builds hierarchical summaries of a repository so that an agent can work with the context of the whole codebase without reading all of it, which Google report cuts token overhead by more than 85%.
The number that matters most in that post is the starting point. Naive AI code scanning, by their own account, ran at under 7% true positives. The multi-agent structure is what raised it to something an engineer will act on, and Google say the core skills behind Mantis are published on GitHub, though the full internal version is not.
They describe two other systems worth knowing about. One writes and repairs fuzzing harnesses without human maintenance. The other is a patching pipeline that reproduces a crash, maps the failure path, drafts a fix and puts it through a regression loop before any human reviews it. Their design review process, separately, checks product designs against a catalogue of more than 200 security requirements and escalates only the risky ones.
Notice what is doing the work in each case. A catalogue of requirements, a written threat model, a rule set. In both companies the artefact carrying the value is a document, and the agents are interchangeable machinery pointed at it. Figma runs two frontier models in parallel; Google swaps agents in and out by role. Neither company’s advantage lies in which model they called.
What a rule actually looks like
“Sixty-eight precedent-based rules” is the kind of phrase that sounds like a large project and tells you nothing, so here is the shape of one, written the way we would write it for a client.
Rule 12. Authorisation checked at the route, not the record. In
api/, any handler that loads a record by an ID taken from the request must check that the authenticated user owns or may access that specific record, in the same function, before returning it. A route-level middleware check that only proves the user is logged in does not satisfy this rule. Precedent: HDS-2025-04, where/api/invoices/:idreturned any customer’s invoice to any authenticated account.
That is the whole trick. A rule names the location, states the condition precisely enough to be checked, disqualifies the common near-miss, and cites the incident that earned it a place. Ten of those, written from your own incidents and near misses, will outperform a hundred generic ones taken from a framework, because generic rules produce generic findings and generic findings are how you arrive at Google’s 7%.
If you have no incidents to draw on yet, Anthropic’s guide for CISOs on agentic AI offers a reasonable starting structure. It suggests assessing any AI-enabled system with four questions: what untrusted content does it ingest, what actions can it take and under whose identity, what is the blast radius if it behaves unexpectedly, and what observability exists over what it did. Those questions are aimed at agents rather than at ordinary application code, so they will not populate a rule set for your API on their own, but they are a sound way to structure the first conversation.
The cost, at a realistic scale
Take Figma’s figure and apply it to a smaller company. A team of ten engineers opening 200 pull requests in a month, reviewed at roughly 50 cents each, comes to about £75 a month. A one-off audit of an entire normal-sized application, the equivalent of the sweep that found Figma more than 100 latent issues, is a few hundred pounds of inference at most.
That is less than a single seat of most commercial security tooling, and it is well within reach of a company that has no security team at all. Whatever is stopping smaller organisations from doing this, it is not the bill for the machinery.
Where we would push back
Both posts are written from a position most readers do not occupy, and it is worth being explicit about where their advice does not transfer.
Figma could afford a first week at 15% precision because their developers already trusted the security team and had reason to be patient. A five-person startup does not have that credit, and a fortnight of noisy findings will teach the team to ignore the tooling permanently. We would run the first two weeks entirely in private, with one person reading every finding before any developer sees one, and only widen once precision is above half.
Google’s architecture is the right shape and the wrong size. Six agent roles, hierarchical summarisation and a self-improving knowledge store are answers to the problem of scanning an enormous codebase continuously. Copying that structure onto a single application adds cost and failure modes without adding accuracy. The transferable idea is narrower: run cheap filtering before expensive judgement, and measure your true positive rate honestly from the first day.
Neither company addresses what happens when the threat model is wrong rather than incomplete, which we think is the more common failure in smaller organisations. A rule set built from the vulnerabilities you happened to notice will faithfully keep finding that same class while missing the one that eventually hurts you. We have no clean answer to this beyond reviewing the rule set against incidents from outside your own history.
What to do first
We have run this sequence on client codebases and on our own, and the step that consistently takes longest is the first one. Writing ten precise rules from real incidents is a day of work for someone who knows both the codebase and what has gone wrong in it, and it is the step people skip.
The gap between what Figma and Google can do and what a forty-person company can do is no longer a technology gap. Both published the method, the inference costs pennies, and the models are available to anyone this afternoon. What the larger companies have is someone whose job was to write down what insecure means in their particular system, and that is a cheaper problem to fix than the one most people think they have.
We work with SaaS teams and startups to turn real incidents into a written rule set your tooling can act on, then help you get agentic review running against it. Free 30-minute call, no commitment.
Book a free call