The Model You Point At Is a Security Decision
Date: 2026-09-28
We shipped ClawHunter several weeks ago. It's been running against real codebases since, and it's held up. Recently, we enhanced a feature most teams would treat as a performance optimization: offloading the reasoning-heavy phases of scans and fix recommendations to a different model than the one running the scan, and not just cloud-hosted models.
When you point a security analysis at a model, you're making important data-flow decisions. The code under analysis, including any secrets it contains, goes wherever the model runs. A local model means the code stays in your network. A cloud model means the code crosses the wire to a vendor's servers. That's not a performance tradeoff. It's a risk decision, and you should make it deliberately.
While building out our own infrastructure, we forced ourselves to answer questions many security teams never ask of their own tooling.
The problem with a feature
ClawHunter is an attacker-first static analyzer. It starts at entry points, traces data flows forward to dangerous sinks, and then actively tries to disprove what it finds before reporting anything. The method comes from Capital One's VulnHunter, and it works. See our original ClawHunter writeup for more on the approach.
Phase 2, the hunt and falsification phase, is where most of the reasoning lives. You're holding a data-flow chain in context across multiple files, following every path, checking each control, and deciding whether a finding survives. Initially, that meant one of three cloud providers: Anthropic, OpenAI, or xAI. All three required an API key. All three took the full prompt, which includes the code under analysis, and ran it on their infrastructure.
That may be fine for public code, but it introduces risks for proprietary code.

This month, we added a fourth provider class: keyless, OpenAI-compatible endpoints on private networks. Reached through a tunnel, the model runs securely on your own hardware, and the code never leaves your environment.
The moment you can route a scan to a local endpoint, you must answer: which code goes to which class of provider? Answers depend on a code's sensitivity, not just a model's capability.
Where the code actually goes
Here's the thing about "local vs. cloud" that gets lost when it's described as a configuration option.
The model isn't only a function you call. It's a place your data goes.
When you send a scan to a cloud provider, you're making the same decision you'd make when you paste a proprietary stack trace into a public tool. The code is the payload. The secrets in it are the part you didn't mean to share. The difference is that a scanner does this to the entire codebase, on every run, by default, and then hands you a report that's at least as sensitive as the code itself.
A local provider inverts that. The model lives inside your boundary. The data-flow path is contained. The secret that was riding along with the code stays riding along with it, inside your network, where your controls already apply.
I don't say this to argue that local is always right. It isn't. For a half-million-line microservice, the gap between a model that can hold the full data-flow chain in context and one that can't is still real. Local models have closed most of that gap, but not all of it, and the calculus changes as the capability floor keeps rising. The choice should be a data-flow decision first and a capability decision second. Many teams decide it in the reverse order, or not at all.
The governance gap
This is a small example of a larger pattern I keep seeing. AI-powered security tools make data-flow decisions that security teams never explicitly consider. The scanner is a tool. The tool routes code to a model. The model runs somewhere. That "somewhere" is a data-flow boundary, and it should appear in the architecture diagram, the risk assessment, and the incident response plan.
Most teams have a data-flow diagram for their application. Few have one for their security tooling. But the scanner is part of the stack. Where it sends data is part of the risk model, and treating it as a performance setting is how a proprietary codebase ends up in a vendor's inference pipeline on the first scan, in the moment, under time pressure.
The question I'd put to any security team isn't "should you use local inference?" It's "do you have a policy for which code goes to which class of provider, and is that policy documented?" If the answer is no, then the first scan of a proprietary codebase is the first time you're making that decision. And that's when the wrong answer is most likely.
Next steps for your strategy
A few things to tackle when strategizing AI-powered analysis, in the order I'd approach them:
- Draw the data-flow path for the tooling. Where does the code under analysis go? If the honest answer is "to a model" and you can't name where that model runs, you don't have a data-flow diagram yet. Draw it.
- Classify the code before you classify the model. Public open-source code is low-risk to send to a cloud provider. Proprietary code is intellectual property and high-risk. Regulated code, as defined by financial, healthcare, government, etc. regulations, is the highest risk. Match the provider class to the code class, not the other way around.
- When using cloud providers for security analysis, document it. The data-flow decision belongs in the risk assessment, not in the tool's default config. "We sent the proprietary codebase to a cloud API for analysis" is a fact that belongs in the incident response plan, because someone will ask it after an incident.
- Treat the local endpoint as a security boundary. The endpoint's address, the tunnel, and the host it runs on are now part of the architecture. A compromised local inference endpoint is a compromised security tool.
- Review the provider config the way I'd review an egress policy. In ClawHunter's config, the
phase2_modelfield is a routing decision. It determines where your code goes. It deserves the same review as an outbound rule. - Protect the output like the code it describes. A vulnerability report with exploit traces is a security artifact. Restrict access to it, don't commit it to public repos without redaction, and treat it in the incident response plan the same way you'd treat a penetration test report.
The point
The interesting part of adding a local inference provider wasn't the capability. It was the questions it raised: which code goes to which provider class, and are the decisions documented? Most teams don't have such a policy. The first proprietary scan is the first time they're making it, in the moment. That's the governance gap.
ClawHunter is a tool, not a verdict. The findings it produces start another round of review, not an automated determination. But you should answer the data-flow questions before running scans, not after.
Repository: github.com/isbitski-klm/clawhunter · Upstream methodology: Capital One VulnHunter
