On Teslas and Tensors
Mitigating AI hallucinations and misalignment with architecture
Last week, one of my AI assistants told me a Tesla Roadster reservation deposit was $5,000. That number wasn't simply false. Tesla allowed a $5,000 partial deposit for 24 hours, then you had to pay the other $45,000. The total deposit was still $50,000. The model treated $5,000 as the whole deposit and wouldn't know the difference. That distinction comes from human experience and other domain knowledge, not from the model. And the more unsettling part: it's confident about the incomplete figure. It had repeated the number across three separate messages, each one more certain than the last. Confidence and repetition hid the gap. When I caught and flagged the mistake, the agent said "you're right" and moved on, as if nothing had happened.
Three days later, the same assistant was telling me Starship Flight 14 had already launched. It hadn't. The date had slipped, and the model had anchored on an early search result and kept repeating it. Every time I asked about SpaceX, the answer was more confident. Every time I checked the daily log, the wrong date was there, reinforcing the next wrong answer.
Models compute over tensors: structured arrays of values that look coherent even when a single entry is wrong. Architects design systems that decide which values to trust, check, and act on. This isn't a story about one bad search result. It's a story about a structural failure mode that every team using AI for critical decisions needs to understand, and an architecture pattern for mitigating hallucinations and misalignment.
The mechanism: self-reinforcing loops
Here's what happens inside the model, in simplified terms:
- A fact enters the context. "Flight 14 is September 22."
- The model generates a response that includes that fact. "Starship Flight 14 is tomorrow."
- That response is logged. The next prompt includes the previous response as context.
- The model now has two sources for the same fact: the original search result and its own previous output. It treats them as corroboration.
- The next response is even more confident. "Flight 14 was today. Here's what happened."
Each repetition makes the next one more confident. The bad value gets broadcast through context the way a corrupted tensor entry propagates through a computation: the shape still looks right, so nothing fails loudly. The model is checking its own previous output, which is circular. And it looks like verification.
This is not a bug. It's the natural behavior of a model optimized for coherence and fluency. The confident-sounding wrong answer flows better than the qualified, uncertain one. The model doesn't have a well-developed "I don't know" muscle. It has a very well-developed "here's the answer" muscle.
The failure mode isn't just wrong answers. It's wrong answers delivered with the confidence of right ones.
Why this is worse with AI than with humans
Humans hedge. They say "I think," "I believe," or "I'm not sure." The hedge is a signal: "this is my best guess, verify it." AI doesn't hedge the same way. It's optimized for fluency and completeness. The confident-sounding wrong answer flows better than the uncertain one.
And there's a second layer. AI can generate a plausible-sounding wrong answer faster and more fluently than a human can. A human might say "I don't remember the exact deposit amount." An AI will say "$5,000" because it's a fluent completion. The confidence is real in the sense that the model is genuinely producing that token with high probability. But high token probability isn't the same as high factuality.
Fluency is a risk factor. Treat AI fluency as a risk signal, not an accuracy signal. The smoother the answer, the more deliberately you should verify it before you act.
The security angle
Potential pitfalls we can run into with security use cases include:
| # | Use Case | Pitfalls |
|---|---|---|
| 1 | Triage CVEs and prioritize patches | A wrong CVSS score or wrong affected-version range means you patch the wrong thing. A confident wrong answer about which version is vulnerable means you leave the vulnerable one exposed. |
| 2 | Analyze threat intelligence | A wrong actor attribution or wrong TTP means you defend against the wrong attack. A confident-sounding false positive about an APT group means you waste weeks investigating a ghost. |
| 3 | Generate incident reports | A wrong timeline or wrong affected-system list means the post-mortem is wrong. The next incident will be handled with the wrong playbook. |
| 4 | Write security policies | A wrong compliance requirement means you're either over-building (wasting budget) or under-building (leaving gaps). |
| 5 | Review code for vulnerabilities | A false negative (AI says "this is fine" when it isn't) is the most dangerous failure mode. The AI's confidence makes you skip the manual review. |
In every one of these cases, the cost of a confident wrong answer is higher than the cost of a qualified "I'm not sure, let me verify." The AI's fluency is working against you. That's an architecture problem, not a prompting problem.
The three-layer verification architecture
The fix isn't "be careful with AI." It's an architecture pattern for mitigating hallucinations and misalignment. Three layers, each with different failure modes, each catching what the others miss. Think of it as reshaping the claim space before a decision ships: generate, independently check, then apply human judgment that no second model can fake.
| # | Layer | Role |
|---|---|---|
| 1 | Primary AI (generator) | Writes the initial analysis; fast, broad synthesis across sources |
| 2 | Verification AI (independent check) | Independently checks the primary AI's output before it reaches a human |
| 3 | Human-in-the-loop (independent verifier) | Applies domain judgment with genuine independence no second model can fake |
Layer 1: Primary AI (the generator)
This model writes the initial analysis, the CVE triage, the threat brief, the policy draft. It's fast, broad, and synthesizes well. It's also the layer most prone to the self-reinforcing loop described above. Once a fact enters its context, it becomes the "source of truth" for every subsequent mention.
What it does well: speed, breadth, pattern recognition across large document sets, synthesis of multiple sources into a coherent narrative.
What it fails at: confidence without verification, circular self-checking, narrative drift over time, anchoring on early search results.
Layer 2: Verification AI (the independent check)
This is a second model from a different provider with different training data and blind spots. It should verify the primary AI's output before it reaches a human.
This is the pattern the research community calls LLM-as-a-judge. The idea: instead of asking one model "is this right?" you ask a different model to evaluate the output. It's a real technique with published results, and it works. UC Berkeley's MT-Bench and Chatbot Arena work found that a strong LLM judge like GPT-4 agrees with human preferences at over 80% on open-ended tasks, about the same rate human judges agree with each other. For a security analyst, that's a meaningful signal. You're not asking the model to guess. You're asking a second, independent model to check the work.
What it adds: independent model, different training corpus, different search behavior. If the primary AI hallucinates a CVE number, the verification AI will likely generate a different one. If both agree, that's meaningful. If they disagree, you've caught the hallucination.
What it still gets wrong: the LLM-as-a-judge literature documents several known biases. Position bias: the judge favors whichever response appears first, regardless of quality. Verbosity bias: longer answers get higher scores, even when they're less accurate. Self-enhancement bias: a model scores its own outputs higher than others, a pattern the Berkeley team explicitly identified as a limitation of the approach. And the ICLR 2025 work by Chen et al. goes further: they show that high agreement between two judges does not imply accurate scores. Two models can agree while both are wrong, and the agreement can look like verification while masking a shared error.
The failure mode: correlated hallucinations. Both models make the same error because they have the same blind spots. The confidence of agreement masks the shared error. "Grok confirmed it" sounds like verification, but if both models drew from the same search results, they're corroborating the same bias, not checking it.
The practical implication: Layer 2 is a real improvement over nothing. It's not a substitute for Layer 3.
Layer 3: Human-in-the-loop (the independent verifier)
This is the only layer with genuine independence. And it's not just a safety net. It's where domain knowledge does work that no model can do.
The Roadster test. My AI told me the Tesla Roadster reservation deposit was $5,000. That figure wasn't simply false. Tesla allowed a $5,000 partial payment for 24 hours, then you owed the other $45,000. The total was still $50,000. The model treated $5,000 as the whole deposit and wouldn't know the difference. I caught the gap because I know what a $200,000+ supercar costs. Treating $5,000 as the full ask on that asset class doesn't fit, and my intuition flagged it immediately. No model needed to tell me that. I didn't know the exact deposit. I didn't search for it. I knew the vehicle's speculated price range, and the incomplete figure didn't fit. That kind of judgment separates a domain expert from a generalist, and a second AI can't replicate it. A second AI might also say $5,000 and stop there, because it doesn't have the same mental model of how that deposit actually works.
What only a human catches:
- Institutional context ("that vendor doesn't actually use that library")
- Client relationships ("I talked to their CISO last month, and they already patched this")
- Retracted advisories ("that CVE was a false positive that was pulled last week")
- Narrative detection ("that outlet has a pattern of inflating severity for clicks")
- Price-sanity checks ("that number doesn't fit the asset class")
- The "that doesn't sound right" intuition that no model has
Why this layer is non-optional in critical systems: the first two layers are both AI. They have the same structural biases. The human layer is the only one with genuine epistemic independence. Without it, you have two AIs checking each other, which is corroboration, not verification. The research is clear: LLM judges are good, but they're not human. They don't have the institutional memory, client relationships, or "that doesn't fit" intuition a domain expert brings. They're a useful check, not a substitute.
That's why we need architects more than ever. Not people who write prompts. People who design a verification architecture pattern for hallucinations and misalignment so fluent wrong answers don't become operational truth.
The framework: five rules for the three layers
These are established principles that most builders and operators have never been taught. They're not new. They're not theoretical. They're the same checks that a seasoned security analyst does instinctively before trusting a piece of information, grounded in critical thinking, and formalized for the AI era. The Roadster deposit and the Flight 14 date aren't edge cases. They're examples of the exact failure modes this framework prevents.
| # | Rule | What it means |
|---|---|---|
| 1 | Two-source minimum | Any fact that drives a critical decision needs two independent sources. One source is a claim. Two is data. Independent means different organizations and incentives; two articles from the same outlet are one source. |
| 2 | Authority hierarchy | Weight by authority, not by search ranking: vendor advisory > CERT/CC > NVD > news outlet > blog post > AI training data. A vendor's own advisory is more authoritative than a news article about it. |
| 3 | Temporal decay | The longer a fact sits unverified, the more stale it gets. Any fact in context for more than 24 hours gets re-verified before being used again. |
| 4 | Narrative detection | Ask who benefits before you repeat. Media (clicks, bias), vendors (sales, liability), lobbyists (regulatory capture), and AI fluency all push toward the confident answer. |
| 5 | Reverification gate | Before you act on a fact, re-verify it against the primary source right now (vendor advisory, CERT/CC, NVD), not a news article or an AI summary. |
Fluency is a risk factor
AI can generate a plausible-sounding wrong answer faster and more fluently than a human. A human might say "I don't remember the exact CVE number." An AI will say "CVE-2026-12345" because it's a fluent completion. The confidence is real in the sense that the model is producing that token with high probability. But high token probability doesn't automatically mean factuality.
And the self-reinforcing loop makes it worse. Once the AI has said "CVE-2026-12345" once, it will say it again. And again. Each repetition makes the next one more confident. The model is checking its own previous output, which is circular. And it looks like verification.
Treat fluency as a risk factor in your review process:
- Prefer a hedged answer you can verify over a polished answer you can't.
- When an answer is especially smooth, raise the verification bar, don't lower it.
- Don't let model confidence substitute for source authority.
- Don't let multi-turn repetition count as corroboration. It's still one claim, restated.
The fix isn't a better model. It's a better architecture pattern for mitigating hallucinations and misalignment. Three layers, each with different failure modes, each catching what the others miss. The primary AI generates. The verification AI checks. The human verifies. And the five rules apply across all three.
The practice
This is a practice problem, not a technology problem. The tools will get better. The verification discipline is on you. Build the architecture now, before the confident wrong answer costs you an incident.
The Roadster total deposit was $50,000, not the $5,000 partial deposit the model treated as the whole ask. Starship Flight 14 hadn't launched yet, though SpaceX did launch it successfully on 9/28/2026. Both gaps were in my AI's context before I caught them. The model was confident. The repetition was smooth. Confidence and repetition hid a factor-of-10 deposit gap in the first instance and a week-off launch date in the second instance.
In a security context, those errors aren't factors of 10. They're the difference between patched and exposed. Between the right playbook and the wrong one. Between catching the real threat and chasing a ghost.
Tensors will keep getting larger. Models will keep getting more fluent. What we need more of isn't bigger models. We need architects who design robust architecture patterns for hallucinations and misalignment so fluency can't outrun truth.
Build the three layers. Apply the five rules. Treat every confident-sounding answer as a claim, not a fact. And treat AI fluency as a risk factor, not a signal of accuracy.
Tesla, Roadster, and SpaceX are trademarks of their owners. KLM Innovation isn't affiliated with them.
