Wispr Flow turns your voice into clean, ready-to-send writing — speak naturally, it strips the filler and fixes the punctuation. I've used it daily since February 2026 to build this newsletter. Read the full review →

The headline you probably saw about this story said a startup is using math to make AI hallucination-proof.

We went looking for who actually said that. It was not the startup.

What happened

Pramaana Labs raised $27 million in June, led by Khosla Ventures, with Accel, Boldcap, Nexus Venture Partners, Premji Invest and Unbound joining. Founded in 2025, based in Palo Alto. The three co-founders are IIT Madras alumni who came from Google Maps, Glean, and Google DeepMind — the DeepMind one was a core contributor to Gemini. The tax work is advised by Danny Werfel, the former IRS Commissioner, with researchers from Yale Law School and Stanford.

This is not a thin wrapper with a good deck.

Introducing The First Agentic CRM

Get revenue agents, workflows, and automations across every stage of your motion. Access customer data in real time through Attio's web app, MCP, API, and SDK.

Then Ask Attio anything about your business and get instant answers.

It's the CRM that runs the work behind every win.

What it actually does

Normal AI answers by predicting what a good answer looks like. Pramaana adds a step that has nothing to do with prediction.

Domain experts encode the real rules of a field — the tax code, clinical protocols — into Lean, a formal proof language that accepts a statement only when every logical step checks out. Your question gets translated into that same language, and a proof engine tries to connect the rules to an answer.

If it succeeds, you get a machine-checkable proof. If it fails, the system tells you which rule breaks — instead of writing you a confident paragraph.

That last part is the genuinely unusual decision. Almost every AI product ever shipped is built to always return something. This one is built to say I cannot prove that and stop.

The phrase they never use

We hunted specifically for an absolute claim. There is exactly one, in the press release: "It has never produced a confidently wrong verified answer."

Every word there is load-bearing. Not never wrong. Not never hallucinates. It is a claim about answers that came out with a proof attached — a narrower set than everything the product says. And it is a company self-report, with no benchmark, no error bar, and no independent auditor behind it.

"We have never seen it happen" and "it cannot happen" are different statements. Only the second would earn the word hallucination-proof. The company does not use it. We are not going to hand it to them.

Why we are this careful

The same promise was already made in the same domain, and it was already measured.

In 2024, a Stanford team ran the first preregistered evaluation of commercial legal-research AI. Casetext had said its method would eliminate hallucinations. Thomson Reuters said avoid. LexisNexis advertised "hallucination-free" legal citations.

Measured result: those tools each hallucinated 17% to 33% of the time — Lexis+ AI and Ask Practical Law AI at 17%, Westlaw AI-Assisted Research at 33%. The paper's verdict on the marketing was blunt — the providers' claims are overstated.

The underlying problem is not shrinking. Researcher Damien Charlotin maintains a public database of court decisions involving AI-hallucinated filings. On 6 August 2026 it listed 1,847 cases.

The real weak point

It is not the proof engine. Lean is about as trustworthy as software gets.

The risk sits one step earlier, in translating messy human rules into formal logic. If that translation misses an exception or a jurisdiction-specific definition, the engine will rigorously prove the wrong thing — and hand you a certificate of correctness for a rule that was never the real rule.

To Pramaana's real credit, they are not hiding this. Their own blog published a post titled "How Lean Handles Ambiguity, Vagueness, and Gaps in Law." A company selling a miracle does not publish a post about the three things blocking the miracle. That is the most reassuring thing in this entire story.

One result shows the upside clearly: formalising Indian income tax law in Lean, they found a bug in the law itself — a marginal-relief flaw where earning more could reduce take-home pay — and proved a fix. The system did not get caught making an error. It caught an error in the rules.

What to do with this

You cannot buy it — there is no consumer product. So take the test instead of the tool.

Split your AI questions into two piles. Pile one has a written rule underneath it: tax treatment, eligibility, contract terms, deadlines, dosage limits. For those, always ask "Which specific rule produces that answer, and quote it." If the model cannot name the rule, you got a plausible sentence, not an answer.

Pile two runs on judgment: is this good writing, is this a sound strategy. No proof engine will ever help there, and confident AI is genuinely useful — just never mistake the confidence for verification.

The habit, starting today: when a wrong answer would cost you real money, make the AI show you the rule. Not a citation — citations get fabricated, which is what those 1,847 court cases are about. The rule, quoted, so you can go read it yourself.

And when a company tells you its AI cannot be wrong, notice that the most credible outfit in this space is the one carefully declining to say so.