Wispr Flow

Wispr Flow turns your voice into clean, ready-to-send writing — speak naturally, it strips the filler and fixes the punctuation. I've used it daily since February 2026 to build this newsletter. Read the full review →

Same reference document. Same twelve questions. Same AI model. Asking them one way costs $0.53. Asking them the other way costs $0.18 — for the exact same answers.

The difference isn't a better prompt. It's whether you started a new chat for every question, or asked all twelve in the same sitting.

How AI pricing actually works

Every major AI API — Anthropic's Claude, OpenAI's GPT, Google's Gemini — bills three different rates for the text you send, not one:

  • Fresh input, the normal rate, for text the model hasn't seen yet.

  • Cache write, a premium rate, the first time you ask the model to hold a block of text in short-term memory.

  • Cache read, a steep discount, every later request that reuses that same held text instead of re-processing it.

Paste the same reference material into a new chat every time, and you only ever pay the expensive rates — you never stick around long enough to collect the read discount. Keep one session open and ask your follow-ups there, and you write the cache once and read it repeatedly.

1,000+ Claude Prompts Top Professionals Actually Use at Work

Claude can be your analyst, editor, and strategist.

But most professionals are using it to fix grammar.

These 1,000+ Claude prompts take it from grammar tool to your most powerful AI work assistant.

Sign up for Superhuman AI and get:

  • 1,000+ ready-to-use Claude prompts to get real work done in minutes — researched, tested, and used by professionals at Google, Microsoft, and NASA

  • Superhuman AI newsletter (4 min daily) so you keep learning new AI tools and skills to stay ahead in your career — the prompts are just the beginning

Verified as of August 2026, and the rules genuinely differ by vendor:

  • Anthropic (Claude Sonnet 5): cache writes cost 1.25x the base input rate on the default 5-minute cache, or 2x on a paid 1-hour cache; reads cost 0.1x base input (90% off). The 5-minute clock resets every time the cache is reused, so a fast-moving session never pays the write premium twice.

  • OpenAI (GPT-5.6): same 1.25x write premium and 0.1x read discount, but the cache lifetime on GPT-5.6+ is a fixed 30 minutes — not adjustable — and busy traffic can occasionally route your request to a different server than the one holding your cache, causing a miss even inside that window.

  • Google (Gemini 3.1 Pro Preview): no published write premium at all — Google's own docs describe writing the cache as billed like ordinary input. Reads still get the same 0.1x discount. But Google charges separately for holding the cache: a per-hour storage fee, billed the whole time it's alive, whether you use it again or not.

The worked example

A consultant spends an afternoon working from one 20,000-token reference document (about 15,000 words) and asks 12 follow-up questions about it, using Claude Sonnet 5 at its published August 2026 rates: $2/MTok input, $10/MTok output, $4/MTok to write a 1-hour cache (2x input), $0.20/MTok to read from it (0.1x input).

Starting a fresh chat for each of the 12 questions means re-pasting the whole document 12 times: 241,800 fresh input tokens plus 4,800 output tokens. Total: $0.5316.

Asking all 12 in one session means writing the document to cache once ($0.08), then reading it from cache for the other 11 questions (a combined $0.044), plus the small per-turn question text and the output. Total: $0.1756.

Same task, same model, same answers — restarting the chat costs just over 3x as much. Do that once a workday for a month and the gap is $11.16 versus $3.69: about $7.50 a month for one person's one recurring habit, and the multiplier only gets bigger with a longer document or a heavier user.

What to actually do

You don't need to touch an API setting. Claude.ai, ChatGPT, and Gemini's consumer apps already keep your conversation's earlier turns in context, so the discount kicks in automatically as long as you stay in one conversation. The behavior change is entirely on you: when you're working from the same source material across several questions, set it up once and keep going in that same chat — don't start over for each new question.

Watch the clock, though. How long you can step away before the discount resets ranges from 5 minutes to an hour by default, depending on the vendor. A lunch break in the middle of a long research session can quietly reset you to square one.

The full breakdown — including a side-by-side vendor table, a copy-paste prompt to check your own habits, and an interactive tool where you can plug in your own document size, question count, and session pacing to watch the cost move — is on the site.

Same sales team. More booked meetings

Aimfox Avatars allows you to rent dedicated LinkedIn profiles that start conversations and hand interested replies to your existing sellers. Expand outbound capacity without adding another SDR seat.