Wispr Flow turns your voice into clean, ready-to-send writing — speak naturally, it strips the filler and fixes the punctuation. I've used it daily since February 2026 to build this newsletter. Read the full review →
If your custom AI assistant has started feeling vaguer since you set it up — not slower, just less sharp — you're not imagining it, and it's not a worse model. It's the same model, doing more work to find your actual question.
Here's the mechanism nobody explains: your assistant has no memory of its own between messages. OpenAI says as much in its own developer docs — "each text generation request is independent and stateless." What feels like an ongoing conversation is an app quietly rebuilding everything relevant and resending it with every call: your custom instructions, your message history, and — depending on the product and how big it's gotten — your Project files.
That "depending on" matters, because Claude and ChatGPT don't handle it identically. Anthropic's help documentation says Claude reads a small Project's files in full, but automatically switches to retrieval-based search once the knowledge base "approaches the context window limit" — no exact threshold published. ChatGPT's Custom GPT "Knowledge" files work differently from the start: OpenAI's help center describes them as chunked and turned into embeddings for retrieval regardless of size. Two real products, two different answers to "does it reload the whole thing" — worth knowing rather than assuming either way.
What doesn't get that same protection, in either product, is your custom instructions and your connected apps. Those are short enough that nobody builds retrieval around them — they're resent whole, every time. Anthropic's own engineering team says as much about connected tools specifically: "most... clients load all tool definitions upfront directly into context," and in extreme cases, agents wired to thousands of tools "process hundreds of thousands of tokens before reading a request." Nobody's running thousands of connectors, but every one you've switched on adds its name, description, and full parameter list before your question gets read — whether you used it today or not.
Unify Your Teams and Tech Stack With HubSpot
Connect your customer data, teams, and tools without the hassle of complex integrations or lengthy setups. One easy platform gives marketing, sales, and service teams a unified customer view and the tools to turn it into growth.
Why HubSpot and what's new
• Generate leads and automate marketing with Marketing Hub
• Build your pipeline and close more deals with Sales Hub
• Scale customer support and drive retention with Service Hub
• Keep customer data clean, connected, and actionable with one, unified platform
Join 306,000+ in over 135 countries using HubSpot to grow their businesses.
See what a more connected approach to growth can do for you and your team. Get setup quickly and start checking off your hardest tasks.
Stack a long instruction block, a Project's worth of files, and a handful of connectors, and you're spending real, finite room. Anthropic's current flagship Claude models publish a 1-million-token context window (its smaller Haiku model: 200,000); OpenAI's current flagship line publishes roughly 1.05 million. Big numbers — but the room they eat isn't "used later." It's gone before your actual question is read.
And it's not only about running out of space. Anthropic's own engineers describe something they call "context rot": as token count in the context window rises, "the model's ability to accurately recall information from that context decreases" — a pattern they say shows up "across all models," to varying degrees. Independent researchers at Chroma tested it directly across 18 models and found the same thing: consistent, if uneven, accuracy loss as input length grew, even with task difficulty held constant. It's a gradient, not a cliff — but it's a real cost against the quality of an answer, not just its speed.
None of this means custom instructions, Projects, or connected apps are mistakes. It means they're not free, and the bill shows up as a duller answer, not a slower one.
On the site, we built something you can actually run on your own account: a ten-minute audit checklist covering exactly what to check and what "too much" looks like, a copy-paste prompt that interviews you and hands back a verdict, and an interactive model you can toggle yourself to see how a light setup compares to a loaded one. Illustrative, not a live measurement — but it makes the shape of the problem visible in a way a paragraph can't.
Free email without sacrificing your privacy
Gmail tracks you. Proton doesn’t. Get private email that puts your data — and your privacy — first.



