I had just made coffee. We had been in a long conversation the night before and I picked it up the next morning. Claude had no idea time had passed.
This has happened more than once. And it got me curious about why.
Claude can reason about time beautifully: timezones, durations, historical events. It just can’t feel it passing. Like someone who can explain how a clock works but has no idea what time it is.
Here’s what’s actually going on
There’s no clock. No persistent state between turns. The date gets injected into the system prompt at conversation start, but in a long conversation, that date is now thousands of tokens away from where the model is currently generating.
Transformer attention isn’t uniform. There’s a well-documented effect where information buried in the middle of a long context gets significantly less attention weight than what’s at the beginning or end. So the date just becomes one more signal competing in a noisy landscape, and it loses. It gets worse as the conversation grows.
Training data is a flat bag of patterns, not a timeline. A sentence from 2019 and one from 2023 live in the same weight space with no clear temporal index. No episodic memory. No subjective “now” to anchor from.
So when you say “yesterday” or “last week” deep in a long conversation, Claude has to reconcile that against a date it can barely locate anymore. It’s not ignoring you. It’s re-deriving where it is in time, every single time, from context that keeps getting noisier.
A few things that actually help
Repeat the date near your message when it matters. Recency dominates attention, so putting the date close to your question works better than hoping it remembers the system prompt from 50 turns ago.
Be explicit with time references. This is the one most people miss. Instead of “since last time we talked about this,” say “on March 14 we discussed X, today is March 18.” Instead of “it’s been a few days,” say “on March 12 we decided Y, I want to revisit that now.” Instead of “what should I do tomorrow,” say “today is Tuesday March 17, what should I do Wednesday March 18.”
The pattern is always the same: name the dates, name the anchors. Relative words like “recently,” “last time,” “tomorrow” require Claude to do inference across a long noisy context, and that’s exactly where it fails.
Think of it like cc’ing a colleague into a thread halfway through. The more you spell out the timeline, the better it performs.
It’s extra work. But once you understand why, it stops feeling annoying and starts feeling like just knowing how to use the tool.
What quirks have you found in your AI tools that changed how you work with them?
