Mori shows you exactly what's in her context window and roughly what the next turn costs — then lets you compress it in one click. Signed-in users also get an always-on plan meter that never lets managed usage surprise you.
Every chat carries a live context gauge in the composer bar. It reads the fraction of the model's context window in use right now, as a percentage and a token count — 48% full · 96k, for example — and its color warms from accent to amber to red as the window fills.
This isn't a decoration. It's measured from exactly what will be sent on your next turn, refreshed as the conversation grows. Click the gauge — or type /usage — to open the full breakdown.
The popover splits the window into the seven things that actually consume it, each as a colored segment on one bar and a labeled line with its token count:
The denominator is the model's own context window — 200k for Claude, more for some others — so the same conversation reads as a different percentage on a bigger-window model. The breakdown is assembled from the real sources on every open; if any one source is briefly unavailable, its bucket just reads zero rather than blocking the popover.
The popover footer shows an estimated dollar cost for your next turn — the current context as input, plus a typical reply as output, priced by model family.
Be clear about what this is: an estimate, not a bill. Token counts are a fast characters-over-four approximation, and the model's own tokenizer will differ slightly; the per-token prices are sensible per-family figures, not a live rate card. It's there to give you an honest sense of scale and to catch a runaway context before it gets expensive — not to reconcile an invoice.
The popover says as much in its own footnote. Treat the number as a ballpark, and lean on your provider's dashboard for exact spend.
When the window gets heavy, hit Compress now in the /usage popover — or type /compress. Mori keeps the opening framing and your most recent turns verbatim, folds everything in the middle into one dense summary, and keeps that summary in play so every later turn inherits it. You get the room back without losing the thread.
A toast reports what happened — how many older turns were folded and roughly how many tokens were freed. Compression needs a bit of history to work on; on a short conversation there's simply nothing in the middle to compress yet, and Mori tells you so instead of pretending.
You rarely have to think about any of this. Once a conversation crosses roughly 70% of the window, Mori compacts the middle on her own — the same fold-the-middle, keep-the-edges move as the manual button, run automatically and only once per new stretch of turns.
The effect is that long sessions stay fast and coherent instead of slamming into a wall. The manual /compress and the /usage button are there for when you want to reclaim space early or deliberately; the automatic pass handles the rest quietly in the background.
Signed-in users get a second meter beside the context gauge: a small plan chip showing your tier and a bar. Every signed-in user sees it — including bring-your-own-key users, because every tier still has managed features with their own caps. Signed out, or if the status can't be fetched, the chip simply hides.
Managed usage flows in three windows at once — a rolling Session, a Today total, and a longer Period. The bar deliberately shows the most-constrained of the three: the window that will pause you first. It warms to amber and then red as that window fills. Hover the chip and it spells out all three with their reset times; click it to open the full set of meters. It refreshes on its own every few minutes, so it stays live even while you're not looking.
Hitting a cap never locks the app. Managed usage pauses gracefully until that window resets on its own — and if you have your own API keys, they keep working instantly, right through the pause.