Token Forecaster puts a range on an LLM reply before you press Enter: the usual length, and a worst case that held for 90.6% of 4,146 unseen calls. While the reply streams, it shows whether it is running long. It only watches and never changes the request. Learns from your local history. Open source, MIT.
Comparing token budgets? Alongside Token Forecaster, try opencode for an in-terminal AI agent, Monica for help on any site, and TypingMind - Chat UI for LLMs to use your own API keys. Need speed? Claude Haiku 4.5 is fast and affordable for coding, while Pretty Prompt polishes prompts for more accurate forecasts.