LLM Context Window Calculator

See if your system, context and user messages fit the model window.

Works Offline
Privacy First
No Login
No API

Messages

Max output: 16,384

Context budget

Estimating while the tokenizer loads

System
7
Context
0
User
0
Chat overhead
7
Total input
14
Window
128,000
Window usage0.01%

Fits — 126,962 tokens left after reserving 1,024 for the reply.

Same prompt on other models

  • GPT-4o0.0% of 128,000
  • GPT-4o mini0.0% of 128,000
  • GPT-4.10.0% of 1,047,576
  • GPT-4.1 mini0.0% of 1,047,576
  • o30.0% of 200,000
  • GPT-4 Turbo0.0% of 128,000
  • GPT-3.5 Turbo0.1% of 16,385
  • Claude Sonnet0.0% of 200,000
  • Claude Haiku0.0% of 200,000
  • Claude Opus0.0% of 200,000
  • Gemini Flash0.0% of 1,048,576
  • Gemini Pro0.0% of 1,048,576
  • Llama 3.3 70B0.0% of 128,000
  • Mistral Large0.0% of 131,072

About the LLM Context Window Calculator

Break a request into system prompt, retrieved context and user message, add realistic chat overhead, reserve room for the reply and instantly see whether the request fits GPT, Claude, Gemini, Llama or Mistral context windows.

Examples

RAG request

system + 8 chunks + question

Output

12,480 / 128,000 tokens — fits

Keyboard shortcuts

  • Copy the main outputCtrl / ⌘ + Shift + C
  • Download the resultCtrl / ⌘ + S
  • Share this toolCtrl / ⌘ + Shift + S
  • Reset the inputsAlt + R
  • Open the tool search paletteCtrl / ⌘ + K

Related tools

Frequently asked questions

Version 1.0.0 · Updated 2026-08-14 · Runs entirely in your browser