How many PDF pages fit in ChatGPT, Claude, and Gemini?
A practical page estimate for model context windows, and why tokens still matter more than page count.
This question shows up every week: how many PDF pages can I paste into ChatGPT before it starts ignoring part of the file? The honest answer is that pages are a rough proxy. Tokens are what models actually read.
Still, a page estimate is useful for planning. It helps you decide if your file will fit in one shot or if you should split it before you send it to a model.
A practical baseline
If your PDF is dense, contracts, technical reports, policy docs, a safe planning baseline is around 500 tokens per page. Lighter documents can be closer to 250 to 350 tokens. Scans with little text can be much lower.
- 128K context: around 250 to 500 pages
- 200K context: around 400 to 800 pages
- 1M context: around 2 000 to 4 000 pages
Most current flagship models take around a million tokens (GPT-6 on OpenAI's models page, Claude Opus 5.5, and Gemini 3.8 Flash), so through the API a single PDF rarely hits the ceiling anymore. The Claude bar is shorter for a reason: Anthropic says its current tokenizer fits about 555K words in 1M tokens, against about 750K on older models. Same document, roughly 35% more tokens.
Chat apps are a different story. They can cap file size or context on their own terms, so the API number isn't a promise about what happens when you drop a PDF into a chat window. The full model-by-model table is in PDF token limits by model.
Why page count drifts so much
Two PDFs with the same page count can differ by 3x in tokens. Tables, legal definitions, repeated headers, and code blocks all inflate token usage. Non-English text also tends to tokenize heavier than plain English prose.
OCR output can push counts up fast when a scan has noise, broken words, and duplicated text lines.
Best workflow before sending to AI
First, run the file through the PDF Token Counter. Then decide if you need to split by chapter, by section, or by page ranges.
If you are over the model limit, do not hope the interface will warn you. Many tools truncate silently and still produce a confident answer.
For long recurring docs, clean and convert to Markdown first. That makes token use more predictable and easier to chunk for RAG or prompt chains.