The same PDF in four languages: how many more tokens it costs
I ran the same UN document in English, Portuguese, German, and Spanish through two OpenAI tokenizers. The newer one cut the extra cost for Portuguese and German by more than half.
If you send PDFs to an AI model in a language other than English, you probably pay more for the same content. The question is how much. I wanted a real number instead of the vague “non-English text tokenizes heavier” line you see everywhere, including, until now, on this site.
How I tested it
A fair test needs the same text in each language, so only the language changes. The UN publishes the Universal Declaration of Human Rights as a PDF in hundreds of languages. I took the English, Portuguese, German, and Spanishversions and ran each one through the same pipeline the PDFShore Token Counter uses: pdf.js to pull out the text, then two of OpenAI's public encodings from tiktoken. o200k_base is the newer one, used by GPT-4o and GPT-5.x. cl100k_base is the older one, used by GPT-4. I ran it on October 5, 2026.
The results
| Language | Words | o200k tokens | vs English | cl100k tokens | vs English |
|---|---|---|---|---|---|
| English | 1,780 | 2,190 | baseline | 2,189 | baseline |
| Portuguese | 1,842 | 2,583 | +18% | 3,237 | +48% |
| German | 1,679 | 2,668 | +22% | 3,425 | +56% |
| Spanish | 2,344 | 3,116 | +42% | 3,510 | +60% |
The newer tokenizer cut the language tax
On the older cl100k_base encoding, the Portuguese version cost almost half again as many tokens as the English one, and German more than half. On o200k_base the gap shrank to 18% and 22%. English barely moved between the two (2,190 against 2,189), so the whole improvement went to the other languages.
If you sized a Portuguese or German workload back when GPT-4 was the default, your old estimate is probably too high for current OpenAI models.
Two different reasons a language costs more
Spanish looks like the worst case at +42%, but the table hides something. The Spanish PDF has 2,344 words against 1,780 in English: the translation itself is about a third longer. Per word, Spanish is actually the cheapest of the three, at 1.33 tokens per word against 1.23 for English. German goes the other way. It uses the fewest words, because it packs ideas into long compounds, and then pays for them at 1.59 tokens per word, the most of the four.
So two separate things are going on. Some languages need more words to say the same thing. Some need more tokens per word. Your bill depends on both, which is why a tokens-per-word rule from English doesn't carry over.
What this test doesn't tell you
- It's one document of about 2,000 words. Contracts, manuals, or chat logs may behave differently.
- Word counts come from splitting on spaces, and each PDF carries a little header and footer text, so treat them as approximate.
- It only covers OpenAI's public encodings. Claude and Gemini use their own tokenizers, which they don't publish for local use; their token-counting APIs give exact numbers. OpenAI also hasn't documented which encoding GPT-6 uses, so o200k is an estimate there.
Check your own file
The quickest way to know is to count your actual PDF. The PDF Token Counter shows both encodings side by side, and the gap between them is a decent hint of how much the language is costing you: near zero for English, much wider for the others. To turn the count into dollars for a batch, use the Token Cost Calculator. Both run in your browser, so the file stays on your device.
For page-level numbers on English documents, see how many tokens are in one PDF page.