LLMs Don’t Read Words. Here’s What They Actually See

Founder & CEO, Contrive Solutions 15+ years in software engineering, SaaS, AI automation and enterprise application development

A plain guide to how LLMs turn your text into numbers, why context windows fill up, and what pushes token usage up.

A token is the basic unit of text that a large language model reads and generates. It is usually a whole word or part of a word. In English, one token is roughly four characters, or around three-quarters of a word. An LLM does not see your text as letters and words. It sees a sequence of token numbers.

Almost everything involved in working with an LLM is measured in tokens. Your API costs, context window usage and even response speed are all tied to how many tokens a request uses.

Most teams start paying attention to tokens when the API bill arrives and is much higher than expected. This guide explains what tokens are, how text is converted into tokens, how tokens use up a model’s context window, and which patterns in prompts, chat applications, RAG systems and AI agents drive token usage up.

The token counts in this post use OpenAI’s o200k_base tokenizer, which is used by GPT-4o and newer OpenAI models. Claude, Gemini and Llama use different tokenizers, so the exact count varies between models. The general patterns apply across LLMs.

What is a token?

A token is a piece of text that exists in a model’s vocabulary. Common words are usually a single token, while longer or less common words are split into several pieces.

Spaces and punctuation are part of tokenization too. In English, many tokens include the space that comes before the word.

Here is a simple example:

How an LLM reads a sentence: the text 'Tokenization is how language models read text.' is split into 9 tokens, each token becomes a token ID such as 4421 and 2860, and each ID becomes a vector of numbers

The sentence has 7 words, 46 characters and 9 tokens. The word “Tokenization” is split into two tokens, “Token” and “ization”. The other words are one token each, with the leading space included.

For English text, these rough conversions are useful:

  • 1 token is about 4 characters
  • 100 tokens is about 75 words
  • A 1,500-word article is about 2,000 tokens

These are only estimates. The actual number depends on the text, tokenizer, language, formatting and model.

How an LLM sees text

A language model does not process your prompt as individual characters. Before the text reaches the model, it goes through three basic steps:

  1. Split. The tokenizer breaks the text into tokens from its vocabulary.
  2. Number. Each token is replaced with its ID in that vocabulary. In o200k_base, “Token” is 4421 and “ization” is 2860.
  3. Embed. Each token ID is looked up in a table and converted into a long list of numbers called a vector or embedding. The neural network works with these numbers, not the original characters.

Output works in the opposite direction. The model predicts the next token ID, converts that ID back into text, predicts the next token, and continues until the response is complete. This is why model speed is described in tokens per second.

Tokenization also explains some of the strange mistakes LLMs make. Ask a model how many r’s are in “strawberry” and it may give you the wrong answer. On its own, “strawberry” is three tokens: “st”, “raw” and “berry”. Inside a sentence, with a space before it, it is a single token.

Either way, the model is working with token IDs, not individual letters. It cannot count letters reliably unless it spells the word out first.

How tokens are created: byte pair encoding

Many modern LLMs build their vocabulary with a method called byte pair encoding, or BPE. The tokenizer is created before the model is trained, by processing a very large collection of text.

The basic process looks like this:

  1. Start with individual characters, or more precisely bytes, as the initial tokens.
  2. Find the pair of neighboring tokens that occurs together most often.
  3. Merge that pair into a new token and add it to the vocabulary.
  4. Keep merging pairs until the vocabulary reaches its target size.

Byte pair encoding example for the word lowest: six single-character tokens become five, four, three and finally two tokens, low and est, as the most frequent pairs are merged

After enough merges, common words and word fragments become individual tokens. Something frequent such as ” the” or ” model” is one token, while a rare word is assembled from several common pieces.

Token vocabularies have grown over time. GPT-2 had roughly 50,000 tokens, GPT-4’s cl100k_base has about 100,000, and o200k_base has about 200,000. A larger vocabulary represents the same text with fewer tokens, particularly in languages other than English.

Because the vocabulary is learned from training data, text that appeared often in that data is represented efficiently. Less common text needs more tokens. You can see the difference in these examples:

Text Tokens Why
How much does it cost to build an AI agent? 11 Common English words, one token each
The same question in Urdu 14 Urdu words are split into more pieces. The older cl100k_base tokenizer needed 35
The meeting is on 2026-10-01 at 3:45 PM. 18 Dates and times split into several smaller tokens
1234567890 4 Long numbers are split into chunks of up to 3 digits
if (user.isActive) { return true; } 11 Symbols and camelCase names need extra tokens

If your product serves users in Urdu, Arabic, Hindi or another non-Latin script, the tokenizer belongs in your model selection. The same conversation can use more tokens, and fill the context window faster, than its English equivalent.

Tokens and the context window

The context window is the maximum number of tokens a model can process in a single request. It includes both the information you send to the model and the output it generates. Current models offer context windows from around 128,000 tokens to more than 1 million.

Everything in a request shares that space:

One request in a 128K-token context window: system prompt 2,000 tokens, tool definitions 3,000, retrieved documents 24,000, chat history 18,000, user message 500 and 4,000 reserved for the reply, leaving 76,500 free

  • System prompt: instructions, rules and persona
  • Tool definitions: the name, description and schema of every tool an agent can use
  • Retrieved documents: content retrieved by a RAG search
  • Chat history: earlier messages included in the conversation
  • The new message: the user’s latest request
  • The reply: output tokens also use the context window

Three things follow from this.

The model has no memory between requests. A chat application sends the conversation history again with each new request. What looks like memory is the application sending the previous context back to the model.

When the context window fills up, something has to be removed or compressed. An application may drop older messages, summarize them, or return an error. Without a clear strategy for long conversations, the problem tends to appear exactly when the conversation has become most useful.

A larger context window is not automatically better. Models use information near the beginning and end of a long prompt more reliably than information buried in the middle, and longer prompts take more time to process. Sending the most relevant 5,000 tokens is often more useful than sending 100,000 just in case.

What causes high token consumption?

1. Resending chat history

Take a customer support chatbot with a 1,500-token system prompt, user messages averaging 100 tokens and replies averaging 300 tokens.

If you only count the messages in a 20-turn conversation, you get around 8,000 tokens. But because the complete history is sent with every request, the model receives about 108,000 input tokens over the conversation, plus 6,000 output tokens.

Input tokens per turn in a 20-turn chat: with full history the request grows from 1,600 to 9,200 tokens, 108,000 in total; keeping the last 3 turns plus a 300-token summary caps it at 3,100, 58,400 in total

At GPT-4o’s $2.50 per million input tokens and $10 per million output tokens, the conversation costs about $0.33. Counting only the messages, you might estimate around $0.07. At 10,000 conversations a month, that is roughly $650 estimated against $3,300 billed.

When you calculate LLM costs, count what the model receives on every request, not just the new message.

2. AI agents that loop

An AI agent calls the model, runs a tool, adds the tool result to the conversation, and then calls the model again. The previous context is sent along with each step.

Say an agent has 15 tools, with each tool description averaging 200 tokens. Add a 1,000-token system prompt and a 200-token task. Each step then adds a 150-token tool call and a 1,500-token tool result.

The first step sends around 4,200 tokens. By the tenth step, the request has grown to about 19,050 tokens. A single ten-step task uses around 116,000 input tokens.

Multi-agent systems push this further, because agents pass information and tool results between one another. This is one reason agent-based applications use far more tokens than a simple request-and-response chatbot.

3. Large system prompts and tool lists

System prompts and tool definitions are sent with every request.

A 3,000-token system prompt used across 50,000 requests a month is 150 million tokens before users even start typing. At $2.50 per million input tokens, that is about $375 a month.

The same applies to tools. If an agent receives descriptions for 30 tools but only uses three of them for a task, the unused descriptions still consume tokens on every request.

4. Retrieval that pulls in too much

A RAG pipeline gets expensive fast if it retrieves more content than the model needs.

If your system retrieves the top 20 chunks of 1,000 tokens each, that puts 20,000 tokens into every request. In many cases, three well-ranked chunks give the model more useful context than twenty loosely related ones.

Better retrieval cuts both cost and noise. We cover the retrieval side in our Laravel vector search postmortem.

5. Long outputs

Output tokens cost more than input tokens. At the prices above, one output token costs four times as much as one input token. Output is also generated one token at a time, so longer responses take longer to produce.

A prompt that says “Explain your reasoning in detail” on every request increases both cost and response time.

Reasoning models add another layer. They generate internal thinking tokens before the final answer. You may not see those tokens in the response, but they are billed as output.

6. Data formats

The way you format data also affects token usage. Here is the same customer record in three formats:

Format Tokens
One plain line: Name: Sara Khan, email sara@example.com, plan Pro, seats 12, renews 2026-11-01 26
Compact JSON, no spaces 38
Pretty-printed JSON with indentation 63

Repeated key names, quotation marks, braces and indentation all add tokens.

For large datasets, compact JSON or plain CSV uses fewer tokens than heavily formatted JSON. The exact savings depend on the structure and tokenizer, but the difference adds up when you send thousands of records.

7. Retries and failed calls

Failed requests are an easy source of unexpected token costs.

If a request times out or returns invalid JSON and your application retries it automatically, the model processes the same input again. You pay twice for the same task.

Output schemas, sensible timeouts, validation and retry limits keep failed calls from quietly eating a large part of your AI budget.

How your prompt affects token usage

The way you write a prompt makes a noticeable difference. Here is a typical conversational prompt:

Hello! I hope you are doing well today. I was wondering if you could possibly help me out with something. I have a customer support email here and I would really like you to please read through it carefully and then let me know what you think the main problem is that the customer is having. Also, if it's not too much trouble, could you maybe also tell me whether the customer seems happy or upset? And it would be great if you could suggest which team should handle it. Please explain your reasoning in detail. Thank you so much in advance for your help, I really appreciate it!

That prompt uses 120 tokens. The same task, stated directly:

Read the support email below. Return JSON with: issue (one sentence), sentiment (positive/neutral/negative), team (billing/technical/sales).

That version uses 31 tokens.

Saving 89 input tokens on one request will not change your bill much by itself. The bigger difference comes from the response. The first prompt asks for a detailed explanation, so the model returns several paragraphs. The second asks for a fixed structure and gets something like this:

{"issue":"Customer was charged twice for the March invoice.","sentiment":"negative","team":"billing"}

That response is only 21 tokens. It is also easier for your application to process, because the data arrives in a predictable structure instead of buried inside a paragraph.

A few prompt habits reduce token usage without losing useful context:

  • Remove padding and repeated instructions. State each requirement once.
  • Specify the format and length. For example, ask for “3 bullets”, “JSON with these fields” or “under 100 words”.
  • Put fixed instructions before changing content. This makes prompt caching more effective.
  • Send only the relevant part of a document. There is little value in sending an entire file when one section matters.
  • Use examples carefully. One good example often does the job of several.

Don’t over-optimize, though. If removing context makes the model misunderstand the task, the errors and retries cost more than the tokens you saved.

How to cut token costs in production

  1. Use prompt caching. Major model providers offer lower pricing for repeated prompt prefixes, such as system instructions, tool definitions or a frequently reused document. Keep the reusable part identical and at the beginning of the request, following your provider’s caching rules.
  2. Trim or summarize chat history. Keep the most recent turns along with a short running summary. In the chatbot example above, keeping the last 3 turns and a 300-token summary cuts input from 108,000 to 58,400 tokens per conversation, a 46% reduction.
  3. Route requests by difficulty. Send simple classification and routing tasks to a smaller model and complex requests to a larger one. For example, GPT-4o Mini was priced at $0.15 per million input tokens in the pricing used for this guide.
  4. Load tools when they are needed. Instead of giving an agent every available tool on every request, provide the tools relevant to the current task.
  5. Set an output limit. Use a maximum output length and ask for a concise format when the task does not need a long answer.
  6. Measure token usage before shipping. Use the provider’s tokenizer or token-counting endpoint and log token usage for production requests. Unexpected AI bills usually come from a feature whose token usage nobody was measuring.

If you are planning an AI feature and want to estimate its token budget before development, our AI development cost guide covers the development side. You can also tell us what you are building for an estimate.

Founded
2014
Years in business
12
People
66
Headquarters
Lahore, Pakistan
Projects delivered
250+
AI systems in production
10
Posted in AI