Types of tokens in LLMs
Tokens are the basic units of text processing in language models. They can be divided by role, visibility, and function.
1. Classification by processing flow
| Type | Description | Counted toward costs? | Example |
|---|---|---|---|
| Input | Tokens from the request or prompt | Yes | “Process this text:” |
| Output | Tokens generated by the model | Yes | “Here is the answer.” |
| Context | The sum of input, conversation history, and output in the context window | Yes | The entire conversation |
| Cached | Input tokens that the system recognizes as previously processed and can reuse | Yes, but usually at a lower price | The same system prompt or an unchanged part of a long context |
2. Special and system tokens
- System — define the model's role, for example: “You are an expert.”
- Special — technical markers such as
<|endoftext|>,<|im_start|>, and<|im_end|>, which describe the data structure. - Thinking / Reasoning — hidden tokens used by the model during reasoning.
- Planning / Critical — tokens supporting planning or key stages of the response logic.
3. Classification by tokenization
| Method | Granularity | Advantages | Disadvantages |
|---|---|---|---|
| Words | Whole words | Simple approach | Weaker with new and rare words |
| Subwords (BPE) | Parts of words | Flexible, often used in LLMs | More complex |
| Characters | Individual letters | Universal | Generate far more tokens |
The economy rule
Input + Cached + Output ≤ the model's context window
for example, 128K tokens.
What do cached tokens mean in practice?
Cached tokens are parts of the input that have not changed between successive model calls.
Most often, these are:
- the same system prompt
- the same conversation history
- the same instructions or documents attached to multiple requests
This means the system does not always need to recompute everything from scratch, and such a portion may be billed at a lower price. Not all models support this.
A practical rule
The more fixed, unchanging context there is between requests, the greater the chance of using cached tokens and reducing the cost.