Back to archive

Types of tokens in LLMs

Tokens are the basic units of text processing in language models. They can be divided by role, visibility, and function.

1. Classification by processing flow

TypeDescriptionCounted toward costs?Example
InputTokens from the request or promptYes“Process this text:”
OutputTokens generated by the modelYes“Here is the answer.”
ContextThe sum of input, conversation history, and output in the context windowYesThe entire conversation
CachedInput tokens that the system recognizes as previously processed and can reuseYes, but usually at a lower priceThe same system prompt or an unchanged part of a long context

2. Special and system tokens

  • System — define the model's role, for example: “You are an expert.”
  • Special — technical markers such as <|endoftext|>, <|im_start|>, and <|im_end|>, which describe the data structure.
  • Thinking / Reasoning — hidden tokens used by the model during reasoning.
  • Planning / Critical — tokens supporting planning or key stages of the response logic.

3. Classification by tokenization

MethodGranularityAdvantagesDisadvantages
WordsWhole wordsSimple approachWeaker with new and rare words
Subwords (BPE)Parts of wordsFlexible, often used in LLMsMore complex
CharactersIndividual lettersUniversalGenerate far more tokens

The economy rule

Input + Cached + Output ≤ the model's context window
for example, 128K tokens.

What do cached tokens mean in practice?

Cached tokens are parts of the input that have not changed between successive model calls.

Most often, these are:

  • the same system prompt
  • the same conversation history
  • the same instructions or documents attached to multiple requests

This means the system does not always need to recompute everything from scratch, and such a portion may be billed at a lower price. Not all models support this.

A practical rule

The more fixed, unchanging context there is between requests, the greater the chance of using cached tokens and reducing the cost.

Token Economics

Token