Back to archive

Token

Token

You paste one sentence into a chat, and the counter shows a dozen or so units. Why not as many as there are words? A computer needs a defined way to divide text that also handles rare names, punctuation, and code fragments.

That unit is a token: a piece of text from the vocabulary used by a particular model. It can comprise an entire word, part of one, a character, or a space with the following letters. A program called a tokenizer divides text and converts the pieces into numbers; only those numbers enter the model.

For example, the Polish sentence “Lubię czytać książki.” (“I like reading books.”) can illustratively be split into Lubię, czy, tać, książ, ki, .. That gives six tokens although the sentence has three words. The real split depends on the tokenizer — this example is not a particular model's output. A model creating a response predicts successive pieces, which the program assembles back into text.

Tokens help explain conversation length limits and how many services calculate fees. Word count alone cannot estimate them. A token is a unit of representation, rather than a unit of meaning: part of a word need not have a separate meaning. Byte Pair Encoding describes how this segmentation arises. The Hugging Face documentation explains tokenization and its variants.

Connections

Types of Tokens in LLMs