token
A token is the smallest unit of data processed and generated by natural language processing (NLP) systems and language models (LLMs), typically produced by a tokenizer that segments text into words, subwords, characters, or bytes.
Tokens are mapped to integer IDs from a fixed vocabulary so that models can process sequences efficiently. Tokens are distinct from words: for English text, a token averages roughly 3.5 to 4 characters, so a long or unusual word can span several tokens. Practical limits, costs, and context windows for LLMs are measured in tokens. Multimodal models tokenize images, audio, and video, so those inputs are counted and billed in tokens too.
Related Resources
Tutorial
Natural Language Processing With Python's NLTK Package
In this beginner-friendly tutorial, you'll take your first steps with Natural Language Processing (NLP) and Python's Natural Language Toolkit (NLTK). You'll learn how to process unstructured data in order to be able to analyze it and draw conclusions from it.
For additional information on related topics, take a look at the following resources:
Have a question about this? Mentor AI can show you examples, compare related terms, and point you to tutorials.
By Leodanis Pozo Ramos • Updated Sept. 23, 2026