Skip to content

one-hot encoding

One-hot encoding is a way of representing a categorical value as a vector of zeros with a single one marking which category is present. A color feature holding red, green, or blue becomes three columns, so a green row encodes as [0, 1, 0].

The scheme exists because most machine learning models expect numbers rather than labels, and numbering the colors 0, 1, and 2 would imply an ordering and a spacing that the categories don’t have. One-hot vectors keep every category equidistant. In Python, scikit-learn’s OneHotEncoder and get_dummies() in pandas both produce this representation.

The two encodings below place the same three colors at different distances from each other:

Interactive diagram — enable JavaScript to view.

The cost is width. A vocabulary of 50,000 tokens needs 50,000 dimensions, and no two of those vectors say anything about how similar their categories are. Embeddings replace them with dense learned representations for that reason, though one-hot targets are still the reference distribution for cross-entropy loss during training.

Linear Regression in Python

Tutorial

Linear Regression in Python

Use Python to build a linear model for regression, fit data with scikit-learn, read R2, and make predictions in minutes.

intermediate data-science machine-learning

For additional information on related topics, take a look at the following resources:

Have a question about this? Mentor AI can show you examples, compare related terms, and point you to tutorials.


By Martin Breuss • Updated Sept. 18, 2026