Skip to content

cross entropy loss

Cross entropy loss, also called log loss, is the standard loss function for classification. It measures the distance between the probability distribution a model predicts over the available classes and the distribution the labeled data implies, which puts all of its probability on the one correct class.

The value falls toward zero as the model puts more probability on the correct class, and it grows without bound as a confident prediction turns out to be wrong. That asymmetry makes it a useful teaching signal, because a model is punished far more for being certain and wrong than for being merely undecided.

That penalty curve is clearest when setting the probability assigned to the correct class directly and watching where the loss lands:

Interactive diagram — enable JavaScript to view.

Computing it takes two steps. A softmax turns the model’s raw output scores, called logits, into probabilities that sum to one, and the loss is the negative logarithm of the probability assigned to the correct class. Frameworks such as PyTorch fold both steps into one operation that consumes logits directly, which is more numerically stable than applying the softmax and the logarithm separately.

It dominates large language model work because autoregressive generation is itself a classification problem. At each position the model picks the next token out of a fixed vocabulary that runs to tens or hundreds of thousands of entries, so training minimizes the average cross entropy loss of the true next token across a corpus.

The same quantity resurfaces in evaluation, since perplexity is the exponential of the average cross entropy loss per token. Common variants adapt the signal to the data, including binary cross entropy for two-class problems, per-class weighting for imbalanced labels, and label smoothing, which softens the target distribution to discourage overconfidence.

Python AI: How to Build a Neural Network & Make Predictions

Tutorial

Python AI: How to Build a Neural Network & Make Predictions

In this step-by-step tutorial, you'll build a neural network from scratch as an introduction to the world of artificial intelligence (AI) in Python. You'll learn how to train your neural network and make accurate predictions based on a given dataset.

intermediate ai data-science machine-learning

For additional information on related topics, take a look at the following resources:

Have a question about this? Mentor AI can show you examples, compare related terms, and point you to tutorials.


By Martin Breuss • Updated Sept. 19, 2026