confusion matrix
A confusion matrix is a table that compares a classifier’s predictions against the true labels, with one row per actual class and one column per predicted class. Every prediction falls into exactly one cell, so the table shows how often a model is wrong and which classes it mistakes for which.
For a binary classifier, the table has four cells, and their names recur throughout classification evaluation:
- True positive (TP): A positive case the model labeled positive.
- False positive (FP): A negative case the model labeled positive, also called a false alarm.
- False negative (FN): A positive case the model labeled negative, also called a miss.
- True negative (TN): A negative case the model labeled negative.
Accuracy, precision, recall, specificity, and the F1 score are all ratios of those four counts, which makes the matrix strictly more informative than any single score drawn from it. The gap shows up on imbalanced data, where a classifier that never predicts the rare class still posts high accuracy while its true-positive cell sits at zero.
The decision threshold is the probability cutoff above which the model calls a case positive. Sweeping that threshold over a deliberately imbalanced sample moves each scored example between the four cells. At the top of the threshold range, the true-positive cell empties out while accuracy stays high:
The layout generalizes to any number of classes as an N by N table, where correct predictions land on the diagonal and every off-diagonal cell names one specific mistake. Row and column conventions vary across the literature, so a matrix is only readable once its axes are labeled. The confusion_matrix() function in scikit-learn puts true labels on the rows and predicted labels on the columns.
Confusion matrices carry over to LLM work wherever a model emits a label rather than free text: intent routing, content moderation and other guardrails, hallucination detectors, and LLM-as-judge graders scored against a human-annotated set.
Related Resources
Tutorial
Logistic Regression in Python
In this step-by-step tutorial, you'll get started with logistic regression in Python. Classification is one of the most important areas of machine learning, and logistic regression is one of its basic methods. You'll learn how to create, evaluate, and apply a model to make predictions.
For additional information on related topics, take a look at the following resources:
- Practical Text Classification With Python and Keras (Tutorial)
- Split Your Dataset With scikit-learn's train_test_split() (Tutorial)
- Learn Text Classification With Python and Keras (Course)
- Splitting Datasets With scikit-learn and train_test_split() (Course)
- Build Your Own Face Recognition Tool With Python (Tutorial)
- Split Your Dataset With scikit-learn's train_test_split() (Quiz)
- Build Your Own Face Recognition Tool With Python (Quiz)
Have a question about this? Mentor AI can show you examples, compare related terms, and point you to tutorials.
By Martin Breuss • Updated Sept. 21, 2026