unsupervised learning
Unsupervised learning is a machine learning paradigm in which a model finds structure in unlabeled data, discovering groupings, compressed representations, or unusual examples without being shown the correct answer during training.
In supervised machine learning, every training example arrives paired with the answer the model is meant to produce. Unsupervised learning drops that pairing, so the objective comes from the shape of the data rather than from human annotation. That makes unsupervised learning useful wherever labels are expensive or absent, which describes most raw text, log, and image collections. Three families of methods dominate:
- Clustering: Grouping similar examples with algorithms such as k-means, DBSCAN, and hierarchical clustering.
- Dimensionality reduction: Compressing many features into a smaller vector space using principal component analysis (PCA), t-SNE, or an autoencoder.
- Density estimation: Modeling how the data is distributed with kernel density estimation or a Gaussian mixture model, so that rare points surface as outliers or anomalies.
The stepper below runs k-means over a scatter of unlabeled points, one half-step at a time. Each pass reassigns every point to its nearest centroid, then recenters each centroid on the points that chose it, until no point changes cluster.
Evaluation is harder than in labeled settings, since there is no ground truth to score against. Results are judged with internal measures such as the silhouette score, or by how much the learned representation improves a downstream task.
Large language model pretraining is more precisely called self-supervised, because the next token is the label that the corpus supplies for itself, though the two terms are often used interchangeably. Purely unsupervised methods still run alongside it, clustering embeddings to deduplicate training data or map its topics.
Related Resources
Tutorial
K-Means Clustering in Python: A Practical Guide
In this step-by-step tutorial, you'll learn how to perform k-means clustering in Python. You'll review evaluation metrics for choosing an appropriate number of clusters and build an end-to-end k-means clustering pipeline in scikit-learn.
For additional information on related topics, take a look at the following resources:
- Build a Recommendation Engine With Collaborative Filtering (Tutorial)
- Embeddings and Vector Databases With ChromaDB (Tutorial)
- Python AI: How to Build a Neural Network & Make Predictions (Tutorial)
- Split Your Dataset With scikit-learn's train_test_split() (Tutorial)
- NumPy Tutorial: Your First Steps Into Data Science in Python (Tutorial)
- K-Means Clustering in Python: A Practical Guide (Quiz)
- Vector Databases and Embeddings With ChromaDB (Course)
- Embeddings and Vector Databases With ChromaDB (Quiz)
- Building a Neural Network & Making Predictions With Python AI (Course)
- Python AI: How to Build a Neural Network & Make Predictions (Quiz)
- Splitting Datasets With scikit-learn and train_test_split() (Course)
- Split Your Dataset With scikit-learn's train_test_split() (Quiz)
- NumPy Tutorial: Your First Steps Into Data Science in Python (Quiz)
Have a question about this? Mentor AI can show you examples, compare related terms, and point you to tutorials.
By Martin Breuss • Updated Sept. 20, 2026