random forest
A random forest is an ensemble machine learning model that combines many decision trees into a single predictor, with each tree trained on a randomized view of the same data. Leo Breiman introduced the method in 2001, and it remains a common baseline for classification and regression on tabular data.
A single deep decision tree fits its training set closely and generalizes poorly, a failure mode known as overfitting. A forest counteracts that by growing trees that make different mistakes, drawing on two sources of randomness:
- Bootstrap sampling, where each tree learns from a sample drawn with replacement from the training set
- Feature subsampling, where each split considers only a random subset of the available features
Because the trees are decorrelated, averaging their outputs cancels much of the error any one of them makes. Scikit-learn averages the predicted class probabilities across trees, while Breiman’s original formulation had each tree vote for a single class.
The simulation below varies the number of trees on a two-feature dataset, and switching off the randomness shows how little a forest gains when every tree comes out a copy of the same one.
The rows left out of a tree’s bootstrap sample form its out-of-bag set, which yields an evaluation estimate of how well the forest generalizes to unseen data, without a separate validation split. A forest also ranks features by how much each one decreases impurity across the trees, though that measure skews toward features with many distinct values.
The trade-off is cost: a forest needs more memory and more inference time than a single tree. Because every prediction is an average over piecewise-constant trees, a forest also can’t extrapolate beyond the range of target values seen during training.
Related Resources
Tutorial
Linear Regression in Python
Use Python to build a linear model for regression, fit data with scikit-learn, read R2, and make predictions in minutes.
For additional information on related topics, take a look at the following resources:
- Split Your Dataset With scikit-learn's train_test_split() (Tutorial)
- Splitting Datasets With scikit-learn and train_test_split() (Course)
- Data Version Control With Python and DVC (Tutorial)
- Setting Up Python for Machine Learning on Windows (Tutorial)
- Starting With Linear Regression in Python (Course)
- Linear Regression in Python (Quiz)
- Split Your Dataset With scikit-learn's train_test_split() (Quiz)
- Setting Up Python for Machine Learning on Windows (Quiz)
Have a question about this? Mentor AI can show you examples, compare related terms, and point you to tutorials.
By Martin Breuss • Updated Sept. 22, 2026