scikit-learn Workflow Cheat Sheet

This page contains a condensed overview of the scikit-learn machine learning workflow. It covers data shapes, splitting train and test sets, preprocessing features, building pipelines, fitting and predicting, evaluating models, cross-validation, and hyperparameter tuning. You can also download the information as a printable cheat sheet:

Free Bonus: scikit-learn Workflow Cheat Sheet

Get a scikit-learn Workflow Cheat Sheet (PDF) and keep the full machine learning workflow at hand, from splitting data and building pipelines to cross-validation and hyperparameter tuning:

scikit-learn Workflow Cheat Sheet

Practice with hands-on coding exercises, quizzes, and guided learning paths. Not sure where to begin? Start here.

New to machine learning with scikit-learn?

Data Shapes

  • Features X are 2D: (n_samples, n_features)
  • The target y is 1D: (n_samples,)
  • From pandas: X = df[["a", "b"]], y = df["c"]
  • pip install scikit-learn, then import sklearn
  • as_frame=True makes loaders return pandas objects
Language: Python Filename: Features Must Be 2D
>>> import numpy as np
>>> x_lr = np.array([5, 15, 25, 35, 45, 55])
>>> x_lr = x_lr.reshape(-1, 1)  # 1 column
>>> y_lr = np.array([5, 20, 14, 32, 22, 38])
>>> x_lr.shape, y_lr.shape
((6, 1), (6,))
Language: Python Filename: Load a Toy Dataset
>>> from sklearn.datasets import load_iris
>>> X, y = load_iris(return_X_y=True)
>>> X.shape, y.shape
((150, 4), (150,))

Tripped up by array shapes before?

Split Train and Test Sets

  • Split first, before any preprocessing, to avoid data leakage
  • random_state makes the split reproducible
  • stratify=y keeps class proportions
  • test_size takes a fraction (0.25) or a row count (38)
  • Unpack in order: X_train, X_test, y_train, y_test
Language: Python Filename: train_test_split()
from sklearn.model_selection import (
    train_test_split,
)
X_train, X_test, y_train, y_test = \
    train_test_split(
        X, y, test_size=0.25,
        random_state=42, stratify=y,
    )
# X_train.shape → (112, 4)
# X_test.shape  → (38, 4)

Want to go deeper on splitting data?

Preprocess Features

  • .fit() on training data only, then .transform() both sets
  • Scale for kNN, SVMs, and logistic regression; trees don’t need it
  • OneHotEncoder(handle_unknown="ignore") survives unseen categories
Language: Python Filename: Scale Numeric Features
from sklearn.preprocessing import (
    OneHotEncoder, StandardScaler,
)
scaler = StandardScaler()
X_train_s = scaler.fit_transform(X_train)
X_test_s = scaler.transform(X_test)
# Train columns now: mean 0, std 1
Language: Python Filename: One-Hot Encode Categories
>>> colors = [["red"], ["green"], ["red"]]
>>> enc = OneHotEncoder(sparse_output=False)
>>> enc.fit_transform(colors)
array([[0., 1.],
       [1., 0.],
       [0., 1.]])
>>> enc.categories_
[array(['green', 'red'], dtype=object)]
Language: Python Filename: Mix Column Types
from sklearn import compose
pre = compose.make_column_transformer(
    (StandardScaler(), ["age"]),
    (OneHotEncoder(), ["city"]),
)
X_ready = pre.fit_transform(df)

Why fit on the training data only?

Build a Pipeline

  • A pipeline is an estimator: .fit(), .predict(), .score()
  • It refits preprocessing inside every CV fold, so no leakage
  • make_pipeline() names steps after the lowercased class
  • pipe[-1] is the last step; pipe.named_steps has them all
Language: Python Filename: Chain Steps
from sklearn.linear_model import (
    LinearRegression, LogisticRegression,
)
from sklearn.pipeline import make_pipeline

pipe = make_pipeline(
    StandardScaler(),
    LogisticRegression(),
)
pipe.fit(X_train, y_train)

Still fuzzy on why pipelines matter?

Fit and Predict

  • Learned attributes end in _: coef_, intercept_
  • .predict() needs 2D input, even for one sample
  • Classifiers add .predict_proba()
  • Swap in KNeighborsRegressor, RandomForestClassifier, …
Language: Python Filename: Linear Regression
>>> model = LinearRegression()
>>> model.fit(x_lr, y_lr)
LinearRegression()
>>> model.coef_  # Slope
array([0.54])
>>> model.intercept_.round(2)
np.float64(5.63)
>>> model.predict([[60]])
array([38.03333333])
>>> model.score(x_lr, y_lr)  # R²
0.7158756137479542
Language: Python Filename: Classification
>>> pipe.predict(X_test[:3])
array([0, 1, 1])
>>> pipe.predict_proba(X_test[:1]).round(2)
array([[0.98, 0.02, 0.  ]])
>>> pipe.score(X_test, y_test)  # Accuracy
0.9210526315789473

Think you’ve got fit and predict down?

Evaluate Models

  • Report scores on the test set, not the training set
  • High train score, low test score means overfitting
  • With imbalanced classes, look past accuracy
Language: Python Filename: Regression Metrics
from sklearn.metrics import (
    mean_absolute_error, r2_score,
    root_mean_squared_error,
)
y_fit = model.predict(x_lr)
r2_score(y_lr, y_fit)                 # 0.716
mean_absolute_error(y_lr, y_fit)      # 5.467
root_mean_squared_error(y_lr, y_fit)  # 5.81
Language: Python Filename: Classification Metrics
>>> from sklearn.metrics import (
...     accuracy_score, confusion_matrix,
... )
>>> y_pred = pipe.predict(X_test)
>>> accuracy_score(y_test, y_pred)
0.9210526315789473
>>> confusion_matrix(y_test, y_pred)
array([[12,  0,  0],
       [ 0, 12,  1],
       [ 0,  2, 11]])

Is accuracy telling the whole story?

Cross-Validate

  • Get a mean and a spread instead of one lucky split
  • Pass the whole pipeline, not pre-scaled data
  • Default is 5 folds, stratified for classifiers
  • Pick the metric with scoring="f1_macro", "r2", …
Language: Python Filename: cross_val_score()
>>> from sklearn.model_selection import (
...     cross_val_score,
... )
>>> scores = cross_val_score(pipe, X, y)
>>> scores.round(2)
array([0.97, 1.  , 0.93, 0.9 , 1.  ])
>>> print(f"{scores.mean():.3f}")
0.960

Pipeline or pre-scaled data?

Tune Hyperparameters

  • Address pipeline params as step__param (two underscores)
  • Search on the training set; touch the test set once, at the end
  • grid.best_estimator_ is refit on all training data
Language: Python Filename: GridSearchCV
from sklearn.model_selection import (
    GridSearchCV,
)
from sklearn.neighbors import (
    KNeighborsClassifier,
)
knn = make_pipeline(
    StandardScaler(), KNeighborsClassifier()
)
grid = GridSearchCV(knn, {
    "kneighborsclassifier__n_neighbors":
        [3, 5, 7, 9, 11],
}, cv=5)
grid.fit(X_train, y_train)
Language: Python Filename: Inspect the Winner
>>> grid.best_params_
{'kneighborsclassifier__n_neighbors': 7}
>>> round(grid.best_score_, 3)  # Mean CV
0.955
>>> grid.score(X_test, y_test)  # Test
0.9473684210526315

Ready to test yourself on tuning?

Ready to go beyond the cheat sheet?

You can download this information as a printable cheat sheet:

Free Bonus: scikit-learn Workflow Cheat Sheet

Get a scikit-learn Workflow Cheat Sheet (PDF) and keep the full machine learning workflow at hand, from splitting data and building pipelines to cross-validation and hyperparameter tuning:

scikit-learn Workflow Cheat Sheet