Skip to content

feature engineering

Feature engineering is the process of transforming raw data into the input variables, called features, that a machine learning model learns from during training. That transformation runs as a pipeline:

Raw data fields fan out into combine, encode, scale, and select steps that build a feature matrix for training and inference.
From Raw Fields to Features, Replayed at Inference

The work combines domain knowledge with iteration over a handful of recurring operations:

  • Creating features by combining, decomposing, or aggregating existing fields, such as splitting a timestamp into hour and weekday
  • Encoding categorical values and text as numbers that a model can consume
  • Scaling numeric ranges so that no single feature dominates the updates made by gradient descent
  • Selecting a subset of features, or reducing dimensionality with methods like principal component analysis (PCA), to cut noise and redundancy

Deep neural networks absorb much of this work, learning their own representations such as embeddings from raw text, images, or audio, so hand-built features matter most on tabular and time-series problems.

Recurring pitfalls include data leakage, where a feature encodes information that wouldn’t be available at prediction time, and applying different transformations during training and inference.

Pythonic Data Cleaning With Pandas and NumPy

Tutorial

Pythonic Data Cleaning With pandas and NumPy

A tutorial to get you started with basic data cleaning techniques in Python using pandas and NumPy.

intermediate data-science numpy

For additional information on related topics, take a look at the following resources:

Have a question about this? Mentor AI can show you examples, compare related terms, and point you to tutorials.


By Martin Breuss • Updated Sept. 17, 2026