Skip to content

training

Training is the process of fitting a model’s parameters to data by minimizing or optimizing a carefully chosen objective (loss or surrogate) via gradient-based (or other) optimization, typically using forward and backward passes.

In practice, training workflows include:

  • Specifying the objective (loss) and evaluation metrics
  • Splitting the data into training, validation, and test sets
  • Iterating over mini-batches across multiple epochs or streaming passes
  • Updating weights with optimizers
  • Applying regularization techniques, such as L2 weight decay, L1 penalties, dropout, data augmentation, and early stopping
  • Optionally staging the run as pretraining followed by post-training, which covers supervised fine-tuning, preference optimization, RLHF-style alignment, reinforcement learning with verifiable rewards for reasoning, and distillation in a pipeline
  • Enhancing efficiency and stability through batching, normalization, mixed precision arithmetic, gradient accumulation, gradient clipping, learning rate scheduling, checkpointing, distributed/parallel training, and use of accelerated hardware (GPUs, TPUs)

Modern training often operates in overparameterized regimes, where the implicit regularization of optimization dynamics, alongside explicit regularizers such as weight decay and early stopping, can influence generalization.

Split Your Dataset With scikit-learn's train_test_split()

Tutorial

Split Your Dataset With scikit-learn's train_test_split()

In this tutorial, you'll learn why splitting your dataset in supervised machine learning is important and how to do it with train_test_split() from scikit-learn.

intermediate data-science machine-learning numpy

For additional information on related topics, take a look at the following resources:

Have a question about this? Mentor AI can show you examples, compare related terms, and point you to tutorials.


By Leodanis Pozo Ramos • Updated Sept. 17, 2026