Skip to content

transformer

A transformer is a neural network model that processes sequences using self-attention to capture long-range dependencies without recurrence or convolutions.

The transformer architecture stacks encoder and decoder blocks, each wrapping multi-head attention and position-wise feed-forward layers in residual connections and layer normalization. Decoder blocks add a sub-layer that attends over the encoder’s output, and their own self-attention is masked so each position can’t see the tokens that come after it.

Because self-attention is order-agnostic, token embeddings also carry positional encodings before they enter the stack.

The design comes in three main variants:

  • Encoder-only: Builds representations for classification and search, as in BERT.
  • Decoder-only: Handles autoregressive generation, as in GPT and most modern large language models.
  • Encoder-decoder: Maps one sequence to another for tasks such as translation.

Transformers form the foundation of modern language and vision models because of efficient parallel training, strong scaling behavior, and adaptable decoding for tasks such as classification, translation, and text or image generation.

Hugging Face Transformers: Leverage Open-Source AI in Python

Tutorial

Hugging Face Transformers: Leverage Open-Source AI in Python

As the AI boom continues, the Hugging Face platform stands out as the leading open-source model hub. In this tutorial, you'll get hands-on experience with Hugging Face and the Transformers library in Python.

intermediate ai

For additional information on related topics, take a look at the following resources:

Have a question about this? Mentor AI can show you examples, compare related terms, and point you to tutorials.


By Leodanis Pozo Ramos • Updated Sept. 25, 2026