07 · Architectures · 5 min read · Interactive · updated
What is a transformer in machine learning and how does the architecture work?
In short
The transformer is an architecture of self-attention and MLP blocks with residual connections. It processes whole sequences at once and powers modern LLMs.
What it is
The transformer is a neural network architecture for processing sequences, built from repeated blocks, each with two parts: multi-head self-attention (tokens exchange information) and an MLP applied to each token separately (each token "processes" what it has gathered). Both parts are wrapped in residual connections and layer normalization. The architecture was introduced by Vaswani and colleagues in the paper "Attention Is All You Need" (2017).
Intuition: each block is a round of a meeting. First, all the participants (tokens) listen to one another and note down what from whom matters to them — that is self-attention. Then each of them quietly thinks over their notes — that is the MLP. After a dozen or several dozen rounds, every token's representation holds rich knowledge of its role in the whole text.
The transformer started out as a translation model (the encoder reads the source sentence, the decoder generates the translation). Today three variants are in use: encoder-only (BERT — understanding and classifying text), decoder-only (GPT and most large language models — generation) and encoder–decoder (T5, translation models). The same architecture also works on images (ViT), audio and proteins.
Mechanism — why it works this way
The flow in a GPT-style model: the text is split into tokens, each token becomes a vector (an embedding), and position information is added. The vectors then pass through N identical blocks. In the pre-norm variant, standard today, a block is: x = x + Attention(LayerNorm(x)), then x = x + MLP(LayerNorm(x)). The MLP is usually two dense layers with an expansion to 4·d and a nonlinearity (GELU or SwiGLU). At the end, a final normalization and a linear layer with a softmax give the probability distribution over the next token.
Why does it work so well? First, self-attention gives every token direct access to every other token, so long-range dependencies do not have to squeeze through dozens of recurrent steps as in an RNN. Second, the whole sequence is computed in parallel — training is largely multiplication of big matrices, which GPUs do very efficiently. This made it possible to train models on unprecedented amounts of text, and quality improved predictably with scale (scaling laws).
Third, a division of labour: attention moves information between positions, while the MLP transforms it within a position. Interpretability research suggests that much of a model's factual knowledge "lives" in the MLP layers, which account for about two-thirds of a block's parameters. Fourth, residual connections form a shared "stream" of representations to which each block merely adds corrections — which is why dozens of blocks can be stacked.
A decoder-style model learns to predict the next token with a causal mask: the token at position t sees only positions 1…t. Thanks to the mask, a single pass over a sentence of n tokens yields n training examples at once.
Limitations: the cost of self-attention grows quadratically with context length; during generation text is produced token by token, so inference is sequential; the transformer has a weaker inductive bias than a CNN or RNN, so it needs a lot of data. On small tabular or image datasets it often loses to simpler models.
By example
Let us count the parameters of GPT-2 small (12 blocks, d = 768, 12 heads, a vocabulary of 50,257 tokens, context of 1,024). Token embeddings: 50,257·768 = 38,597,376. Positions: 1,024·768 = 786,432. One block: self-attention 4·768² + 4·768 = 2,362,368, the MLP with expansion to 3,072: 768·3,072 + 3,072 + 3,072·768 + 768 = 4,722,432, two normalizations 3,072 — in total 7,087,872. Twelve blocks make 85,054,464, the final normalization 1,536. Sum: 124,439,808, i.e. the well-known "124 million" (the output layer shares its weights with the embeddings, so it adds no parameters). We checked this by counting the model's parameters in the transformers library.
A curiosity from this calculation: the MLP accounts for 67% of a block's parameters, attention for 33%, and the embedding table alone for 31% of the whole model. In larger models the share of embeddings shrinks quickly. The original 2017 transformer had 65 million parameters in the base version and 213 million in the big version; the latter achieved a BLEU score of 28.4 on English→German translation on the WMT 2014 dataset, with training that took 3.5 days on 8 GPUs.
In practice
- Ready-made models: the Hugging Face
transformerslibrary (AutoModel.from_pretrained("gpt2"),AutoModelForSequenceClassification). - PyTorch also offers the building blocks
nn.TransformerEncoderLayer(d_model, nhead, norm_first=True, batch_first=True)andnn.TransformerEncoder. - Typical training hyperparameters: AdamW, a learning-rate warm-up followed by cosine decay; gradient clipping to a norm of 1.
- For your own data it is usually better to fine-tune an existing model (fine-tuning, LoRA) than to train from scratch.
- Common mistake: no causal mask in a generative model — the training loss falls suspiciously low, and the generated text is nonsense.
Frequently asked questions
- How is a transformer different from an RNN?
- An RNN reads text step by step and stores everything in a single hidden state. A transformer looks at the whole sequence at once through self-attention, so it connects distant words more easily and trains in parallel, which allows enormous scale.
- What does the "T" in GPT stand for?
- GPT stands for Generative Pre-trained Transformer: a generative model, pretrained on a large text corpus, based on the decoder-only variant of the transformer.
- Do transformers work only on text?
- No. Images are split into patches (ViT), audio into frames, proteins into amino acids — anything that can be represented as a sequence of tokens can be processed by a transformer.
Sources
- Vaswani et al. "Attention Is All You Need", NeurIPS 2017, arXiv:1706.03762.
- Radford et al. "Language Models are Unsupervised Multitask Learners", OpenAI technical report, 2019.
- Devlin, Chang, Lee, Toutanova "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding", NAACL 2019.
- Zhang et al. "Dive into Deep Learning", d2l.ai, ch. 11 ("Attention Mechanisms and Transformers").
- Jurafsky, Martin "Speech and Language Processing", 3rd ed. draft, chapter on transformers and large language models, https://web.stanford.edu/~jurafsky/slp3/