ML Atlas

08 · LLMs · 5 min read · Interactive · updated

What are word embeddings and why do similar words have similar vectors?

In short

A word embedding is a vector of numbers in which words used in similar contexts lie close together. It is the foundation on which an LLM builds meaning.

What it is

A word embedding (word vector) is a dense vector of a few hundred numbers assigned to a word or token, learned so that words used in similar ways have similar vectors. Similarity is usually measured by the cosine of the angle between vectors. In an LLM the very first layer is exactly such an embedding table: the token's ID selects a row from it.

Before embeddings, words were encoded as one-hot vectors: a vocabulary of 50,000 words meant a vector of 50,000 zeros with a single one. Such an encoding carries no information about meaning — "cat" is just as far from "dog" as from "parliament". An embedding squeezes a word into a few hundred dimensions and thereby forces the model to generalize: what it knows about "cat" partly carries over to "dog".

Individual dimensions usually have no names. Meaning lives in directions and distances across the whole space, not in separate coordinates.

Mechanism — why it works this way

The distributional hypothesis. Words that occur in similar contexts have similar meanings — "you shall know a word by the company it keeps". If "cat" and "dog" both appear next to "feed", "vet" and "fur", a model that predicts the context from a word (or the word from its context) will reach a lower loss by giving them similar vectors. Vector similarity is thus a side effect of a predictive task, not something we teach directly.

word2vec. Mikolov et al. (2013) proposed two simple models: CBOW predicts a word from the average of its neighbours' vectors, and skip-gram predicts the neighbours from the word. Training pulls together the vectors of words that genuinely co-occur and pushes them away from randomly sampled "negative" words. GloVe (Pennington et al., 2014) arrives at similar vectors by a different route — by factorizing the co-occurrence matrix. At heart, both approaches compress the statistics of "who appears with whom".

The arithmetic of meaning. The famous result: vector(king) − vector(man) + vector(woman) lies close to vector(queen). This works because the difference "man → woman" is approximately a constant direction that recurs across many pairs. In fairness, though, the standard evaluation excludes the input words from the answers; without that exclusion the nearest word is often simply "king" (Nissim et al., 2020). Analogies are therefore a tendency, not a law.

Static versus contextual embeddings. word2vec gives one vector per word, so a river "bank" and a savings "bank" share a vector. In an LLM the embedding table is static too, but after the first few self-attention layers a token's vector depends on its surroundings — this is a contextual embedding. It is these contextual representations that are used in semantic search.

Limitations. Embeddings inherit biases from text: Bolukbasi et al. (2016) showed that in popular word vectors, occupations line up along a gender axis. Rare words have poorly learned vectors. Distributional similarity does not distinguish synonyms from antonyms — "hot" and "cold" occur in very similar contexts.

By example

Let us build toy 4-dimensional vectors (dimensions: power, masculinity, femininity, food): king = (0.9, 0.8, 0.1, 0), queen = (0.9, 0.1, 0.8, 0), man = (0.1, 0.9, 0.1, 0.1), woman = (0.1, 0.1, 0.9, 0.1), apple = (0, 0.1, 0.1, 0.9). The cosine for "king–man" is 0.74, for "king–queen" 0.66, and for "king–apple" only 0.08. Words with shared features are close; unrelated ones are almost perpendicular.

Now the analogy: king − man + woman = (0.9, 0, 0.9, 0). The cosine of this vector with "queen" is 0.995, with "woman" 0.77, with "king" 0.59. Subtracting "masculinity" and adding "femininity" took us exactly where we needed to go. In a real model the dimensions have no names and there are several hundred of them, but the geometry works similarly. GPT-2 small shows the scale: a table of 50,257 tokens × 768 dimensions is 38.6 million parameters, close to a third of the whole model.

In practice

  • Classic word vectors are trained in gensim (Word2Vec, FastText); fastText builds a word's vector from character n-grams, which works well for highly inflected languages such as Polish.
  • In PyTorch an embedding table is torch.nn.Embedding(num_embeddings, embedding_dim) — an ordinary weight matrix trained by gradient descent.
  • For searching sentences and documents, today's practice is sentence embedding models (e.g. via sentence-transformers) rather than averaged word vectors.
  • Normalize vectors to unit length before comparing them; the dot product then equals the cosine.
  • Common mistake: visualizing embeddings in 2D (t-SNE, UMAP) and drawing conclusions from the distances between far-apart clusters — these methods mainly preserve local neighbourhoods.

Frequently asked questions

How many dimensions should an embedding have?
Classic word vectors usually have 100–300 dimensions, sentence models 384–1,024, and the hidden states of large LLMs several thousand. More dimensions mean more capacity, but also more data needed for training and more memory in the index.
How is an embedding different from one-hot encoding?
One-hot is sparse, huge and contains no similarities. An embedding is dense, short and learned so that vector similarity reflects similarity in how words are used.
Do embeddings understand meaning?
They encode co-occurrence statistics, which partly overlap with meaning. That is why they capture topic and similarity well, but negation, numbers and logical relations less well.

Sources

  • Mikolov T. et al., 2013, "Efficient Estimation of Word Representations in Vector Space", arXiv:1301.3781.
  • Mikolov T. et al., 2013, "Distributed Representations of Words and Phrases and their Compositionality", NeurIPS 2013.
  • Pennington J., Socher R., Manning C., 2014, "GloVe: Global Vectors for Word Representation", EMNLP 2014.
  • Bolukbasi T. et al., 2016, "Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings", NeurIPS 2016.
  • Nissim M., van Noord R., van der Goot R., 2020, "Fair Is Better than Sensational: Man Is to Doctor as Woman Is to Doctor", Computational Linguistics 46(2).

See also