ML Atlas

06 · Neural nets · 5 min read · updated

What is a perceptron and how is it different from logistic regression?

In short

The simplest artificial neuron: it takes a weighted sum of inputs plus a bias and outputs 1 when it exceeds zero, splitting feature space with a hyperplane.

What it is

The perceptron (Rosenblatt 1958) is a single artificial neuron: it multiplies each input by a weight, sums the results, adds an offset (the bias) and outputs 1 when the total is greater than zero, and 0 otherwise. It is a linear classifier: its decision boundary is a straight line in two dimensions and a hyperplane in general.

Every neural network and every large language model is built from such neurons, supplemented with a nonlinear activation function. The perceptron itself is a close relative of logistic regression, an ancestor of the SVM and the template for the neurons in modern networks.

Intuition: a doctor assesses a tumour from a few measurements. Each measurement "votes" with some strength — a large radius for malignant, smooth edges for benign. The perceptron adds up these votes and compares the sum with a threshold. Learning means choosing how strong each vote should be.

Mechanism — why it works this way

The formula: y = 1 if w·x + b > 0, otherwise 0. The set of points where w·x + b = 0 is a hyperplane; the weight vector w is perpendicular to it, and the bias b shifts it away from the origin. Each weight therefore says how strongly, and in which direction, a feature votes for the answer 1.

The learning rule is simple: after every mistake, w ← w + η (y − ŷ) x and b ← b + η (y − ŷ). An example wrongly classified as 0 pulls the hyperplane in the direction that lets it through; one wrongly classified as 1 pushes it away. When training starts from zero weights, the step size η does not affect the course of learning (it only rescales the weights) — unlike in networks trained by gradient descent.

Novikoff's theorem (1962): if the classes are linearly separable with margin γ and all examples lie inside a ball of radius R, the perceptron makes at most (R/γ)² corrections and then stops. The other half of the story is often left out: when the classes cannot be separated by a hyperplane, the algorithm never stops — the weights cycle forever. Minsky and Papert (1969) showed that XOR is one such problem, which is what motivates hidden layers.

The perceptron outputs only "yes/no", with no degree of confidence. Logistic regression is the same neuron with a sigmoid in place of the threshold: it returns a probability between 0 and 1 and can be trained by following the gradient of a smooth loss. That is why, in modern networks, "neuron" almost always means a weighted sum with a smooth activation rather than a hard threshold. The biological analogy is loose: real neurons do not compute an exact weighted sum.

By example

On the Iris dataset (150 flowers, 4 measurements) I ran scikit-learn's Perceptron(random_state=0). The task "setosa or not" is linearly separable: the perceptron reaches 100% accuracy after 7 epochs and stops. The task "versicolor or virginica" is not separable: with the stopping criterion turned off, the algorithm used up its whole limit of 1,000 epochs, and in a separate run on standardized data the number of mistakes (out of 100 flowers) over 50 consecutive epochs jumped between 2 and 9 and never once reached zero — exactly as the theory predicts.

On Breast Cancer Wisconsin (569 tumours, 30 features, standardized, 75/25 stratified split, random_state=0) the perceptron is right on 94.4% of test cases, while logistic regression on the same data reaches 95.8%. The difference is small, but logistic regression gives you something extra: a probability, e.g. 0.996 for one tumour and 0.00003 for another, instead of a bare 0/1.

In practice

  • scikit-learn: Perceptron (defaults eta0=1.0, max_iter=1000), in practice replaced by LogisticRegression or SGDClassifier(loss='log_loss').
  • PyTorch: nn.Linear(n_in, 1) is exactly a weighted sum plus bias; the threshold is replaced by sigmoid and BCEWithLogitsLoss.
  • A single linear neuron cannot solve XOR or any rule that needs two cuts — you need a hidden layer.
  • Common mistake: training a perceptron on non-separable data and waiting for it to "settle down". It won't — watch the number of mistakes per epoch.
  • Weights are comparable only after standardizing the inputs; otherwise a large weight may simply reflect a feature with a small scale.

Frequently asked questions

How is a perceptron different from logistic regression?
Both compute the same weighted sum. The perceptron compares it with zero, returns a hard 0/1 and learns only from mistakes. Logistic regression passes it through a sigmoid, returns a probability and is trained on the gradient of cross-entropy, so it also improves its confidence, not just which side of the boundary a point falls on.
Why can't a perceptron solve the XOR problem?
XOR requires the answer 1 for (0, 1) and (1, 0), and 0 for (0, 0) and (1, 1). No straight line separates these two pairs of points — you need two cuts, i.e. a hidden layer with at least two neurons (Minsky and Papert 1969).
Will a perceptron always learn if the data are separable?
Yes — Novikoff's theorem guarantees a finite number of corrections, no more than (R/γ)². It does not guarantee that the boundary it finds will generalize well, though: the perceptron stops at the first boundary that fits, not necessarily the one with the largest margin (that is what an SVM does).

Sources

  • Rosenblatt, F. (1958). "The perceptron: a probabilistic model for information storage and organization in the brain". Psychological Review 65(6), 386–408. doi:10.1037/h0042519
  • Novikoff, A. (1962). "On convergence proofs on perceptrons". Proc. Symposium on the Mathematical Theory of Automata 12, 615–622.
  • Minsky, M., Papert, S. (1969). Perceptrons. MIT Press.
  • Bishop, C. (2006). Pattern Recognition and Machine Learning, ch. 4.1.7 "The perceptron algorithm".
  • Hastie, Tibshirani, Friedman (2009). The Elements of Statistical Learning, 2nd ed., ch. 4.5.1 "Rosenblatt's perceptron learning algorithm".

See also