06 · Neural nets · 5 min read · updated
What is a perceptron and how is it different from logistic regression?
In short
The simplest artificial neuron: it takes a weighted sum of inputs plus a bias and outputs 1 when it exceeds zero, splitting feature space with a hyperplane.
What it is
The perceptron (Rosenblatt 1958) is a single artificial neuron: it multiplies each input by a weight, sums the results, adds an offset (the bias) and outputs 1 when the total is greater than zero, and 0 otherwise. It is a linear classifier: its decision boundary is a straight line in two dimensions and a hyperplane in general.
Every neural network and every large language model is built from such neurons, supplemented with a nonlinear activation function. The perceptron itself is a close relative of logistic regression, an ancestor of the SVM and the template for the neurons in modern networks.
Intuition: a doctor assesses a tumour from a few measurements. Each measurement "votes" with some strength — a large radius for malignant, smooth edges for benign. The perceptron adds up these votes and compares the sum with a threshold. Learning means choosing how strong each vote should be.
Mechanism — why it works this way
The formula: y = 1 if w·x + b > 0, otherwise 0. The set of points where w·x + b = 0 is a hyperplane; the weight vector w is perpendicular to it, and the bias b shifts it away from the origin. Each weight therefore says how strongly, and in which direction, a feature votes for the answer 1.
The learning rule is simple: after every mistake, w ← w + η (y − ŷ) x and b ← b + η (y − ŷ). An example wrongly classified as 0 pulls the hyperplane in the direction that lets it through; one wrongly classified as 1 pushes it away. When training starts from zero weights, the step size η does not affect the course of learning (it only rescales the weights) — unlike in networks trained by gradient descent.
Novikoff's theorem (1962): if the classes are linearly separable with margin γ and all examples lie inside a ball of radius R, the perceptron makes at most (R/γ)² corrections and then stops. The other half of the story is often left out: when the classes cannot be separated by a hyperplane, the algorithm never stops — the weights cycle forever. Minsky and Papert (1969) showed that XOR is one such problem, which is what motivates hidden layers.
The perceptron outputs only "yes/no", with no degree of confidence. Logistic regression is the same neuron with a sigmoid in place of the threshold: it returns a probability between 0 and 1 and can be trained by following the gradient of a smooth loss. That is why, in modern networks, "neuron" almost always means a weighted sum with a smooth activation rather than a hard threshold. The biological analogy is loose: real neurons do not compute an exact weighted sum.
By example
On the Iris dataset (150 flowers, 4 measurements) I ran scikit-learn's Perceptron(random_state=0). The task "setosa or not" is linearly separable: the perceptron reaches 100% accuracy after 7 epochs and stops. The task "versicolor or virginica" is not separable: with the stopping criterion turned off, the algorithm used up its whole limit of 1,000 epochs, and in a separate run on standardized data the number of mistakes (out of 100 flowers) over 50 consecutive epochs jumped between 2 and 9 and never once reached zero — exactly as the theory predicts.
On Breast Cancer Wisconsin (569 tumours, 30 features, standardized, 75/25 stratified split, random_state=0) the perceptron is right on 94.4% of test cases, while logistic regression on the same data reaches 95.8%. The difference is small, but logistic regression gives you something extra: a probability, e.g. 0.996 for one tumour and 0.00003 for another, instead of a bare 0/1.
In practice
- scikit-learn:
Perceptron(defaultseta0=1.0,max_iter=1000), in practice replaced byLogisticRegressionorSGDClassifier(loss='log_loss'). - PyTorch:
nn.Linear(n_in, 1)is exactly a weighted sum plus bias; the threshold is replaced bysigmoidandBCEWithLogitsLoss. - A single linear neuron cannot solve XOR or any rule that needs two cuts — you need a hidden layer.
- Common mistake: training a perceptron on non-separable data and waiting for it to "settle down". It won't — watch the number of mistakes per epoch.
- Weights are comparable only after standardizing the inputs; otherwise a large weight may simply reflect a feature with a small scale.
Frequently asked questions
- How is a perceptron different from logistic regression?
- Both compute the same weighted sum. The perceptron compares it with zero, returns a hard 0/1 and learns only from mistakes. Logistic regression passes it through a sigmoid, returns a probability and is trained on the gradient of cross-entropy, so it also improves its confidence, not just which side of the boundary a point falls on.
- Why can't a perceptron solve the XOR problem?
- XOR requires the answer 1 for (0, 1) and (1, 0), and 0 for (0, 0) and (1, 1). No straight line separates these two pairs of points — you need two cuts, i.e. a hidden layer with at least two neurons (Minsky and Papert 1969).
- Will a perceptron always learn if the data are separable?
- Yes — Novikoff's theorem guarantees a finite number of corrections, no more than (R/γ)². It does not guarantee that the boundary it finds will generalize well, though: the perceptron stops at the first boundary that fits, not necessarily the one with the largest margin (that is what an SVM does).
Sources
- Rosenblatt, F. (1958). "The perceptron: a probabilistic model for information storage and organization in the brain". Psychological Review 65(6), 386–408. doi:10.1037/h0042519
- Novikoff, A. (1962). "On convergence proofs on perceptrons". Proc. Symposium on the Mathematical Theory of Automata 12, 615–622.
- Minsky, M., Papert, S. (1969). Perceptrons. MIT Press.
- Bishop, C. (2006). Pattern Recognition and Machine Learning, ch. 4.1.7 "The perceptron algorithm".
- Hastie, Tibshirani, Friedman (2009). The Elements of Statistical Learning, 2nd ed., ch. 4.5.1 "Rosenblatt's perceptron learning algorithm".