ML Atlas

03 · Supervised · 5 min read · Interactive · updated

How does linear regression work and how do you interpret its coefficients?

In short

Linear regression predicts a number as a weighted sum of features plus a constant. Least squares picks the weights, and each one describes a feature's effect.

What it is

Linear regression is a model that predicts a numeric value as a weighted sum of features plus a constant: ŷ = b₀ + b₁·x₁ + b₂·x₂ + … + bₚ·xₚ. The weights (coefficients) are chosen to make the sum of squared differences between predictions and true values as small as possible. It is the oldest supervised learning model and still one of the most widely used.

With a single feature, the model is a straight line through a scatter plot: b₀ is the intercept and b₁ the slope, i.e. how much y increases on average when x goes up by one unit. With many features the line becomes a plane (or a hyperplane), but the interpretation stays similar: bⱼ tells you how much the prediction changes when xⱼ increases by one, holding all other features fixed.

That last phrase is crucial and often overlooked. A coefficient in multiple regression is not the same as the feature's correlation with the outcome — it is the "net" contribution, after removing what the other features already explain.

Mechanism — why it works this way

Fitting means minimising the residual sum of squares (RSS). For linear regression this problem has a closed-form solution, the normal equations: b = (XᵀX)⁻¹Xᵀy. No iteration is needed — linear algebra is enough. The details, and the justification for squaring, are covered in the article on least squares.

The model is "linear in the parameters", not necessarily in the features. If you add x² or log(x) as new columns, you still have linear regression, just with transformed features. That is how the same machinery handles curves (polynomial regression) and interactions.

The classical assumptions under which the coefficients have good statistical properties are: a linear relationship, independent observations, constant residual variance and — for confidence intervals — normally distributed residuals. Normality is not needed for prediction alone; what is needed is that the relationship really is approximately linear in the features you use.

The biggest interpretive trap is collinearity. When two features are strongly correlated, the model may give one a large positive weight and the other a large negative one — their effects nearly cancel and the predictions are fine. The individual coefficients then become unstable and must not be read as "effects". Regularisation (ridge) or dropping redundant features helps.

The second trap: a coefficient depends on the feature's units. A weight of 68 for a feature ranging from 3 to 6 and a weight of 0.04 for age in years say nothing about which feature matters more. For comparisons, standardise the features (mean 0, standard deviation 1).

By example

The Diabetes dataset (Efron et al., 2004): 442 diabetes patients, 10 features measured at baseline (age, sex, BMI, blood pressure and six blood-test measurements, s1–s6), and the target is a numeric measure of disease progression one year later (ranging from 25 to 346, mean 152). BMI alone explains a lot: on the full dataset the slope of the line is 10.2 — each extra BMI point means a progression score 10.2 units worse on average — and R² = 0.34.

Regression on all 10 features, trained on 331 patients and tested on 111 (random 75/25 split, random_state=0), reaches a test R² of 0.36 (0.56 on training), and RMSE drops from 70.5 (predicting the mean) to 56.4. This particular split is rather unlucky: 5-fold cross-validation gives a mean R² of 0.49. After standardising the features, the collinearity trap shows up: s1 (total cholesterol) gets a weight of −37.7 and s2 (the LDL fraction) +22.7, even though both measure almost the same thing (correlation 0.90). Together these two weights make sense; separately, none at all.

In practice

  • sklearn.linear_model.LinearRegression().fit(X, y); coefficients are in coef_, the intercept in intercept_, R² via score(X, y).
  • For inference (standard errors, p-values, confidence intervals) use statsmodels.OLS — scikit-learn does not report them.
  • Standardise features before comparing weights (StandardScaler); encode categorical variables as dummies (OneHotEncoder(drop="first")).
  • Always look at residuals plotted against predictions: curvature means missing non-linearity, a "funnel" means non-constant variance.
  • Many correlated features, or more features than observations? Reach for Ridge or Lasso.

Frequently asked questions

What does R² mean and what is a "good" value?
R² is the share of the variance of y explained by the model: 0 means "no better than the mean", 1 means a perfect fit. On test data R² can even be negative. What counts as good depends on the field: physics expects 0.99, while in medicine and the social sciences 0.3–0.5 can be a very respectable result.
Does a regression coefficient imply causation?
No. It describes how the prediction changes when a feature changes and the other features are held fixed, in this data. If an important variable that affects both the feature and the outcome has been left out, the coefficient will be biased. Causal conclusions require an experiment or a carefully designed study.
Can linear regression be used to predict categories?
Technically you can code the classes as 0 and 1, but predictions will fall outside [0, 1] and are hard to interpret as probabilities. Classification is the job of logistic regression, which passes the same weighted sum through a sigmoid function.

Sources

  • James G., Witten D., Hastie T., Tibshirani R. "An Introduction to Statistical Learning", 2nd ed., 2021, ch. 3.
  • Hastie T., Tibshirani R., Friedman J. "The Elements of Statistical Learning", 2nd ed., 2009, ch. 3.
  • Efron B., Hastie T., Johnstone I., Tibshirani R. "Least Angle Regression", Annals of Statistics 32(2), 2004 — source of the Diabetes dataset.
  • scikit-learn documentation, "Linear Models": https://scikit-learn.org/stable/modules/linear_model.html

See also