Logistic regression is often the first classifier students learn and the last one experts abandon. It is fast, interpretable, well-calibrated and hard to beat as a baseline. It is also a single neuron with a sigmoid activation โ so understanding it deeply prepares you for neural networks.
Why not linear regression for classification?#
For a binary label $y \in \{0, 1\}$, linear regression predicts values outside $[0, 1]$ that cannot be interpreted as probabilities, and it is badly distorted by points far from the boundary. We want a model that outputs a valid probability $P(y = 1 \mid \mathbf{x})$.
From log-odds to the sigmoid#
The odds of an event with probability $p$ are $p/(1 - p)$, ranging over $(0, \infty)$. The log-odds (logit) range over all real numbers โ so we can model them linearly:
Solving for $p$ gives the sigmoid (logistic) function:
The sigmoid maps any real score to $(0, 1)$, with $\sigma(0) = 0.5$, $\sigma(z) \to 1$ as $z \to \infty$, and the symmetry $\sigma(-z) = 1 - \sigma(z)$.
Training by maximum likelihood#
Each label is Bernoulli with probability $p_i = \sigma(\mathbf{w}^\top\mathbf{x}_i)$. The negative log-likelihood is the binary cross-entropy:
Its gradient has a beautiful form:
โ identical in form to linear regression's gradient, with predictions passed through a sigmoid. There is no closed-form solution, but the loss is convex, so gradient descent, Newton's method (known here as iteratively reweighted least squares) or L-BFGS find the global optimum.
The decision boundary#
We predict class 1 when $p \ge 0.5$, i.e. when $\mathbf{w}^\top\mathbf{x} \ge 0$. The boundary $\mathbf{w}^\top\mathbf{x} = 0$ is a hyperplane โ logistic regression is a linear classifier. To obtain curved boundaries, add polynomial or other basis features.
The threshold 0.5 is not sacred. If false negatives are costly (missing a disease), lower the threshold; if false positives are costly, raise it. Choose it using the costs of errors and the validation set.
Interpreting coefficients#
Each coefficient is a change in log-odds per unit of the feature. Exponentiating gives an odds ratio:
If $w_{\text{smoker}} = 0.9$, smokers have $e^{0.9} \approx 2.46$ times the odds of the outcome, holding other features fixed. This interpretability is why logistic regression dominates in medicine, epidemiology and credit scoring.
Implementation from scratch#
import numpy as np
def sigmoid(z):
return np.where(z >= 0, 1 / (1 + np.exp(-z)), np.exp(z) / (1 + np.exp(z))) # stable
def train_logreg(X, y, lr=0.1, epochs=3000, lam=1e-3):
Xb = np.column_stack([np.ones(len(X)), X])
w = np.zeros(Xb.shape[1])
for _ in range(epochs):
p = sigmoid(Xb @ w)
grad = Xb.T @ (p - y) / len(y)
grad[1:] += lam * w[1:] # L2, intercept not penalised
w -= lr * grad
return w
rng = np.random.default_rng(0)
n = 500
hours = rng.uniform(0, 10, n)
attendance = rng.uniform(0.3, 1.0, n)
logit = -6 + 0.8 * hours + 4 * attendance
passed = (rng.random(n) < sigmoid(logit)).astype(float)
X = np.column_stack([hours, attendance])
w = train_logreg((X - X.mean(0)) / X.std(0), passed)
print("weights (standardised features):", w.round(3))
print("odds ratio per 1 SD of study hours:", np.exp(w[1]).round(2))Calibration#
Because it is trained by maximum likelihood, logistic regression tends to produce well-calibrated probabilities: among cases predicted at 0.7, roughly 70% are positive. Many more powerful models (boosted trees, deep networks) are less well calibrated. Check with a reliability diagram (sklearn.calibration.calibration_curve).
Strengths and limitations#
Strengths: fast, convex, interpretable, calibrated, works well with many sparse features (text), strong baseline.
Limitations: linear decision boundary unless features are engineered; sensitive to strongly correlated features (use regularisation); cannot capture complex interactions automatically.