๐Ÿ“ˆ Machine Learning ยท Lecture 7 of 47

Logistic Regression: Probabilistic Classification Done Right

Despite its name, logistic regression is a classifier โ€” and one of the most reliable. We derive it from log-odds, train it by maximum likelihood, interpret its coefficients and understand its decision boundary.

Logistic regression is often the first classifier students learn and the last one experts abandon. It is fast, interpretable, well-calibrated and hard to beat as a baseline. It is also a single neuron with a sigmoid activation โ€” so understanding it deeply prepares you for neural networks.

Why not linear regression for classification?#

For a binary label $y \in \{0, 1\}$, linear regression predicts values outside $[0, 1]$ that cannot be interpreted as probabilities, and it is badly distorted by points far from the boundary. We want a model that outputs a valid probability $P(y = 1 \mid \mathbf{x})$.

From log-odds to the sigmoid#

The odds of an event with probability $p$ are $p/(1 - p)$, ranging over $(0, \infty)$. The log-odds (logit) range over all real numbers โ€” so we can model them linearly:

$$ \ln\frac{p}{1 - p} = \mathbf{w}^\top\mathbf{x} = z $$

Solving for $p$ gives the sigmoid (logistic) function:

$$ p = \sigma(z) = \frac{1}{1 + e^{-z}} $$

The sigmoid maps any real score to $(0, 1)$, with $\sigma(0) = 0.5$, $\sigma(z) \to 1$ as $z \to \infty$, and the symmetry $\sigma(-z) = 1 - \sigma(z)$.

Training by maximum likelihood#

Each label is Bernoulli with probability $p_i = \sigma(\mathbf{w}^\top\mathbf{x}_i)$. The negative log-likelihood is the binary cross-entropy:

$$ J(\mathbf{w}) = -\frac{1}{n}\sum_{i=1}^{n}\Big[y_i\ln p_i + (1 - y_i)\ln(1 - p_i)\Big] $$

Its gradient has a beautiful form:

$$ \nabla_{\mathbf{w}}J = \frac{1}{n}\sum_{i=1}^{n}(p_i - y_i)\,\mathbf{x}_i = \frac{1}{n}\mathbf{X}^\top(\mathbf{p} - \mathbf{y}) $$

โ€” identical in form to linear regression's gradient, with predictions passed through a sigmoid. There is no closed-form solution, but the loss is convex, so gradient descent, Newton's method (known here as iteratively reweighted least squares) or L-BFGS find the global optimum.

The decision boundary#

We predict class 1 when $p \ge 0.5$, i.e. when $\mathbf{w}^\top\mathbf{x} \ge 0$. The boundary $\mathbf{w}^\top\mathbf{x} = 0$ is a hyperplane โ€” logistic regression is a linear classifier. To obtain curved boundaries, add polynomial or other basis features.

The threshold 0.5 is not sacred. If false negatives are costly (missing a disease), lower the threshold; if false positives are costly, raise it. Choose it using the costs of errors and the validation set.

Interpreting coefficients#

Each coefficient is a change in log-odds per unit of the feature. Exponentiating gives an odds ratio:

$$ e^{w_j} = \text{factor by which the odds multiply when } x_j \text{ increases by 1} $$

If $w_{\text{smoker}} = 0.9$, smokers have $e^{0.9} \approx 2.46$ times the odds of the outcome, holding other features fixed. This interpretability is why logistic regression dominates in medicine, epidemiology and credit scoring.

Implementation from scratch#

python
import numpy as np

def sigmoid(z):
    return np.where(z >= 0, 1 / (1 + np.exp(-z)), np.exp(z) / (1 + np.exp(z)))  # stable

def train_logreg(X, y, lr=0.1, epochs=3000, lam=1e-3):
    Xb = np.column_stack([np.ones(len(X)), X])
    w = np.zeros(Xb.shape[1])
    for _ in range(epochs):
        p = sigmoid(Xb @ w)
        grad = Xb.T @ (p - y) / len(y)
        grad[1:] += lam * w[1:]                      # L2, intercept not penalised
        w -= lr * grad
    return w

rng = np.random.default_rng(0)
n = 500
hours = rng.uniform(0, 10, n)
attendance = rng.uniform(0.3, 1.0, n)
logit = -6 + 0.8 * hours + 4 * attendance
passed = (rng.random(n) < sigmoid(logit)).astype(float)

X = np.column_stack([hours, attendance])
w = train_logreg((X - X.mean(0)) / X.std(0), passed)
print("weights (standardised features):", w.round(3))
print("odds ratio per 1 SD of study hours:", np.exp(w[1]).round(2))

Calibration#

Because it is trained by maximum likelihood, logistic regression tends to produce well-calibrated probabilities: among cases predicted at 0.7, roughly 70% are positive. Many more powerful models (boosted trees, deep networks) are less well calibrated. Check with a reliability diagram (sklearn.calibration.calibration_curve).

Strengths and limitations#

Strengths: fast, convex, interpretable, calibrated, works well with many sparse features (text), strong baseline.

Limitations: linear decision boundary unless features are engineered; sensitive to strongly correlated features (use regularisation); cannot capture complex interactions automatically.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ“ˆ Machine Learning

Softmax Regression and Multiclass Classification Strategies

How do we classify into more than two classes? We generalise logistic regression to softmax regression, derive its gradient, and compare it with one-vs-rest and one-vs-one strategies.

Beginnerโฑ 5 min#057
๐Ÿ“ˆ Machine Learning

k-Nearest Neighbours: Learning by Similarity

The simplest learning algorithm stores the data and asks the neighbours. We analyse k-NN's biasโ€“variance behaviour, distance choices, scaling, efficient search structures and its surprising theoretical guarantees.

Beginnerโฑ 5 min#064
๐Ÿ“ˆ Machine Learning

Support Vector Machines: Maximum-Margin Classification

Among all separating hyperplanes, SVMs pick the one with the widest margin. We derive the hard- and soft-margin formulations, the hinge loss, support vectors and the role of the C parameter.

Intermediateโฑ 5 min#072