[home]   [coding projects]   [research projects]


Linear Regression with Mini-Batch Gradient Descent

Mini-batch gradient descent, learning-rate selection, batching, and weight decay.

[full source]

This project is a from-scratch NumPy study of gradient descent for a known linear regression problem. The point is less the regression model itself and more the optimization behavior: how initialization, learning rate, batch size, and \(L_2\) weight decay change convergence.

1. The model and loss

The synthetic data use five Gaussian features and the true weight vector \(w^\star=[4,-3,2.5,-1,0.5]\). Targets are generated from the linear model plus Gaussian noise. The optimized objective is

\[ L(w)=\frac{1}{2B}\|Xw-y\|_2^2+\lambda\|w\|_2^2, \qquad \nabla L(w)=\frac{X^\top(Xw-y)}{B}+2\lambda w. \]
Christopher_Housholder_HW1.py
def compute_gradient(X, y, w, lam=0.0):
    B = len(X)
    r = X @ w - y
    return (X.T @ r) / B + 2 * lam * w

def compute_loss(X, y, w, lam=0.0):
    B = len(X)
    r = X @ w - y
    return (r @ r) / (2 * B) + lam * (w @ w)

2. One routine for SGD, mini-batch, and full-batch descent

A nice implementation detail is that the three common gradient-descent regimes are not separate algorithms here. They are the same routine with different batch sizes: \(1\) for SGD, an intermediate integer for mini-batch descent, and \(N\) for full-batch descent.

Christopher_Housholder_HW1.py
batch_size = len(X_train) if batch_size == 'N' else batch_size

for i in range(0, n, batch_size):
    X_batch = X_epoch[i:i+batch_size]
    y_batch = y_epoch[i:i+batch_size]

    grad = compute_gradient(
        X_batch, y_batch, w, weight_decay
    )
    w = w - lr * grad

3. Experiments

The writeup compares three initializations, four learning rates, batch sizes from stochastic to full-batch, and weight decay. The broad conclusion is that, on this deliberately well-behaved linear problem, the hyperparameters mostly alter the route and speed of convergence rather than the final solution.

Weight initialization comparison.
Weight initialization comparison.
Learning-rate comparison.
Learning-rate comparison.
Batch-size comparison.
Batch-size comparison.
Training/test MSE with and without weight decay.
Training/test MSE with and without weight decay.
Squared parameter norm under weight decay.
Squared parameter norm under weight decay.

These pages are selective technical manuals: enough source to expose the mechanism, not a mirror of the entire repository.
Last updated: September 14, 2026.