[home] [coding projects] [research projects]
Mini-batch gradient descent, learning-rate selection, batching, and weight decay.
This project is a from-scratch NumPy study of gradient descent for a known linear regression problem. The point is less the regression model itself and more the optimization behavior: how initialization, learning rate, batch size, and \(L_2\) weight decay change convergence.
The synthetic data use five Gaussian features and the true weight vector \(w^\star=[4,-3,2.5,-1,0.5]\). Targets are generated from the linear model plus Gaussian noise. The optimized objective is
def compute_gradient(X, y, w, lam=0.0):
B = len(X)
r = X @ w - y
return (X.T @ r) / B + 2 * lam * w
def compute_loss(X, y, w, lam=0.0):
B = len(X)
r = X @ w - y
return (r @ r) / (2 * B) + lam * (w @ w)
A nice implementation detail is that the three common gradient-descent regimes are not separate algorithms here. They are the same routine with different batch sizes: \(1\) for SGD, an intermediate integer for mini-batch descent, and \(N\) for full-batch descent.
batch_size = len(X_train) if batch_size == 'N' else batch_size
for i in range(0, n, batch_size):
X_batch = X_epoch[i:i+batch_size]
y_batch = y_epoch[i:i+batch_size]
grad = compute_gradient(
X_batch, y_batch, w, weight_decay
)
w = w - lr * grad
The writeup compares three initializations, four learning rates, batch sizes from stochastic to full-batch, and weight decay. The broad conclusion is that, on this deliberately well-behaved linear problem, the hyperparameters mostly alter the route and speed of convergence rather than the final solution.





These pages are selective technical manuals: enough source to expose the mechanism, not a mirror of the entire repository.
Last updated: September 14, 2026.