SKlearn and Statsmodels
Choosing the Right Link Function
Generalised Linear Models (GLM) offer a flexible framework for regression analysis. When building a GLM for a practical problem, the first and most important decision is selecting an appropriate link function — and this choice is driven by the nature of the response variable. A binary response calls for the logit link, count data calls for the log link, a proportion response calls for the probit or complementary log-log link, and so on. Once the link function is fixed, the model is ready to be fitted with data.
Fitting the Model: Iterative Optimisation Algorithms
GLMs are not fitted in one step. Since Maximum Likelihood Estimation for GLMs generally has no closed-form solution, parameter estimates are arrived at through numerical iterative optimisation algorithms. Starting from an initial guess, the algorithm updates the parameter estimates step by step, each iteration bringing the estimates closer to the values that maximise the likelihood.
Commonly used iterative algorithms in GLM fitting include:
- Newton-Raphson — uses both the gradient and the Hessian (matrix of second derivatives of the log-likelihood). Converges fast and accurately. Naturally handles predictors on different scales due to the curvature information in the Hessian.
- Fisher Scoring — a variant of Newton-Raphson that replaces the observed Hessian with the expected (Fisher) information matrix. Equivalent to Newton-Raphson for canonical link functions such as logit and log.
- Iteratively Reweighted Least Squares (IRLS) — the classical GLM fitting algorithm. At each step it solves a weighted least squares problem. Mathematically equivalent to Fisher Scoring for GLMs.
- Gradient Descent — uses only the first derivative (gradient). Computationally cheap per iteration but slow to converge and sensitive to feature scale.
- Stochastic Gradient Descent (SGD) — a variant that updates parameters using one or a small batch of observations at a time. Useful for very large datasets but introduces noise in convergence.
- L-BFGS (Limited-memory Broyden–Fletcher–Goldfarb–Shanno) — a quasi-Newton method that approximates the Hessian using gradient history from recent iterations. Memory-efficient but sensitive to differences in predictor scale.
- Conjugate Gradient (CG) — an iterative method that uses successive conjugate directions to minimise the objective. More efficient than plain gradient descent but still first-order in nature.
The essential aspect across all these algorithms is the stopping rule — the criterion that decides when the iteration ends. Two common approaches are:
- A predetermined number of steps — the algorithm runs for a fixed number of iterations and stops
- A tolerance limit — the algorithm compares parameter estimates from two consecutive steps and stops when the difference falls below a pre-set threshold, indicating convergence
Python Libraries: An Important Practical Difference
When fitting GLMs in Python, two libraries are widely used — sklearn and statsmodels. Both solve the same underlying optimisation problem, but their behaviour in practice differs in a way that matters to the analyst.
sklearn — An Unnatural Restriction
sklearn‘s LogisticRegression uses L-BFGS as its default solver. This solver is sensitive to the relative scale of predictor variables. When predictors are on very different scales — which is the normal situation in most real datasets — L-BFGS struggles to converge within its iteration limit and throws a ConvergenceWarning:
ConvergenceWarning: lbfgs failed to converge (status=1):
STOP: TOTAL NO. of ITERATIONS REACHED LIMIT.
To avoid this, sklearn mandates scaling of predictors before fitting. The official sklearn documentation states this explicitly:
“‘sag’ and ‘saga’ fast convergence is only guaranteed on features with approximately the same scale. You can preprocess the data with a scaler from sklearn.preprocessing.” — sklearn LogisticRegression API Documentation
This is an unnatural restriction. Scaling is not a requirement of the GLM or its likelihood — it is a workaround imposed by the solver choice. Even for a modest dataset of around 4,000 observations with a handful of predictors, this warning appears if the data is not manually scaled beforehand.
statsmodels — No Such Restriction
statsmodels uses Newton-Raphson as its default solver for logistic regression. Newton-Raphson is a second-order method that uses the actual Hessian, giving it natural robustness to differences in predictor scale. No manual scaling is required.
The formula-based interface is straightforward:
import statsmodels.formula.api as smf
model = smf.logit(
"y_response ~ age + balance + day + duration + campaign + pdays + previous",
data=df
).fit()
print(model.summary())
The output includes coefficients, standard errors, z-statistics, p-values, confidence intervals, log-likelihood, AIC, and BIC. The official statsmodels documentation shows convergence in just 7 iterations with no warnings and no preprocessing required. (statsmodels Discrete Models Documentation)
Summary
| Aspect | statsmodels | sklearn |
|---|---|---|
| Default solver | Newton-Raphson | L-BFGS |
| Scaling required | No | Yes (practically mandatory) |
| p-values, std errors | Yes | No |
| Formula syntax | Yes | No |
| Convergence warnings | Rare | Common without scaling |
For GLM fitting in Python, statsmodels handles the iterative optimisation naturally — with no additional preprocessing burden on the analyst. sklearn‘s default solver imposes a scaling requirement that is external to the GLM framework itself.
References
- Nelder, J.A. and Wedderburn, R.W.M. (1972). Generalized Linear Models. Journal of the Royal Statistical Society, Series A, 135(3), 370–384.
- McCullagh, P. and Nelder, J.A. (1989). Generalized Linear Models (2nd ed.). Chapman and Hall.
- scikit-learn. LogisticRegression API Documentation. https://scikit-learn.org/stable/modules/generated/sklearn.linear_model.LogisticRegression.html
- scikit-learn. ConvergenceWarning. https://scikit-learn.org/stable/modules/generated/sklearn.exceptions.ConvergenceWarning.html
- statsmodels. Regression with Discrete Dependent Variable. https://www.statsmodels.org/stable/discretemod.html
- statsmodels. Logit.fit() API. https://www.statsmodels.org/stable/generated/statsmodels.discrete.discrete_model.Logit.fit.html
All web links cited in this write-up were accessed and the content verified during the preparation of this document in May 2026.