Variable Treatment and Coefficient Interpretation in Logistic Regression
A Case Study Using the Bank Marketing Dataset
About the Data
The dataset used in this analysis is the Bank Marketing Dataset from the UCI Machine Learning Repository, contributed by Sérgio Moro, P. Cortez, and P. Rita (2014). It is related to direct marketing campaigns — conducted via phone calls — of a Portuguese banking institution. The outcome of interest is whether a client subscribed to a term deposit (y: yes/no).
- Source: UCI Machine Learning Repository — Bank Marketing
- Total instances: 45,211 (full dataset); the working subset used here has 4,521 observations
- Features: 16 input variables covering client demographics, campaign contact details, and previous campaign outcomes
- Response variable:
y— binary (yes = 1, no = 0) - DOI: https://doi.org/10.24432/C5K306
Variables Used in the Model
From the available 16 features, the following were selected for the logistic regression model:
| Variable | Type | Description |
|---|---|---|
age | Numeric | Client age in years |
balance | Numeric | Average yearly account balance (euros) |
duration | Numeric | Last contact duration in seconds |
campaign | Numeric | Number of contacts during this campaign |
pdays | Numeric | Days since last contact from previous campaign |
previous | Numeric | Number of contacts before this campaign |
month | Categorical | Last contact month → recoded to quarter |
day | Numeric | Last contact day of month → recoded to fortnight |
job | Categorical | Occupation → recoded to job_C |
Derived Variables — Recoding for Better Interpretability
Three variables were recoded before model fitting. The rationale and mapping for each is given below.
1. month → quarter
The original month variable has 12 levels, producing 11 dummy variables in the model. This adds noise without meaningful gain. Collapsing into quarters aligns with business campaign cycles and reduces dummy explosion.
Q1 — jan, feb, mar
Q2 — apr, may, jun
Q3 — jul, aug, sep
Q4 — oct, nov, dec
df$quarter <- ifelse(df$month %in% c("jan","feb","mar"), "Q1",
ifelse(df$month %in% c("apr","may","jun"), "Q2",
ifelse(df$month %in% c("jul","aug","sep"), "Q3", "Q4")))
2. day → fortnight
The day of the month as a raw numeric predictor carries no meaningful linear interpretation in logistic regression. Splitting into fortnights captures early versus late month contact timing in a practically meaningful way.
F1 — day 1 to 15
F2 — day 16 to 31
df$fortnight <- ifelse(df$day >= 1 & df$day <= 15, "F1", "F2")
3. job → job_C
The original job variable has 12 levels, producing 11 dummy variables. Six business-meaningful groups were formed to reduce model noise while preserving interpretable occupational distinctions.
Employed — services, management, blue-collar, technician, admin.
Self-Employed — self-employed, entrepreneur
Student — student
Unemployed — unemployed
Others — retired, housemaid, unknown
df$job_C <- ifelse(df$job %in% c("services","management","blue-collar",
"technician","admin."), "Employed",
ifelse(df$job %in% c("self-employed","entrepreneur"), "Self-Employed",
ifelse(df$job == "student", "Student",
ifelse(df$job == "unemployed", "Unemployed", "Others"))))
The Problem with balance in Its Raw Form
The balance variable records the average yearly account balance in euros. Its distribution is:
| Statistic | Value |
|---|---|
| Minimum | −3,313 |
| Maximum | 71,188 |
| Median | 444 |
| Mean | 1,423 |
| SD | 3,010 |
When balance is entered into the logistic regression in its raw form, the estimated coefficient across all models comes out as approximately 0.0000 — appearing as zero in any table rounded to four decimal places.
This is not a modelling failure and it has nothing to do with significance. It is a straightforward arithmetic consequence of scale. The coefficient represents the change in log-odds for a one euro increase in balance. Given that balance spans a range of over 74,000 euros, a one-euro change carries negligible weight — and so the coefficient shrinks to a very small number. When rounded to four decimal places, it displays as zero.
Why the Naive Shift-and-Log Is Not Appropriate
A common workaround for applying log transformation to a variable with negative values is to shift the entire variable upward by adding a constant — typically abs(min(x)) + 1 — before taking the log. While this makes the log computable, it is artificial. The shift distorts the original meaning of the variable, since the zero point and negative values in balance carry real economic information (a negative balance indicates debt, zero indicates no savings, positive indicates savings). Artificially relocating this scale misrepresents the variable.
The Appropriate Transformation: Inverse Hyperbolic Sine (asinh)
The inverse hyperbolic sine transformation is defined as:
asinh(x) = ln(x + √(x² + 1))
Its key properties make it well suited for a variable like balance:
- For large positive values, it behaves like a log transformation — compressing the right tail
- For large negative values, it mirrors the same compression symmetrically
- At zero, it passes through naturally — no adjustment needed
- It handles negative, zero, and positive values in a single expression without any manual constant
This is increasingly used in econometrics for income, wealth, and balance-type variables precisely because of these properties.
df$balance_asinh <- asinh(df$balance)
What Changes After Transformation
After applying asinh, the coefficient of balance_asinh in the logistic regression model becomes interpretable and visible. The estimated log-odds ratio (ln OR) across all fitted models is consistent:
| Model | ln(OR) for balance_asinh |
|---|---|
Base glm() | 0.0449 |
brglm (Firth bias reduction) | 0.0449 |
brglm2 (extended bias reduction) | 0.0449 |
speedglm | 0.0449 |
biglm | 0.0449 |
glmnet — Lasso (penalised) | 0.0383 |
The consistency across glm, brglm, brglm2, speedglm, and biglm is expected — all five use maximum likelihood with Newton-Raphson or IRLS, and for this dataset size (4,521 observations), bias correction makes negligible difference to point estimates.
The glmnet Lasso estimate is slightly lower at 0.0383 — this is the effect of L1 regularisation, which shrinks coefficients toward zero as part of its penalisation. The shrinkage is mild here, consistent with balance_asinh being a moderately informative predictor.
The ln(OR) of 0.0449 means that for each unit increase in asinh(balance) — roughly corresponding to a doubling of balance in the positive range — the log-odds of subscribing to a term deposit increases by 0.0449, holding all other variables constant.
Full R Code
# ============================================================
# GLM Fitting on Bank Marketing Data — All R Packages
# Derived Variables: quarter, fortnight, job_C
# balance treated with asinh transformation
# Response: y (yes/no → 1/0)
# ============================================================
# ---- Install packages if not already installed ----
packages <- c("glmnet", "MASS", "brglm", "brglm2", "speedglm",
"biglm", "knitr", "kableExtra", "dplyr")
install.packages(setdiff(packages, rownames(installed.packages())))
# ---- Load libraries ----
library(glmnet)
library(MASS)
library(brglm)
library(brglm2)
library(speedglm)
library(biglm)
library(knitr)
library(kableExtra)
library(dplyr)
# ============================================================
# 1. READ DATA
# ============================================================
df <- read.csv("Bank_Mktg_Data.csv", stringsAsFactors = FALSE)
# ============================================================
# 2. DERIVED COLUMNS
# ============================================================
# month → quarter
df$quarter <- ifelse(df$month %in% c("jan","feb","mar"), "Q1",
ifelse(df$month %in% c("apr","may","jun"), "Q2",
ifelse(df$month %in% c("jul","aug","sep"), "Q3", "Q4")))
# day → fortnight
df$fortnight <- ifelse(df$day >= 1 & df$day <= 15, "F1", "F2")
# job → job_C
df$job_C <- ifelse(df$job %in% c("services","management","blue-collar",
"technician","admin."), "Employed",
ifelse(df$job %in% c("self-employed","entrepreneur"), "Self-Employed",
ifelse(df$job == "student", "Student",
ifelse(df$job == "unemployed", "Unemployed", "Others"))))
# balance → asinh transformation
df$balance_asinh <- asinh(df$balance)
# Response variable: y → binary 0/1
df$y_bin <- ifelse(df$y == "yes", 1, 0)
# Convert derived cols to factors
df$quarter <- factor(df$quarter)
df$fortnight <- factor(df$fortnight)
df$job_C <- factor(df$job_C)
# ============================================================
# 3. FORMULA
# ============================================================
formula_glm <- y_bin ~ age + balance_asinh + duration + campaign +
pdays + previous + quarter + fortnight + job_C
# ============================================================
# 4. FIT ALL MODELS
# ============================================================
# --- 4a. Base glm() ---
model_glm <- glm(formula_glm, data = df, family = binomial(link = "logit"))
# --- 4b. brglm ---
model_brglm <- brglm(formula_glm, family = binomial(link = "logit"), data = df)
# --- 4c. brglm2 ---
model_brglm2 <- glm(formula_glm, family = binomial(link = "logit"),
data = df, method = "brglmFit")
# --- 4d. speedglm ---
model_speed <- speedglm(formula_glm, data = df,
family = binomial(link = "logit"))
# --- 4e. biglm ---
model_big <- bigglm(formula_glm, data = df,
family = binomial(link = "logit"), chunksize = 500)
# --- 4f. glmnet (Lasso at optimal lambda) ---
X <- model.matrix(formula_glm, data = df)[, -1]
Y <- df$y_bin
set.seed(42)
cv_glmnet <- cv.glmnet(X, Y, family = "binomial", alpha = 1)
coef_glmnet <- as.matrix(coef(cv_glmnet, s = "lambda.min"))
# --- 4g. MASS glm.nb() — Negative Binomial (count response: previous) ---
formula_nb <- previous ~ age + balance_asinh + duration + campaign +
quarter + fortnight + job_C
model_nb <- glm.nb(formula_nb, data = df)
# ============================================================
# 5. HELPER — EXTRACT COEF TABLE
# ============================================================
extract_coef <- function(model, model_name) {
s <- summary(model)
co <- as.data.frame(s$coefficients)
names(co) <- c("Estimate", "Std.Error", "z_or_t", "p.value")[seq_len(ncol(co))]
co$Variable <- rownames(co)
co$Model <- model_name
co <- co[, c("Model", "Variable", "Estimate", "Std.Error", "p.value")]
rownames(co) <- NULL
co
}
# ============================================================
# 6. COLLECT RESULTS
# ============================================================
res_glm <- extract_coef(model_glm, "glm")
res_brglm <- extract_coef(model_brglm, "brglm")
res_brglm2 <- extract_coef(model_brglm2, "brglm2")
res_speed <- extract_coef(model_speed, "speedglm")
res_nb <- extract_coef(model_nb, "MASS glm.nb\n(response: previous)")
# biglm — extract manually
coef_big <- summary(model_big)$mat
res_big <- data.frame(
Model = "biglm",
Variable = rownames(coef_big),
Estimate = round(coef_big[, "Coef"], 4),
Std.Error = round(coef_big[, "SE"], 4),
p.value = NA_real_,
stringsAsFactors = FALSE
)
# glmnet — no SE or p-values
res_glmnet <- data.frame(
Model = "glmnet (Lasso)",
Variable = rownames(coef_glmnet),
Estimate = round(as.numeric(coef_glmnet), 4),
Std.Error = NA_real_,
p.value = NA_real_,
stringsAsFactors = FALSE
)
res_glmnet <- res_glmnet[res_glmnet$Estimate != 0, ]
# ============================================================
# 7. COMBINE ALL RESULTS
# ============================================================
all_results <- bind_rows(res_glm, res_brglm, res_brglm2,
res_speed, res_big, res_glmnet, res_nb)
all_results$Estimate <- round(all_results$Estimate, 4)
all_results$Std.Error <- round(all_results$Std.Error, 4)
all_results$p.value <- round(all_results$p.value, 4)
all_results$Sig <- ifelse(is.na(all_results$p.value), "",
ifelse(all_results$p.value < 0.001, "***",
ifelse(all_results$p.value < 0.01, "**",
ifelse(all_results$p.value < 0.05, "*",
ifelse(all_results$p.value < 0.1, ".", "")))))
# ============================================================
# 8. WIDE FORMAT & KABLEEXTRA TABLE
# ============================================================
format_cell <- function(est, se, sig) {
ifelse(is.na(est), "—",
paste0(est,
ifelse(is.na(se), "", paste0("\n(", se, ")")),
ifelse(sig == "", "", paste0(" ", sig))))
}
all_results$Cell <- format_cell(all_results$Estimate,
all_results$Std.Error,
all_results$Sig)
wide <- tidyr::pivot_wider(
all_results[, c("Model", "Variable", "Cell")],
names_from = "Model",
values_from = "Cell",
values_fill = "—"
)
wide %>%
kable(
format = "html",
caption = "GLM Results — All Packages | Bank Marketing Data<br>
<small>Estimate (Std. Error) | Significance: *** p<0.001 ** p<0.01 * p<0.05 . p<0.1<br>
biglm: p-values not directly available | glmnet: penalised — no SE or p-values<br>
MASS glm.nb: Negative Binomial fitted on count response (previous)</small>",
align = c("l", rep("c", ncol(wide) - 1)),
escape = FALSE
) %>%
kable_styling(
bootstrap_options = c("striped", "hover", "condensed", "bordered"),
full_width = TRUE,
font_size = 13
) %>%
column_spec(1, bold = TRUE, background = "#f0f0f0") %>%
row_spec(0, bold = TRUE, background = "#2c3e50", color = "white") %>%
add_header_above(
c(" " = 1,
"Standard GLM" = 1,
"Bias Reduction" = 2,
"Large Data" = 2,
"Penalised" = 1,
"Neg. Binomial" = 1)
) %>%
scroll_box(width = "100%", height = "600px")
Key Takeaway
The balance variable, when entered raw into logistic regression, produces an estimated coefficient that is arithmetically near zero — not because it is unimportant, but because its unit (one euro) is too small relative to its actual range. This is a variable treatment issue, not a modelling issue.
Applying the inverse hyperbolic sine transformation resolves this cleanly — without any artificial shifting, without discarding negative values, and without losing the economic meaning of the variable. The resulting coefficient of 0.0449 (ln OR) is consistent across all five maximum likelihood based models, with a mild shrinkage to 0.0383 under Lasso penalisation in glmnet — which is expected behaviour for regularised estimation.
This finding underlines a broader point: before interpreting any regression coefficient as negligible, the scale and distribution of the predictor must be examined. A zero-looking coefficient is not always a zero effect.
Reference
Moro, S., Rita, P., & Cortez, P. (2014). Bank Marketing [Dataset]. UCI Machine Learning Repository. https://doi.org/10.24432/C5K306
All web links cited in this write-up were accessed and the content verified during the preparation of this document in May 2026.