Transformation in GLM

Variable Treatment and Coefficient Interpretation in Logistic Regression

A Case Study Using the Bank Marketing Dataset


About the Data

The dataset used in this analysis is the Bank Marketing Dataset from the UCI Machine Learning Repository, contributed by Sérgio Moro, P. Cortez, and P. Rita (2014). It is related to direct marketing campaigns — conducted via phone calls — of a Portuguese banking institution. The outcome of interest is whether a client subscribed to a term deposit (y: yes/no).

  • Source: UCI Machine Learning Repository — Bank Marketing
  • Total instances: 45,211 (full dataset); the working subset used here has 4,521 observations
  • Features: 16 input variables covering client demographics, campaign contact details, and previous campaign outcomes
  • Response variable: y — binary (yes = 1, no = 0)
  • DOI: https://doi.org/10.24432/C5K306

Variables Used in the Model

From the available 16 features, the following were selected for the logistic regression model:

VariableTypeDescription
ageNumericClient age in years
balanceNumericAverage yearly account balance (euros)
durationNumericLast contact duration in seconds
campaignNumericNumber of contacts during this campaign
pdaysNumericDays since last contact from previous campaign
previousNumericNumber of contacts before this campaign
monthCategoricalLast contact month → recoded to quarter
dayNumericLast contact day of month → recoded to fortnight
jobCategoricalOccupation → recoded to job_C

Derived Variables — Recoding for Better Interpretability

Three variables were recoded before model fitting. The rationale and mapping for each is given below.

1. month → quarter

The original month variable has 12 levels, producing 11 dummy variables in the model. This adds noise without meaningful gain. Collapsing into quarters aligns with business campaign cycles and reduces dummy explosion.

Q1 — jan, feb, mar
Q2 — apr, may, jun
Q3 — jul, aug, sep
Q4 — oct, nov, dec
df$quarter <- ifelse(df$month %in% c("jan","feb","mar"), "Q1",
              ifelse(df$month %in% c("apr","may","jun"), "Q2",
              ifelse(df$month %in% c("jul","aug","sep"), "Q3", "Q4")))

2. day → fortnight

The day of the month as a raw numeric predictor carries no meaningful linear interpretation in logistic regression. Splitting into fortnights captures early versus late month contact timing in a practically meaningful way.

F1 — day 1 to 15
F2 — day 16 to 31
df$fortnight <- ifelse(df$day >= 1 & df$day <= 15, "F1", "F2")

3. job → job_C

The original job variable has 12 levels, producing 11 dummy variables. Six business-meaningful groups were formed to reduce model noise while preserving interpretable occupational distinctions.

Employed      — services, management, blue-collar, technician, admin.
Self-Employed — self-employed, entrepreneur
Student       — student
Unemployed    — unemployed
Others        — retired, housemaid, unknown
df$job_C <- ifelse(df$job %in% c("services","management","blue-collar",
                                  "technician","admin."), "Employed",
            ifelse(df$job %in% c("self-employed","entrepreneur"), "Self-Employed",
            ifelse(df$job == "student",    "Student",
            ifelse(df$job == "unemployed", "Unemployed", "Others"))))

The Problem with balance in Its Raw Form

The balance variable records the average yearly account balance in euros. Its distribution is:

StatisticValue
Minimum−3,313
Maximum71,188
Median444
Mean1,423
SD3,010

When balance is entered into the logistic regression in its raw form, the estimated coefficient across all models comes out as approximately 0.0000 — appearing as zero in any table rounded to four decimal places.

This is not a modelling failure and it has nothing to do with significance. It is a straightforward arithmetic consequence of scale. The coefficient represents the change in log-odds for a one euro increase in balance. Given that balance spans a range of over 74,000 euros, a one-euro change carries negligible weight — and so the coefficient shrinks to a very small number. When rounded to four decimal places, it displays as zero.


Why the Naive Shift-and-Log Is Not Appropriate

A common workaround for applying log transformation to a variable with negative values is to shift the entire variable upward by adding a constant — typically abs(min(x)) + 1 — before taking the log. While this makes the log computable, it is artificial. The shift distorts the original meaning of the variable, since the zero point and negative values in balance carry real economic information (a negative balance indicates debt, zero indicates no savings, positive indicates savings). Artificially relocating this scale misrepresents the variable.


The Appropriate Transformation: Inverse Hyperbolic Sine (asinh)

The inverse hyperbolic sine transformation is defined as:

asinh(x) = ln(x + √(x² + 1))

Its key properties make it well suited for a variable like balance:

  • For large positive values, it behaves like a log transformation — compressing the right tail
  • For large negative values, it mirrors the same compression symmetrically
  • At zero, it passes through naturally — no adjustment needed
  • It handles negative, zero, and positive values in a single expression without any manual constant

This is increasingly used in econometrics for income, wealth, and balance-type variables precisely because of these properties.

df$balance_asinh <- asinh(df$balance)

What Changes After Transformation

After applying asinh, the coefficient of balance_asinh in the logistic regression model becomes interpretable and visible. The estimated log-odds ratio (ln OR) across all fitted models is consistent:

Modelln(OR) for balance_asinh
Base glm()0.0449
brglm (Firth bias reduction)0.0449
brglm2 (extended bias reduction)0.0449
speedglm0.0449
biglm0.0449
glmnet — Lasso (penalised)0.0383

The consistency across glm, brglm, brglm2, speedglm, and biglm is expected — all five use maximum likelihood with Newton-Raphson or IRLS, and for this dataset size (4,521 observations), bias correction makes negligible difference to point estimates.

The glmnet Lasso estimate is slightly lower at 0.0383 — this is the effect of L1 regularisation, which shrinks coefficients toward zero as part of its penalisation. The shrinkage is mild here, consistent with balance_asinh being a moderately informative predictor.

The ln(OR) of 0.0449 means that for each unit increase in asinh(balance) — roughly corresponding to a doubling of balance in the positive range — the log-odds of subscribing to a term deposit increases by 0.0449, holding all other variables constant.


Full R Code

# ============================================================
# GLM Fitting on Bank Marketing Data — All R Packages
# Derived Variables: quarter, fortnight, job_C
# balance treated with asinh transformation
# Response: y (yes/no → 1/0)
# ============================================================

# ---- Install packages if not already installed ----
packages <- c("glmnet", "MASS", "brglm", "brglm2", "speedglm",
              "biglm", "knitr", "kableExtra", "dplyr")
install.packages(setdiff(packages, rownames(installed.packages())))

# ---- Load libraries ----
library(glmnet)
library(MASS)
library(brglm)
library(brglm2)
library(speedglm)
library(biglm)
library(knitr)
library(kableExtra)
library(dplyr)

# ============================================================
# 1. READ DATA
# ============================================================

df <- read.csv("Bank_Mktg_Data.csv", stringsAsFactors = FALSE)

# ============================================================
# 2. DERIVED COLUMNS
# ============================================================

# month → quarter
df$quarter <- ifelse(df$month %in% c("jan","feb","mar"), "Q1",
              ifelse(df$month %in% c("apr","may","jun"), "Q2",
              ifelse(df$month %in% c("jul","aug","sep"), "Q3", "Q4")))

# day → fortnight
df$fortnight <- ifelse(df$day >= 1 & df$day <= 15, "F1", "F2")

# job → job_C
df$job_C <- ifelse(df$job %in% c("services","management","blue-collar",
                                  "technician","admin."), "Employed",
            ifelse(df$job %in% c("self-employed","entrepreneur"), "Self-Employed",
            ifelse(df$job == "student",    "Student",
            ifelse(df$job == "unemployed", "Unemployed", "Others"))))

# balance → asinh transformation
df$balance_asinh <- asinh(df$balance)

# Response variable: y → binary 0/1
df$y_bin <- ifelse(df$y == "yes", 1, 0)

# Convert derived cols to factors
df$quarter   <- factor(df$quarter)
df$fortnight <- factor(df$fortnight)
df$job_C     <- factor(df$job_C)

# ============================================================
# 3. FORMULA
# ============================================================

formula_glm <- y_bin ~ age + balance_asinh + duration + campaign +
                       pdays + previous + quarter + fortnight + job_C

# ============================================================
# 4. FIT ALL MODELS
# ============================================================

# --- 4a. Base glm() ---
model_glm <- glm(formula_glm, data = df, family = binomial(link = "logit"))

# --- 4b. brglm ---
model_brglm <- brglm(formula_glm, family = binomial(link = "logit"), data = df)

# --- 4c. brglm2 ---
model_brglm2 <- glm(formula_glm, family = binomial(link = "logit"),
                    data = df, method = "brglmFit")

# --- 4d. speedglm ---
model_speed <- speedglm(formula_glm, data = df,
                        family = binomial(link = "logit"))

# --- 4e. biglm ---
model_big <- bigglm(formula_glm, data = df,
                    family = binomial(link = "logit"), chunksize = 500)

# --- 4f. glmnet (Lasso at optimal lambda) ---
X <- model.matrix(formula_glm, data = df)[, -1]
Y <- df$y_bin
set.seed(42)
cv_glmnet   <- cv.glmnet(X, Y, family = "binomial", alpha = 1)
coef_glmnet <- as.matrix(coef(cv_glmnet, s = "lambda.min"))

# --- 4g. MASS glm.nb() — Negative Binomial (count response: previous) ---
formula_nb <- previous ~ age + balance_asinh + duration + campaign +
                         quarter + fortnight + job_C
model_nb <- glm.nb(formula_nb, data = df)

# ============================================================
# 5. HELPER — EXTRACT COEF TABLE
# ============================================================

extract_coef <- function(model, model_name) {
  s  <- summary(model)
  co <- as.data.frame(s$coefficients)
  names(co) <- c("Estimate", "Std.Error", "z_or_t", "p.value")[seq_len(ncol(co))]
  co$Variable <- rownames(co)
  co$Model    <- model_name
  co <- co[, c("Model", "Variable", "Estimate", "Std.Error", "p.value")]
  rownames(co) <- NULL
  co
}

# ============================================================
# 6. COLLECT RESULTS
# ============================================================

res_glm    <- extract_coef(model_glm,    "glm")
res_brglm  <- extract_coef(model_brglm,  "brglm")
res_brglm2 <- extract_coef(model_brglm2, "brglm2")
res_speed  <- extract_coef(model_speed,  "speedglm")
res_nb     <- extract_coef(model_nb,     "MASS glm.nb\n(response: previous)")

# biglm — extract manually
coef_big <- summary(model_big)$mat
res_big  <- data.frame(
  Model     = "biglm",
  Variable  = rownames(coef_big),
  Estimate  = round(coef_big[, "Coef"], 4),
  Std.Error = round(coef_big[, "SE"],   4),
  p.value   = NA_real_,
  stringsAsFactors = FALSE
)

# glmnet — no SE or p-values
res_glmnet <- data.frame(
  Model     = "glmnet (Lasso)",
  Variable  = rownames(coef_glmnet),
  Estimate  = round(as.numeric(coef_glmnet), 4),
  Std.Error = NA_real_,
  p.value   = NA_real_,
  stringsAsFactors = FALSE
)
res_glmnet <- res_glmnet[res_glmnet$Estimate != 0, ]

# ============================================================
# 7. COMBINE ALL RESULTS
# ============================================================

all_results <- bind_rows(res_glm, res_brglm, res_brglm2,
                         res_speed, res_big, res_glmnet, res_nb)

all_results$Estimate  <- round(all_results$Estimate,  4)
all_results$Std.Error <- round(all_results$Std.Error, 4)
all_results$p.value   <- round(all_results$p.value,   4)

all_results$Sig <- ifelse(is.na(all_results$p.value), "",
                   ifelse(all_results$p.value < 0.001, "***",
                   ifelse(all_results$p.value < 0.01,  "**",
                   ifelse(all_results$p.value < 0.05,  "*",
                   ifelse(all_results$p.value < 0.1,   ".", "")))))

# ============================================================
# 8. WIDE FORMAT & KABLEEXTRA TABLE
# ============================================================

format_cell <- function(est, se, sig) {
  ifelse(is.na(est), "—",
         paste0(est,
                ifelse(is.na(se), "", paste0("\n(", se, ")")),
                ifelse(sig == "", "", paste0(" ", sig))))
}

all_results$Cell <- format_cell(all_results$Estimate,
                                all_results$Std.Error,
                                all_results$Sig)

wide <- tidyr::pivot_wider(
  all_results[, c("Model", "Variable", "Cell")],
  names_from  = "Model",
  values_from = "Cell",
  values_fill = "—"
)

wide %>%
  kable(
    format  = "html",
    caption = "GLM Results — All Packages | Bank Marketing Data<br>
               <small>Estimate (Std. Error) | Significance: *** p<0.001  ** p<0.01  * p<0.05  . p<0.1<br>
               biglm: p-values not directly available | glmnet: penalised — no SE or p-values<br>
               MASS glm.nb: Negative Binomial fitted on count response (previous)</small>",
    align  = c("l", rep("c", ncol(wide) - 1)),
    escape = FALSE
  ) %>%
  kable_styling(
    bootstrap_options = c("striped", "hover", "condensed", "bordered"),
    full_width        = TRUE,
    font_size         = 13
  ) %>%
  column_spec(1, bold = TRUE, background = "#f0f0f0") %>%
  row_spec(0, bold = TRUE, background = "#2c3e50", color = "white") %>%
  add_header_above(
    c(" " = 1,
      "Standard GLM" = 1,
      "Bias Reduction" = 2,
      "Large Data" = 2,
      "Penalised" = 1,
      "Neg. Binomial" = 1)
  ) %>%
  scroll_box(width = "100%", height = "600px")

Key Takeaway

The balance variable, when entered raw into logistic regression, produces an estimated coefficient that is arithmetically near zero — not because it is unimportant, but because its unit (one euro) is too small relative to its actual range. This is a variable treatment issue, not a modelling issue.

Applying the inverse hyperbolic sine transformation resolves this cleanly — without any artificial shifting, without discarding negative values, and without losing the economic meaning of the variable. The resulting coefficient of 0.0449 (ln OR) is consistent across all five maximum likelihood based models, with a mild shrinkage to 0.0383 under Lasso penalisation in glmnet — which is expected behaviour for regularised estimation.

This finding underlines a broader point: before interpreting any regression coefficient as negligible, the scale and distribution of the predictor must be examined. A zero-looking coefficient is not always a zero effect.


Reference

Moro, S., Rita, P., & Cortez, P. (2014). Bank Marketing [Dataset]. UCI Machine Learning Repository. https://doi.org/10.24432/C5K306

All web links cited in this write-up were accessed and the content verified during the preparation of this document in May 2026.

Scroll to Top