Plugging the Leak: Gradient Boosting and Survival Analysis for Customer Retention

Why Churn Modeling Matters

  • The “Leaky Bucket” metaphor was first popularized by subscription marketers using CRM in the ‘90’s and I see it finally making its way into B2B marketing as a result of XaaS and Cloud technology.
  • Coming from subscriber marketing, the idea of customer churn and retention were always paramount for B2C marketers. With the rise of XaaS and Cloud technology, the ‘Leaky Bucket’ metaphor—a staple of 90s CRM—has evolved from a B2C concept into a mission-critical B2B framework.
  • Churn modeling fits nicely with Propensity-to-Buy and Customer Lifetime Value because it employs many of the same techniques. There are two ways to look at churn which require methodologies we’ve seen in my past articles:
  • What is the likelihood of a customer leaving in the next 30 days? This requires a statistical or machine-learning classifier such as XGBoost (which I used in my article on propensity to buy).  Here we want a list of customers that have a high churn likelihood, so the target for prediction is changed from purchase to churn (yes/no).
  • When will a customer churn? Knowing the timing of churn means turning to survival modeling, which I employed in my article on Customer Lifetime Value analysis, and here the recommended python library is called Lifelines.  There are also some great visualizations of survival probability of different customer segments that can be generated, and I think that is an excellent way to identify and understand the drivers (and profile) the customers who churn – whether they are people (B2C) or businesses (B2B accounts).

The Dataset

The IBM Telco Customer Churn dataset is a famous dataset used by data scientists to predict churn and work with marketers to develop customer retention strategies. It is a fictional telecommunications customer database, providing a mix of demographic, service, and financial data. It consists of data on 7,043 customers with 21 features including:

Demographics: Information about the customer’s gender, age range (Senior Citizen), and whether they have partners or dependents.

Account Information: How long they’ve been a customer (tenure), their contract type (Month-to-month, One year, Two year), payment method, paperless billing, and charges (MonthlyCharges and TotalCharges).

Services: Specific services the customer has signed up for, including Phone, Multiple Lines, Internet (DSL, Fiber Optic, etc.), Online Security, Online Backup, Device Protection, Tech Support, and Streaming TV/Movies.

The Target: The Churn column, indicating whether the customer left within the last month (Yes/No).


Exploratory Data Analysis (EDA) and Data Quality

An analysis of the dataset showed that it was very complete, with only 11 missing Total Charges so other than a few data transformations (strings à integers) I wasn’t concerned about doing a lot of cleaning and data manipulation.

Since the data had a field for Churn (Yes/No) a customer profile could be generated which provided a lot of insight into churned customers:

From these charts, we can see that locking customers in with long term contracts and automatic withdrawal seems a good retention strategy, as churners tend to be on monthly contracts and paying by check.  Perhaps Senior Citizens are more price-sensitive and have bandwidth to shop for discounts, but that would have to be tested.  I also generated a correlation matrix which showed the correlation of the features to customer churn as another way of looking at the characteristics of churned customers:


The Classification Question: Will they leave? (XGBoost): Real-world tradeoffs in modeling the probability and timing of customer churn.

There are a lot of features that can be used together to predict churn from this dataset. The first question is “will a customer churn”?  So, I think of this as a propensity to churn model which classifies churners based on the statistical probability of churn (essentially a propensity to buy model with the target changed from “purchase” to “churn”).

Typically, in sales and marketing models we have to decide when we target:

  1. Do we have a conservative model that is relatively accurate in predicting purchases or churn, but misses a lot of customers because it is so conservative? From a marketing expense perspective this is efficient.
  2. Do we tune more aggressively and cast a wider net, but target a lot of prospects or customers that will not purchase or churn (false positives)? This is less cost-efficient but will uncover more absolute revenue or churn “by knocking on more doors.”

In my experience, I have always leaned towards the more aggressive model (within reason), and sacrifice precision to hit more potential purchasers/churners.

Looking at the confusion matrix below for my baseline (aggressive) model, it is great at capturing churn within a 30 day window (Recall = 0.82 means that it captures 82% of the customers who churned) but at a high cost of false positives (Precision = 0.50 which means that any marketing or sales effort will be inefficient because it will cast a very wide net).  Further tuning improved precision, but at the cost of rejecting a lot of churning customers who had lower probability scores based on the available data, so I would not go with the conservative model or would continue to tune.

In short: a false positive will increase marketing or discounting, while a false negative will lose a customer.

In the high-risk customer base, we have two groups:

  1. High risk and high monetary value customers that are on month-to-month contracts that should be encouraged to sign long term contracts.
  2. High risk and lower monetary value customers that are locked into one- to two-year contracts that are targets for long term retention and customer satisfaction programs.

For some more input on sales and marketing program design, XGBoost also produces a list of the features (variables below) that are most influential in the model, which can be used to test tactical adjustments such as discounting for two-year contracts and content rebalancing.


The Time-based Question: When will they leave? (Lifelines)

While the XGBoost classifier was able to tell us “This customer is at-risk” a survival model can tell us “This customer has a 60% chance of making it a year” so that a marketer can time retention efforts. To find out when a customer will churn, we need a model that is designed to predict the timing of events.  Miller finds that “a good example of a duration or survival model in marketing is customer lifetime estimation” (Miller, 2015). These are generally categorized as Survival Models:

“… medical researchers are often interested in the effects of certain drugs on the timing of death (or recovery) among a sample of patients. In fact, these statistical models are known most as survival models because they are used often by biostatisticians, epidemiologists, and other researchers to study the time between diagnosis and death. … social and behavioral scientists have adopted these models for a variety of purposes.”

Hoffmann, J. P. (2016)

In Python there is a package called Lifelines (lifelines import KaplanMeierFitter) which I used for this task. The survival probability curve by contract type highlights the value of two-year contracts.


Financial Impact

By combining two churn modeling approaches sales and marketing can evolve from a reactive strategy to proactively developing retention programs that improve financial performance. An aggressive XGBoost model gives the business the ability to identify high-risk/ high-monetary value customers in advance, while Survival Analysis improves marketing effectiveness by guiding the timing of retention and customer satisfaction campaigns.

The addition of churned customer profiling can be used to develop relevant marketing communications and pricing strategies.


Citations:

Chen, T., & Guestrin, C. (2016). XGBoost: A Scalable Tree Boosting System. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785–794). ACM. https://doi.org/10.1145/2939672.2939785

Davidson-Pilon, C. (2023). lifelines: survival analysis in Python (Version 0.27.8) [Software]. Zenodo. https://doi.org/10.5281/zenodo.8259706

Hoffmann, J. P. (2016). Generalized Linear Models: An Applied Approach (2nd ed.). Routledge.

IBM Sample Data Sets. (n.d.). Telco Customer Churn: Focused customer retention programs. Retrieved from https://community.ibm.com/community/user/businessanalytics/viewdocument/telco-customer-churn

Miller, T. W. (2015). Marketing Data Science: Modeling Techniques in Predictive Analytics with R and Python. Pearson Education.


Technical Keywords & Methodology Index

Methodology & Strategy: Churn Modeling, Time-to-Event Analysis, Predictive Classification, Customer Retention Strategy, “Leaky Bucket” Lifecycle Framework, Customer Lifetime Value (CLV) Alignment.

Statistical Concepts: Gradient Boosting (XGBoost), Survival Analysis (Kaplan-Meier), Confusion Matrix (Precision-Recall trade-offs), Propensity Scoring, Hazard Modeling, Feature Importance Analysis.

Data Engineering & Diagnostics: Imbalanced Dataset Handling, Feature Engineering for Churn, Data Quality Auditing, Temporal Logic (Survival Modeling), Outlier Detection.

Revenue Operations (RevOps) Logic: Retention Forecasting, High-Risk/High-Value Segmentation, Marketing Efficiency vs. Coverage (Recall), Reactive vs. Proactive Retention Design.

Python Libraries & Documentation

For data scientists and revenue operations engineers looking to replicate this hybrid retention architecture, the following stack integrates classification modeling with duration-based survival analytics.

LibraryRole in PipelineStrategic Purpose
XGBoostPropensity ClassifierIdentifies high-risk churn candidates (Yes/No) using gradient boosting; crucial for immediate “30-day” intervention lists.
LifelinesSurvival AnalysisModels the timing of churn events, providing survival probabilities that help marketers time retention campaigns effectively.
Scikit-LearnPerformance MetricsGenerates confusion matrices, precision/recall trade-offs, and classification reports to balance marketing efficiency.
Pandas / NumPyData WranglingFacilitates data normalization, outlier management, and feature transformation for the IBM Telco dataset.
Matplotlib / SeabornVisualizationProduces survival curves, correlation matrices, and churn profiles for stakeholder visibility and executive alignment.

Posted in

One response to “Plugging the Leak: Gradient Boosting and Survival Analysis for Customer Retention”

  1. […] Temporal Consistency: Date ordering must be validated (e.g., ensuring MQL < SAL < SQL) to catch data entry errors or CRM sync issues before performing survival analysis. […]

Leave a Reply

Discover more from The Marketing Science Signal

Subscribe now to keep reading and get access to the full archive.

Continue reading