The Demand Generation Workhorse

Why Propensity-to-Buy Targeting Matters

  • Precision Resource Allocation: It identifies the specific customers and prospects with the highest likelihood to purchase, allowing marketing and sales teams to focus their budget and energy where it will yield the highest return.
  • Churn Mitigation: By predicting which contacts or companies are at risk of leaving for a competitor, P2B allows for proactive retention strategies before the customer actually departs.
  • Optimized Marketing Frequency: It prevents “over-indexing” or inundating low-potential customers with excessive outreach, identifying the “tipping point” where increased contact frequency actually decreases the probability of a sale.
  • Foundation for Advanced Metrics: P2B generates the critical probability scores required to calculate more complex financial models, such as Customer Lifetime Value (CLV), effectively bridging the gap between marketing activity and long-term profitability.
  • Strategic Profiling: It isolates high-propensity populations to uncover the influential variables (e.g., communication channel, past success, or financial stability) that define the “ideal” customer profile for future targeting.

Introduction

The “workhorse” of modern demand generation is a family of binary classification models. These models are designed to predict a specific response, such as “…whether a customer buys, whether a customer stays with the company or leaves to buy from another company, and whether the customer recommends a company’s products to another customer” (Miller, 2015). Bruce Ratner (2017), in his work on machine learning and data mining, calls this approach the “workhorse of response modeling.”

There are many statistical and machine-learning techniques for building Propensity-to-Buy (P2B) models—whether predicting churn, response rates, or identifying CRM opportunities that will convert to sales.  The concept has been around since the early days of direct mail (1980s), when direct mail marketers relied on Logistic Regression (still a powerful technique). Now, fueled by big data and high-performance computing, one of the most popular and effective techniques is called eXtreme Gradient Boosting (XGBoost) and I am using that for my examination of how it can be applied in marketing.

Propensity-to-buy has been the first technique we have used in any company I’ve had the pleasure of working in; it has evolved to become the foundation of modern digital marketing.  Technically efficient and scalable, once the base code is set, a P2B model can be trained to predict different targets with limited recalibration. 

Key Applications of P2B:

  1. Purchase Likelihood for a Product: Identifying which customers and prospects are most likely to purchase specific products. This applies to both new product launches and up-sell/cross-sell for existing products.
  2. Churn Risk: Predicting which companies or individual contacts are at risk of leaving to competitors.
  3. Lead Scoring and Conversion Probability: Forecasting which CRM opportunities are most likely to convert to “Closed-Won” to optimize the marketing and sales pipeline.
  4. CLV Integration: Generating the probability scores required to calculate a Customer Lifetime Value (CLV) model.
  5. Campaign Response: Determining which contacts are most likely to engage with or respond to a specific marketing program.
  6. Customer Valuation: Identifying which customers will be the most valuable overall across their entire purchase history.  For example, comparing a company’s Total IT Spend with P2B.
  7. Strategic Profiling: Isolating a “high-propensity” population to build ideal demographic and firmographic profiles for future targeting.

For the following example, XGBoost was perfect (an open-source machine learning ensemble method that builds multiple decision trees to correct previous errors). First introduced to me six years ago by data scientist Fuqiang Shi, it has become the gold standard for structured data. XGBoost gained global fame around 2014–2016 for dominating Kaggle competitions (often outperforming popular methods like Random Forest, Support Vector Machine, Bayesian Classifier, etc.).  That said, in practice a data scientist should test several methods to find the best model by comparing performance metrics.


The Data

To demonstrate how to build and score a model, I utilized the Bank Marketing Dataset from the well-known UC Irvine Machine Learning Repository. This dataset consists of 45,211 rows and 17 columns, representing a real-world scenario of a direct telemarketing campaign from a Portuguese banking institution (2008-2010). Here are the first five rows:

Bank Marketing Dataset from the well-known UC Irvine Machine Learning Repository.

UCI Machine Learning Repository: Bank Dataset. University of California, Irvine.


Model Development: A High-Level Overview

Propensity model development process.

This code snippet demonstrates how to optimize an XGBoost propensity-to-buy model by using GridSearchCV to systematically evaluate different combinations of hyperparameters, such as max_depth and learning_rate. By performing 3-fold cross-validation and selecting the model with the highest roc_auc score, this approach helps ensure the final model is tuned for better predictive accuracy in identifying high-intent customers. Running a grid search has replaced manual tuning as the standard for modelers.

Predictive Performance

The model achieved a predictive performance of 81% (ROC AUC), which was within range for a real-world application in my experience (although I have seen between 65% for prospecting to 90% for customer models). While I initially hard-coded the parameters, I followed up with a grid search (GridSearchCV) to ensure optimization. Since both approaches achieved the same predictive performance, I stopped tuning at that point. [Environment: Python Jupyter Notebook (Anaconda)].


Using Model for Decision Support: Rebalancing Marketing Frequency.

From a targeting perspective, we now have a list of customers and can develop a profile based on the characteristics of that population – either by analyzing the segment directly or by examining the most influential variables used in the P2B model.

The model reveals a ‘tipping point’ in telemarketing outreach. The campaign variable shows that as the number of contacts increases (red dots moving left in the SHAP plot), the propensity to buy drops. This suggests the bank is currently over-indexing on low-potential customers, essentially ‘inundating’ them—while missing the opportunity to focus that energy on high-potential segments that require lower frequency to convert.

Scatterplot showing customers by number of telemarketing contacts and propensity to purchase with size = account balance.

The Negative Correlation: In the summary plot, the high values for campaign/ telemarketing calls (dark red dots) shift to the left of the center line. This indicates that a high number of contacts during a single campaign actually decreases the probability of a purchase.

Further examination (and data) is required here to determine whether messaging, media mix, brand awareness or other factors also come into play during execution, but this is definitely a red flag since typically effective reach is around ~3X+.

SHAP Insights:

SHAP (SHapley Additive exPlanations) visualization showing how variables change the outcome (positive or negative) and bar chart of feature importance.

Based on the SHAP (SHapley Additive exPlanations) visualizations provided, we can determine exactly which “levers” drive purchase propensity. This is the “explainability” phase that translates a black-box model like XGBoost into actionable business insights.

  • The Bar Chart (Left) – Importance. It shows variables prioritized by the model (e.g., “Cellular contact is the most important piece of information”).  One caveat here: this is an older dataset from the ML Library and for illustration only; interpret with caution!
  • The Summary Plot (Right) – Direction. Shows how those variables change the outcome (e.g., “Being contacted via cell phone increases propensity, while a housing loan decreases it”).

Summary of Feature Influence

The model shows that a mix of communication channels, past behavior, and economic stability are the primary drivers of a “Yes” prediction.

  1. Primary Driver: Communication Method (contact_cellular). The most influential variable and  the strongest predictor of a purchase.
  2. The “Momentum” Effect (poutcome_success). Success in previous marketing campaigns is a powerful indicator of future success. This validates my previous blog’s assertion that RFM (Recency/Frequency/Monetary Value) are highly influential in P2B models. 
  3. Financial Stability (housing no and balance). Customers without housing loans (housing_no) show a higher propensity to purchase. Further, higher bank balance levels correlate positively with conversion.
  4. Timing and Outreach (day_of_week, month_jun, campaign). The specific timing of the outreach (months like June or March) influences the model, though to a lesser degree than the contact method.  So, a telemarketing group and marketing programs should be adjusted for seasonality.

Conclusion

I’ll be returning to propensity-to-buy in future articles, since like RFM analysis this technique is foundational to successful quantitative marketing.  Both techniques trace their origins to the 1980s as statistical tools for direct mailers and have evolved over time to become the foundational “workhorse” of modern digital marketing.


Citations

Miller, Thomas W. Marketing Data Science: Modeling Techniques in Predictive Analytics with R and Python. FT Press, 2015.

Moro, S., Rita, P., & Cortez, P. (2014). Bank Marketing [Dataset]. UCI Machine Learning Repository. https://doi.org/10.24432/C5K306.

Ratner, Bruce. Statistical and Machine-Learning Data Mining: Techniques for Better Predictive Modeling and Analysis of Big Data. 3rd ed., CRC Press, 2017.


Technical Keywords & Methodology Index

Methodology & Strategy: Propensity-to-Buy (P2B) Modeling, Binary Classification, Lead Scoring/Pipeline Conversion Probability, Campaign Frequency Optimization (Tipping Point Analysis), Model Explainability (XAI).

Statistical Concepts: eXtreme Gradient Boosting (XGBoost), SHAP (SHapley Additive exPlanations), Feature Importance Analysis, Grid Search (Hyperparameter Optimization), ROC AUC (Receiver Operating Characteristic).

Data Engineering & Analytics: Imbalanced Dataset Handling, Feature Engineering (Firmographics vs. Demographics), Telemetry/Behavioral Data Integration, Seasonal Adjustment Modeling, Lead Velocity Metrics.

Revenue Operations (RevOps) Logic: Demand Generation Strategy, Churn Mitigation, CLV Integration, Account Prioritization, Resource Allocation Efficiency.

Python Libraries & Documentation

For data scientists and revenue operations engineers building a propensity framework, the following stack bridges the gap between predictive accuracy (XGBoost) and executive-level transparency (SHAP).

LibraryRole in PipelineStrategic Purpose
XGBoostPredictive EngineExecutes the ensemble learning to classify high-propensity vs. low-propensity targets.
SHAPExplainability (XAI)Decodes the “black box” of gradient boosting to identify the specific variables—and their directions—driving conversion (e.g., the frequency tipping point).
Scikit-LearnTuning & ValidationManages Grid Search (GridSearchCV) to optimize parameters and calculates ROC AUC for baseline model performance.
Pandas / NumPyData WranglingManages the ingestion and standardization of the UCI Bank Marketing Dataset for model training.


Posted in

2 responses to “The Demand Generation Workhorse”

  1. […] Part 2, we will dive deeper into Propensity-to-Buy Modeling. We will explore how to codify these insights into a functional AI Agent that can predict velocity […]

  2. […] propensity-to-buy (P2B) model scores the likelihood that a lead, prospect, or customer will convert, respond, or churn. It is […]

Leave a Reply

Discover more from The Marketing Science Signal

Subscribe now to keep reading and get access to the full archive.

Continue reading