Why Propensity-to-Buy Prospecting Matters
- Precision Resource Allocation: It identifies the prospects with the highest likelihood to purchase, allowing marketing and sales teams to focus their budget and energy where it will yield the highest return.
- Relevant Communication: By profiling high-potential prospects with high P2B scores marketers can craft relevant communications to overcome the anonymity associated with a “cold start”.
- Optimized Marketing Frequency: It prevents “over-indexing” or inundating low-potential prospects with excessive outreach, identifying the “tipping point” where increased contact frequency decreases the probability of a sale.
- Foundation for Advanced Metrics: Propensity-to-Buy (P2B) generates the critical probability scores required to calculate more complex financial models, such as Customer Lifetime Value (CLV), effectively bridging the gap between marketing activity and long-term profitability.
- Strategic Profiling: It isolates high-propensity populations to uncover the influential variables (e.g., job title, industry, demographics, or financial stability) that define the “ideal” customer profile for future targeting.
| GTM Bottom Line 1. The gap is real, and it’s bigger than it looks. A blended model says relationship history is worth almost nothing (+0.0125 AUC). Split the file in two and the true gap is 0.113 AUC — nine times larger. 2. Prospects still convert — just not for the reasons you’d expect. Without purchase history, the model leans on contact timing and channel, not demographics. The top 30% of the list still captures 64% of every conversion. 3. The metric you lead with changes the story. ROC-AUC undersells this problem by a factor of three; Average Precision is the number that actually predicts whether a campaign pays for itself. 4. What to do with it: score prospects separately from customers, and treat the accuracy ceiling as the dollar-figure case for buying third-party enrichment — not as a model that needs fixing. |

The pay-off is a dashboard that shows the accuracy and impact of both customer and prospect models to use as decision support for marketing data scientists.
A propensity-to-buy (P2B) model scores the likelihood that a lead, prospect, or customer will convert, respond, or churn. It is typically a statistical or machine-learning binary classifier that uses everything you know about a lead, prospect, or customer to make that prediction. Every propensity model I have built in twenty+ years of doing this has started with the assumption that the people or businesses we are scoring have a past. However, when we look at the data, we often find that the most influential information needed to predict purchase or response lies with customers and is largely absent from prospects.
Recency, frequency, monetary value. I wrote about the power of RFM Segmentation. Then we have marketing responsiveness (prior campaign response). Product ownership. Customer support tickets. The entire apparatus of behavioral scoring rests on the premise that this person, buying group, company or account has done something before, and that what they did predicts what they will do next.
Then marketing asks you to score the prospect file for a new customer acquisition program, and the assumption collapses and uncertainty increases.
This article asks a specific question: should never-contacted prospects be scored with their own model, separate from customers with contact, marketing and purchasing history? Instead of running one model across the entire database and hoping the statistics sort it out, we can separate people with history from those without it and measure exactly what that history is worth.
The answer can make a significant difference in GTM effectiveness, and the reason offers a lesson in metric selection that applies well beyond this dataset. In my experience, past purchase recency, frequency, and monetary value (RFM) are such powerful predictors that high-potential prospects critical for new customer acquisition will remain buried in the lower deciles of a blended prospect + customer scored list. They will be overlooked, and new customer acquisition will take a backseat to existing-customer installed-base cross-sell/up-sell marketing.
A Note on Data Integrity and Ethics
To maintain the highest ethical standards and ensure zero overlap with proprietary information from past or current employers, the analysis in this series is conducted on a publicly available open-source dataset from the University of California at Irvine’s Machine Learning Archive.
In this article I use the UCI Bank Marketing dataset — 45,211 records from a Portuguese retail bank’s term-deposit telemarketing campaigns, published by Moro, Rita, and Cortez (2014). Readers of my past articles will know that I have found this dataset excellent for a wide range of marketing data science use. Generally speaking, I always use public or synthetic data and have used many of the UCI datasets because they are (a) extremely popular for use by academic institutions and (b) a good fit for a wide range of modeling.
The Dataset and Target Variable
The target is binary: did the customer subscribe to a term deposit (y = yes) or not. Across the full file the base rate is 11.7%. The dataset was trained on the full file. However, 81% of that file had never been contacted before. So, for most of the population the three key features: promotion days, previous contact and outcome were all null or zero. So we have zero variance and zero information from the most powerful features for predicting purchase in my experience. Here is a sample of the data to illustrate:

Deliberate exclusion: duration. The dataset includes the length of the call that produced the outcome. A long call is a very strong signal that someone is about to say yes — and it is completely unusable, because call length does not exist until after the call happens, and the entire purpose of this model is to decide who to call before calling them. This is a textbook example of temporal/look-ahead data leakage (also called “future information leakage”) where the data is not available at the moment the prediction must be made.
This is the same error I removed from the pipeline model in Part 2 From Signal to Score: Building a Propensity Model to Predict Closed-Won, where Forecast_Category was updated by reps as deals progressed and would have handed the model an answer it could not have at prediction time. Different dataset, identical failure mode. Leaving duration in produces a model that looks considerably more accurate. However, this is not because it’s a better predictor, but because duration is a proxy for the outcome that doesn’t exist until after the decision has already been made.
The Population Split
Three fields in this dataset encode prior relationship: pdays (days since last contact), previous (number of prior contacts), and poutcome (outcome of the last campaign). A value of [pdays = -1] flags a record that has never been contacted before
| Population | Records | Share | Conversion |
| Never contacted (prospects) | 36,954 | 81.7% | 9.2% |
| Previously contacted | 8,257 | 18.3% | 23.1% |

Two things stand out. The first is the conversion gap: previously contacted customers convert roughly two and a half times the rate of cold prospects, which is unsurprising and consistent with what I have seen during my real-world work and is the reason relationship marketing exists.
The second is more consequential for modeling. Within the prospect population, all three relationship fields are constants; pdays are -1 for every one of the 36,954 records, previous is 0, poutcome is “unknown.” Zero variance, zero information because a column holding the same value for every row is a constant and without variability it cannot help a model distinguish a likely buyer from an unlikely one. It is the modeling equivalent of segmenting a list by a field that is blank for everyone in it.
And yet a single model trained on the full file has those columns available. In this case, it learns what they mean from the 18% who have history, then applies that learning across a population where 82% of records carry no signal in them at all.

index on 200 records is not the same as a high index on 8,000 because the error is much higher.
The Ablation, and the Trap Inside the Feature Exclusion Test
Why this matters: the standard way of measuring a feature’s value can make an important signal look nearly worthless — here’s the test that got it wrong, and the fix.
The obvious way to measure what relationship history is worth is an ablation test: train the identical model on the identical rows twice, once with the relationship features and once without, and compare.
| Feature set | ROC-AUC | Average Precision |
| Full population, with history | 0.8075 | 0.4691 |
| Full population, without history | 0.7950 | 0.3947 |
An AUC difference of +0.0125. Read at face value, this says relationship history barely matters, it is not influential in prediction.
Methodological Note: AUC-ROC is the area under the Receiver Operating Characteristic curve measuring the ability of the model to rank a random positive case higher than a random negative case.
That conclusion is wrong, and the reason why is the methodological core of this article.
The relationship features are constant for 82% of the rows in this test. They can only contribute predictive value to the 18% minority, and whatever they contribute there gets averaged across the whole file. The aggregate ablation measures the effect of a coupon on a mailing list where four out of five recipients never received one.
Aggregate ablation understates the value of a feature that is only present for a subpopulation. This is not specific to this dataset. Any time you ablate a feature that is structurally missing for a large share of your records — third-party enrichment that only matched on some accounts, product telemetry that only exists for one SKU, intent data that only fires on a fraction of the file — the aggregate test will tell you the feature is worth less than it is.
The fix is to stop blending the populations.
Two Models, Two Populations
| Model | Records | Base rate | ROC-AUC | Avg Precision |
| Prospect (never contacted) | 36,954 | 9.2% | 0.7464 | 0.3114 |
| Contacted (has history) | 8,257 | 23.1% | 0.8592 | 0.6676 |
Now the gap is visible: 0.113 ROC-AUC, nine times the diluted ablation figure.
The prospect model is stable. Five-fold stratified cross-validation on the prospect population returns 0.7567 ± 0.0101 — the single hold-out was not a lucky slice.
Methodological Note: in k-fold cross-validation we split the prospect data into x folds (5) , train on all but one fold (4) and test on the hold-out fold (1) , rotate through all five combinations and look at how much the score moves across the five runs.

Why Average Precision Is the Number That Matters Here
Why this matters: two ways to grade the same model gave answers three times apart — and only one of them tells you if the campaign pays for itself.
Look again at the two columns.
The ROC-AUC gap between the models is 0.113. The Average Precision gap is 0.356 — more than three times larger, describing the same two models on the same data.

Both metrics are correct. They are answering different questions, however only one of them is the question a campaign manager is asking.
ROC-AUC asks: if I hand the model two prospects, one who will convert and one who will not, how often does it correctly rank the converter higher? It is a pure ranking question, and it treats false positives as a rate against the negative class. When the negative class is 91% of the file, that denominator is enormous, and the false-positive rate stays flattering even when the model is generating a great deal of waste.
Average Precision asks: of the people the model tells me to call, how many convert? That is the question that determines whether a campaign pays for itself.
Methodological Note: Average Precision (AP) summarizes a classifier’s precision-recall curve into a single number — roughly, the average precision achieved across all recall levels, where recall is the share of all actual positives (e.g., true converters) that the model successfully identifies — and unlike AUC it isn’t diluted by the (often huge) true-negative class, which is why it tells a sharper story on imbalanced data like a low-base-rate prospect file.
This matters because ROC-AUC is the default metric in most propensity work, including my own in Part 2 From Signal to Score: Building a Propensity Model That Predicts Closed Won. On a near-balanced dataset — the 55/45 split in the pipeline model — the two metrics tell similar stories. At a 9% base rate, they diverge sharply, and reporting only AUC would have led me to conclude that relationship history was worth about a third of what it is worth.
Rule I now apply: when the positive class is under roughly 20%, lead with Average Precision and treat ROC-AUC as the secondary number.

The Scores Are a Ranking, Not a Probability
Why this matters: a model can rank leads perfectly and still mislead you about their odds — here’s why your pipeline forecast might be inflated 4x.
This is the section I owe to a misinterpretation in my own earlier work.
In Part 2 I assigned marketing actions using probability thresholds — scores above 0.8 routed as High Priority, above 0.5 as Nurture. That framing is incorrect, and the reason is scale_pos_weight.
scale_pos_weight is the standard XGBoost correction for class imbalance. It tells the model to pay disproportionate attention to the rare positive cases during training, which improves ranking substantially. It also inflates the predicted probabilities, because the model has been trained on an artificially rebalanced view of the world.
On the prospect model:
| Measure | Value |
| Mean predicted probability | 0.4110 |
| Actual conversion rate | 0.0916 |
| Inflation factor | 4.49× |
| Brier score (uncalibrated) | 0.1845 |
| Brier score (isotonic calibrated) | 0.0726 |
A score of 0.41 does not mean a 41% chance of conversion. This is a common misinterpretation when scored lists are put to use by marketing in my experience. It means “this record ranks in a particular position relative to the others.” Both numbers live between 0 and 1, which is precisely why the error is easy to make and hard to notice.
Isotonic calibration fixes it. Training a second model without the class weighting and calibrating it against held-out data brings the mean prediction to 0.0926 against an actual rate of 0.0916 and cuts the Brier score from 0.1845 to 0.0726 — while ROC-AUC moves only from 0.7464 to 0.7497. The ranking is essentially unchanged. Only the scale is corrected.

The operational distinction:
- Ranking, decile targeting, call-list construction — raw scores are fine. Order is all that matters.
- Revenue projection, expected-value math, probability-based thresholds — use the calibrated column, or your forecast will be inflated by roughly the same factor as your scores.
If you have ever built a propensity-weighted pipeline forecast on top of an imbalance-corrected model, this is worth checking before your next planning cycle.

What the Model Delivers Operationally
AUC is a modeling metric. Decile lift is what you hand to a campaign manager.
| Decile | Records | Converters | Rate | Lift | Cumulative capture |
| 1 | 740 | 262 | 35.4% | 3.87× | 38.7% |
| 2 | 739 | 97 | 13.1% | 1.43× | 53.0% |
| 3 | 739 | 73 | 9.9% | 1.08× | 63.8% |
| 4 | 739 | 45 | 6.1% | 0.66× | 70.5% |
| 5 | 739 | 46 | 6.2% | 0.68× | 77.3% |
| 6–10 | 3,695 | 154 | 4.2% | 0.45× | 100% |

The top decile converts at 35.4% against a 9.2% base rate and contains 38.7% of every converter in the file. Working the top three deciles — 30% of the list — reaches 63.8% of all available conversions.
That is a real call list, built on a population where the model has no relationship history to work with at all.
On thresholds: the default 0.50 cut is arbitrary and, at a 9% base rate, a poor default. Two defensible alternatives:
- F2 optimization, which weights recall twice as heavily as precision — appropriate when a missed converter costs more than an unnecessary call, which is usually true for outbound phone. F2 peaks at threshold 0.525: precision 0.232, recall 0.573, 1,673 contacts.
- Capacity, which is the real constraint in most organizations. If the team can work 1,000 leads, take the top 1,000 scores and let the threshold fall where it falls. Call-center hours are fixed regardless of what an optimal F-beta says.

What Actually Drives Cold Prospect Conversion
Why this matters: without purchase history to lean on, the model ends up telling you more about when and how to call than who to call — which changes what you do with it.
Here the findings are uncomfortable, and I think it is the most useful thing in this analysis.
With relationship history unavailable, SHAP shows the prospect model leaning on contact month, contact channel, and housing-loan status. It is learning when to call and how to reach people far more than who is worth calling.

contribute less than a marketer would expect.
That is not a model failure. It is an accurate description of the information available. Strip out behavioral history and what remains in a standard CRM record is demographics and campaign operations — and campaign operations turn out to carry more signal than demographics do.
Two implications follow, and they point in opposite directions.
The optimistic one: campaign timing and channel are levers you control. If the model says March and October outperform May by a wide margin, that is directly actionable in a way that “single people convert better than married people” is not.
The cautionary one: a model built mostly on operational levers is measuring your campaign, not your market. Rerun it after a calendar change and the feature importances will move.
Audit the top feature before you trust it
contact = unknown ranks high in the SHAP output. Before treating that as a customer insight, I checked what it actually is:
| Contact channel | Records | Conversion | Index vs. base |
| Cellular | 21,729 | 12.0% | 1.31 |
| Telephone | 2,275 | 11.1% | 1.21 |
| Unknown | 12,950 | 4.0% | 0.44 |

“Unknown” converts at 44% of base rate across nearly 13,000 records. But “unknown” is not a channel. It is an absence of data — most likely an older batch of records where contact method was never logged, or accounts with no phone number on file. This is not uncommon in real-life datasets.
This is a data quality issue. The model is not learning that unreachable people do not convert. It is learning that badly maintained records do not convert, which is true, is predictive, and will evaporate the moment someone cleans up the CRM.
This is the category of finding that survives validation and dies in production. A feature that predicts because of how the record was created, rather than because of who the customer is, has no durability. It belongs in a data-quality remediation ticket, not a targeting strategy.
Recommendations
1. Split the file before you model it. If a meaningful share of your database has no relationship history, a single blended model is quietly applying learned patterns to a population those patterns do not describe. Segment on history availability, then model each population on its own terms.
2. Do not ablate across populations where the feature is structurally missing. The aggregate test will understate the feature’s value by roughly the proportion of records where it is absent. Measure within the population that has it.
3. Lead with Average Precision below a 20% base rate. ROC-AUC will make a mediocre model look respectable when negatives dominate. Report both; interpret AP.
4. Calibrate before any number leaves the model as a probability. If you use an imbalance correction, your scores are ranks. Isotonic calibration costs essentially nothing in ranking performance and makes the scores mean what they appear to mean.
5. Audit your top features for data-quality artefacts. Ask of every high-importance feature: is this predictive because of who the customer is, or because of how the record was created? The second kind does not survive a CRM cleanup.
6. Treat the prospect model’s ceiling as a business case, not a modeling failure. The 0.113 AUC and 0.356 AP gaps quantify exactly what relationship history is worth. That number is the argument for third-party enrichment — intent data, firmographics, technographics — with a figure attached rather than a vendor’s assertion.
What Comes Next
The prospect model answers who is most likely to convert. It does not answer the question a marketer with a fixed budget actually needs: who converts because we contacted them.
Those are different populations. Some prospects would have converted anyway; contacting them spends budget to purchase an outcome already in hand. Others are actively pushed away by outreach. Separating the persuadable from the already-decided requires a randomized control group, which this dataset does not have.
Part 5 will move to a dataset with one.
Academic & Technical Citations
Michael E. Foley (2026). The Cold Start Problem: Why Prospects Need Their Own Propensity Model. The Marketing Science Signal.
@article{foley2026coldstart,
author = {Foley, Michael E.},
title = {The Cold Start Problem: Why Prospects Need Their Own Propensity Model},
journal = {The Marketing Science Signal},
year = {2026},
url = {https://mikesdatamarketing.com/},
keywords = {Cold Start Problem, Prospect Scoring, Propensity Modeling,
Probability Calibration, Average Precision, Class Imbalance,
Ablation Testing, Data Leakage}
}
Further Reading & Technical References
Chen, T., & Guestrin, C. (2016). “XGBoost: A Scalable Tree Boosting System.” Proceedings of the 22nd ACM SIGKDD International Conference.
[Foundational paper for the gradient boosting implementation used throughout this series.]
Moro, S., Rita, P., & Cortez, P. (2014). Bank Marketing [Dataset]. UCI Machine Learning Repository. https://doi.org/10.24432/C5K306
[Source data for this analysis.]
Niculescu-Mizil, A., & Caruana, R. (2005). “Predicting Good Probabilities with Supervised Learning.” Proceedings of the 22nd International Conference on Machine Learning.
[The reference treatment of isotonic and Platt calibration, and why classifier scores are not probabilities by default.]
Saito, T., & Rehmsmeier, M. (2015). “The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets.” PLOS ONE.
[Empirical basis for the Average Precision argument in this article.]
Kaufman, S., Rosset, S., & Perlich, C. (2012). “Leakage in Data Mining: Formulation, Detection, and Avoidance.” ACM Transactions on Knowledge Discovery from Data.
[Formal treatment of the leakage class that includes duration and Forecast_Category.]
Lundberg, S. M., & Lee, S. (2017). “A Unified Approach to Interpreting Model Predictions.” Advances in Neural Information Processing Systems.
[SHAP methodology used for the feature attribution above.]
Peppers, D., & Rogers, M. (2016). Managing Customer Experience and Relationships: A Strategic Framework. Wiley.
Technical Keywords & Methodology Index
- Methodology: Cold Start Problem, Population Segmentation, Ablation Testing, Propensity Modeling, Binary Classification, XGBoost, SHAP Explainability, Decile Lift Analysis.
- Statistical Concepts: ROC-AUC, Average Precision, Gini Coefficient, Kolmogorov-Smirnov Statistic, Brier Score, Log Loss, Stratified Cross-Validation, Class Imbalance (scale_pos_weight), Precision-Recall Tradeoff, F-beta Optimization.
- Calibration: Isotonic Regression, Platt Scaling, Reliability Diagrams, Probability Inflation, Rank-vs-Probability Distinction.
- Feature Engineering & Data Integrity: Data Leakage Prevention, Temporal Feature Validity, Zero-Variance Feature Detection, Data-Quality Artefact Auditing, Structural Missingness.
- Business Intelligence: Prospect Scoring, Call List Prioritization, Capacity-Constrained Targeting, Third-Party Data Enrichment, Campaign Timing Effects.
- MLOps: Assertion-Based Leakage Checks, Hold-Out Scoring Discipline, Model Reproducibility.

















































































































































