Lead Quality Scoring That Survives Contact With Reality
Most lead scoring models are elaborate ways of restating what the sales team already knew. Building one that changes decisions.
Most lead scoring models fail in the same way. They get built, they produce a number, the sales team ignores the number, and eighteen months later someone proposes building a new one.
The failure is rarely statistical. It is that the score was never wired to a decision.
Start from the decision, not the data
Before any modelling, answer this: what will change based on this score?
Valid answers are concrete. Leads scoring above X get called within five minutes; below X go to the nurture queue. Sources whose median score falls below Y get paused. Bids increase on keywords producing above-median scores.
Invalid answers are the common ones. "It will give the team visibility." "We'll use it for reporting." A score nobody acts on is a dashboard ornament, and building it costs the same as building a useful one.
Score against the outcome you are paid for
The most common modelling error is scoring against an intermediate event because it is easier to observe. Contact rate is easy. Appointment set is easy. Revenue is hard, because it arrives weeks later through a different system.
Score against revenue anyway, or the closest proxy you have to it. A model optimised for contact rate will faithfully find you people who answer the phone and buy nothing.
Three features usually do most of the work
Elaborate models tend to underperform simple ones in this domain, because the signal is concentrated in a handful of variables and the rest is noise that varies by season.
In call-based lead generation, three features carry most of the weight almost everywhere:
- Source, at the granularity of keyword or placement, not channel
- Time-to-contact, which is more predictive than nearly anything about the lead itself
- A single qualification flag — the one hard criterion for your vertical
Start there. Add features only when you can show they change a decision, not just improve a fit statistic.
Recalibrate on a schedule, not when it breaks
Source mix drifts. Seasonality shifts. A model trained on last November's Medicare traffic describes a world that no longer exists by February.
Set a recalibration cadence — quarterly is usually right — and hold to it. Models that get revisited only when someone complains have already been making bad decisions for months by the time anyone notices.
Publish the score to the people it constrains
If a score determines that certain leads go to a nurture queue, the person working that queue should be able to see the score and the top reasons for it.
Two things follow. Sales trusts a model they can interrogate. And they will catch the failure modes you did not — the ones where a score is confidently wrong for a reason no one thought to encode.
The test that matters
Six months after launch, ask one question: what decision is different because this exists?
If nobody can point to a routing rule, a paused source, or a changed bid, the model is not underperforming. It was never in the loop to begin with.