Back to News
Marketing Operations
Demand Generation
Revenue Operations

Lead Scoring That Reflects Real Buyer Readiness

Jonathan Martins
June 2, 2026
14 min read
TL;DR

Build a B2B lead scoring model that accurately predicts pipeline readiness — combining firmographic fit, behavioral intent signals, and machine learning to prioritize the right leads.

Lead scoring is one of the most widely deployed and most frequently miscalibrated systems in B2B marketing. The concept is straightforward: assign numerical values to leads based on their characteristics and behavior, rank them by score, and focus sales energy on the highest-ranked leads. In practice, most lead scoring models produce scores that correlate loosely with MQL volume but poorly with actual pipeline conversion — which is the outcome they were built to predict.

The gap between a lead scoring model that looks right and one that actually works comes down to a single question: what signals, in your specific data, actually predict closed-won? Not what signals predict engagement. Not what signals predict MQL volume. What signals predict that a lead becomes an opportunity, advances through the pipeline, and closes?

Answering this question rigorously — rather than building a model on intuition, industry benchmarks, or the defaults in your MAP — is the difference between a scoring system that prioritizes your best leads and one that sends your reps to the wrong conversations.

The Two Dimensions of an Effective Scoring Model

Every functional B2B lead scoring model has two distinct dimensions that it combines: fit and intent. Treating these as a single composite score, or focusing exclusively on one, is the most common cause of scoring model failure.

Fit scoring measures how closely a lead's firmographic and demographic profile matches your ideal customer profile. The variables typically include company size (revenue or headcount), industry or vertical, geography, technology stack, and role seniority. A lead who works at a 300-person SaaS company in your target vertical, holds a Director-level role, and uses the technology stack that integrates with your product is a high-fit lead. A lead who is a student filling out a form for research is a low-fit lead.

Fit scoring is relatively stable — a company's size and industry do not change overnight, and a person's role changes infrequently. This makes fit a good baseline filter but a poor predictor of timing. A high-fit company may be a perfect eventual customer but have no active buying intent for months or years.

Intent scoring measures what a lead is actively doing — the behaviors that signal active evaluation rather than passive interest. High-intent signals include: visiting pricing pages, viewing competitor comparison content, reading implementation or integration documentation, engaging with customer case studies in the same industry, attending a product-focused webinar, using a self-serve trial or demo environment, or explicitly requesting a sales conversation. Low-intent signals include: downloading a top-of-funnel guide, opening an email newsletter, following the company on social media.

Intent scoring is highly time-sensitive — a lead who visited the pricing page six months ago and has not engaged since is very different from a lead who visited it yesterday. Intent scores should decay over time, reducing a lead's score if they have been inactive for 30, 60, or 90 days.

Why Most Models Get Calibrated Wrong

The typical scoring model is built during a MAP implementation, often by someone working from a template or a vendor recommendation. Points are assigned to signals based on intuition ("pricing page visits seem important, let's give them 20 points") rather than analysis of which signals have actually predicted conversion in historical data. The model goes live, generates MQL volume, and is assumed to be working until someone runs a cohort analysis and discovers that high-scored MQLs are converting to opportunities at roughly the same rate as mid-scored ones.

Machine learning analytics gradient infographic lead scoring data
Fit scoring uses stable firmographic data to establish a baseline; intent scoring uses time-sensitive behavioral signals to identify accounts in an active buying window.

Specific calibration errors show up repeatedly across B2B organizations:

Over-weighting content downloads: Downloading a white paper or e-book is easy, free, and frequently done by students, job seekers, and researchers with no buying intent. Many models assign 10-20 points to any content download, which inflates scores for leads with no real purchasing interest. In most B2B datasets, content downloads are weak predictors of closed-won compared to bottom-of-funnel behaviors.

Under-weighting negative signals: A lead who visits the careers page is probably evaluating you as an employer, not a vendor. A lead using a personal email address (gmail, yahoo) at a company that uses a corporate domain probably gave a fake email. A lead from a country you do not serve is unlikely to convert. These negative signals should reduce scores, but most models simply ignore them.

No score decay: A lead who was highly engaged six months ago and has since gone completely silent is not the same as a lead who was highly engaged last week. Without time-decay applied to behavioral signals, old engagement continues to inflate scores long past its predictive relevance.

Demographic over-fitting: Some models weight title or seniority so heavily that they consistently elevate VP and C-suite leads above all others regardless of behavioral intent. In reality, in many B2B buying processes the practitioner who will use the product is often the most active researcher and the most reliable intent signal — ignoring them in favor of executives who have shown no engagement wastes sales resources on unresponsive contacts.

Building a Calibrated Model from Your Own Data

The foundation of an accurate scoring model is a closed-won analysis: looking back at the leads who became customers and identifying which signals most consistently appeared in their history before the sales conversation. This analysis typically reveals a handful of signals that are genuinely predictive and a larger group that seemed important but have little correlation with outcomes.

The process has four steps:

Step 1 — Pull your closed-won history. Export the last 12-24 months of closed-won opportunities and the contact records associated with them. For each contact, pull their engagement history before the first sales contact: pages visited, content consumed, email interactions, form fills, event attendance.

Step 2 — Identify the common signals. Across your closed-won population, which engagement behaviors appeared most frequently in the 30-60 days before the sales conversation? Which pages were visited? Which content was consumed? Which actions preceded an inbound demo request? These are your highest-weight intent signals.

Step 3 — Validate with your lost and disqualified population. Check whether the signals you identified also appear frequently in leads that did not convert. If a signal appears equally in your closed-won and closed-lost populations, it is not predictive — it is just common. True predictive signals appear disproportionately in closed-won compared to closed-lost.

Step 4 — Assign weights proportionally. Weight signals based on their predictive strength in your data. Signals that appeared in 80% of closed-won but only 20% of closed-lost deserve high weight. Signals that appeared in both populations equally deserve low or zero weight.

Machine Learning: When It Helps and When It Does Not

Machine learning-based lead scoring, offered by platforms like Salesforce Einstein, Marketo Predictive Content, and standalone tools like MadKudu, can outperform manually calibrated models when the data conditions are right. ML models can identify non-obvious signal combinations — the interaction between company size and specific page visit patterns, for example — that manual analysis might miss.

Data analytics graph diagram halftone business intelligence vintage abstract
Score decay prevents old engagement from inflating lead priority — a lead who visited the pricing page six months ago and went silent is not equivalent to one who visited yesterday.

However, ML scoring models require sufficient training data to be reliable. MadKudu's own published guidance suggests a minimum of 500 closed-won opportunities to train a model with meaningful predictive accuracy. Companies below this threshold often find that ML models trained on limited data overfit to noise and perform no better than a well-calibrated rule-based model. For those teams, a rigorous manual model built on closed-won analysis is more reliable than an ML model built on insufficient history.

For companies that do have the data volume, the combination approach — a manually calibrated fit scoring layer combined with an ML-based intent scoring layer — tends to outperform either approach alone. Fit scores are transparent, stable, and easy to audit; intent scores benefit from the pattern recognition that ML provides at scale.

Maintaining the Model: The Quarterly Review

A lead scoring model built once and never revisited will degrade as buyer behavior, product positioning, and market dynamics evolve. Teams that recalibrate scoring models quarterly — comparing current MQL score distributions with closed-won outcomes — consistently outperform teams that treat scoring as a one-time implementation task.

The quarterly review should answer four questions: Are high-scored MQLs converting to opportunities at a significantly higher rate than mid-scored ones? Which signals have the highest correlation with closed-won in the most recent quarter's data? Are there new behaviors (new content types, new pages, new product features) that should be added to the model? Are any signals that were previously predictive no longer performing?

Integrating Third-Party Intent Data Into Your Scoring Model

First-party behavioral data — the signals generated by a lead's direct interactions with your website, emails, and product — is the most reliable input for lead scoring because it reflects engagement with your brand specifically. But first-party data has a structural limitation: it only captures what happens within your owned channels. A buyer who is actively researching your category on G2, reading competitor reviews, and engaging with industry content on LinkedIn generates no first-party signal until they visit your site or fill out a form.

Machine learning AI concept abstract technology business illustration
ML-based scoring models require sufficient training data — typically 500+ closed-won opportunities — before they outperform a rigorously calibrated manual model.

Third-party intent data fills this gap by tracking research behavior across the broader web — identifying companies that are actively searching for solutions in your category before they have made direct contact with you. Providers like Bombora, TechTarget, and G2 aggregate content consumption, search behavior, and review activity across large publisher networks to identify accounts showing elevated interest in specific topics.

Integrating third-party intent data into lead scoring requires care. Intent data signals that a company is researching a category — it does not indicate which vendor they are leaning toward, how far along they are in their evaluation, or whether the research is driven by a genuine near-term purchase intent or a background awareness activity. Used as a standalone trigger, intent data produces a high false-positive rate. Used as a score multiplier for accounts that already have high first-party engagement, it can meaningfully improve the model's ability to identify accounts in an active buying window.

The most reliable integration model treats third-party intent as a context signal rather than a primary scoring dimension. When a high-fit account shows elevated third-party intent in your category, the model increases the urgency of acting on existing first-party signals from that account — flagging it for outbound prospecting if no contact has been made, or accelerating the nurturing sequence for contacts already in the system. This approach uses intent data to amplify what the model already knows rather than as a standalone ranking criterion.

Bombora's published data on intent-based scoring integrations suggests that companies using third-party intent data to prioritize outbound prospecting within their ICP see higher connect rates and shorter sales cycles compared to companies that rely on firmographic data alone — because they are reaching accounts during an active evaluation rather than at a random point in the buying cycle. The key constraint is list size: running intent scoring effectively requires a defined target account list and a clear process for acting on elevated intent signals within a short enough window that the research activity is still ongoing.

Communicating Scores Across the Go-to-Market Team

A lead scoring model that marketing trusts but sales ignores is not generating value. Getting cross-functional adoption requires presenting scores in context — not just as a number, but as a ranked list with the specific signals that drove each score visible to the rep at the point of decision. When a rep sees that a lead scored 87 because they visited the pricing page twice, completed a product trial, and work at a company that matches the ICP on all five criteria, they have the context to act on the score confidently. When they see only the score, adoption depends on institutional trust that often is not there.

Building score transparency into CRM views — a dedicated scoring panel on the lead and contact record that shows the score, the top three signals driving it, and the score change over the last 30 days — closes the adoption gap more reliably than training sessions or mandate-driven rollouts. The information is available at the moment the rep needs it, without requiring them to navigate to a separate system or report.

Frequently Asked Questions

How many points should the top score be?
The specific maximum score does not matter — what matters is that the distribution is meaningful. If 80% of your MQLs cluster within a 10-point range at the top of your scale, the model is not differentiating effectively. A good scoring model produces a distribution where there is a clear break between the top tier (the leads worth immediate sales attention) and the rest. Adjust point values until that differentiation is visible in your actual MQL population.

Should we score leads or accounts in ABM programs?
Both. In an account-based motion, account-level signals (multiple contacts from the same company engaging, a target account visiting pricing pages, company-level intent data spiking) are often more predictive than individual lead scores. The best practice is to run both models: individual lead scoring for contact-level routing and sequence assignment, and account scoring that aggregates signals across the entire buying committee for prioritizing account-level sales and marketing investment.

How do we handle leads from very small companies that match our ICP?
Company size is a valid scoring criterion, but it should not be treated as an automatic disqualifier. If your data shows that companies below a headcount threshold rarely convert to revenue, reduce their fit score accordingly. But if you have meaningful closed-won history with smaller companies, ensure the scoring model reflects that — blanket penalties for company size without data support create false negatives that waste legitimate pipeline opportunities.

What is score decay and how should we implement it?
Score decay is the reduction of a lead's behavioral score over time to account for the fact that old engagement is less predictive of current buying intent than recent engagement. A simple decay model reduces the behavioral score by a fixed percentage — typically 25-50% — every 30 or 60 days of inactivity. More sophisticated decay models reduce scores at different rates for different signal types: high-intent signals like pricing page visits decay faster than low-intent signals like email opens.

Should we share our scoring model with the sales team?
Yes, with appropriate context. Sales teams who understand why a lead has a high score — "this lead visited pricing twice, read a case study in your industry, and attended the product webinar" — can use that information to personalize their outreach. Sales teams who see a score with no explanation often distrust or ignore it. Transparency in how scores are calculated, and surfacing the key signals alongside the score in the CRM, is one of the highest-leverage steps for improving sales adoption of marketing's scoring model.

How do we know when our scoring model needs to be rebuilt versus recalibrated?
Recalibration (adjusting point values for existing signals) is appropriate when the model's predictive power has drifted slightly but the underlying signals are still relevant. A rebuild is warranted when the correlation between your top-scored MQLs and closed-won has dropped significantly — typically when less than 30% of your top-tier MQLs are converting to opportunities — or when your ICP, product, or market has changed substantially enough that the signals that predicted conversion 18 months ago no longer reflect current buyer behavior.

Key Takeaways

  • Lead scoring often fails to predict actual pipeline conversion.
  • Fit and intent are the two essential dimensions of scoring models.
  • High-fit leads may lack immediate buying intent.
  • Calibration errors often lead to inflated scores from low-intent signals.

Frequently Asked Questions

What is lead scoring?
Lead scoring assigns numerical values to leads based on their characteristics and behavior. This helps prioritize sales efforts on the highest-ranked leads.
Why do most lead scoring models fail?
Many models are built on intuition rather than data analysis. They often correlate poorly with actual conversion rates.
What is the difference between fit and intent scoring?
Fit scoring assesses how closely a lead matches your ideal customer profile. Intent scoring measures active behaviors indicating buying interest.
How should intent scores be managed over time?
Intent scores should decay if a lead remains inactive for 30, 60, or 90 days. This reflects the lead's current engagement level.

See Where Your Business Stands in Search

Get a free site audit. We identify what is holding you back and what to fix first.

Ready to Transform Your Marketing Operations?

Join mid-market teams transforming their marketing operations with RankWorks AI. Get unified workflows, predictable execution, and measurable growth.

4.9/5 Rating
Google Certified
Enterprise Ready