How to Score Buying Intent Without Black-Box AI: A Transparent Scoring Framework

A scoring model you can't explain to a rep in one sentence per point is a black box, even if it's called AI. Here's a rule-based framework you can build, defend, and audit like code, not rent from a vendor.

Anshul
Anshul Bhatia
Founder
July 16, 2026 · 14 min read

A rep pulls up a lead scored 87 in the pipeline review. Someone asks why. The answer is a shrug: "that's what the model gave it." And nobody in the room can trace the number to a single thing the account did. That's the moment you need a way to score buying intent without a black box, because a number nobody can explain isn't intelligence. It's a liability wearing a dashboard.

We build these scoring systems for clients as part of larger outbound builds. We don't sell scoring software, so there's no vendor angle in what follows, just the framework we'd hand you if you hired us to build one.

The actual definition of a black box isn't the marketing one. A score is a black box the moment nobody in the room can trace it to a specific, verifiable signal. That covers plenty of "AI-powered" scoring tools. But it also covers a spreadsheet nobody has touched since Q1 that still has three people's names buried in the formula. The fix isn't a more explainable vendor. It's a scoring function you own: rule-based, weighted on purpose, simple enough to explain to a rep in one sentence per point.

This is that framework: a fit gate, signal weights tied to buyer cost, decay logic, and a routing threshold you can re-derive from your own closed-won data. And it covers where AI belongs in the pipeline, which isn't as the scoring mechanism itself.

Why black-box intent scores stop getting used

The failure mode isn't abstract mistrust. It's operational, and it shows up in three specific places.

First, reps stop acting on the number. A score a rep can't defend upward to their manager or downward to a skeptical prospect gets quietly ignored. But not rejected outright, just routed around. The rep builds their own mental model of who's actually ready and treats the official score as background noise. Within a few months the team has an informal scoring system living in Slack threads, and the official one is decoration on the CRM record.

Second, marketing can't debug it. When conversion on "high-intent" leads drifts down, someone needs to answer why. With a rule-based rubric, that's a research question: which signal category degraded, and did the weight need updating. But with a black-box model, it's not answerable at all. You can't open the model and ask which input mattered most this month versus last, not without a data science team and a support ticket to the vendor.

Third, and this is the one people skip past: a vendor's model was trained on somebody else's closed-won data. Even a well-built AI scoring product is learning patterns from its training customers, not yours. And your buyer's behavior, your sales cycle, your definition of sales-ready don't automatically transfer. A model that's accurate on the vendor's aggregate dataset can still be wrong for your specific motion, and you have no way to check, because you can't see the training data or the weights.

None of this is about trust as a vague feeling. It's about whether a person in your building can look at the number and do something useful with it: defend it, debug it, or override it with a reason. So a black box fails at all three. This is the same failure mode that shows up anywhere teams buy an outcome instead of building the system behind it, which is most of what GTM engineering actually looks like in practice done right instead of skipped.

The two things any intent score is actually made of

Most "intent scores" quietly merge two different questions into one number, and that's a problem before AI even enters the picture.

But this is the part that causes real damage. When a vendor blends the two into a single number, a bad-fit account with a lot of activity can outscore a great-fit account that's quiet. A 50-person agency binge-reading your blog can out-rank an enterprise account where two stakeholders just quietly visited your pricing page.

Rule-based intent scoring keeps fit and behavior separate on purpose. Fit decides whether an account is even in the game. Behavior decides how urgently, within that game, they need attention. So one number does one job.

A transparent scoring framework you can build and defend

Five decisions, in order. Skip one and the rest of the rubric is decoration.

Step 1: Define your fit gate, yes or no, not points

Most scoring models blend fit and intent into one weighted number, 50/50 or whatever split feels reasonable. That's the mistake. A points-blend lets fit and intent quietly cancel each other out, so a bad-fit account with high activity slips through wearing a good score.

Run fit as a gate instead: a yes-or-no check that happens before any point math starts. Define the two or three firmographic facts that actually predict a closed-won deal in your business, not the ten you'd like to be true. If the account fails the gate, it doesn't get scored at all. And it doesn't matter how many demo requests it racks up after that.

This is the single biggest structural difference between a rubric that survives contact with real pipeline data and one that quietly rots. A gate is binary and auditable. But a blended weight is neither.

Step 2: Weight signals by commitment, not by volume

Once an account clears the gate, weight what it does by what the behavior costs the buyer, not by how easy the signal is to collect.

A demo request costs a real calendar slot and a stated reason for wanting one. A pricing page revisit means someone's doing math, not browsing. A single blog read, by contrast, costs the visitor nothing. And it can happen from a paid social click they don't even remember making. Treating those as remotely comparable is where most scoring rubrics quietly go wrong, and it's a flaw a flat category-bucket model can't fix, because it never asks what the action cost the buyer in the first place.

A starting rubric, built on that logic, looks like the table below. Tune the point values against your own funnel; the ordering is the part that should survive contact with your data.

Signal weighting by buyer commitment
SignalPointsWhy
Demo or consultation request40Costs a real calendar slot and a stated reason for asking
Pricing page revisit (after an earlier visit)25Means the buyer is doing math, not browsing
Competitor or alternatives page read20Signals active evaluation against other options, not idle research
Webinar or live session attended15Real time spent, but often still in the research phase
Repeat blog reads (same account, multiple articles)10A pattern worth noticing, still cheap for the visitor to fake
Review-site visit (G2, Capterra, and similar)8Passive research signal with no commitment behind it
Single blog read3Costs the visitor nothing, easy to trigger by accident

The enrichment layer that pulls this signal data together, wherever it lives (a Clay table, a CRM field, a reverse-IP tool), matters less than whether the weighting logic sitting on top of it can be inspected. So if you're still deciding what fills that pipe, see our breakdown of the enrichment stack that feeds a rubric like this.

Step 3: Apply decay asymmetry, not a flat decay rate

Most published scoring guides apply one decay rate to every signal. Understory's model, for instance, applies a flat weekly decay across its entire point system, which means a demo request loses value at the same rate as a blog read. That's backwards. But a demo request signals a real decision underway; it should hold its value across a long window. A blog read signals mild curiosity; it should decay fast, because curiosity from a while back tells you almost nothing about intent today.

You don't need a universal decay percentage. So you need a relative ordering instead: rank your signals from Step 2 by how slowly they should lose value, and make sure high-commitment signals decay meaningfully slower than low-commitment ones. If your rubric applies the same decay curve to every row in that table, you've rebuilt the blended-score problem one layer down, just with a fancier name.

Step 4: Set a routing threshold you can re-derive, not just declare

A threshold picked in a meeting because "70 felt right" isn't a threshold. It's a guess with a number attached. Validate it instead: pull your closed-won and closed-lost accounts, run each one's last known score, and look for the point where the two distributions actually separate.

Where the two groups start to diverge is your real threshold, and it will almost never land on a round number. That's fine. Round numbers are for slide decks.

Once an account clears that threshold, it needs somewhere to go. That's also the trigger point for a research-first outreach motion (a piece for another day), where the account has already earned attention instead of landing on a cold list. Practically, that's the handoff to the infrastructure that sends: the sequence and the sender built for an account that already cleared a bar. And revisit the threshold every time you update the weights above it, not on a fixed schedule you'll forget to keep.

Step 5: Version and audit it like code

This is the part that separates an owned system from a marketing-automation setting nobody remembers editing. Treat the scoring logic as a function with a change log, not a checkbox buried three tabs deep in your MAP.

// scoring_v4, changed 2026-06-02, Anshul
// Dropped repeat-blog-read weight from 10 to 6.
// Q2 review showed it was over-scoring a cluster of
// paid-social retargeting clicks that never came back.

function score(account):
    if not fit_gate(account):
        return "NO SCORE: fails fit gate"
    points = 0
    for signal in account.signals:
        points += weight(signal.type) * decay(signal.type, signal.age)
    return points

Every change gets a reason attached, in writing, next to the change itself. Not because anyone's going to read the changelog for fun, but because months from now, when conversion on "high-intent" leads drifts and someone asks why, the answer is a two-minute look back instead of a guessing game. That's the whole argument in one sentence: inspectable logic beats agent-count or model sophistication, every time someone in the business needs to trust the number. And if you want the exact Clay build behind a rubric like this, we walk through the implementation separately; this piece is the argument and the portable version of the rubric, not the formula syntax.

What this looks like running

Take a hypothetical account moving through this rubric. Call it a 60-person logistics-software company, picked because that description is deliberately unmemorable. Nothing here is a real client. It's a walkthrough.

  1. The fit gate runs first. The account is in the right industry, the right employee-count band, and using a tech stack the rubric flags as compatible. It passes. If it hadn't, none of the next steps would matter. No amount of activity earns a score once the gate fails.
  2. A marketing contact at the account reads a blog post. Three points, decaying fast.
  3. And a while later, someone at the same account (could be the same person, could be a different one) revisits the pricing page. Twenty-five points, decaying slowly. The blog-read points from step two have already faded to almost nothing by now, because low-commitment signals shouldn't linger.
  4. A different contact at the account reads the competitor-comparison page. Twenty points, decaying slowly.
  5. A demo gets requested. Forty points, barely decaying at all. A demo request stays meaningful for a long stretch.
  6. Running total, weighted and decayed: comfortably past the threshold the team validated against its own closed-won data. The account routes to a rep, and the routing note isn't "the model liked it." It's three lines: pricing revisit, competitor read, demo request, in that order.

That last part is the entire point.

A rep opening this record for the first time can read the reasoning in ten seconds, without asking anyone what an 87 means.

Where AI still earns a place in this system

This isn't an anti-AI piece, whatever the framing so far might suggest. And we run Claude Code as the intelligence layer across most of what we build for clients, including systems that feed this exact rubric. OneGTM's 2026 State of GTM Engineering report, a self-selected survey of 228 GTM engineers across 30-plus countries (meaningful, by the authors' own description, but not statistically representative), found that 72% of respondents report direct, measurable revenue impact from their work. An auditable rubric, where you can trace which signal drove which outcome, is what makes that kind of attribution possible in the first place. But a black-box score forecloses exactly this kind of tracing.

The distinction is where AI sits. Enrichment, summarization, signal extraction: solid uses. Have a model read a call transcript and flag "asked about pricing tiers" as a structured signal. Have it pull technographic data off a job posting. Have it summarize a long research thread into three lines a rep can scan before a call. All of that is AI doing work that feeds into the rubric.

What it shouldn't do is replace the rubric. The objection was never to AI in the pipeline. It's to AI as the unaccountable scoring mechanism itself, the step where a number comes out and nobody can trace it back to a signal a human could verify. A model that extracts "mentioned budget authority" from a transcript and hands that fact to your fit gate is doing useful work. A model that reads the same transcript and quietly outputs "score: 74" with no visible reasoning is the exact problem this piece argues against, just wearing a newer badge.

Common mistakes when building this yourself

Blending fit and intent into one number causes the most damage, because it hides a bad-fit account with high activity behind a score that looks strong. Step 1 above exists specifically to stop that.

Teams also build the rubric once and never touch it again. A scoring model built in January and untouched by summer isn't a system anymore. It's a fossil, and the version-log habit from Step 5 exists to make revisiting cheap instead of painful.

Applying the same decay rate to every signal is the third one, and it's sneaky because it looks reasonable on a spreadsheet. And if a demo request and a blog read lose value at the same speed, you've built one flat category bucket wearing a rubric's clothing.

Scoring in a spreadsheet with no history is the fourth. A formula that changed three times with no record of when or why isn't auditable just because it lives in a cell instead of a vendor's model. But same problem, different wrapper.

And the one that trips up even good rubrics: letting the number replace judgment entirely on buying-committee accounts. Multiple weak signals spread across several people at one account often beat one strong signal from a single person, and a purely additive point system can miss that pattern if nobody's watching for it. The rubric should inform the call. But it shouldn't make it alone.

A rubric you can explain beats a probability you can't, every time someone in the room asks why. That's not a slogan, it's the operational test: can a rep defend this number to a prospect, can marketing debug it when it drifts, can you point to the exact signal that moved it. If the honest answer is no, the fix isn't a better vendor demo. It's building the five pieces above (gate, weights, decay, threshold, version log) and owning what comes out the other end.

Frequently asked questions

What's the difference between a fit score and an intent score?

Fit measures whether an account is the kind that becomes a customer at all, based on firmographic and technographic data that barely changes week to week. Intent, or behavior, measures whether that specific account is acting like a buyer right now: pricing visits, demo requests, content consumption. But blending the two into one number hides bad-fit accounts with high activity behind a score that looks strong. Keep them separate.

Why do reps stop trusting AI-generated lead scores?

Because they can't defend the number. A score a rep can't explain upward to a manager or downward to a skeptical prospect gets quietly ignored, not rejected outright, just routed around. So reps build an informal scoring system of their own instead, usually living in Slack threads, while the official score becomes decoration on the CRM record nobody checks before a call.

Does a transparent scoring rubric mean you can't use AI at all?

No. AI is useful for enrichment, signal extraction, and summarization work that feeds into the rubric: reading a call transcript for a budget mention, pulling technographic data, summarizing research for a rep. The objection is narrower than "no AI." But it's to AI replacing the scoring mechanism itself, where a number comes out with no visible reasoning a human can check against the underlying signal.

How often should you update signal weights and the routing threshold?

Whenever you have enough new closed-won and closed-lost data to check whether the current weights still separate the two groups, not on a fixed calendar you'll forget to follow. So treat each change like a code change: write down what moved, and why, next to the change itself. That log is what turns a future conversion drift into a research question instead of a guess.

What counts as a black-box scoring model, exactly?

Any score where nobody in the room can trace the number back to a specific, verifiable signal the account took. That includes plenty of AI-powered vendor tools, and it includes an old spreadsheet formula three people have edited with no record of why. The label on the tool doesn't matter. Whether a human can audit the logic behind the number does.

Supporting

  1. OneGTM: The 2026 State of GTM Engineering
  2. Understory: Mastering B2B Intent Signals to Accelerate SaaS Pipeline Growth
Written by
Anshul

Anshul Bhatia

Founder
IIT Kharagpur. Builds GTM systems for B2B SaaS.

Anshul builds the outbound systems behind Lead Line Partners. Clay workflows, AI enrichment, and research-first sequencing for teams that want more with less.

More posts
GTM EngineeringGuide · 13 min read

The GTM Engineer Job Description: What to Actually Put in the Req (With a Real Scorecard)

Most GTM engineer job descriptions are tool lists wearing a job title. Here's a req that scopes to company stage, cites a real two-source comp range, and ships the interview scorecard nobody else in the field publishes.

By Anshul Bhatia
GTM EngineeringGuide · 19 min read

GTM Engineering FAQ: 20 Questions Founders and RevOps Leaders Actually Ask

Straight, sourced answers to the questions founders and RevOps leaders actually ask about GTM engineering: cost, hiring stage, tools, RevOps turf, and how results get measured.

By Anshul Bhatia
GTM EngineeringComparison · 13 min read

GTM Engineering vs Marketing Ops vs RevOps: Where Each Role Starts and Stops

Three job titles, one org chart, and no agreement on where one role ends and the next begins. Here's the boundary line for each, and what actually breaks when one is missing.

By Anshul Bhatia

Ready to engineer your GTM motion?

Tell us how your motion runs today. We'll show you what we'd engineer.

Contact us