AI in GTM

Black-box scoring

Black-box scoring is any lead or intent score that nobody in the room can trace back to a specific, verifiable signal, whether it comes from an opaque AI model or an old spreadsheet formula nobody remembers, making the number impossible to defend, debug, or override.

The real test for a black box has nothing to do with whether AI is involved. A score is a black box the moment nobody can trace it to a specific fact the account actually did. That covers plenty of "AI-powered" scoring tools, but it also covers a spreadsheet formula three people have quietly edited since Q1 with no record of why.

Why it stops getting used

A score nobody can defend gets quietly ignored rather than rejected outright; reps build their own informal read of who's ready and treat the official number as background noise. When conversion on "high-intent" leads drifts, a rule-based rubric gives you a research question, which signal category degraded, while a black-box model gives you nothing to open and inspect. And a vendor's model was trained on someone else's closed-won data, which may simply not describe how your buyers behave.

The fix is a scoring function built to be traced, not a more explainable vendor. Fit runs as a yes-or-no gate before any points get added, so a bad-fit account with high activity can't outscore a good-fit account that's quiet. Signals get weighted by what the action cost the buyer, not by how easy the data was to collect. And the routing threshold gets set by finding where an account's closed-won and closed-lost outcomes actually diverge, not by a number that felt right in a meeting.

In practice

AI still earns a real place in this system, just not as the score itself. Having a model read a call transcript and flag "asked about pricing tiers" as a structured signal is useful work feeding the rubric. A model that reads the same transcript and outputs a bare number with no visible reasoning is the exact problem this whole approach exists to avoid, just wearing a newer label.

What people get wrong

The most common mistake is blending fit and behavior into one weighted number, which lets a bad-fit account with lots of activity outrank a great-fit account that's quiet. Applying one flat decay rate to every signal type is another: a demo request and a single blog read shouldn't lose value at the same speed. And picking a routing threshold because it "felt right" instead of checking where real closed-won and closed-lost accounts actually separate leaves the whole rubric ungrounded.

Related terms
Where we use this
Updated July 25, 2026

Ready to engineer your GTM motion?

Tell us how your motion runs today. We'll show you what we'd engineer.

Contact us