I’ve had a version of the same conversation with a handful of revenue leaders over the last few months. They don’t want to look at another AI tool. When you ask why, the answer comes back as a count. Eleven tools. Fourteen. Somewhere in that range.
The count is real. What it gets blamed for is the interesting part. The complaint always arrives as some version of “we have too much AI,” and I don’t think that’s what any of them mean.
You don’t walk into your house and complain about having too much electricity. You complain about the bill, and about the four appliances nobody has touched since 2019.
Electricity is one input. Appliances are many, and each one either finishes a job or hands it back to you. Your dishwasher finishes. The bread machine that needs three hours of babysitting does not, which is exactly why its in the basement.
Your AI stack is appliances.
How sales tools got here
The team at Airspeed has the cleanest description of the progression I’ve seen. Call recording captured what was said, and somebody still had to go watch it. Conversational intelligence explained what it meant, and the work still landed on the rep. Each generation got closer to the work being finished, then stopped an inch short.
Their word for the third stage is execution, and its the right word. Recording captured. Intelligence explained. Execution finishes the job: it reads the signals, decides what the next move is, and then goes and makes that move inside the systems where work actually happens. Slack, the inbox, the CRM record. No rep in the middle, copying a recommendation out of one window and into another.
That last clause is the hard part, and its worth being specific about why. A system that takes action needs three things a dashboard never needed. Permission to write and not only to read. Memory of the account, so the action it takes lines up with the last six months of the deal rather than with the last call. And somewhere to act into that the team already lives in, because an action delivered inside a tool nobody opens is a notification.
Miss any one of the three and you’re back to producing recommendations, with extra steps.
Let me be the skeptic for a second, because I’ve sat through enough of these evaluations. When a vendor tells you their agents take action, the next question is what happens when the action is wrong. Who sees it, how fast, and how do you undo it. Anything with permission to write to your CRM and email your buyers can be wrong in public. That’s not a reason to stay away from it. It is a reason to ask, and the serious ones have an answer ready before you finish the question.
That inch between explaining and finishing is the whole thing.
Any tool that produces a recommendation and stops has created a to-do for a human. A deal risk score is a to-do. A flagged coaching moment on a call recording is a to-do. A dashboard is a pile of them, sorted by column.
Which means a stack that generates insight faster than your team can act on it is adding work to your team. That is the feeling underneath “we have too much AI.” Unlike the feeling, it can be counted.
Try it for one week. How many recommendations did your stack produce, and how many did somebody act on? The ratio tends to be worse than a leader wants to say out loud. You’re allowed to keep the number to yourself.
Worth saying clearly: the individual tools are usually fine. The call recorder records calls. The intelligence layer is intelligent. Each one does what it said it would do on the demo. The failure shows up at the seams, where output from one tool becomes somebody’s manual input into the next, and nobody owns the handoff because it was never anybody’s job. Eleven good tools with ten unowned seams between them is a worse system than four tools that finish things.
The first three questions in the audit below are how you test an execution claim without taking anybody’s word for it. Which brings me to the exercise.
The audit
Sit down with your stack, one row per tool, and score each one out of 12. It takes about 20 minutes if you already know what you pay for.
Two points for the good answer, zero for the bad one, one point for anything in between.
- Does the output end in information, or in a finished action?
Zero for a summary, a score, or a dashboard. Two for a drafted email, an updated record, a booked next step. The test is whether a human has to do something after the tool is done. If the answer is yes, the tool handed the work back to you and called it insight.
- Who does the last mile?
Zero if the answer is your reps, during selling hours. Two if the answer is nobody, or somebody whose time isn’t your constraint. Every hour a rep spends working a recommendation is an hour they spent away from a buyer, and you’re paying for that hour twice.
- Does it remember?
Zero if every session starts over and somebody has to re-upload the context. Two if it knows the history of your biggest deal without being told. This is Airspeed’s argument and its a fair question to put to everything in the stack, including the general-purpose assistants your team pastes call notes into. A tool with no memory of your deals is a very fast intern on their first day, every day.
- How many days pass between the thing happening and a human knowing about it?
Zero if you find out at pipeline review. Two for same day. If pipeline review is where deal risk gets discovered at your company, you’ve built a postmortem and put it on the calendar every Monday.
- What percentage of licenses got used last week without anybody being told to?
Zero under 40%. Two over 75%. Usage that only happens when you ask for it stops the week you stop asking, and every leader reading this has watched that movie. Pull the actual seat data before you answer this one. Leaders guess high on their own stack more often than not, and the gap between the guess and the report is the most honest thing in this whole exercise.
- What breaks if you turn it off Monday morning?
Zero if nobody can name anything. Two if somebody can name a ritual or a number. This is the most useful question on the list and the one people skip, because a fair number of tools survive on nothing but an auto-renewal and the fact that the champion who bought it left in March.
Reading the score
10 to 12. Keep it, then go find out who set it up. Whatever they did is repeatable and you’re probably not doing it anywhere else.
6 to 9. Worth keeping, with an owner and a date on it. Most tools in this band are one integration away from a much higher score.
Under 6. Kill it, or fold it into something that scored higher. Expect resistance, and expect almost all of it to be about sunk cost.
One note on the kill list. The instinct is to cut the lowest scores and stop there. The better move is to look at which of your 10 to 12 scorers could absorb the job of a 4, because consolidation is usually available and rarely considered. Two tools scoring 5 each, doing adjacent work, are a single tool nobody built yet.
While you’re in the spreadsheet, change the number you report upward. Cost per seat tells you what you spent. Cost per completed action tells you what you got.
Take a tool’s annual cost and divide it by the finished actions it produces in a year. Say a platform runs you $40,000 and it drafts and sends 200 follow-ups a week. That’s roughly 10,000 a year, so $4 per completed action. Now take a different $40,000 tool that finishes nothing and surfaces 200 recommendations a week. Its cost per completed action has no denominator. What it does have is 10,000 to-dos a year, handed to the same team you’re asking to hit a number.
That is the slide your CFO understands, and its a much better conversation than defending a line item called “AI.”
The rule for the next one
The next tool you evaluate has to remove a step that exists today. Make somebody name the step out loud in the first meeting, and write it down.
If nobody can name it, what you’re buying is a report, a login, and a quarterly usage review.
For what its worth, I’m not against any of this. I said on a podcast a while back that the reason some of the newer AI tools are working is that they’re built on buyer behavior, what buyers say on calls and type in emails, rather than sellers marking activities in a CRM. I DO think that’s where this goes, and the tools that get there will be worth every dollar.
The complaint usually survives the audit, by the way. What changes is the sentence. It stops being a feeling about a category and turns into a list of four tools with a name and a date next to each one.
This may not be the right call for your org, but its the one I believe. Twenty minutes and one spreadsheet gets you a keep list, a kill list, and a number finance recognizes. Worth running it yourself before somebody in finance runs it for you.

