The 58% vs 35% Problem: Why Generic AI Fails at Real Sales Conversations

Here’s an uncomfortable question nobody’s asking in revenue enablement meetings right now.

If your reps are practicing cold calls, discovery, and objection handling on ChatGPT, do you know how often it gets it right?

I didn’t, until I went looking. And the answer isn’t good.


What “Multi-Turn” Actually Means

 

Quick bit of translation, because the AI research world loves a bit of jargon.

“Single turn” means one question in, one answer out. You ask it something, it answers, done. That’s the easy test. It’s also the test most AI vendors quote when they talk about accuracy.

“Multi-turn” means an actual conversation. The AI doesn’t have all the information up front. It must ask a follow-up question, take in a partial answer, ask another one, and build up to a correct response over several exchanges.

Have a think about which one of those sounds like a sales call.

Discovery isn’t one question. Objection handling isn’t one rebuttal. A buyer very rarely hands your rep every fact they need in the opening thirty seconds. Sales is a multi-turn sport, always has been. So, the only score that matters here is the multi-turn one.


The Study: 58% Down to 35%

 

Salesforce’s own AI research team ran both tests, back-to-back, on a stack of leading models, GPT-4o, Gemini, Llama, and others, across nineteen distinct CRM tasks pulled from real customer service, sales, and pricing workflows. Over 83,000 synthetic records were validated by working CRM professionals so the scenarios held up to scrutiny.

On the single-turn version, the best agents landed around 58% success.

Not brilliant, but you could squint and call it promising.

Then they ran the same tasks as a proper conversation, the kind where the agent has to ask before it can act. Every model’s score fell. Sharply.

THE NUMBER THAT SHOULD WORRY YOU Leading AI agents: 58% success on single-turn tasks. 35% on multi-turn ones. Same models, same underlying skill. The only thing that changed was whether the conversation ran more than one exchange.

That’s not a rounding error. That’s the difference between a coin flip and a system you’d really trust with a rep’s practice time.


The One Thing AI Is Genuinely Good At

 

In fairness, it wasn’t all bad news, and it’s worth being straight about that.

One skill category stood well clear of the rest. Workflow execution, the rule-based stuff, routing a lead to the right rep, following a fixed policy step by step, hit over 83% success in the single-turn setting. If the task is “follow this exact rule,” AI is properly reliable. It still lost ground once the conversation had more than one exchange, everything does, but it stayed the strongest category by a clear margin.

The trouble is that’s not what a sales conversation is. Nobody closes a deal by following a fixed rule. Objection handling, reading a buyer’s real priorities, and adjusting your pitch mid-call, that’s judgment, not workflow. And judgment is precisely where the numbers fell apart.


Why the Score Collapses When the Conversation Continues

 

Here’s the part I found most telling.

The researchers went back and manually reviewed twenty failed conversations from their best-performing model. In nine of those twenty, the agent had simply failed to ask for information it needed. It didn’t ask the follow-up question. It guessed, or answered with what it had, incomplete as that was.

Only one failure out of twenty came down to the simulated buyer being unhelpful.

Read that again. The agent usually had the reasoning power to get to the right answer. What it lacked was the instinct to stop and ask for the missing piece before plowing ahead.

WHY IT ACTUALLY FAILS In nine out of twenty failed multi-turn conversations, the AI never asked the question it needed to ask. It just moved forward on incomplete information and hoped for the best.

Now picture a new rep on the other side of that. The “buyer” gives a vague, partial answer, exactly like a real one would. A generic AI won’t reliably push back, dig deeper, or clock the gap. It tends to plow on regardless, because proactive clarification is precisely the skill its weakest at.

That’s the exact opposite of the habit you want a rep building.


The Quieter (Bigger) Problem: Confidence Without Competence

 

There’s a second issue sitting underneath all this, and it might be the more dangerous one.

Research into how these models’ express confidence keeps finding the same pattern. They’re frequently most confident exactly when they’re wrong. Not a little overconfident, genuinely miscalibrated, sounding certain about answers that turn out to be false, with nothing built in to flag the doubt.

A generic AI doesn’t put its hand up and say, “I’m not sure about this one, I will go and check.” It just answers, smoothly, whether it’s right or not.

For drafting a follow-up email it is mildly annoying. For a rep rehearsing how to handle a technical objection or a competitive question, that’s actively dangerous. They absorb the confident delivery and the wrong content in the same breath, with no signal telling them which parts to trust.

THE QUIET RISK It’s not that generic AI gets things wrong sometimes. Every tool does. It’s that it rarely warns you when it has. Reps walk out of a practice session having rehearsed confidence in the wrong answer.


Reasoning Models Don’t Fix This Either

 

Worth heading off the obvious objection here. Surely the newer, smarter “reasoning” models close this gap?

They help, genuinely. The strongest reasoning models in the study outperformed the lighter, faster ones by a wide margin, in some cases over 20 points. But even the best of them still landed around the mid-thirties on multi-turn tasks. Better than the pack. Nowhere near good enough to build a training program on.

This isn’t a “wait for GPT-6” problem. It’s a design problem. The same study found that the models most willing to ask clarifying questions performed best in multi-turn settings, clarification-seeking correlated directly with success. That’s a behavior you build in on purpose. It’s not something that shows up automatically because the underlying model got bigger.


What Purpose-Built Systems Do Differently

 

So what closes the gap? Not a bigger model. A different design brief.

A generic model is simulating a buyer from general knowledge about how conversations tend to sound. It’s guessing at plausible dialogue. A purpose-built system is grounded in something narrower and far more useful, your actual captured sales calls, your actual objections, your actual product, your actual buyer personas. When it doesn’t have enough information, it’s built to ask, because asking is the job, not an edge case the model happened to pick up.

Put a coach in that loop too, not just a model, reviewing what “good” looks like against real outcomes, and you get a closed feedback loop instead of an open-ended guess. That’s the whole reason we built the SenPro Gym around real call data and real coaches, rather than a general-purpose chat window with a sales persona bolted on.


Stop Training Reps Against a Coin Flip

 

Here’s where I land on this.

Generic AI is a genuinely useful tool for a hundred things. Drafting an email, summarizing a call, knocking out a first-pass objection list. Nobody’s saying don’t use it for any of that.

But roleplay is different. Roleplay is reps, in the gym sense as much as the sales sense. It’s the thing you do fifty times so that it turns into muscle memory. Nobody builds strength lifting a weight that gives way at random. If the tool your team is repeating against is right 58% of the time on the easy version of the test, and 35% on the version that resembles a sales call, you’re not training reps. You’re training them on noise and hoping the good habits outnumber the bad ones by accident.

We built SentientPro because reps deserve better odds than that. Real coaching, from people who’ve carried a real commercial number. Roleplay grounded in your own calls, not a generic script pulled from the wider internet. That’s what Educate and Enable really means to us. Not a tagline. A design decision.

If your team’s currently practicing on a generic model and it feels like it’s working, fair enough, keep going. Just know the number you’re working with.


58% on the easy test. 35% on the one that matters.


Source: Huang, K-H. et al., “CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions,” Salesforce AI Research, 2025. arxiv.org/abs/2505.18878

SentientPro Revenue Gym Easy Lift

Want to learn more?

Take a look at the SenPro Gym

Looking for more? Keep on reading...