Back to all articles
AI coaching
pilot programmes
scaling
change management
adoption

Beyond the Pilot: Why Most AI Coaching Programmes Stall and How to Scale Yours

James Mitchell
12 min read
Share

The pilot went well. Twenty reps used the AI coaching tool for eight weeks. Engagement was high. Feedback was positive. The training team produced a slide deck showing improved confidence scores and enthusiastic qualitative comments. Leadership nodded approvingly and said "let's roll it out."

Then nothing happened.

Six months later, the tool is still live for the original pilot group (half of whom have stopped using it) and the broader rollout is perpetually "in planning." The champion who ran the pilot has moved on to other priorities. IT has flagged a security review that nobody has scheduled. The budget for enterprise licensing is stuck in next year's planning cycle. The AI coaching programme has not failed, exactly. It has simply stopped moving forward.

This story is so common that it has become the default outcome for AI coaching pilots in life sciences. I have watched it happen at companies of every size, with tools from every vendor. The pattern is remarkably consistent: the pilot succeeds on its own terms, but the organisation fails to convert that success into scaled adoption.

Understanding why this happens is the first step toward preventing it.

Why pilots succeed but scaling stalls

Pilots have structural advantages that disappear at scale. Recognising these advantages helps explain why a successful pilot does not automatically predict a successful rollout.

The pilot team was hand-picked. Most pilot groups consist of willing participants. They volunteered, or their manager volunteered them because they are the type of people who embrace new tools. This selection bias means the pilot tests the tool's effectiveness with enthusiastic users, not with the sceptical majority who will encounter it during a broader rollout. Scaling means reaching the reps who did not volunteer and the managers who are not naturally curious about new technology.

The champion provided energy that does not scale. Every successful pilot has a champion, someone who sends reminder emails, troubleshoots problems, celebrates wins, and personally ensures that adoption stays high. This person invested significant time and attention in making the pilot work. During a scaled rollout, no single person can provide this level of attention to hundreds of users. If the programme depends on one person's energy, it will not survive the transition to a broader audience.

The pilot had no integration burden. During a pilot, it is acceptable for the tool to sit outside normal workflows. Reps log into a separate platform, use it when they remember, and report their usage back to the champion. At scale, this standalone approach creates friction. Reps will not maintain a separate login and a separate workflow indefinitely. The tool needs to fit inside how they already work.

Nobody challenged the business case. A twenty-person pilot does not attract serious financial scrutiny. The cost is small enough to sit within a discretionary training budget. But scaling to 500 users means a significant line item that requires proper justification. Suddenly, the "improved confidence scores" from the pilot are not sufficient. Finance wants to know what the return on investment is, and "people felt better about their conversations" does not satisfy that question.

The six common stalling points

In my experience, AI coaching programmes stall for specific, identifiable reasons. Usually more than one is present simultaneously.

1. The pilot was too small to prove ROI. Twenty reps over eight weeks does not generate statistically significant commercial outcomes. You cannot demonstrate revenue impact from a sample this small, and without revenue impact, the business case for scaled investment is weak. The pilot proves the tool works. It does not prove the tool is worth the cost of enterprise deployment.

2. Leadership changed or shifted priorities. The executive who sponsored the pilot moved to a different role. Their replacement has different priorities and does not feel ownership over the initiative. Or a reorganisation shifted the AI coaching programme from one department to another, and the receiving team does not see it as their project.

3. The champion left or moved on. The person who made the pilot work day-to-day was promoted, changed roles, or simply ran out of bandwidth as other projects demanded their attention. Without someone actively pushing the programme forward, momentum drops to zero. Programmes in motion stay in motion. Programmes without a driver stop.

4. IT raised security or procurement concerns. During the pilot, the tool was approved under a limited-use exemption or a trial agreement. Scaling requires full security review, data processing agreements, SSO integration, and potentially a new procurement process. Each of these steps takes time, and any one of them can introduce a months-long delay.

5. Managers did not adopt it. The pilot focused on reps. Managers were informed but not involved. When the programme tries to scale, managers become the bottleneck. They do not understand the tool well enough to coach their teams on using it. They do not review the data it produces. They do not reinforce its use in one-to-ones. Without manager involvement, rep adoption decays once the initial novelty fades.

6. Reps stopped using it after the novelty wore off. Even within the pilot group, usage often follows a predictable curve: high initial engagement, a plateau, then a decline. If the programme offers the same scenarios repeatedly, if scoring does not feel meaningful, if the connection between practice and performance is not visible, reps deprioritise it in favour of more immediately urgent tasks.

What successful scaling looks like

The organisations that successfully move from pilot to full deployment share common characteristics. None of these are particularly surprising, but they require deliberate planning that most pilot teams do not do.

Executive sponsorship that survives the pilot phase. The sponsor needs to remain engaged through implementation, not just approve the initial investment. This means regular updates to leadership on progress, blockers, and results. It means the sponsor actively removes obstacles rather than delegating everything to the training team. Successful programmes have a named executive who considers the AI coaching initiative part of their personal portfolio.

Integration into existing workflows. At scale, the tool cannot be a separate destination. It needs to show up where reps already spend their time. That might mean integration with CRM, with the existing LMS, or with whatever communication tool the team uses daily. The principle is simple: reduce the friction of access to near zero. Every additional click, login, or context switch reduces usage.

Manager enablement before rep rollout. Before scaling to reps, train the managers. Show them how to read coaching data. Teach them how to use session transcripts in their one-to-ones. Help them understand what the scores mean and how to act on them. A manager who can say "I noticed your objection handling scores dropped this week, let's talk about what's happening" is a manager who reinforces the programme daily without any additional effort from the training team.

Clear success metrics agreed before the pilot starts. This is the most commonly missed step. The time to define success criteria for the pilot is before it begins, not after it ends. What specifically would need to be true for the organisation to commit to a full rollout? What data would finance need to see? What level of adoption would demonstrate genuine value? Write these criteria down. Get stakeholder agreement. Then design the pilot to produce exactly that evidence.

Designing a pilot that sets up successful scaling

If you have not yet started your pilot, or if your current pilot is still early enough to adjust, here is what a well-designed pilot looks like.

Team size: 40-60 reps. This is large enough to generate meaningful data but small enough to manage closely. Twenty reps is too few for statistical significance on any outcome metric. One hundred is too many to support with the attention a pilot requires.

Duration: 12-16 weeks. Eight weeks is not long enough to observe behaviour change in the field. You need at least twelve weeks to see the practice-to-performance connection in commercial metrics. Sixteen weeks gives you buffer for the inevitable slow-start period where adoption is building.

Include both new hires and tenured reps. If your pilot only includes new hires, you cannot demonstrate value for the majority of your salesforce. If it only includes tenured reps, you miss the most obvious use case (onboarding acceleration). Include both, and segment your results.

Include the managers. From day one, managers of pilot participants should be involved. They should see the data, reference it in coaching conversations, and provide feedback on whether the skills they see in simulation are translating to the field. This does two things: it generates field validation data, and it builds manager capability for the scaled rollout.

Pre-agree the decision criteria. Before the pilot begins, sit down with every stakeholder who would need to approve a full rollout: finance, IT security, commercial leadership, compliance. Ask each one: "What would you need to see from this pilot to support a full deployment?" Document their answers. Design the pilot to produce that evidence. If finance needs revenue correlation, build the measurement approach into the pilot from week one. If IT needs a security assessment, start it in parallel rather than waiting until the pilot ends.

Measure leading indicators weekly. Do not wait until week twelve to look at data. Track practice frequency, score trajectories, scenario completion rates, and manager engagement weekly. This lets you intervene early if adoption is dropping and gives you a story of progressive improvement rather than a single endpoint measurement.

The scaling timeline

Successful programmes typically follow a predictable timeline from pilot completion to full deployment.

Weeks 1-4 post-pilot: Business case development. Compile pilot results against the pre-agreed success criteria. Build the financial model for full deployment. Present to leadership with a specific request and timeline.

Weeks 5-8: Procurement and security. Run the full security review. Negotiate enterprise licensing. Complete the data processing agreement. Establish SSO integration requirements.

Weeks 9-12: Manager enablement. Train all managers who will oversee the programme at scale. Give them access to the tool and their team's data. Run workshop sessions on how to interpret scores and integrate AI coaching data into their existing management rhythms.

Weeks 13-16: Phased rep rollout. Deploy to the broader population in cohorts rather than all at once. This allows the training team to support each cohort during their first two weeks, address questions, and maintain quality of implementation.

Weeks 17-20: Optimisation. Monitor adoption across all cohorts. Identify teams or regions where usage is low and diagnose why. Adjust scenarios, introduce new content, and publish early success stories from the scaled deployment.

The uncomfortable truth

Some pilots should not scale. If the pilot produced genuinely disappointing results, if reps did not improve, if managers found the data useless, if the tool did not fit the workflow, then the honest conclusion is that this particular solution is not right for your organisation.

The mistake is treating every pilot as a success because people "liked" it. Satisfaction is not the same as impact. If you set rigorous success criteria before the pilot and the pilot did not meet them, that is valuable information. It saved you the cost of a failed full deployment.

But if the pilot met its criteria and the programme is stalling for organisational reasons rather than performance reasons, then the problem is solvable. It requires deliberate attention to the structural factors that kill momentum: sponsorship, integration, manager enablement, and a clear business case that speaks to what finance actually cares about.

Most programmes do not fail because the technology does not work. They fail because nobody planned for what comes after the pilot ends.

Frequently Asked Questions

Share this article