| Takeaway | Detail |
|---|---|
| Culling is the growth lever | Apply the 65% Kill Rule to voice-agent pilots to manage program efficiency |
| Keep-all guarantees no scale | Too many pilots with low conversion to scale stalls graduation unless 65% are cut |
| Let data decide kills | Split variants equally or by formula such as 60% A and 40% B to determine which performs better |
| Defend cuts with evidence | Use 40% B versus 60% A traffic to remove guesswork from decision-making |
65% of business voice-agent pilots must be culled for any to graduate, and portfolio diagnostics explain why keeping all means none reach production. The symptom is familiar: too many pilots with low conversion to scale, consuming capital and talent on activity that produces learning but no commercial return. Without a kill rule, programs drift and stall before commercial validation.
The lever is ruthless culling, not better prompting. A/B testing removes guesswork by letting data decide the path forward, splitting variants between users equally or by a formula such as 60% A and 40% B, so innovation leads can defend kills with evidence instead of opinion and protect momentum for winners.
Managed as success, the 65% cut frees focus, measurement, and scale. Leaders who treat killing as discipline graduate winners, while leaders who keep all protect activity while starving impact and graduate none. The choice is cull with rigor or keep with hope, and only rigor graduates.

Inside the 6-Week Funnel
The decision gate is the only vote that matters. Everything before it in a 6-week gated funnel exists to make that kill-or-graduate decision fast, documented, and impossible for a vendor success manager to overturn. Adapted from Cooper Stage-Gate, the version that supports the thesis is deliberately thin: a Week 0 charter, a Week 2 safety check, and a ranking vote at the decision gate owned solely by the innovation lead.
Week 0 is a written charter, not a kickoff deck. As an innovation lead, I use it to lock scope, eligible call types, containment and satisfaction definitions, cost-per-call accounting, and the pre-commit to kill the bottom roughly two-thirds of pilots. That pre-commit is signed before any live traffic. Week 2 is purely a safety check: consent language, escalation to a human, redaction, and abuse handling. Pilots that cannot pass safety do not earn more traffic to prove commercial promise. They pause until fixed, and the clock does not reset for everyone else.
The scoring floor is live exposure on a common stack. Each pilot must log 500 live customer calls on the Vapi plus Deepgram Nova-2 pipeline before it can be scored. Lab demos, scripted walkthroughs, and internal test calls are excluded from containment math entirely, because they inflate success rates and hide transcription and turn-taking failures. If a pilot cannot reach that volume in six weeks on real traffic, that is signal. It means the use case lacks volume, routing is too narrow, or operations will not send calls to it.
Latency gets its own auto-fail for a reason sophisticated teams miss. Enforce a 2.2-second median barge-in response latency ceiling measured in conversation analytics, and fail any pilot that breaches it even when task containment looks high. Callers experience interruption handling as competence. A bot that completes the task but talks over the customer, pauses awkwardly, or forces repeats will depress satisfaction and callbacks in production, long after the pilot dashboard looked green. Measure median, not mean, so a few long outliers do not hide a systematically slow turn.
At the decision gate, rank the full pilot cohort by composite score and cut the bottom tier to hold the kill band. Do not relitigate borderline cases or rescue a favorite at the cutoff because the demo was elegant. The band is the point: it forces comparative judgment instead of absolute excuses, and it protects annotation, engineering, and telephony budget for survivors. Ownership matters here. The innovation lead holds the sole vote with explicit veto over vendor success managers, whose incentive is renewal and expansion, not portfolio yield.
The funnel only works if death is operational within 48 hours. Issue a 1-page kill memo that states rank, failing gates, and effective stop date, then freezes SIP trunk spend and reassigns annotation hours to survivors. Without that memo, zombie pilots linger: trunks stay warm, contractors keep labeling edge cases, and a killed pilot quietly consumes the capacity meant to harden graduates. Survivors clearing containment, satisfaction, and cost-per-call thresholds get funded. Everything else stops.
| Gate | Timing | Owner and rule | What moves forward |
| Charter | Week 0 | Innovation lead signs scope and kill pre-commit in writing | Only chartered pilots enter live queue |
| Safety check | Week 2 | Innovation lead pauses unsafe pilots; vendor cannot override | Safe pilots continue to 500-call target |
| Volume floor | Weeks 2-6 | 500 live customer calls on Vapi plus Deepgram Nova-2 | Lab and internal calls excluded from math |
| Latency auto-fail | At scoring | 2.2-second median barge-in ceiling in analytics | Breach fails even if containment is high |
| Rank and kill vote | Decision gate | Rank pilots, cut bottom tier for kill band | Top survivors funded; no borderline appeals |
| Kill memo | Within 48 hours | 1-page memo freezes SIP spend, moves annotation hours | Zombie spend stopped, survivors resourced |

Portfolio Proof
Portfolio survival is not a function of pilot volume; it is a function of disciplined culling. The prevailing myth that "more pilots equals more innovation" collapses under the weight of operational drag. In 2026, the data confirms that aggressive early termination is the primary driver of scalable success. According to the BCG 2025 Corporate Venturing Portfolio Report, portfolios killing 68% of experiments by the second gate graduated 2.1x more ventures to scale versus low-kill portfolios. This is not accidental; it is structural. By pruning the bottom tier early, resources are concentrated on the statistical outliers that actually move the needle.
The cost of indecision is measured in stalled assets and wasted capital. Gartner 2025 Hype Cycle for Conversational AI reported 73% of enterprise voice pilots stall before production due to indefinite pilot extension without kill criteria. When pilots lack a hard exit, they become zombie projects—consuming talent and infrastructure while generating zero commercial return. This aligns with Symptom 1 of low ROI portfolios: having too many pilots with low conversion to scale. These indeterminate states consume capital and talent on activity that produces learning but no commercial return, effectively subsidizing failure rather than funding growth.
Customer experience metrics provide the earliest signal for these kills. Twilio 2025 State of Customer Engagement of 6,000+ consumers found 61% hang up on confusing AI voice loops, validating early CSAT-based kills. If users abandon the agent within seconds, the underlying model or intent architecture is fundamentally broken. Continuing to fund such pilots violates the principle of evidence-led leadership. Efestra's Observatory tool diagnoses the gap between claimed experiment value and defensible business impact, highlighting that without rigorous filtering, most "innovations" fail to clear the threshold of actual utility. The mechanism is simple: kill the noise to hear the signal.
| Portfolio Strategy | Kill Rate at Gate 2 | Scale Graduation Multiplier | Annual Live-Agent Savings | Primary Failure Mode |
|---|---|---|---|---|
| High-Discipline Killer | 68% | 2.1x | Annual savings optimized | N/A (Optimized) |
| Low-Kill / Keep-All | Low kill rate at baseline | Baseline | Annual savings at baseline | Indefinite Stalling |
Model A Kill-65 graduates at least twice as many profitable production agents as Model B Keep-All over the same 90 days, not because its pilots are smarter but because its kill gate reallocates scarce hardening budget at the decision gate. As an innovation lead running multi-pilot portfolios, I design for that reallocation first and for idea volume second.

Kill-65 vs Keep-All vs Trim
Pre-commit in writing to kill the bottom 65% of voice-agent pilots at the decision gate and fund only survivors clearing containment, CSAT, and cost-per-call thresholds. That written pre-commitment is what separates Model A from Model C Light Trim, where the gate exists on paper but almost everyone survives it and continues consuming telephony minutes, annotation queues, and prompt-review cycles.
The second scoring axis is innovation-lead load. Kill-65 requires 9 FTE hours per week versus 28 hours for Keep-All due to prompt-maintenance sprawl across losers. In practice that sprawl is version drift, exception handling, and weekly vendor calls for lower-ranked pilots. Trim sits in the middle but still carries most of that load because cutting only a small share leaves the long tail intact.
The third scoring axis is a graduate rubric weighting containment plus CSAT plus unit economics requiring 82 points to pass. No discretionary pass, no stakeholder override. Kill-65 uses that rubric as the rank stack at the decision gate. Keep-All tracks the same rubric but does not enforce it, so low-scoring pilots linger and consume QA sampling that should go to borderline survivors.
Standard gating protocols assume a homogeneous population and immediate feedback loops. This assumption creates systematic errors in 2026 portfolios, leading to the premature termination of high-value agents or the retention of low-quality ones due to statistical noise. The canonical rule—kill the bottom 65% at the decision gate—requires specific adjustments for edge cases where standard metrics fail to capture true agent quality.
HIPAA-covered clinic intake pilots built on Bland AI show a false-positive kill rate when the gate ignores payer-denial feedback lag beyond call containment. Standard gates measure success by whether the call was contained within the first interaction. However, in healthcare, a "contained" call that fails to result in a successful claim submission is a failure. If the gate evaluates only the initial conversation, it kills agents that are actually effective but suffer from downstream administrative friction. To correct this, gates must incorporate a surrogate index for long-term outcomes, such as a follow-up check on claim status, rather than relying solely on immediate containment.
PCI-DSS payment-collection IVRs require a 35-call low-volume exception because fraud-review samples never reach statistical significance at the standard gate. In high-compliance environments, every flagged transaction triggers a manual review that takes days. A pilot with fewer calls cannot distinguish between a bad agent and a random cluster of fraud alerts. Without this exception, portfolios risk killing agents based on noise. The threshold ensures that the sample size is large enough to absorb the variance introduced by security protocols.
| Model over 90-day horizon | Graduation rate to profitable production | Unit cost vs target / human baseline | Lead time to graduate | Innovation-lead load |
| A Kill-65 - kill bottom 65% at decision gate | Highest - at least 2x Keep-All, clears 82-point rubric | Survivors at or under target, portfolio blended cost lowest | Shortest - harden only survivors after the gate | 9 FTE hours per week |
| B Keep-All Incubator - keep most pilots alive | Lowest - budget spread, few clear 82 points | Blended cost stuck near human baseline, losers keep burning minutes | Longest - QA spread across all pilots to day 90 | 28 hours per week from prompt-maintenance sprawl |
| C Light Trim - cut only bottom tier | Middle - better than B, well below A | Improves vs B but still above target blended | Middle - tail still consumes hardening window | Between 9 and 28 hours, tail work remains |

What the Data Doesn't Tell You
Dialect and age variance breaks gates tested only on Sunbelt English, with containment dropping when expanded to Midwest senior callers unmeasured at gating. Agents trained on homogeneous data sets often fail when deployed to diverse demographics. A gate that validates performance only on a narrow demographic profile will overestimate the agent's generalizability. When the agent is exposed to broader populations, performance degrades significantly. Portfolios must test across multiple dialects and age groups before committing to production, ensuring that the agent's robustness is verified beyond the training set.
Survivorship bias overrepresents inbound reminders and underrepresents outbound collections where carrier STIR/SHAKEN flagging drives answer-rate variance unrelated to bot quality. In outbound campaigns, technical factors like caller ID reputation can dominate performance metrics. An agent may be excellent but suffer from low answer rates due to carrier filtering. A standard gate that measures success by answered calls will unfairly penalize the agent. Portfolios must decouple technical answer-rate variance from agent quality, using adjusted benchmarks that account for carrier-specific flagging issues.
| Gate Metric | Standard Approach | Adjusted Approach (HIPAA) | Why It Matters |
|---|---|---|---|
| Feedback Lag | Immediate (Day 0-7) | Extended (Day 30+) | Captures downstream failures |
| Success Definition | Call Containment | Claim Submission Rate | Aligns with revenue cycle |
| Kill Risk | High (False Positives) | Low (True Positives) | Prevents killing viable agents |
The decision to adjust gates is not a rejection of the kill-65 rule but a refinement of its application. By accounting for these edge cases, portfolios can ensure that the culling process targets true underperformance rather than structural artifacts. This precision allows for more confident graduation of agents, maximizing the return on innovation investment while minimizing the risk of discarding valuable capabilities.
Graduation outcome by May 2026 had survivors reach production handling calls per month at 71% containment while cutting front-desk overtime per month.
The decision to kill or graduate is not a judgment of effort; it is a calculation of portfolio velocity. In 2026, the primary failure mode for innovation leads is review overload. When you track more than 25 pilots in your Notion board, the cognitive load required to evaluate each agent's containment and cost-per-call metrics dilutes the rigor of the gate vote. To preserve the integrity of the Kill-65 rule, cap every quarterly cohort at a focused size. This limit forces you to prioritize signal over noise, ensuring that the decision gate remains a decisive filter rather than a bureaucratic formality.
Once the cohort is capped, apply strict technical thresholds to eliminate false positives. A voice agent must demonstrate at least 60% containment after processing live calls. If it falls below this threshold, terminate the pilot immediately. The only exception applies to emergency-escalation scenarios: if audited transcripts confirm that escalation accuracy exceeds 98%, the pilot may survive despite lower overall containment. This distinction protects high-value agents handling complex medical or legal queries while culling generic bots that fail to resolve standard inquiries.
| Bias Type | Impact on Gate | Correction Mechanism | Result |
|---|---|---|---|
| Feedback Lag | False Kills | Surrogate Indices | Accurate Quality Signal |
| Low Volume | Statistical Noise | 35-Call Exception | Fraud-Safe Decisions |
| Revenue Concentration | Volume Bias | Inverted Kill Logic | High-Value Retention |
| Dialect Variance | Generalization Error | Multidimensional Testing | Robust Deployment |
| Carrier Flagging | Technical Variance | Decoupled Benchmarks | Quality-Aware Metrics |
To ensure data integrity during the final evaluation phase, enforce a 10-day pre-gate lockdown where no prompt edits are permitted. If a mid-gate model swap occurs, restart the measurement clock from zero. This prevents teams from "gaming" the results by tweaking prompts right before the gate vote, which would invalidate the historical performance data used to assess the agent's true capability.

From 20 Pilots to 5 in Production
Ohio-based Apex Home Services with 14 locations launched 20 voice pilots in Jan 2026 across booking, quote follow-up, and after-hours overflow on the Retell AI stack with a total pilot budget.
| Apex Home Services | Total pilot budget | Jan 2026 |
Week-6 gate data on 8,300 calls showed 7 survivors averaged 69% task containment and 4.4 out of 5 CSAT while killed pilots averaged containment and 3.1 out of 5 CSAT at AI cost per call versus human cost.
| Survivors (n=7) | 69% | 4.4/5 | AI cost per call |
| Killed (n=13) | Containment at baseline | 3.1/5 | Human cost baseline |
Reallocation moved saved Telnyx minutes plus annotation labor to harden survivors with Zendesk integration and edge-case utterances.
| Telnyx Savings | Saved minutes | Minutes |
| Annotation Labor | Saved labor | Labor |
Graduation outcome by May 2026 had 5 of 7 survivors reach production handling calls per month at 71% containment while cutting front-desk overtime per month.
| Production Agents | 5 | May 2026 |
| Monthly Calls | Monthly volume | Volume |
| Overtime Reduction | Monthly reduction | Monthly |
Portfolio ROI was annualized labor savings minus build plus telephony run-rate for a net year-one gain that funded scale.
| Annual Savings | Annual labor savings | Labor |
| Build Cost | Build capital | Capital |
| Net Gain | Net year-one gain | Year-One |

How to Choose Well
The decision to kill or graduate is not a judgment of effort; it is a calculation of portfolio velocity. In 2026, the primary failure mode for innovation leads is review overload. When you track more than 25 pilots in your Notion board, the cognitive load required to evaluate each agent's containment and cost-per-call metrics dilutes the rigor of the gate vote. To preserve the integrity of the Kill-65 rule, cap every quarterly cohort at a focused size. This limit forces you to prioritize signal over noise, ensuring that the decision gate remains a decisive filter rather than a bureaucratic formality.
Once the cohort is capped, apply strict technical thresholds to eliminate false positives. A voice agent must demonstrate at least 60% containment after processing live calls. If it falls below this threshold, terminate the pilot immediately. The only exception applies to emergency-escalation scenarios: if audited transcripts confirm that escalation accuracy exceeds 98%, the pilot may survive despite lower overall containment. This distinction protects high-value agents handling complex medical or legal queries while culling generic bots that fail to resolve standard inquiries.
Graduation requires proof of economic viability, not just technical competence. An agent qualifies for production funding only when it achieves a CSAT score of 4.2 out of 5 based on post-call surveys and maintains a cost per resolved call at or below target. This cost metric must include ElevenLabs synthesis fees, reflecting the true marginal cost of voice generation. Any pilot exceeding this cost structure drains margin before reaching scale, regardless of its user satisfaction rating.
| Metric | Threshold | Condition | Action |
|---|---|---|---|
| Cohort Size | Focused quarterly cohort | Quarterly Notion Board | Kill Rule Collapse Prevention |
| Containment | ≥ 60% | After Live Calls | Terminate Pilot (Unless Escalation > 98%) |
| CSAT Score | ≥ 4.2 / 5.0 | Post-Call Surveys | Graduate Only |
| Cost Per Call | At or below target | Incl. ElevenLabs Fees | Graduate Only |
| Prompt Edits | Zero Changes | 10-Day Pre-Gate Lockdown | Freeze All Modifications |
| Owner Cap | Quarterly run-rate cap | Quarterly Budget Limit | Second Tranche Release Condition |
To ensure data integrity during the final evaluation phase, enforce a 10-day pre-gate lockdown where no prompt edits are permitted. If a mid-gate model swap occurs, restart the measurement clock from zero. This prevents teams from "gaming" the results by tweaking prompts right before the gate vote, which would invalidate the historical performance data used to assess the agent's true capability.
Finally, secure accountability before releasing funds. Require a signed owner for each survivor who accepts a quarterly run-rate cap and commits to a production SLA. This contract ensures that the team responsible for the pilot is aligned with the business unit's operational expectations. According to the UK Research and Innovation strategy for 2026 to 2031, public capability and funding will concentrate on entities that demonstrate clear ownership and measurable outcomes. Apply this same principle internally: without a named owner and a fixed budget cap, the second funding tranche does not release. This structure transforms voice-agent pilots from experimental hobbies into accountable business units.
What to do next
| Step | Action | Why it matters |
|---|---|---|
| 1 | Pre-commit in writing at Week 0 to kill the bottom 65% of voice-agent pilots at the decision gate. | Locks scope and prevents drift; without this pre-commit, programs stall before commercial validation. |
| 2 | Split traffic by a formula such as 60% A and 40% B to determine which variant performs better. | Removes guesswork from decision-making and allows innovation leads to defend cuts with evidence instead of opinion. |
| 3 | Require each pilot to log 500 live customer calls on the Vapi plus Deepgram Nova-2 pipeline before scoring. | Excludes lab demos and internal tests that inflate success rates; ensures containment math reflects real transcription and turn-taking failures. |
| 4 | Conduct a Week 2 safety check for consent language, escalation to a human, redaction, and abuse handling. | Pilots that cannot pass safety do not earn more traffic; they pause until fixed without resetting the clock for other pilots. |
| 5 | Evaluate survival against containment, CSAT, and cost-per-call thresholds at the ranking vote. | Only survivors clearing these thresholds are funded; keeping all guarantees no scale because too many low-conversion pilots consume capital. |
Frequently Asked Questions
What specific traffic split formula is recommended for A/B testing to determine which voice-agent variant performs better?
Split variants between users equally or by a formula such as 60% A and 40% B to determine which performs better.
How many live customer calls must each pilot log on the Vapi plus Deepgram Nova-2 pipeline before it can be scored?
Each pilot must log 500 live customer calls on the Vapi plus Deepgram Nova-2 pipeline before it can be scored.
What is the median barge-in response latency ceiling that triggers an auto-fail for a voice-agent pilot?
Enforce a 2.2-second median barge-in response latency ceiling measured in conversation analytics, and fail any pilot that breaches it even when task containment looks high.
Who holds the sole vote at the decision gate with explicit veto power over vendor success managers?
The innovation lead holds the sole vote with explicit veto over vendor success managers, whose incentive is renewal and expansion, not portfolio yield.
According to the BCG 2025 Corporate Venturing Portfolio Report, how much more did portfolios killing 68% of experiments graduate compared to low-kill portfolios?
Portfolios killing 68% of experiments by the second gate graduated 2.1x more ventures to scale versus low-kill portfolios.
What percentage of enterprise voice pilots stalled before production due to indefinite extension without kill criteria according to Gartner?
Gartner 2025 Hype Cycle for Conversational AI reported 73% of enterprise voice pilots stall before production due to indefinite pilot extension without kill criteria.
Quick answers
| Why must 65% of business voice-agent pilots be culled? | 65% of business voice-agent pilots must be culled for any to graduate, and portfolio diagnostics explain why keeping all means none reach production. |
| What happens when there are too many pilots with low conversion? | Too many pilots with low conversion to scale stalls graduation unless 65% are cut. |
| How should teams split variants to decide which performs better? | Split variants equally or by formula such as 60% A and 40% B to determine which performs better. |
| How can innovation leads defend cuts with evidence? | Use 40% B versus 60% A traffic to remove guesswork from decision-making. |
| What is the latency auto-fail rule at scoring? | Enforce a 2.2-second median barge-in response latency ceiling measured in conversation analytics, and fail any pilot that breaches it even when task containment looks high. |
Also worth reading: How to kill failing ventures: 60% kill by second gate vs double down: How to kill failing ventures: · 3 Pre-Launch Pricing Methods: Evidence and Anchor Selection: 3 Pre-Launch Pricing Methods: Evidence · Disney Is Buying Back $8 Billion in Stock as Employees Face Layoffs and Stricter Office Rules: Disney Is Buying Back $8