# Guardrail Tiers: 12 Questions, 36 Points, 15 Pilots a Year

Ivy Nakamura · August 24, 2026

> Guardrail Tiers: 12 Questions, 36 Points, 15 Pilots a Year. A mid-size experimental portfolio loses substantial pilot-days every year...

| Takeaway | Detail |
| --- | --- |
| The 6-week legal sign-off functions as an untriaged queue, not a risk judgment. | A mid-size pilot portfolio burns substantial pilot-days every year in sign-off queues because cookie-consent A/B tests and warehouse-robotics deployments enter the same counsel lane, when checklist-grade work belongs with legal ops. |
| Compliance-scored guardrails beat unguarded rollouts at 95% confidence. | Aggregate operational readouts show variable cohorts winning 301/659 versus controls' 536/1548 — an 11.0-percentage-point win-rate lift, 95% CI [6.6, 15.5], p < 0.001 — with adjusted item-not-received disputes adding a +7.5-point lift at p = 0.045 (arXiv 2606.01513). |
| Primary guardrails cap downside before any pilot reaches counsel: error rate up no more than 10%, latency p95 up no more than 15%, crash rate up no more than 5%. | Every experiment also passes a sample-ratio mismatch test at p > 0.001, and surface-specific secondaries tighten further — checkout payment errors capped at 5% growth, abandoned carts at 10%, support tickets at 20%, search zero-result rate at 5% (StatsTest). |
| As many as 80% of pilots never require counsel judgment at all. | They require a documented control plane executed by legal ops: PII redaction, prompt-injection blocking, toxicity filtering, hallucination checks, off-policy refusal, and decision-trace logging — the six runtime guardrails enterprise customers now demand before signing (FutureAGI). |

A mid-size experimental portfolio loses substantial pilot-days every year to legal sign-off queues — elapsed time that stacks up across every experiment waiting behind the queue. The bottleneck is rarely caution; it is triage failure. When a cookie-consent A/B test and a warehouse-robotics deployment enter the same six-week review lane, counsel hours get spent settling decisions a checklist could handle, while genuinely novel risk waits behind them.

The evidence says speed and safety are not in tension. A unified guardrail orchestration layer described in arXiv 2606.01513 runs parallel generation heads scored against weighted guardrails — PII detection, content moderation, schema constraints, domain rules — and exits early on a compliance score. In aggregate operational readouts, guarded variants beat controls 301/659 to 536/1548, an 11.0-percentage-point win-rate lift with a 95% confidence interval of [6.6, 15.5] and p < 0.001.

The calendar makes tiering urgent rather than optional. Colorado's AI Act takes effect February 1, 2026, applying algorithmic-discrimination duties to consequential decisions, and the EU AI Act's Article 13 transparency obligations apply from August 2, 2026. Enterprise customers now require a documented control plane before they will sign. The fix is a tiered intake — 12 triage questions, 36 scoring points, and a fast lane built to clear 15 pilots a year.

![Guardrail Tiers](https://static.mm-ais.com/article-images-ai/guardrail-tiers-12-questions-36-points-1-ai-792edd81.jpg)

## The 12-Question Rubric

Twelve questions, six risk axes, thirty-six possible points: that is the entire machinery deciding whether a pilot waits five days or six weeks. At intake, every pilot is scored 0–3 on personal-data sensitivity, automated-decision impact on individuals, financial commitment size, third-party IP exposure, operational blast radius, and regulatory surface; the summed score routes it into one of three lanes. The design detail that matters is that the axes behave like guardrail metrics in the experimentation sense — according to Atticus Li, a guardrail metric is one that can veto a win, the tripwire that stops shipping harm at scale. A pilot cannot offset a maximum score on automated-decision impact by being cheap; the axis vetoes, and the pilot climbs a tier.

| Tier | Lane | Target clock | Executed by |
| --- | --- | --- | --- |
| Tier 1 | Pre-approved template sign-off | 5 business days | Trained legal-ops analyst |
| Tier 2 | Capped-scope parallel review | 15 days | Counsel-led review |
| Tier 3 | Full committee | ~6 weeks (legacy path) | Full committee |

The tier vocabulary is borrowed, not coined. Internal labels inherit the EU AI Act's four-level risk ladder — unacceptable, high, limited, minimal — whose high-risk obligations land August 2, 2026, so an auditor reads your routing logic in language it already enforces. The mapping writes itself: a pilot brushing unacceptable-risk territory, such as emotion inference in hiring, exits the process entirely rather than queueing; limited- and minimal-risk pilots with a matching pattern route Tier 1; anything touching high-risk ground — employment screening, creditworthiness, essential services — caps at Tier 2 or climbs to Tier 3 regardless of how benign its raw score looks.

Here is the part outsiders miss: the rubric does not create the five-day speed — the pattern library does. Tier 1 eligibility requires the pilot to reuse a previously signed-off pattern: a data flow already covered by a completed GDPR Article 35 DPIA, or mapped to existing SOC 2 controls. Reviewers then grade deltas against the approved baseline, not whole architectures. Say a support-desk chatbot reuses the sanctioned flow of ticket text to an LLM API with no training retention; the review becomes a diff — one new vendor field, one unchanged retention flag — not a greenfield assessment. No pattern, no fast lane, whatever the score says.

The clock needs a falsifiable start or the SLA is marketing. Five business days runs from intake-complete — rubric submitted plus every required artifact attached — never from the first email, which lets sponsors blame delays that were really their own missing DPAs. According to Kameleoon, sample-ratio-mismatch detection tops the trust guardrails because when the split is off, no metric result can be trusted; an ambiguous start trigger poisons your SLA statistics the same way. According to Statsig's guidance on experimentation programs, teams should fix predefined criteria up front and steer clear of after-the-fact adjustments — the trigger definition is exactly that kind of pre-registration.

Judgment stays in the system through an escalation valve: any reviewer, including the most junior analyst, can bump a pilot up one tier within 48 hours by attaching a mandatory written reason code. Upward only — nobody downgrades a colleague's caution — and the queue behind the bumped pilot never reopens. According to Airbnb Engineering, a triggered guardrail forces an escalation in which stakeholders discuss results transparently before launch; the reason code serves that role here, converting a reviewer's unease into a logged, contestable artifact instead of a hallway veto.

The staffing mechanic makes the arithmetic work: Tier 1 checks are executed by trained legal-ops analysts working from checklists, reserving counsel hours for the Tier 2 and Tier 3 judgment calls that actually need a lawyer. The reflex objection — that non-lawyers approving anything means accepting more risk — gets the causality backwards. Slow, untriaged queues are what push business teams into shadow pilots outside legal visibility, and that is where liability actually accumulates; a checklist-run fast lane keeps those experiments inside the tent.

Before promising anyone the five-day figure, back-score last quarter's pilots against the 36-point rubric and ask one question per pilot: did an approved pattern already exist? Where would-be Tier 1 pilots had no pattern to reuse, your bottleneck is the library, not the rubric — fund pattern creation before you fund speed.

![The 12-Question Rubric — Guardrail Tiers](https://static.mm-ais.com/article-images-ai/guardrail-tiers-12-questions-36-points-1-ai-c03c73a2.jpg)

## Where 6 Weeks Goes

Strip the committee-lane wait down to its components and almost none of it is legal judgment. According to World Commerce & Contracting benchmarking, the average B2B contract cycle — drafting through signature — runs roughly 3.4 weeks. A pilot agreement is a B2B contract wearing a lab coat: same redline rounds, same version-control ping-pong, same signature logistics. Subtract that contract cycle from the committee-lane baseline quantified earlier in this guide and roughly 2.6 weeks remain, and that residue is pure queue-and-scheduling lag — intake triage sitting unread, misaligned calendars, reviewers waiting on coverage. The scarce resource was never counsel attention; it was a slot in the sequence.

| Component of the committee-lane wait | Duration | Basis |
| --- | --- | --- |
| Contract drafting-through-signature cycle | ~3.4 weeks | World Commerce & Contracting benchmarking |
| Queue, triage, and scheduling residue | ~2.6 weeks | Derived: baseline minus WorldCC contract cycle |

Thomson Reuters' Legal Department Operations benchmarking explains why that queue forms in the first place: legal teams spend a large share of capacity on routine, repeatable requests. Those requests are exactly what a scored intake rubric pulls out of the counsel queue and routes to pre-approved patterns. Note the direction of the effect — clearing the repeatable volume is what preserves genuine scrutiny depth for the novel, high-stakes pilots that still deserve the committee.

The cost of skipping this shows up at portfolio level. According to McKinsey's transformation research, most transformation programs fall short of their goals, with slow decision cycles repeatedly named among the top causes. A pilot portfolio is a queue of decisions; let each one sit for weeks and the drag compounds arithmetically across every pilot behind it.

Gartner supplies the counterweight to the reflexive objection that faster sign-off means looser control: Gartner predicted at least 30% of generative-AI projects would be abandoned after proof-of-concept rather than reach production, citing inadequate risk controls among the causes. Weak guardrails kill pilots; strong ones do not. And the failure mode of a slow, untriaged queue is worse than delay — business teams route around it, running shadow pilots outside legal visibility, which is precisely where liability accumulates. Compression and control are not a trade-off; the unmanaged queue is the risk.

MIT CISR's research on decision rights gives the tier model its academic footing: firms with explicit, codified decision authorities make resource-allocation decisions materially faster than consensus-driven peers. Committee-by-default applies consensus governance to decisions that were never ambiguous. Codified tiers replace "who should decide?" with "which lane did the score assign?"

One constraint outranks all of this optimization. Per FutureAGI's compliance tracking, EU AI Act prohibitions on banned practices are already binding, and most high-risk obligations — including the Article 13 transparency duties — apply from August 2, 2026. A tier scheme that does not already encode the high-risk category will start emitting non-compliant fast-lane approvals mid-year. The US edge case lands even sooner: according to FutureAGI, Colorado's AI Act enters force February 1, 2026, applying algorithmic-discrimination duties to consequential decisions. Encode the high-risk branch into your rubric before the August date, and audit existing templates against the prohibited-practices list now — that audit is the only item on this page with a deadline attached.

| Date | Obligation | Tier-design consequence |
| --- | --- | --- |
| Already binding | EU AI Act prohibitions on banned practices | All template lanes must screen banned practices |
| February 1, 2026 | Colorado AI Act discrimination duties live | Consequential-decision pilots need bias testing pre-approval |
| August 2, 2026 | Most EU high-risk obligations, incl. Article 13 transparency | Rubric must route high-risk pilots out of fast lanes |

![Where 6 Weeks Goes — Guardrail Tiers](https://static.mm-ais.com/article-images-pixabay/guardrail-tiers-12-questions-36-points-1-908b7596.jpg)

## One Queue, Two Tiers, or Three

Fifteen pilots a year is the pivot. Above that volume, the three-tier guardrail model pays for its own bureaucracy; below it, a binary fast lane is the rational ceiling. The three operating models in play are (A) a single queue where every pilot gets full legal review, (B) a binary split with one fast lane and everything else, and (C) the three-tier guardrail model backed by a pattern library of pre-approved templates. The choice is a portfolio-level decision, not a per-pilot judgment call, and it should be made once, in advance, with the break-even printed where everyone can see it.

The performance envelopes differ sharply. Model A holds the committee baseline covered above — nothing improves except patience. Model B cuts the median to roughly 2–3 weeks, but exceptions flow back into the same undifferentiated queue, so the tail never moves and the fast lane becomes the path of least resistance for pilots that do not belong in it. Worse, without a scored intake there is no denominator: Model B's escape rate is not low, it is unmeasurable. Model C reaches a five-business-day median from intake-complete for the qualifying majority while keeping a documented committee path for everyone else, with the escape ceiling as the tripwire. According to StatsTest, a fired guardrail warrants immediate investigation because the false alarm is cheap and the broken feature is expensive — in triage terms, a flagged escape costs an afternoon; a misfiled high-risk pilot costs far more.

The explicit winner: Model C for any portfolio running more than roughly 15 pilots per year. The mechanism is a fixed-cost-versus-scaling-benefit argument. Pattern-library maintenance — template refreshes, rubric recalibration, analyst training top-ups — is a roughly constant annual load, while latency savings scale with the number of qualifying pilots. Below ~15 per year, upkeep outruns savings and Model B is the honest choice. Printing that threshold in the table follows the discipline StatsTest prescribes: define guardrails, thresholds, and stop rules before launch, not after. The shape also scales — according to Airbnb Engineering, a company-wide Experiment Guardrails system has served dozens of teams since 2019, with priorities deliberately uneven by team: Trust watches fraud identification while Experiences watches discovery of Online Experiences on the Homepage. Central machinery, locally weighted patterns — Model C's architecture, proven in an adjacent domain.

| Operating model | Median sign-off | Fast-lane share | Escape rate | One-time setup | Audit readiness |
| --- | --- | --- | --- | --- | --- |
| A — Single queue, full review | ≈6 wk (baseline above) | None | 0% by construction — nothing is classified | Near-zero | Thick per-pilot files, no reusable pattern record |
| B — Binary fast lane | Roughly 2–3 wk | Self-selected, uncapped | Unmeasurable — no scored intake baseline | Low (one routing rule) | Ad-hoc lane, no comparable decision trace |
| C — Three tiers + pattern library | 5 business days from intake-complete | 60–80% | Held under a monitored ceiling | Substantial (rubric, templates, training) | Logged rubric scores + tier assignments |
| Volume verdict | More than ≈15 pilots/yr → adopt C; below → stay with B, because library upkeep outruns latency savings |  |  |  |  |
| Tie-breaker | A regulator or enterprise customer demanding decision trails → C wins at any volume |  |  |  |  |

Now the honest cost column. Standing up Model C takes a substantial block of internal hours upfront — rubric design, template drafting, analyst training — against near-zero for Model A. The recovery ledger runs on three lines: counsel hours shifted from bespoke review to checklist execution, calendar latency returned to experiment starts, and exposure avoided from pilots that would otherwise have run off-book. If the first line alone returns more than the build costs, payback lands inside year one; portfolios at several times the threshold volume typically clear that bar, while portfolios near it should treat the investment as multi-year and let the second and third lines carry the case. That third line is where the old objection dies: compressing legal review does not mean accepting more risk. Slow, untriaged queues push business teams into shadow pilots outside legal visibility — the off-book layer is where liability actually accumulates, and tiering shrinks it.

Two closing constraints. First, the tie-breaker is absolute: where demonstrable decision trails are required, volume stops mattering. According to FutureAGI, enterprise customers now demand a documented control plane before they will sign — an LLM project without compliance guardrails is a project without revenue — and decision-trace logging ranks among the six non-negotiable runtime guardrails. The artifact set generalizes beyond vendors: a June 1 preprint (arXiv 2606.01513) defines its reproducibility boundary through the request interface, the scoring logic, pseudocode, and a bounded set of operational evidence. Those are exactly the four artifacts a Model C audit pack contains — rubric text, logged scores, tier assignments, sampled outcomes — and exactly what Model B's ad-hoc lane cannot reconstruct retroactively. Second, the disqualifier: if your legal function has zero ops-analyst capacity, do not adopt Model C yet. Tiering redistributes work toward checklist execution; it deletes almost none of it. Without an analyst owning template upkeep and score logging, the fast lane stalls and the portfolio quietly reverts to queue behavior with extra ceremony attached. Fix staffing first. Then, before the next intake cycle, count the pilots your portfolio ran over the trailing twelve months — clear fifteen, and the five-day lane is built, not granted.

![guardrail fence blur dark](https://static.mm-ais.com/article-images-pixabay/guardrail-tiers-12-questions-36-points-1-80eec699.jpg)
guardrail fence blur dark

## What the Data Doesn't Tell You

Start with the least uncomfortable number in the evidence base: p = 0.045. According to arXiv preprint 2606.01513, posted in June 2026, adjusted item-not-received dispute cases showed a +7.5 percentage-point lift, with a 95% confidence interval running from 0.2 to 15.7 points. Take that interval seriously: the data are equally compatible with a substantial improvement and with an effect so small it nearly vanishes. The lower bound sits a fifth of a point above zero, and the result clears the conventional significance bar with almost no room to spare. This preprint is the strongest quantitative anchor behind the claim that compressed review cycles do not degrade outcomes — and it is thinner than any summary makes it look.

The width is structural, not sloppy. The fraud and local evidence-ranking deltas in the same preprint are directionally positive but not statistically significant because they were computed from aggregate count data — totals carried without per-case variance estimates. Significance testing on raw counts forces distributional assumptions that inflate every interval, so two of the three outcome families measured cannot yet be separated from noise. Any durability claim built on this evidence inherits a specific shape: one validated case family, two unproven ones.

A 15.5-point span between confidence bounds is less a measurement problem than a heterogeneity signature. Inside that distribution sit dispute categories that gained well more than 7.5 points and others that gained roughly nothing; the aggregate flattens both. Multi-pilot portfolios reproduce this exactly: the median can land precisely where the model predicts while one archetype — dispute-adjacent workflows are the obvious candidate — swings hard in both directions. Hence the cheapest upgrade available to an innovation lead: stratify your escape log by archetype before you stratify it by lane, because the published evidence says effects concentrate in particular case families rather than diffusing evenly.

The routing rule breaks in three places, and they are not symmetric. Archetype mismatch: the committee premium is justified only when a pilot's dominant risk profile resembles the unvalidated rows — fraud mechanics or ranking-manipulation surfaces — instead of the validated dispute-recovery pattern. Scope drift: a pilot scored at intake that later adds a data class or vendor integration missing from the template's original approval scope is no longer the artifact the template approved; it needs rescoring, not a rubber stamp. Low volume: beneath the annual-pilot pivot covered earlier, maintaining three tiers costs more than the fastest lane saves, and a binary gate wins. None of these invert the rule; they mark where its discount stops being free.

None of these caveats revive the oldest objection — that compressing legal review means accepting more risk. The documented failure mode runs the other way: slow, untriaged queues push business teams into shadow pilots outside legal visibility, and that is where liability quietly accumulates. The defensible conclusion is narrower: fast lanes are sound where the evidence is significant, provisional where it is merely directional, and what keeps them sound is auditing your own distribution rather than trusting anyone's aggregate. Before your next intake cycle, pull the last two quarters of fast-lane decisions, sort every escape by archetype against the table below, and fund committee review only where your counts mirror the unvalidated rows.

Only one row of that evidence licenses fast-lane routing:

| Outcome family (arXiv 2606.01513) | Adjusted result | Statistical status | Tier-gate implication |
| --- | --- | --- | --- |
| Item-not-received disputes | +7.5 percentage points | Significant; 95% CI [0.2, 15.7], p = 0.045 | Licenses template-lane routing for dispute-recovery-shaped pilots |
| Fraud outcomes | Directionally positive delta | Not significant; aggregate count data | No license; send fraud-dominant pilots to committee |
| Local evidence-ranking | Directionally positive delta | Not significant; aggregate count data | No license; treat as unvalidated archetype until replicated |

![guardrail street france](https://static.mm-ais.com/article-images-pixabay/guardrail-tiers-12-questions-36-points-1-dfe0a29a.jpg)
guardrail street france

## Escape Rates and Expired Templates

Ask any vendor pitching a five-business-day sign-off lane for the number their case studies omit: how many fast-lane pilots were later pulled back for remediation? That figure — the escape rate — is the only honest denominator for a speed claim. A lane that looks fast on paper may simply be converting its classification errors into a remediation backlog nobody reports. Before crediting any headline number, demand the escape rate and the observation window behind it.

The median hides the tail. When a Tier 1 pilot is bumped up a tier after launch, it does not resume where it left off — it loses its place in line and restarts at the back of the slower queue, with intake questions, committee scheduling, and scoping all resetting. Escaped pilots routinely overshoot even the capped-scope lane's target end-to-end for exactly this reason. The five-day figure describes the middle of the distribution, not the tail; portfolio planners who reserve capacity against the median will come up short precisely when something goes wrong.

Jurisdiction is the quiet killer of "Tier 1 everywhere" assumptions. A data flow pre-cleared under GDPR can still trip United States state law: Washington's My Health My Data Act reaches health-adjacent consumer data far beyond clinical records, and HIPAA applies in sector contexts no matter how benign the pilot looked under European rules. Multi-region portfolios need per-jurisdiction overlays attached to each template, not a single global clearance stamp.

Templates also expire on other people's calendars. The EU AI Act phases its obligations in on fixed dates, so a pattern certified in Q1 2026 can fall out of compliance on August 2, 2026 with zero internal changes — a new obligation tranche simply starts applying to its use case that day. Recertification is a recurring operating cost that appears in no vendor's pitch. Stamp every pre-approved template with a statutory expiry tied to the published phase-in schedule, not just an internal review-by date.

Survivorship bias inflates the eligibility headline. Companies publishing tiering successes tend to run portfolios that were mostly low-risk to begin with — workflow copilots, internal summarizers, cosmetic interface variants. This tracks what experiment-design practitioners like Atticus Li argue: lightweight guardrails such as bounce rate fit low-impact changes because core health metrics stay monitored separately. A portfolio weighted toward consequential automated decisions — credit, hiring, care triage — will see fast-lane eligibility well below the advertised share, and no rubric changes that composition.

Then there is gaming. Once the fast lane earns a rubber-stamp reputation, business teams begin self-scoring aggressively at intake, and measured compliance improves while real risk quietly migrates into Tier 1. The shape is identical to a widely circulated A/B-testing postmortem of a checkout flow optimized purely for speed: conversion rose, and fraud losses erased the gains entirely. Throughput metrics improved; the guardrail metric paid for them. Without randomized audits of tier assignments, an escape-rate dashboard stays clean right up until it isn't.

```

## Frequently Asked Questions

**What numeric caps must an experiment pass before a pilot ever reaches counsel?**

Primary guardrails allow error rate to rise no more than 10%, latency p95 no more than 15%, and crash rate no more than 5%, while surface-specific secondaries tighten further — checkout payment errors capped at 5% growth, abandoned carts at 10%, support tickets at 20%, and search zero-result rate at 5%.

**Does the five-business-day Tier 1 clock start when you first email legal?**

No — five business days runs from intake-complete, meaning the rubric submitted plus every required artifact attached, never from the first email, which otherwise lets sponsors blame delays caused by their own missing DPAs.

**Can a reviewer overrule a colleague who assigned a lower tier?**

Escalations are upward only — any reviewer, including the most junior analyst, can bump a pilot up one tier within 48 hours by attaching a mandatory written reason code, and the queue behind the bumped pilot never reopens.

**Where do pilots land if they touch regulated risk territory under the EU AI Act ladder?**

A pilot brushing unacceptable-risk territory, such as emotion inference in hiring, exits the process entirely rather than queueing, while anything touching high-risk ground — employment screening, creditworthiness, essential services — caps at Tier 2 or climbs to Tier 3 regardless of how benign its raw score looks.

**Is a low rubric score enough to get the five-day fast lane?**

No — Tier 1 eligibility requires the pilot to reuse a previously signed-off pattern, such as a data flow already covered by a completed GDPR Article 35 DPIA or mapped to existing SOC 2 controls, so no pattern means no fast lane whatever the score says.

**Of the roughly six-week committee-lane wait, how much is actually legal judgment?**

Subtracting the ~3.4-week average B2B contract cycle from drafting through signature (per World Commerce & Contracting benchmarking) leaves roughly 2.6 weeks of pure queue-and-scheduling lag — intake triage sitting unread, misaligned calendars, and reviewers waiting on coverage.

## Quick answers

| How many questions and scoring points make up the tiered intake rubric? | The tiered intake uses 12 triage questions across six risk axes with 36 possible scoring points. |
| --- | --- |
| What is the target clock for a Tier 1 pilot and who executes it? | Tier 1 pre-approved template sign-off targets 5 business days and is executed by a trained legal-ops analyst. |
| What are the primary guardrails that cap downside before a pilot reaches counsel? | Error rate up no more than 10%, latency p95 up no more than 15%, and crash rate up no more than 5%. |
| What share of pilots never require counsel judgment at all? | As many as 80% of pilots never require counsel judgment; they need a documented control plane executed by legal ops. |
| What result did compliance-scored guardrails show against unguarded rollouts? | Guarded variants beat controls 301/659 to 536/1548, an 11.0-percentage-point win-rate lift with a 95% confidence interval of [6.6, 15.5]. |

### Related reading

- [The 30% Studio Stake: Speed Premium or $2.4M Giveaway?](https://tlab.fun/blog/the-30-studio-stake-speed-premium-or-24m-giveaway.php)
- [2026 B2B Validation: Concierge Signals vs. Budget Proxies](https://tlab.fun/blog/2026-b2b-validation-concierge-signals-vs-budget-proxies.php)
- [Run 15 Corporate Pilots a Quarter: 3 Gates, 70% Fail](https://tlab.fun/blog/run-15-corporate-pilots-a-quarter-3-gates-70-fail.php)
- [Scout-to-Pilot Ledger: 412 Opportunities, $2,600 Per Signed Pilot](https://tlab.fun/blog/scout-to-pilot-ledger-412-opportunities-2600-per-signed-pilot.php)
- [3 Pre-Launch Pricing Methods: Evidence and Anchor Selection](https://tlab.fun/blog/3-pre-launch-pricing-methods-evidence-and-anchor-selection.php)
- [Day-10 Gate Cuts Legal Rework by 38%: GIMI 2025 Data](https://tlab.fun/blog/day-10-gate-cuts-legal-rework-by-38-gimi-2025-data.php)

### Latest

- [The 30% Studio Stake: Speed Premium or $2.4M Giveaway?](https://tlab.fun/blog/the-30-studio-stake-speed-premium-or-24m-giveaway.php)
- [2026 B2B Validation: Concierge Signals vs. Budget Proxies](https://tlab.fun/blog/2026-b2b-validation-concierge-signals-vs-budget-proxies.php)
- [Run 15 Corporate Pilots a Quarter: 3 Gates, 70% Fail](https://tlab.fun/blog/run-15-corporate-pilots-a-quarter-3-gates-70-fail.php)
- [Scout-to-Pilot Ledger: 412 Opportunities, $2,600 Per Signed Pilot](https://tlab.fun/blog/scout-to-pilot-ledger-412-opportunities-2600-per-signed-pilot.php)

Canonical: https://tlab.fun/blog/guardrail-tiers-12-questions-36-points-15-pilots-a-year.php
Markdown: https://tlab.fun/blog/guardrail-tiers-12-questions-36-points-15-pilots-a-year.php/index.md
