"We assigned 100 licenses, so the rollout delivered value." That sentence skips most of the work. A license was assigned, someone may have logged in, a scenario may have been repeated, and a business result may have improved. Each claim needs different evidence.
This lesson follows a small service-operations team using Copilot for weekly leadership updates. We will choose the scenario, label the measurements, inspect each seat, and make the expansion decision from the evidence available at the end.
A pilot scenario has to clear three gates
A scenario describes a bounded work outcome. "Use Copilot more" is too vague. "Prepare a fact-checked weekly leadership update from approved operations notes, with the pilot manager verifying names, dates, and figures" names the persona, trigger, approved source, expected output, human decision, and measure of improvement.
Score candidates locally from 0 to 2 across business value, frequency, readiness, adoption potential, and risk control. Add a note showing what supports every score. After ranking them, apply three hard gates separately:
- A named owner exists.
- The source material is approved and accessible.
- The output has a defined human-review step.
A failed gate removes the candidate from the pilot, whatever its total. The weekly leadership update has an owner, approved notes, and a review step, so it remains eligible. The HR policy-question agent may score well, but an unapproved source and unknown participant access defer it. The project-portfolio summary has ready sources, yet nobody accepts ownership and there is no review checklist. It fails two gates. The weekly update enters the pilot because it ranks highest among the candidates that meet all three conditions.
Keep the label attached to every number
Once the pilot starts, the main risk is misreading what each figure can support. Label every number in the measurement pack as one of four types:
- Observed: recorded directly in a dashboard, system, or survey (active users rose from 8 to 14).
- Calculated: derived from observed values with visible arithmetic (an active-user rate of 14 ÷ 20 = 70%).
- Modeled: dependent on an assumption (annual value using a 50% realization rate and an hourly value).
- External: published evidence from another population (a Forrester study's result).
These sit on a chain: readiness, activation, adoption, impact, value. Usage is evidence that Copilot was used. Impact is evidence that work changed. ROI is a financial interpretation of that change. Keep them apart. Assisted hours going up is an impact signal, not cash in the bank. Converting it to money requires an explicit, owner-approved realization rate and an evidence bridge showing how those hours become usable capacity.
Watch how careful a modeled figure has to be. Say the team handles 240 weekly-update artifacts a month and preparation dropped by 11 minutes each. Under a modeled 50% realization rate and an illustrative $40 hourly value (not a Microsoft price), the model gives 240 × 11 ÷ 60 × 50% × $40 × 12 ≈ $10,560 in annual value, and against a $9,600 modeled program cost, about a 10% modeled ROI. Every load-bearing word there is "modeled." Present it as a scenario, not a savings account, and never compare it head-to-head with an external benchmark like Forrester's three-year, risk-adjusted 116% ROI, because the populations, horizons, and methods are different.
License decisions need human evidence
Measurement isn't only a matter of aggregate rates. It's also a per-seat audit that turns "who has a license" into "who has a purpose." Classify each assigned user:
- Keep: a relevant scenario and owner exist, with current use or active-practice evidence.
- Coach: a scenario and owner exist, but access, skill, trust, or reinforcement needs help.
- Reassign: no supported scenario or owner, and a named approved recipient is waiting with a real need.
- Hold: the seat belongs to a documented future cohort with an owner and a future start date.
"We might need the seat later" is not a Hold. A Hold has a name and a date. And when telemetry shows someone stopped using Copilot, resist inventing why. Interview a sample, and record everyone you couldn't interview as unknown—not interviewed. Telemetry locates a possible problem. It never explains one.
When can you expand?
Now the decision. Expansion requires a five-gate check. Mark each gate Pass, Fail, or Unknown. There is no averaging or partial credit:
Then decide. Expand if all five pass. Repair if Scenario passes but another gate fails or is unknown, giving each gap an owner and review date. Pause if Scenario fails or no owner accepts the next cycle.
Back to our team. After guided practice, repeat use held and second-day outputs passed the checklist, but one active-cohort access issue was still open and no support capacity was scheduled for a bigger group. Scenario: Pass. Repetition: Pass. Measurement: Pass. Access: Fail. Support: Fail. The correct decision is Repair: keep the current cohort, resolve access, schedule support, and run the gates again. Three passing gates and a positive mood do not authorize expansion. One more caution while you diagnose: a second-week drop in use is a practitioner warning pattern worth investigating. It does not establish that every pilot must decline in week two.
A note on agent scenarios
When the scenario is an agent rather than a prompt, the same discipline gets a four-word frame: Purpose, Permission, Practice, Proof. Purpose is the work problem. Permission is the access, controls, and approved source boundary. Practice is repeated attempts with feedback. Proof is recorded evidence that it's useful and safe. The order matters. A visible agent is not proof of permission: if participant access is unconfirmed and the source isn't approved, the scenario stays paused until those are recorded, no matter how good the demo looked. Enthusiasm and attendance are not substitutes for confirmed access, repeat practice, and reviewed outputs.