Scenarios, Measurement, and Anti-Patterns

Pick a scenario with hard gates, measure it with disciplined evidence labels, and clear five expansion gates, so you never mistake assigned licenses or a busy dashboard for value.


What you'll learn

  • Prioritize scenarios with a local rubric and three non-negotiable hard gates
  • Label evidence as observed, calculated, modeled, or external without conflating them
  • Audit licenses with Keep, Coach, Reassign, and Hold decisions
  • Decide expand, repair, or pause from five gates. Usage numbers and a calendar date don't settle it.
On this page

"We assigned 100 licenses, so the rollout delivered value." That sentence skips most of the work. A license was assigned, someone may have logged in, a scenario may have been repeated, and a business result may have improved. Each claim needs different evidence.

This lesson follows a small service-operations team using Copilot for weekly leadership updates. We will choose the scenario, label the measurements, inspect each seat, and make the expansion decision from the evidence available at the end.

A pilot scenario has to clear three gates

A scenario describes a bounded work outcome. "Use Copilot more" is too vague. "Prepare a fact-checked weekly leadership update from approved operations notes, with the pilot manager verifying names, dates, and figures" names the persona, trigger, approved source, expected output, human decision, and measure of improvement.

Score candidates locally from 0 to 2 across business value, frequency, readiness, adoption potential, and risk control. Add a note showing what supports every score. After ranking them, apply three hard gates separately:

  1. A named owner exists.
  2. The source material is approved and accessible.
  3. The output has a defined human-review step.

A failed gate removes the candidate from the pilot, whatever its total. The weekly leadership update has an owner, approved notes, and a review step, so it remains eligible. The HR policy-question agent may score well, but an unapproved source and unknown participant access defer it. The project-portfolio summary has ready sources, yet nobody accepts ownership and there is no review checklist. It fails two gates. The weekly update enters the pilot because it ranks highest among the candidates that meet all three conditions.

A score never beats a failed gate

When a stakeholder pushes a favorite scenario that scored 9 out of 10 but has no approved source, the answer is not "close enough." Return it to readiness work and preserve the failed-gate evidence. Lowering a threshold to let it through is how unsafe scenarios enter pilots.

Test a scenario with a visible source boundary· copilot-chat
Bad example

Turn these operations notes into a polished weekly leadership update.

Good example

You are helping prepare a weekly leadership update. Use only the supplied operations notes. Do not infer decisions, sentiment, deadlines, or commitments that aren't stated. Return five sections: Progress, Delays, Customer impact, Decisions needed, Next actions. Preserve named owners and dates, and write "Needs confirmation" for anything the notes don't establish. Then list any statement that isn't traceable to the notes.

Why this works: A grounded draft where every gap shows as "Needs confirmation" instead of a confident guess, plus a list you can check before anyone relies on it.

Keep the label attached to every number

Once the pilot starts, the main risk is misreading what each figure can support. Label every number in the measurement pack as one of four types:

  • Observed: recorded directly in a dashboard, system, or survey (active users rose from 8 to 14).
  • Calculated: derived from observed values with visible arithmetic (an active-user rate of 14 ÷ 20 = 70%).
  • Modeled: dependent on an assumption (annual value using a 50% realization rate and an hourly value).
  • External: published evidence from another population (a Forrester study's result).

These sit on a chain: readiness, activation, adoption, impact, value. Usage is evidence that Copilot was used. Impact is evidence that work changed. ROI is a financial interpretation of that change. Keep them apart. Assisted hours going up is an impact signal, not cash in the bank. Converting it to money requires an explicit, owner-approved realization rate and an evidence bridge showing how those hours become usable capacity.

Watch how careful a modeled figure has to be. Say the team handles 240 weekly-update artifacts a month and preparation dropped by 11 minutes each. Under a modeled 50% realization rate and an illustrative $40 hourly value (not a Microsoft price), the model gives 240 × 11 ÷ 60 × 50% × $40 × 12 ≈ $10,560 in annual value, and against a $9,600 modeled program cost, about a 10% modeled ROI. Every load-bearing word there is "modeled." Present it as a scenario, not a savings account, and never compare it head-to-head with an external benchmark like Forrester's three-year, risk-adjusted 116% ROI, because the populations, horizons, and methods are different.

License decisions need human evidence

Measurement isn't only a matter of aggregate rates. It's also a per-seat audit that turns "who has a license" into "who has a purpose." Classify each assigned user:

  • Keep: a relevant scenario and owner exist, with current use or active-practice evidence.
  • Coach: a scenario and owner exist, but access, skill, trust, or reinforcement needs help.
  • Reassign: no supported scenario or owner, and a named approved recipient is waiting with a real need.
  • Hold: the seat belongs to a documented future cohort with an owner and a future start date.

"We might need the seat later" is not a Hold. A Hold has a name and a date. And when telemetry shows someone stopped using Copilot, resist inventing why. Interview a sample, and record everyone you couldn't interview as unknown—not interviewed. Telemetry locates a possible problem. It never explains one.

Audit seats against their evidence· excel
Bad example

Review this Copilot license list and tell me who should keep a seat.

Good example

Using this fictional seat table (columns user, scenario, owner, use evidence or barrier, future start date, approved alternate), classify each user Keep, Coach, Reassign, or Hold, with a one-line reason. Require a future cohort, owner, and start date before any Hold, and a named approved recipient before any Reassign. For any user with no direct evidence, write "unknown—not interviewed" rather than guessing a cause. [Paste the table.]

Why this works: A seat-by-seat audit where every Hold and Reassign is backed by the evidence its rule requires, not by seniority, politics, or a telemetry guess.

Draft a measurement brief that won't overclaim· copilot-chat
Bad example

Write a positive summary showing that our Copilot pilot delivered value.

Good example

Act as an adoption analyst. Using only the figures and definitions below, draft a one-page measurement brief with headings Business question, Baseline, Observed usage, Business outcome signals, Financial model, External benchmarks, and Limitations. Preserve every number exactly and label each one observed, calculated, modeled, or external. Do not claim causation. If evidence is missing, write "evidence not provided" instead of estimating. [Paste your figures.]

Why this works: A brief where every number carries its evidence label and no correlation is dressed up as cause: exactly what survives a skeptical finance review.

When can you expand?

Now the decision. Expansion requires a five-gate check. Mark each gate Pass, Fail, or Unknown. There is no averaging or partial credit:

Gate Passes when
Scenario A named task, permitted source, expected output, owner, and human-verification checklist exist
Repetition At least two users completed the scenario on two separate days, each later output passing the same checklist
Access No unresolved access issue remains in the active cohort
Support A named contact, coverage, and a maximum supported cohort size are recorded, and the proposed cohort fits within it
Measurement Comparable periods and populations are documented. Telemetry, interviews, and unknowns are kept separate

Then decide. Expand if all five pass. Repair if Scenario passes but another gate fails or is unknown, giving each gap an owner and review date. Pause if Scenario fails or no owner accepts the next cycle.

Back to our team. After guided practice, repeat use held and second-day outputs passed the checklist, but one active-cohort access issue was still open and no support capacity was scheduled for a bigger group. Scenario: Pass. Repetition: Pass. Measurement: Pass. Access: Fail. Support: Fail. The correct decision is Repair: keep the current cohort, resolve access, schedule support, and run the gates again. Three passing gates and a positive mood do not authorize expansion. One more caution while you diagnose: a second-week drop in use is a practitioner warning pattern worth investigating. It does not establish that every pilot must decline in week two.

Don't let a fluent answer skip the source check

The most dangerous Copilot output is a polished one built on an invented fact: a plausible owner the notes never named, a delivery date no source stated. Whether the scenario is a chat prompt or an agent, require a named source, explicit uncertainty ("Needs confirmation"), and a human decision before use.

A note on agent scenarios

When the scenario is an agent rather than a prompt, the same discipline gets a four-word frame: Purpose, Permission, Practice, Proof. Purpose is the work problem. Permission is the access, controls, and approved source boundary. Practice is repeated attempts with feedback. Proof is recorded evidence that it's useful and safe. The order matters. A visible agent is not proof of permission: if participant access is unconfirmed and the source isn't approved, the scenario stays paused until those are recorded, no matter how good the demo looked. Enthusiasm and attendance are not substitutes for confirmed access, repeat practice, and reviewed outputs.

Try it yourself

Audit a stalled cohort and call the decision

Six people hold licenses in a fictional finance pilot. Classify each seat and then judge whether the pilot can expand. Aim for five to ten minutes.

  1. 01

    For each user, assign Keep, Coach, Reassign, or Hold. Example set. Amina: scenario and owner, two verified attempts. Ben: no scenario or owner, but Grace is an approved alternate with a real need. Carla: scenario and owner, unresolved access issue. Farah: future HR cohort with an owner and an August start date.

    Hint: Reassign needs a named approved recipient. Hold needs a future cohort, owner, and date.

  2. 02

    One active-cohort access issue is unresolved and no support capacity is scheduled for a larger group. Mark each of the five expansion gates Pass, Fail, or Unknown.

  3. 03

    State the decision (Expand, Repair, or Pause) and name the owner and review date for each failed gate.

A seat-by-seat classification and a gate-based expansion decision (it should be Repair) that another reviewer could defend from the evidence alone.

Key takeaways

  • A scenario needs a persona, trigger, approved source, output, owner, and review step. A high score never overrides a failed hard gate.
  • Label every figure observed, calculated, modeled, or external. Usage, impact, and ROI are different claims.
  • Assisted hours are an impact signal, not cash. Monetize only through an explicit, owner-approved realization rate.
  • Classify seats Keep, Coach, Reassign, or Hold, and record users you couldn't interview as unknown—not interviewed.
  • Expand only when Scenario, Repetition, Access, Support, and Measurement all pass. Otherwise Repair or Pause.

Check your understanding

  1. 1. Candidate A scores 9 out of 10 but has no approved, accessible source. Candidate B scores 8 out of 10 with a named owner, an approved source, and a defined human-review step. Which may enter the pilot?

  2. 2. Annual value is estimated using a 50% realization rate and an assumed hourly value. How should that figure be labeled?

  3. 3. A sponsor says, "We assigned 100 licenses, so the rollout delivered value." What is the mistake?

  4. 4. A local model produces 85% ROI while an external Forrester TEI study reports 116% ROI. Which interpretation is defensible?

  5. 5. A pilot's Scenario, Repetition, and Measurement gates pass, but one active-cohort access issue is unresolved and no support capacity is approved for a larger cohort. What is the correct decision?

  6. 6. An agent scenario has a useful purpose and an assigned reviewer, but participant access is unconfirmed and the proposed source is unapproved. What should the adoption lead do next?

Frequently asked questions

Terms used in this lesson

hard gate
A pass/fail requirement (named owner, approved and accessible source, defined human review) that a total score cannot override.
realization rate
The share of a modeled benefit, such as assisted hours, that an owner permits you to count as value.
expansion gates
The five checks (Scenario, Repetition, Access, Support, Measurement) that must all pass before a pilot expands.
unknown—not interviewed
The label for a user whose change in use has no direct explanation, kept separate from any inferred cause.

Further reading