13 measures · five dimensions · one disposition
Measure AI change that holds up.
A practitioner scorecard that joins reach, retained behavior, task and service outcomes, human control, evidence quality, and rollback—because usage alone cannot prove useful or responsible adoption.
- Practice
- Public-sector technology · privacy · responsible AI
- Updated
- August 4, 2026
- Use
- Source-backed · printable · adaptable
Use it in a working session, then attach the evidence behind every answer.
Measurement position
Adoption and control share one scorecard.
The scorecard is a decision aid, not a composite maturity score. Report each measure with its definition, numerator, denominator, baseline, threshold, evidence link, owner, and date.
Do not average away a critical failure. A strong usage or efficiency number cannot offset an open safety, rights, privacy, security, accessibility, continuity, or rollback gap.
Every number below is a transparent starting guardrail from Michael's practitioner framework—not a universal benchmark or legal requirement. Replace it with an owner-approved threshold that reflects the workflow's baseline, risk classification, public impact, frequency, sample size, and applicable authority. Keep the original starting threshold in the change history so a weaker target cannot be introduced silently.
Calculation dictionary
Thirteen measures that can be inspected.
Choose the measures that match the decision, but preserve at least one signal from readiness, behavior, outcome, human control, operating health, evidence completeness, and rollback.
Reach + readiness
Did the intended people receive a usable path—not just an announcement?
| Measure | Formula | Starting threshold | Interpretation + evidence |
|---|---|---|---|
| Ready reach | People with required access + role guidance + practice completed ÷ eligible people × 100 | Green ≥95% · Watch 80–94% · Hold below 80% when broad coverage is required | Separate eligible, enabled, trained, and practiced populations; they are not interchangeable.Evidence: Access export, role roster, practice result, and dated eligibility rule. |
| Support readiness | Priority support scenarios with a named owner and tested response ÷ required scenarios × 100 | 100% before unsupervised use; any ownerless critical scenario is a hold | Readiness includes exception and incident support, not only how-to documentation.Evidence: Scenario list, owner, rehearsal record, response target, and escalation path. |
Activation + retained use
Are people using the approved path when the workflow calls for it?
| Measure | Formula | Starting threshold | Interpretation + evidence |
|---|---|---|---|
| Meaningful activation | Enabled people completing the first intended task correctly ÷ enabled people × 100 | Green ≥70% in 30 days · Watch 40–69% · Investigate below 40% | A login, page view, or training attendance does not count as activation.Evidence: Task event, correctness rule, unique person or role, and approved workflow identifier. |
| Four-week retained use | Activated people completing an approved task in week 4 ÷ activated people still eligible in week 4 × 100 | Green ≥60% · Watch 35–59% · Investigate below 35%; adjust for task frequency | For infrequent work, measure the next eligible workflow event instead of an arbitrary week.Evidence: Eligibility snapshot plus approved task events at activation and the return window. |
| Approved-path share | Observed eligible workflow instances using the approved AI or non-AI path ÷ sampled eligible instances × 100 | Green ≥95% · Watch 85–94% · Hold scale below 85% until shadow paths are understood | The goal is observable, supported work—not forcing AI into every eligible task.Evidence: Representative sample, workflow definition, sanctioned alternatives, and shadow-use review. |
Task + service outcome
Did the change improve the work without degrading quality, access, or public value?
| Measure | Formula | Starting threshold | Interpretation + evidence |
|---|---|---|---|
| Task success | Tasks completed correctly without material rework ÷ observed AI-assisted task attempts × 100 | At or above the non-AI baseline; hold if quality drops by more than 2 percentage points | Set the correctness rubric before reviewing results and stratify by task and affected group.Evidence: Blind or consistent review rubric, representative cases, baseline, exceptions, and rework record. |
| Median cycle-time improvement | (Baseline median minutes − AI-assisted median minutes) ÷ baseline median minutes × 100 | Scale signal ≥15% only when task success and control measures stay at or above baseline | Use medians to reduce outlier distortion; report queue time separately from hands-on time.Evidence: Comparable start and stop events, baseline period, case mix, sample size, and quality result. |
| Outcome disparity | Largest group task-success rate − smallest group task-success rate, in percentage points | Investigate any new gap ≥5 points; hold for material legal, civil-rights, or service-access concern regardless of size | Use only lawful, appropriate, privacy-protective group analysis with qualified review.Evidence: Approved grouping method, sample caveats, rates and denominators, impact review, and corrective action. |
Human control + exceptions
Can people recognize, interrupt, correct, and contest the system in real work?
| Measure | Formula | Starting threshold | Interpretation + evidence |
|---|---|---|---|
| Override + escalation rate | Reviewed instances overridden or escalated by a person ÷ reviewed instances × 100 | Review at ≥2× baseline or a ≥5-point increase; do not assume lower is always better | A rise may reveal drift, poor guidance, harder case mix, or a healthy control culture.Evidence: Reason-coded override, decision context, reviewer, outcome, and correction record. |
| Correction / appeal completion | Correction or appeal cases closed with notice and disposition ÷ cases due in the period × 100 | 100% of due high-impact cases; no overdue case without a named owner and notice | Track acknowledgment, resolution time, remedy, and whether the underlying system changed.Evidence: Case intake, owner, due date, notice, disposition, remedy, and systemic follow-up. |
Operating health + evidence
Is the service stable, supportable, auditable, and recoverable?
| Measure | Formula | Starting threshold | Interpretation + evidence |
|---|---|---|---|
| Material incident rate | Material AI-related incidents ÷ production workflow instances × 1,000 | Contain on any severity-1 or severity-2 event; review any increase above the approved baseline | Define severity and attribution before launch. Never wait for a rate threshold to contain serious harm.Evidence: Incident definition, case count, denominator, severity, affected people, containment, and disposition. |
| Six-receipt completeness | Sampled decisions with all six required receipts ÷ sampled decisions × 100 | Green ≥95% · Watch 80–94% · Hold scale below 80% | Completeness does not prove quality; separately review whether the evidence supports the decision.Evidence: Signal, boundary, workflow, readiness, outcome/control, and rollback receipts linked to one decision ID. |
| Rollback readiness | Required rollback steps successfully executed within target time ÷ required steps × 100 | 100% before unsupervised scale; any failed critical step is a hold | A written rollback plan is not readiness. Witness revocation, correction, recovery, and communication.Evidence: Test date, actors, actual steps, elapsed time, gaps, retest, and final disposition. |
Operating disposition
End the review with a verb.
A dashboard without a decision owner becomes observation theater. Record one state, its boundary, evidence, residual risk, approver, date, and next trigger.
Scale
No critical hold; outcome is at or above baseline; evidence completeness is at least 95%; rollback is 100%; owner accepts the residual risk for a named boundary and review period.
Proceed narrowly
No critical hold, but one or more watch signals need a limited population, closer supervision, shorter review window, or additional evidence.
Revise
The service case remains valid but the workflow, guidance, control, model, data boundary, vendor term, or measurement method must change before expansion.
Pause / contain
A material safety, rights, privacy, security, accessibility, continuity, evidence, or ownership gap is open. Stop the affected path while preserving records and support.
Retire
The service value is not sustained, a safer alternative is better, the system cannot remain inside its boundary, or the cost and control burden are not justified.
Measurement hygiene
Make every number reproducible.
A credible scorecard exposes how the number was built and where it can mislead. Write these notes next to the metric—not in a separate methodology no operator sees.
Freeze the boundary first.
Put the workflow, population, time window, model or configuration version, approved data, and connected actions on the scorecard. A changing denominator can make improvement or regression disappear.
Use a real baseline.
Compare the AI-assisted path with the prior safe process or a lawful concurrent comparison. Record case mix, sample size, missing data, and seasonality; a target without a baseline is only a preference.
Keep counts beside percentages.
Show numerator and denominator next to every rate. A 100% result from two cases should not carry the same confidence as 1,000 representative cases.
Stratify material outcomes.
Review task type, channel, role, accessibility path, language, geography, and affected group when lawful and appropriate. Aggregate performance can hide who absorbs the errors or burden.
Measure the safe alternative too.
Approved non-AI work is not adoption failure. Track whether people can choose or return to the safe path without losing service, support, records, or appeal rights.
Protect people in the measurement layer.
Collect the minimum data needed, set access and retention, avoid covert employee surveillance, and involve privacy, labor, civil-rights, records, security, and accessibility owners as applicable.
Reusable artifact
Copy the working scorecard.
One row per measure is enough when the evidence link is real. Add a separate risk or compliance artifact where your organization requires it; do not force every discipline into one spreadsheet.
AI CHANGE MANAGEMENT SCORECARD Workflow / decision ID: Owner: Baseline period: Review period: Population / boundary: METRIC | BASELINE | CURRENT | THRESHOLD | STATUS | EVIDENCE LINK | OWNER Ready reach | | | | | | Support readiness | | | | | | Meaningful activation | | | | | | Retained use | | | | | | Approved-path share | | | | | | Task success | | | | | | Median cycle-time improvement | | | | | | Outcome disparity | | | | | | Override + escalation rate | | | | | | Correction / appeal completion | | | | | | Material incident rate | | | | | | Six-receipt completeness | | | | | | Rollback readiness | | | | | | CRITICAL HOLD OPEN? [ ] No [ ] Yes — name it: DISPOSITION: [ ] Scale [ ] Proceed narrowly [ ] Revise [ ] Pause / contain [ ] Retire Decision owner / date: Residual risk and boundary: Next trigger / review date: Evidence index:
Method, limits + sources
A working aid—not a substitute for accountable review.
Michael’s practitioner synthesis connects operating change, public-sector delivery, and the six receipts before scale. Every organization remains responsible for applying its own authority, expertise, evidence, and risk tolerance.
Limitations
- This is practitioner guidance, not legal, audit, labor-relations, procurement, records, privacy, security, civil-rights, or accessibility advice.
- The starting thresholds are operating guardrails, not universal benchmarks. Replace them with the applicable law, policy, risk classification, service baseline, collective-bargaining obligation, and tolerance approved by your organization.
- A completed template is not evidence by itself. Attach source records, test results, approvals, observed outcomes, and a final disposition.
- Do not average away a critical failure. A material safety, rights, privacy, security, accessibility, or mission-continuity gap remains a stop condition even when the overall score looks strong.
Primary and public sources
- Artificial Intelligence Risk Management Framework (AI RMF 1.0)National Institute of Standards and TechnologyVoluntary, rights-preserving framework for governing, mapping, measuring, and managing AI risk.
- NIST AI RMF PlaybookNational Institute of Standards and TechnologySuggested actions and documentation practices; NIST explicitly describes it as neither a universal checklist nor an ordered set of steps.
- Artificial Intelligence: An Accountability Framework for Federal Agencies and Other EntitiesU.S. Government Accountability OfficeAccountability practices organized around governance, data, performance, and monitoring.
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence ProfileNational Institute of Standards and TechnologyCompanion profile for risks that are distinctive to or intensified by generative AI.
Update history
Versioned in public.
- Initial publication: 13 formulas across five dimensions, explicit starting thresholds, evidence requirements, decision rules, calculation guidance, reusable template, primary sources, and limitations.