Skip to resource
MHBMMichael Hanna-Butros MeyeringComplex systems · human outcomes
Menu

13 measures · five dimensions · one disposition

Measure AI change that holds up.

A practitioner scorecard that joins reach, retained behavior, task and service outcomes, human control, evidence quality, and rollback—because usage alone cannot prove useful or responsible adoption.

Practice
Public-sector technology · privacy · responsible AI
Updated
August 4, 2026
Use
Source-backed · printable · adaptable
Reusable artifact

Use it in a working session, then attach the evidence behind every answer.

Measurement position

Adoption and control share one scorecard.

The scorecard is a decision aid, not a composite maturity score. Report each measure with its definition, numerator, denominator, baseline, threshold, evidence link, owner, and date.

Decision principle
Do not average away a critical failure. A strong usage or efficiency number cannot offset an open safety, rights, privacy, security, accessibility, continuity, or rollback gap.
Threshold status

Every number below is a transparent starting guardrail from Michael's practitioner framework—not a universal benchmark or legal requirement. Replace it with an owner-approved threshold that reflects the workflow's baseline, risk classification, public impact, frequency, sample size, and applicable authority. Keep the original starting threshold in the change history so a weaker target cannot be introduced silently.

Calculation dictionary

Thirteen measures that can be inspected.

Choose the measures that match the decision, but preserve at least one signal from readiness, behavior, outcome, human control, operating health, evidence completeness, and rollback.

01

Reach + readiness

Did the intended people receive a usable path—not just an announcement?

MeasureFormulaStarting thresholdInterpretation + evidence
Ready reachPeople with required access + role guidance + practice completed ÷ eligible people × 100Green ≥95% · Watch 80–94% · Hold below 80% when broad coverage is requiredSeparate eligible, enabled, trained, and practiced populations; they are not interchangeable.Evidence: Access export, role roster, practice result, and dated eligibility rule.
Support readinessPriority support scenarios with a named owner and tested response ÷ required scenarios × 100100% before unsupervised use; any ownerless critical scenario is a holdReadiness includes exception and incident support, not only how-to documentation.Evidence: Scenario list, owner, rehearsal record, response target, and escalation path.
02

Activation + retained use

Are people using the approved path when the workflow calls for it?

MeasureFormulaStarting thresholdInterpretation + evidence
Meaningful activationEnabled people completing the first intended task correctly ÷ enabled people × 100Green ≥70% in 30 days · Watch 40–69% · Investigate below 40%A login, page view, or training attendance does not count as activation.Evidence: Task event, correctness rule, unique person or role, and approved workflow identifier.
Four-week retained useActivated people completing an approved task in week 4 ÷ activated people still eligible in week 4 × 100Green ≥60% · Watch 35–59% · Investigate below 35%; adjust for task frequencyFor infrequent work, measure the next eligible workflow event instead of an arbitrary week.Evidence: Eligibility snapshot plus approved task events at activation and the return window.
Approved-path shareObserved eligible workflow instances using the approved AI or non-AI path ÷ sampled eligible instances × 100Green ≥95% · Watch 85–94% · Hold scale below 85% until shadow paths are understoodThe goal is observable, supported work—not forcing AI into every eligible task.Evidence: Representative sample, workflow definition, sanctioned alternatives, and shadow-use review.
03

Task + service outcome

Did the change improve the work without degrading quality, access, or public value?

MeasureFormulaStarting thresholdInterpretation + evidence
Task successTasks completed correctly without material rework ÷ observed AI-assisted task attempts × 100At or above the non-AI baseline; hold if quality drops by more than 2 percentage pointsSet the correctness rubric before reviewing results and stratify by task and affected group.Evidence: Blind or consistent review rubric, representative cases, baseline, exceptions, and rework record.
Median cycle-time improvement(Baseline median minutes − AI-assisted median minutes) ÷ baseline median minutes × 100Scale signal ≥15% only when task success and control measures stay at or above baselineUse medians to reduce outlier distortion; report queue time separately from hands-on time.Evidence: Comparable start and stop events, baseline period, case mix, sample size, and quality result.
Outcome disparityLargest group task-success rate − smallest group task-success rate, in percentage pointsInvestigate any new gap ≥5 points; hold for material legal, civil-rights, or service-access concern regardless of sizeUse only lawful, appropriate, privacy-protective group analysis with qualified review.Evidence: Approved grouping method, sample caveats, rates and denominators, impact review, and corrective action.
04

Human control + exceptions

Can people recognize, interrupt, correct, and contest the system in real work?

MeasureFormulaStarting thresholdInterpretation + evidence
Override + escalation rateReviewed instances overridden or escalated by a person ÷ reviewed instances × 100Review at ≥2× baseline or a ≥5-point increase; do not assume lower is always betterA rise may reveal drift, poor guidance, harder case mix, or a healthy control culture.Evidence: Reason-coded override, decision context, reviewer, outcome, and correction record.
Correction / appeal completionCorrection or appeal cases closed with notice and disposition ÷ cases due in the period × 100100% of due high-impact cases; no overdue case without a named owner and noticeTrack acknowledgment, resolution time, remedy, and whether the underlying system changed.Evidence: Case intake, owner, due date, notice, disposition, remedy, and systemic follow-up.
05

Operating health + evidence

Is the service stable, supportable, auditable, and recoverable?

MeasureFormulaStarting thresholdInterpretation + evidence
Material incident rateMaterial AI-related incidents ÷ production workflow instances × 1,000Contain on any severity-1 or severity-2 event; review any increase above the approved baselineDefine severity and attribution before launch. Never wait for a rate threshold to contain serious harm.Evidence: Incident definition, case count, denominator, severity, affected people, containment, and disposition.
Six-receipt completenessSampled decisions with all six required receipts ÷ sampled decisions × 100Green ≥95% · Watch 80–94% · Hold scale below 80%Completeness does not prove quality; separately review whether the evidence supports the decision.Evidence: Signal, boundary, workflow, readiness, outcome/control, and rollback receipts linked to one decision ID.
Rollback readinessRequired rollback steps successfully executed within target time ÷ required steps × 100100% before unsupervised scale; any failed critical step is a holdA written rollback plan is not readiness. Witness revocation, correction, recovery, and communication.Evidence: Test date, actors, actual steps, elapsed time, gaps, retest, and final disposition.

Operating disposition

End the review with a verb.

A dashboard without a decision owner becomes observation theater. Record one state, its boundary, evidence, residual risk, approver, date, and next trigger.

Scale

No critical hold; outcome is at or above baseline; evidence completeness is at least 95%; rollback is 100%; owner accepts the residual risk for a named boundary and review period.

Proceed narrowly

No critical hold, but one or more watch signals need a limited population, closer supervision, shorter review window, or additional evidence.

Revise

The service case remains valid but the workflow, guidance, control, model, data boundary, vendor term, or measurement method must change before expansion.

Pause / contain

A material safety, rights, privacy, security, accessibility, continuity, evidence, or ownership gap is open. Stop the affected path while preserving records and support.

Retire

The service value is not sustained, a safer alternative is better, the system cannot remain inside its boundary, or the cost and control burden are not justified.

Measurement hygiene

Make every number reproducible.

A credible scorecard exposes how the number was built and where it can mislead. Write these notes next to the metric—not in a separate methodology no operator sees.

Freeze the boundary first.

Put the workflow, population, time window, model or configuration version, approved data, and connected actions on the scorecard. A changing denominator can make improvement or regression disappear.

Use a real baseline.

Compare the AI-assisted path with the prior safe process or a lawful concurrent comparison. Record case mix, sample size, missing data, and seasonality; a target without a baseline is only a preference.

Keep counts beside percentages.

Show numerator and denominator next to every rate. A 100% result from two cases should not carry the same confidence as 1,000 representative cases.

Stratify material outcomes.

Review task type, channel, role, accessibility path, language, geography, and affected group when lawful and appropriate. Aggregate performance can hide who absorbs the errors or burden.

Measure the safe alternative too.

Approved non-AI work is not adoption failure. Track whether people can choose or return to the safe path without losing service, support, records, or appeal rights.

Protect people in the measurement layer.

Collect the minimum data needed, set access and retention, avoid covert employee surveillance, and involve privacy, labor, civil-rights, records, security, and accessibility owners as applicable.

Reusable artifact

Copy the working scorecard.

One row per measure is enough when the evidence link is real. Add a separate risk or compliance artifact where your organization requires it; do not force every discipline into one spreadsheet.

AI CHANGE MANAGEMENT SCORECARD

Workflow / decision ID:
Owner:
Baseline period:
Review period:
Population / boundary:

METRIC | BASELINE | CURRENT | THRESHOLD | STATUS | EVIDENCE LINK | OWNER
Ready reach | | | | | |
Support readiness | | | | | |
Meaningful activation | | | | | |
Retained use | | | | | |
Approved-path share | | | | | |
Task success | | | | | |
Median cycle-time improvement | | | | | |
Outcome disparity | | | | | |
Override + escalation rate | | | | | |
Correction / appeal completion | | | | | |
Material incident rate | | | | | |
Six-receipt completeness | | | | | |
Rollback readiness | | | | | |

CRITICAL HOLD OPEN? [ ] No [ ] Yes — name it:
DISPOSITION: [ ] Scale [ ] Proceed narrowly [ ] Revise [ ] Pause / contain [ ] Retire
Decision owner / date:
Residual risk and boundary:
Next trigger / review date:
Evidence index:

Method, limits + sources

A working aid—not a substitute for accountable review.

Michael’s practitioner synthesis connects operating change, public-sector delivery, and the six receipts before scale. Every organization remains responsible for applying its own authority, expertise, evidence, and risk tolerance.

Limitations

  • This is practitioner guidance, not legal, audit, labor-relations, procurement, records, privacy, security, civil-rights, or accessibility advice.
  • The starting thresholds are operating guardrails, not universal benchmarks. Replace them with the applicable law, policy, risk classification, service baseline, collective-bargaining obligation, and tolerance approved by your organization.
  • A completed template is not evidence by itself. Attach source records, test results, approvals, observed outcomes, and a final disposition.
  • Do not average away a critical failure. A material safety, rights, privacy, security, accessibility, or mission-continuity gap remains a stop condition even when the overall score looks strong.

Primary and public sources

  1. Artificial Intelligence Risk Management Framework (AI RMF 1.0)National Institute of Standards and TechnologyVoluntary, rights-preserving framework for governing, mapping, measuring, and managing AI risk.
  2. NIST AI RMF PlaybookNational Institute of Standards and TechnologySuggested actions and documentation practices; NIST explicitly describes it as neither a universal checklist nor an ordered set of steps.
  3. Artificial Intelligence: An Accountability Framework for Federal Agencies and Other EntitiesU.S. Government Accountability OfficeAccountability practices organized around governance, data, performance, and monitoring.
  4. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence ProfileNational Institute of Standards and TechnologyCompanion profile for risks that are distinctive to or intensified by generative AI.

Update history

Versioned in public.

  1. Initial publication: 13 formulas across five dimensions, explicit starting thresholds, evidence requirements, decision rules, calculation guidance, reusable template, primary sources, and limitations.