Automation field notes

How to Measure Review Time and Exceptions in an AI Workflow

Time the whole record cohort, not just successful model actions. Separate routine review, escalations, missed errors and recovery work to see whether AI actually releases staff capacity.

What should the measurement answer?

Human review time is the sum of work created by the proposed workflow, not merely the minutes a reviewer spends looking at a clean AI draft. Compare that total with the time the same team needs to complete comparable records manually. Include records the automation rejects, fails to process or sends back for human handling. The difference is released or added staff capacity; it is not automatically a cash saving. A vendor dashboard can start a conversation, but it cannot measure your reviewer burden: Zapier Enterprise Analytics, for example, documents a default estimate of two minutes saved per successful action step and lets an account owner adjust it. That unit is a task, not necessarily a completed business record.

  • Use a fixed observation window and one defined output, such as a verified invoice draft or a classified request; do not mix several workflows in one average.
  • Keep a count of all eligible arrivals and a separate count of completed, failed, skipped and manually handled records. Multiple actions on one record do not enlarge the record denominator.
  • Have the process owner approve the correctness standard and identify which errors would make release unsafe.

Further reading: Use the broader ROI model after timing the work

Fix the denominator and time a manual baseline

For a fair comparison, define N as the number of eligible business records arriving during the observation window, including duplicates, poor-quality inputs and records that need clarification unless the written scope excludes them before either path is timed. Record excluded types and counts separately. Time a representative manual batch from receipt through verified output, including ordinary checking and corrections. Use comparable cases, staff roles and service conditions for the proposed path; otherwise changes in case mix or staffing masquerade as savings. Report totals and minutes per eligible record, not only the median of easy successes.

  • Baseline fields: date or batch, anonymous case ID, input type, eligible or prespecified exclusion, active handling minutes, routine verification minutes, correction minutes, severity of any error, and final disposition.
  • Record elapsed waiting time separately from active staff time. A two-day wait for an approval is not two days of labor, although it may matter for turnaround.
  • If cases vary materially, stratify by type or difficulty and report each group's denominator. Preserve the original mix when combining groups.
  • Choose consecutive arrivals or a prespecified randomized sample across ordinary and busy periods; document any missing observations and do not extrapolate a convenience sample of clean records.

Use one privacy-safe review ledger

Log an event once under the work it actually caused. All-record review means the first routine inspection applied to every AI-handled draft; escalated review is additional investigation for flagged or suspicious cases, not that first inspection again. A missed error is an incorrect output that passed routine review and was found in a separate audit or downstream; count its correction and any remediation without also assigning those minutes to the initial exception. Recovery from a failed run belongs in replay time, while staff time spent locating context, switching tasks or chasing an alert belongs in interruption time. An identifier and category suffice for analysis; do not copy customer names, document contents or sensitive fields into the worksheet.

Blank time-study columns; enter counts and active minutes for the same eligible-record cohort.
Worksheet fieldWhat to recordCounting rule
Cohort and baselineWindow; N eligible; excluded count and reason; total manual minutesManual minutes ÷ N is the comparable baseline per record.
Routine all-record reviewAI-handled record count; first-check minutesCount a first check once per reviewed record.
Exception reviewFlagged count; additional investigation minutes; outcomeFlag rate = flagged records ÷ N; include manually routed cases.
Independent audit and missed errorsAudited count; audit minutes; misses after routine review by severity; correction/remediation minutesMiss rate = misses ÷ audited records; do not present unaudited records as error-free.
Recovery and operating overheadInterruptions, failed replays, manual fallback and administration minutesRecord active staff minutes; assign each activity to one row only.

Find misses that the routine reviewer did not flag

Routine exception counts cannot reveal errors that pass unnoticed. Before testing, define the correct output using a source record and a domain owner, then audit a prespecified sample of apparently accepted cases, plus all high-severity or disputed cases where feasible. Keep audit minutes visible as a distinct cost. Report the audit denominator, miss count and severity; if only 20 of 100 accepted cases were audited, a count of zero misses means zero observed among 20, not proof that the other 80 were correct. Check disagreements against the source with an accountable reviewer rather than treating the model's confidence or another model's answer as ground truth.

  • Classify a miss by consequence: a reversible internal typo, a recoverable wrong route, or a potentially consequential payment, privacy or customer-commitment error require different decisions.
  • Track whether the miss arose from absent source data, extraction, reasoning, reviewer oversight or a downstream write; the owner needs a remedy, not only an accuracy percentage.
  • Set the sampling method and acceptable error severities before seeing the pilot results. Increase scrutiny for rare, costly errors; a small sample cannot establish their absence.

Further reading: See reviewed data-entry boundaries

Work a reproducible hypothetical batch

Illustration only, not a GLCO client result or a forecast: 100 eligible records require four active manual minutes each, including normal manual correction, so baseline staff time is 100 × 4 = 400 minutes. The AI path handles all 100, with 100 routine checks at 1.5 minutes, 20 additional exception investigations at four minutes, an independent audit of all 100 at 0.5 minute, five misses found by that audit corrected at eight minutes, two interruptions at ten minutes and two failed replays at six minutes. These are disjoint time entries: the five miss corrections are not included in exception investigation, and the audit time does not include the corrections.

Hypothetical active staff minutes for one 100-record cohort; replace every input with observed counts and times.
WorkFormulaMinutes
Manual verified baseline100 × 4400
Routine review of all records100 × 1.5150
Extra exception investigation20 × 480
Independent audit100 × 0.550
Correction of missed errors5 × 840
Interruptions2 × 1020
Failed replays2 × 612
Total AI-path staff time150 + 80 + 50 + 40 + 20 + 12352
Net capacity released400 − 352; then ÷ 10048 total; 0.48 minute per eligible record

Test the negative case before scaling

Keep the same hypothetical batch and all other timings, but suppose independent audit finds 12 missed errors rather than five. Correction becomes 12 × 8 = 96 minutes; AI-path staff time rises to 150 + 80 + 50 + 96 + 20 + 12 = 408 minutes. Net capacity is 400 − 408 = −8 minutes: this workflow consumes staff time before any subscription, setup, maintenance or harm from an undetected error. The example is a sensitivity calculation, not evidence that either miss rate is typical. If an error escapes the audit, estimate downstream remediation separately and ask the owner whether its severity permits any deployment.

  • Net active minutes per eligible record = (manual baseline minutes − routine review − exception review − audit − miss correction − interruption − replay − other operating minutes) ÷ N.
  • For sampled audits, retain the observed sample denominator and show uncertainty or a range for unaudited cases; do not silently scale zero observed misses into zero expected work.
  • Compare measured capacity with cash costs in the existing ROI guide. Freed minutes become cash only through a demonstrable reduction in paid work or incremental contribution; do not count the same hours and their output twice.

Further reading: Compare year-one costs in the ROI guide·Consider a no-AI path

Decide who can release, pause or redesign the workflow

The process owner should sign off on the record definition, audit method, severity rubric, minimum acceptable net time and release authority before a live trial. A reviewer handles ambiguous records; an operator monitors failures and replay; the domain owner resolves source disputes and decides whether a missed error triggers a pause. Neither a favorable average nor a platform time-saved estimate overrides a critical escape, an untraceable output or an unresolved data-access restriction. This worksheet measures active labor and observed misses in a chosen case mix; it does not establish future reliability, vendor pricing, legal compliance or cash ROI. Bring an anonymized input/output pair, weekly record volume and a named approver to a scoped conversation if you want help measuring one workflow.

  • Agree in advance on a rollback or manual fallback and who receives the exception queue when the usual reviewer is absent.
  • Recheck after a source format, rule, model or application change; preserve comparable baselines and document why the mix changed.
  • If the negative case or unacceptable severity is observed, pause automation and improve the input or return to the manual/rules path.

Further reading: Discuss a measurable workflow with GLCO·Review owner and handoff responsibilities

Helpful sources