Automation field notes
How to Measure Review Time and Exceptions in an AI Workflow
Time the whole record cohort, not just successful model actions. Separate routine review, escalations, missed errors and recovery work to see whether AI actually releases staff capacity.
What should the measurement answer?
Human review time is the sum of work created by the proposed workflow, not merely the minutes a reviewer spends looking at a clean AI draft. Compare that total with the time the same team needs to complete comparable records manually. Include records the automation rejects, fails to process or sends back for human handling. The difference is released or added staff capacity; it is not automatically a cash saving. A vendor dashboard can start a conversation, but it cannot measure your reviewer burden: Zapier Enterprise Analytics, for example, documents a default estimate of two minutes saved per successful action step and lets an account owner adjust it. That unit is a task, not necessarily a completed business record.
- Use a fixed observation window and one defined output, such as a verified invoice draft or a classified request; do not mix several workflows in one average.
- Keep a count of all eligible arrivals and a separate count of completed, failed, skipped and manually handled records. Multiple actions on one record do not enlarge the record denominator.
- Have the process owner approve the correctness standard and identify which errors would make release unsafe.
Further reading: Use the broader ROI model after timing the work
Fix the denominator and time a manual baseline
For a fair comparison, define N as the number of eligible business records arriving during the observation window, including duplicates, poor-quality inputs and records that need clarification unless the written scope excludes them before either path is timed. Record excluded types and counts separately. Time a representative manual batch from receipt through verified output, including ordinary checking and corrections. Use comparable cases, staff roles and service conditions for the proposed path; otherwise changes in case mix or staffing masquerade as savings. Report totals and minutes per eligible record, not only the median of easy successes.
- Baseline fields: date or batch, anonymous case ID, input type, eligible or prespecified exclusion, active handling minutes, routine verification minutes, correction minutes, severity of any error, and final disposition.
- Record elapsed waiting time separately from active staff time. A two-day wait for an approval is not two days of labor, although it may matter for turnaround.
- If cases vary materially, stratify by type or difficulty and report each group's denominator. Preserve the original mix when combining groups.
- Choose consecutive arrivals or a prespecified randomized sample across ordinary and busy periods; document any missing observations and do not extrapolate a convenience sample of clean records.
Use one privacy-safe review ledger
Log an event once under the work it actually caused. All-record review means the first routine inspection applied to every AI-handled draft; escalated review is additional investigation for flagged or suspicious cases, not that first inspection again. A missed error is an incorrect output that passed routine review and was found in a separate audit or downstream; count its correction and any remediation without also assigning those minutes to the initial exception. Recovery from a failed run belongs in replay time, while staff time spent locating context, switching tasks or chasing an alert belongs in interruption time. An identifier and category suffice for analysis; do not copy customer names, document contents or sensitive fields into the worksheet.
| Worksheet field | What to record | Counting rule |
|---|---|---|
| Cohort and baseline | Window; N eligible; excluded count and reason; total manual minutes | Manual minutes ÷ N is the comparable baseline per record. |
| Routine all-record review | AI-handled record count; first-check minutes | Count a first check once per reviewed record. |
| Exception review | Flagged count; additional investigation minutes; outcome | Flag rate = flagged records ÷ N; include manually routed cases. |
| Independent audit and missed errors | Audited count; audit minutes; misses after routine review by severity; correction/remediation minutes | Miss rate = misses ÷ audited records; do not present unaudited records as error-free. |
| Recovery and operating overhead | Interruptions, failed replays, manual fallback and administration minutes | Record active staff minutes; assign each activity to one row only. |
Find misses that the routine reviewer did not flag
Routine exception counts cannot reveal errors that pass unnoticed. Before testing, define the correct output using a source record and a domain owner, then audit a prespecified sample of apparently accepted cases, plus all high-severity or disputed cases where feasible. Keep audit minutes visible as a distinct cost. Report the audit denominator, miss count and severity; if only 20 of 100 accepted cases were audited, a count of zero misses means zero observed among 20, not proof that the other 80 were correct. Check disagreements against the source with an accountable reviewer rather than treating the model's confidence or another model's answer as ground truth.
- Classify a miss by consequence: a reversible internal typo, a recoverable wrong route, or a potentially consequential payment, privacy or customer-commitment error require different decisions.
- Track whether the miss arose from absent source data, extraction, reasoning, reviewer oversight or a downstream write; the owner needs a remedy, not only an accuracy percentage.
- Set the sampling method and acceptable error severities before seeing the pilot results. Increase scrutiny for rare, costly errors; a small sample cannot establish their absence.
Further reading: See reviewed data-entry boundaries
Work a reproducible hypothetical batch
Illustration only, not a GLCO client result or a forecast: 100 eligible records require four active manual minutes each, including normal manual correction, so baseline staff time is 100 × 4 = 400 minutes. The AI path handles all 100, with 100 routine checks at 1.5 minutes, 20 additional exception investigations at four minutes, an independent audit of all 100 at 0.5 minute, five misses found by that audit corrected at eight minutes, two interruptions at ten minutes and two failed replays at six minutes. These are disjoint time entries: the five miss corrections are not included in exception investigation, and the audit time does not include the corrections.
| Work | Formula | Minutes |
|---|---|---|
| Manual verified baseline | 100 × 4 | 400 |
| Routine review of all records | 100 × 1.5 | 150 |
| Extra exception investigation | 20 × 4 | 80 |
| Independent audit | 100 × 0.5 | 50 |
| Correction of missed errors | 5 × 8 | 40 |
| Interruptions | 2 × 10 | 20 |
| Failed replays | 2 × 6 | 12 |
| Total AI-path staff time | 150 + 80 + 50 + 40 + 20 + 12 | 352 |
| Net capacity released | 400 − 352; then ÷ 100 | 48 total; 0.48 minute per eligible record |
Test the negative case before scaling
Keep the same hypothetical batch and all other timings, but suppose independent audit finds 12 missed errors rather than five. Correction becomes 12 × 8 = 96 minutes; AI-path staff time rises to 150 + 80 + 50 + 96 + 20 + 12 = 408 minutes. Net capacity is 400 − 408 = −8 minutes: this workflow consumes staff time before any subscription, setup, maintenance or harm from an undetected error. The example is a sensitivity calculation, not evidence that either miss rate is typical. If an error escapes the audit, estimate downstream remediation separately and ask the owner whether its severity permits any deployment.
- Net active minutes per eligible record = (manual baseline minutes − routine review − exception review − audit − miss correction − interruption − replay − other operating minutes) ÷ N.
- For sampled audits, retain the observed sample denominator and show uncertainty or a range for unaudited cases; do not silently scale zero observed misses into zero expected work.
- Compare measured capacity with cash costs in the existing ROI guide. Freed minutes become cash only through a demonstrable reduction in paid work or incremental contribution; do not count the same hours and their output twice.
Further reading: Compare year-one costs in the ROI guide·Consider a no-AI path
Decide who can release, pause or redesign the workflow
The process owner should sign off on the record definition, audit method, severity rubric, minimum acceptable net time and release authority before a live trial. A reviewer handles ambiguous records; an operator monitors failures and replay; the domain owner resolves source disputes and decides whether a missed error triggers a pause. Neither a favorable average nor a platform time-saved estimate overrides a critical escape, an untraceable output or an unresolved data-access restriction. This worksheet measures active labor and observed misses in a chosen case mix; it does not establish future reliability, vendor pricing, legal compliance or cash ROI. Bring an anonymized input/output pair, weekly record volume and a named approver to a scoped conversation if you want help measuring one workflow.
- Agree in advance on a rollback or manual fallback and who receives the exception queue when the usual reviewer is absent.
- Recheck after a source format, rule, model or application change; preserve comparable baselines and document why the mix changed.
- If the negative case or unacceptable severity is observed, pause automation and improve the input or return to the manual/rules path.
Further reading: Discuss a measurable workflow with GLCO·Review owner and handoff responsibilities