Automation field notes

When Not to Use AI for a Business Workflow

Keep a process manual when judgment or missing evidence makes mistakes costly; use a rule for stable structured decisions; consider AI only for bounded, reviewable variation that clears owner-set gates.

When is no AI the right answer?

Do not use AI merely because a process is repetitive. Keep the work manual when the necessary source evidence or authority is missing, the consequences of an error cannot be contained, or volume is too low to justify review and operating cost. Simplify the form or use an existing rule when the inputs are structured and the decision can be specified exactly. Consider AI assistance only when variable language or documents defeat those simpler methods, authorized inputs are available, a human can inspect evidence before an external effect, and the net result is better than the alternatives. NIST's AI Risk Management Framework Playbook explicitly calls for considering non-AI alternatives and for a proceed-or-not decision based on benefits and risks; it does not supply a universal numerical cutoff.

  • First ask whether the underlying step is necessary. Delete an unnecessary handoff before choosing a tool.
  • Document the exact output and its effects: internal draft, proposed route, sent message, posted record or payment are not equivalent risk levels.
  • An owner, not a model vendor, decides which errors are tolerable and who may authorize release.

Further reading: Browse process candidates separately

Use three gates before comparing tools

Gate one is authority and impact: identify whose data may be used, which destination may be changed and who approves a consequential decision. Gate two is input fitness: on a recent representative batch, count records with complete, current, attributable source material and those requiring a person to obtain missing facts. Gate three is operating economics: estimate volume, manual minutes, expected routine review, exception handling, correction and ongoing ownership. A missing authority cannot be offset by high volume or attractive projected savings. Percentages should be calculated on all eligible arrivals, not only the easy records an AI system processed.

  • Fail authority gate: no AI ingestion or live writes until permissions, retention, privacy and approver responsibilities are resolved.
  • Fail input gate: repair the intake form or source of truth first; a model must not fill a missing authorization or invent a fact.
  • Fail economic gate: keep manual or choose a native/rules feature if low volume or review/rework overwhelms removed effort.
  • Pass all three provisionally: compare a bounded AI draft against manual and rules on representative cases before any release.

Further reading: Compare AI and deterministic rules

Set severity and data-quality thresholds with the owner

Use a severity rubric tied to this workflow, not a generic accuracy score. A low-severity issue is a reversible internal formatting problem; a moderate issue is a recoverable wrong route or missing field with an identified reviewer; a critical issue could change a payment, expose protected data, deny a service or make an unauthorized customer commitment. The examples are prompts for domain review, not legal classifications. The owner specifies what counts as a correct source, an adequate sample, the minimum fraction of inputs fit for processing, the maximum error counts by severity and the human review capacity. For critical outcomes, lack of a reliable pre-effect approval path is itself a stop condition.

  • For each sampled input, record complete-and-authorized, incomplete, contradictory or inaccessible; do not collapse these into a single model confidence score.
  • Measure both flagged errors and misses that pass routine review; an independent audit of accepted output is needed to find the latter.
  • State thresholds before seeing results and keep severity-specific counts. Zero observed critical misses in a small sample is not proof of zero future risk.
  • Require source traceability and a named fallback owner. If the reviewer cannot inspect the source or is unavailable, route to manual handling rather than quietly sending the output onward.

Further reading: Measure reviewer time and missed errors

Compare three illustrative outcomes, not three promises

The following are separate hypothetical workflows with different volumes; their numbers are planning arithmetic, not GLCO client results, benchmark rates or claims about a product. The manual baseline for each case is the stated number of arrivals times active minutes per completed item. Capacity figures exclude implementation and software cost, so they are not ROI. The correct choice follows from data quality and severity as well as minutes. In the AI case, every drafted output remains subject to human approval, and 12 incomplete inputs never reach the model.

Illustrative monthly alternatives; each row compares its own baseline with its proposed path.
Input profile and gatePath and active-time calculationDecision
Eight unusual refund exceptions; three lack approval evidence; a wrong commitment could be criticalManual: 8 × 6 = 48 minutes; a model cannot supply the missing authorizationKeep manual and repair approval intake; low volume and high consequence do not justify AI.
200 structured routing requests; 196 have valid reason codes and four require clarificationManual baseline: 200 × 3 = 600 minutes. Rule path: 196 × 0.5 checking + 4 × 3 manual + 30 upkeep = 140 minutes; 460 minutes of potential capacitySimplify the form/use a rule; exact reason-code mapping does not need AI. Owner still reviews exceptions and measured errors.
120 mixed-format service requests; 108 have complete authorized source material, 12 do not; output is a reversible internal draftManual baseline: 120 × 5 = 600 minutes. AI path: 108 × 2 review + 12 × 5 manual + 60 audit/replay/operating work = 336 minutes; 264 minutes of potential capacityConsider AI-assisted drafting only if owner-set safety and data gates hold on a representative trial; no automatic external send.

Write a stop criterion before the AI trial

For the hypothetical 120-request draft case, an owner could require at least 108 of 120 arrivals to have complete, authorized and attributable source material; human review of all 108 drafts before release; no critical unsupported claim or unauthorized disclosure observed in an independently checked, prespecified trial sample; and total active staff time below the 600-minute comparable manual baseline after exceptions, audit, correction, interruption and replay. These values describe one illustrative owner's proposed gate, not a general safety standard. A critical observed escape, missing access approval, reviewer backlog that defeats timely review, or non-positive net capacity triggers a pause and return to manual or rules. Any critical-risk requirement may demand stronger evidence, controls or no use at all.

  • Name the process owner who can stop or restart, the reviewer who approves individual drafts and the operator who handles outages; write down the manual fallback.
  • Choose the sample across easy, ambiguous, peak-period and known failure cases before inspecting outputs; do not weaken a gate afterward because an average looks good.
  • If the model changes, inputs shift or exception rates rise, repeat the relevant checks rather than treating the original trial as permanent approval.

Further reading: Time the full review and exception burden

Treat a negative decision as useful work

A no-AI result can still remove work: eliminate duplicate approvals, improve required fields, use a native application feature, or make the correct rule and exception route explicit. Keep the audit trail and handoff owner even if a tool never gets built. If the rule path saves staff time but creates a high-severity misroute, stop that path too; deterministic is not the same as safe. Conversely, a well-reviewed AI draft may be reasonable for variable text without granting the model decision authority. Compare the same output quality and staff workload under each candidate, then examine setup, subscriptions and maintenance before buying. Favor the least complex path that clears the owner's gates.

  • Do not assign a dollar benefit to released minutes unless a real paid cost falls or redeployment produces measurable additional contribution.
  • Do not use an example score, model confidence or a vendor time-saved estimate as proof of actual local accuracy or safety.
  • If no path meets the gate, retain the manual process and address its source-data or policy problem first.

Further reading: Read the ROI model for full costs·Compare build, buy and existing features

Bring a decision record, not a request for a demo

This worksheet cannot establish whether a particular deployment complies with law, resolves privacy obligations, prevents rare harm or pays for itself. It gives the accountable team a way to reject AI early and to specify evidence if AI remains plausible. Bring anonymized examples of an ordinary input, a missing-data input and a costly failure; include monthly volume, manual active minutes, the correct output, prohibited actions and the person authorized to approve changes. GLCO can discuss a bounded workflow and its manual or rules alternatives without presuming that AI is the answer. A scoped conversation is preferable to a broad promise of automation.

  • Confirm that any prospective provider can explain data access, review queues, failed-run recovery and who maintains the workflow after handoff.
  • Write down the decision and the reason for stopping, simplifying or testing; revisit only when the input quality, risk controls or economics materially change.

Further reading: Discuss a scoped workflow with GLCO·Review automation ownership at handoff

Helpful sources