Field note 09

Map workflow failure modes before you automate

A frayed off-white process tape is detected by a compact red inspection light while loose fibers collect in a charcoal recovery tray.
Map failure before automation

A workflow map shows how work is supposed to move. It rarely shows what happens when a source is stale, a required field disappears or a reviewer never signs off. Those gaps become more dangerous after automation because the workflow can move faster while carrying the mistake downstream.

A failure-mode map adds the missing layer. For each step, it records the intended result, a credible failure, the evidence that would reveal it and the safe response. A process owner can then design checks and recovery paths before connecting automation to live tools.

The method here is a MAJLS adaptation for business workflows, not a NASA method for AI operations. NASA's engineering handbook describes FMECA as an analysis of failure modes, likelihood, effects and mitigations, and includes a process-oriented approach that examines procedural steps, human error, detection and mitigation.[1] We borrow that structure without pretending a policy workflow is a spacecraft or importing a made-up precision score.

Build the map in six passes

1. Set the boundary

Write down the trigger, intended outcome, first and last step, system of record and accountable owner. Name what sits outside the workflow.

2. Walk one step at a time

State the intended result in observable terms. "Prepare translation" is vague. "Produce an Arabic draft linked to the approved English source version" can be tested.

3. Name credible failure modes

Ask how the intended result could be absent, wrong, late, duplicated or unauthorized. Keep the failure mode separate from its cause. "Draft uses the wrong source version" is a failure mode; "the shared folder contained two files with similar names" is one possible cause.

4. Trace the effects

Record the local effect, the downstream effect and the affected party. A terminology conflict may look small inside the translation step, yet it can produce two policies with different meanings for employees. Effects determine where a stop is necessary.

5. Define detection as evidence

A check should produce evidence. Useful signals include a version mismatch, an absent section, a failed reconciliation, a missed due time or a missing approval record. Do not route work by an undefined model-confidence score. Test output against the source, required fields and approved terminology.

6. Design prevention, safe state and recovery

Prevention reduces the chance of failure. A safe-state action limits harm once it appears. Recovery says who corrects the record, what they re-run and what evidence permits restart. Consequential approval for policy meaning, legal interpretation or publication stays with the named human owner.

NIST's AI Risk Management Framework Playbook recommends cataloguing failed designs and negative outcomes, and documenting continuous monitoring and feedback plans.[2] Treat the map as a living control record. Add real incidents and near misses, and update it when the workflow, source system or approval policy changes. NASA's handbook likewise says FMECA insights depend on updates after design, operational and test changes.[1]

The workflow failure-mode map

Use one row per step. Split unrelated failures until each has one detection and response path.

Field Question to answer
Step and intended result What must be true when this step finishes?
Failure mode How could that result be absent, wrong, late, duplicated or unauthorized?
Likely cause What credible condition could produce the failure?
Local and downstream effects What changes here or later, and who is affected?
Detection evidence Which record, comparison, reconciliation, time limit or approval state reveals it?
Prevention What reduces the chance of occurrence?
Safe-state action What stops, isolates or remains unchanged when detected?
Recovery and restart evidence Who corrects it, what is re-run and what permits restart?
Human-only decision Which judgment or approval cannot be delegated?

Filled example: an Arabic-English policy update

This is an illustrative example with synthetic details. It is not a client workflow or evidence of a MAJLS deployment.

Scope block

Filled failure rows

Step and intended result Failure mode and effects Detection evidence Prevention, safe state and recovery
Intake: register the approved English source Wrong source version. The Arabic draft and reviewer comments attach to superseded wording; employees could receive inconsistent instructions. Submitted file hash or version ID does not match the policy register. Accept only a registered version. Stop drafting, ask the Policy Owner to resolve the source, then restart with the confirmed file and recorded hash.
Draft: produce an Arabic version linked section by section Required clause omitted. The review pack is incomplete and a later comparison may miss the gap. Section-count and heading reconciliation fails; one source clause has no linked Arabic segment. Create the section map before drafting. Hold the pack, restore the missing segment, then re-run the full reconciliation. The Arabic reviewer decides whether the restored wording preserves meaning.
Terminology check: apply the controlled register Conflicting Arabic term. Two sections use different terms for the same defined concept, creating ambiguity downstream. Term scan returns a non-approved variant or two Arabic terms linked to one controlled English term. Load the current register by version. Mark the conflict unresolved; the qualified Arabic reviewer selects or approves the term and records the decision before restart.
Review assembly: collect source, draft, differences and comments Review incomplete. A comment is unresolved or one language version lacks approval. Open-comment count is not zero, required reviewer field is blank or approval timestamps refer to different package hashes. Lock the review-pack version. Keep status at "review required" until every required reviewer signs the same hashes. Rebuild the pack if either file changes.
Release preparation: stage the approved pair Unauthorized publication. A prepared file is distributed before the Policy Owner approves both versions. No matching final approval record exists for the staged hashes, or the publishing account is not the authorized human owner. The automated role cannot publish. Keep files private, revoke the staged action if possible and route the evidence to the Policy Owner. Only that owner authorizes publication under the organization's policy.

Run a tabletop test before implementation

Use synthetic records to trigger every row. Confirm the signal appears, the workflow enters its safe state, the owner receives the evidence and the process cannot restart early. Then test combinations such as a stale source plus an omitted clause.

The map is ready for implementation when every credible failure has an observable signal, a bounded response, a named owner and explicit restart evidence. It does not guarantee that failure will never happen. It makes failure visible before speed turns it into silent propagation.

Sources

  1. NASA GSFC-HDBK-8004: Guideline for Failure Modes and Effects Analysis and Risk Assessment
  2. NIST AI RMF Playbook: Manage

Discuss the process with MAJLS ↗