An automation exception needs an owner, an affected record, a clear failure description and an authorized recovery action. Decide how failures become visible, which retries are safe and when manual work should take over. Confirm the final result before closing the exception so a successful rerun does not hide duplicate or incomplete downstream work.
- Detect the failureRecord the engagement and the failed action.
- Hold the affected stepShow what already succeeded and what is incomplete.
- Review the causeAn owner checks the data, permissions and safe retry point.
- Resume and verifyConfirm the intended result without duplicate actions.
Verify the external result before retrying. A full restart can repeat a client message or create another record.
Define failure beyond a tool error
A tool can report a successful run while the business task is incomplete. The workflow may have updated the wrong record, created a duplicate or used an input that was already outdated. Define the expected business result as well as the technical completion event.
For a client onboarding workflow, that result could mean one project exists with the correct engagement reference, required owner and visible next action. It is more precise than “the integration ran.” Record what your team will check when a run is questionable.
Prepare an exception record
| Record field | Purpose | Example |
|---|---|---|
| Affected item | Identify work requiring recovery | Engagement or task reference |
| Failure point | Show where progress stopped | Project creation after intake |
| Observed result | Describe what actually happened | Record created without an owner |
| Responsible person | Establish who resolves it | Workflow maintenance owner |
| Recovery action | State what will be done | Correct existing record, then verify |
| Verification | Show the business result | Correct project and owner confirmed |
Include enough context for repair without exposing unrelated client information or credentials in broad notifications.
- Affected itemIdentify the engagement or task requiring recovery.
- Failure pointShow where progress stopped.
- Responsible ownerEstablish who resolves the exception.
- Verified resultConfirm the corrected business state.
Include enough repair context without exposing unrelated client information or credentials.
Decide when a retry is safe
Before repeating a step, check whether it already created an effect. Repeating a read may differ from repeating a message, payment-related action or project creation. Your implementation needs to identify which actions can be repeated safely and how duplicates are detected.
- Inspect prior effectCheck whether the first attempt changed anything.
- Repeat safe actionUse only understood, tested repeat behavior.
- Correct existing recordUpdate partial work instead of starting again.
- Escalate ambiguityHold consequential or unknown outcomes for review.
A platform retry function does not itself repair the business state.
Do not assume that a platform's retry function repairs the business state. Review the tool's current documentation and test the workflow's actual behavior. If a previous attempt partly succeeded, the correct recovery may be to update the existing record rather than start again.
Create a manual route
Name what a person should do when the automated route cannot continue. Identify the information needed, the actions already completed and the steps still outstanding. A vague message saying “please handle manually” leaves the person to reconstruct the workflow.
For an illustrative approval reminder, an error may occur after the reminder was sent but before the status was recorded. The recovery owner should check the communication record before sending again. Then reconcile the status so later runs do not keep repeating the action.
Make alerts actionable
Route an alert to the person who can resolve it, with an escalation path if that person is unavailable. Separate informational logs from exceptions requiring action. If every routine event raises an alert, meaningful failures become harder to notice.
Define the review expectation based on the consequence of delay. A blocked client delivery and a delayed internal summary may need different handling. The workflow owner should be able to explain that choice.
Verify recovery and learn from recurrence
After repair, confirm the expected business state and any downstream effects. Record whether a retry, correction or manual completion was used. Close the exception only after that check.
Review recurring causes such as missing inputs, changed permissions and outdated mappings. Give the fix an owner through the operations review. Use the change control guide when the repair requires altering the workflow itself.
Watch for a missing run as well as a failed step
Define what you expect to happen and how the owner can confirm it. A failure alert can identify an execution that started and encountered a problem. That alone does not establish that an expected scheduled execution began, or that a completed execution produced the correct business result.
- Step failedInspect the record and completed actions.
- Wrong resultCompare expected output with business state.
- Expected run missingInspect trigger, schedule and last verified result.
- Assign review ownerGive each condition accountable follow-up.
An execution alert does not prove an expected run started.
| Condition | What to inspect | Recovery question |
|---|---|---|
| A step failed | The affected record and actions already completed | Which remaining action can be repeated safely? |
| A run completed with an incorrect result | The expected output and actual business state | What needs correction without repeating valid work? |
| An expected run is missing | The trigger, schedule and last verified result | Did work begin, and could another execution still be in progress? |
Give each condition a review owner. Where an external action has an unknown outcome, inspect its destination before retrying. Close the incident when the intended state is verified, and record the cause so the team can improve its checks.
Questions and answers
Is a successful retry proof that an exception is resolved?
It shows the repeated operation completed according to the tool. Verify the business result, including duplicates and downstream work, before closing the exception. A partially successful original attempt can leave effects that a rerun does not reconcile.
Who should receive automation failure alerts?
Route actionable failures to a named owner who can inspect and repair the workflow, with a backup or escalation path. Include the affected item and failure point. Sending every log to the whole team can obscure responsibility.
Should failed actions retry automatically?
Use automatic retries only where the action and recovery behavior are understood and tested. Check whether the first attempt produced a partial effect. For consequential or ambiguous actions, pause for a review rather than repeating them blindly.