Lifestyle

Automation failure, retry and stopping

Use execution records to detect prior effects, avoid duplicates, repair prerequisites and confirm a schedule has stopped.

About 15 min read · Practice 20 min

Original workflow illustration, not a product screenshot.
Image: Mokaair (© Mokaair)
On this page
  1. Goal and prerequisites
  2. Step 1: Identify the run state first
  3. Step 2: Reproduce an answer without completion evidence
  4. Step 3: Handle a real timeout or disconnection
  5. Step 4: Retry one objective while keeping attempts distinct
  6. Stopping and acceptance

Goal and prerequisites

Lessons and resources mentioned here: ·

Step 1: Identify the run state first

A schedule has saved configuration, while each trigger has its own run. Pausing configuration does not cancel an active run, and deleting a schedule does not undo previous changes. A local wrapper timeout only means its wait limit was reached; an offline UI only means the connection is unavailable. Neither proves every remote or child process has stopped. Record the task/schedule name, host, folder, timezone, input revision, last start time and visible run identifier before acting.

SymptomMissing evidenceNext action
Active saved, no new runHost, next time, actual triggerInspect Scheduled and host state
Long-running stateWhether the process/task is activeInspect progress and last records before retry
Timeout/disconnectionRemaining execution and side effectsReconnect and inspect the original run
Nonzero exit or turn.failedSpecific error and complete outputRetain records and correct the cause
Exit 0 but wrong valuesRevision and content verificationReject the result and inspect its source

An old file can already contain a successful answer, so existence alone is not success.

Step 2: Reproduce an answer without completion evidence

Copy successful run-01 to recovery-running using the file manager. Change only the copy's status.json to the content below, retaining its correct final.json and events. This is a labeled simulated interruption; never alter real failure records to manufacture success. Run verify-only from the exercise folder using PowerShell on Windows or a terminal on macOS/Linux. Expect Not accepted: Process did not exit successfully and nonzero exit.

recovery-running/status.json (simulated state) · json
{
  "state": "running"
}
Windows: validate a simulated interruption · powershell
py -3 run_summary.py recovery-running --verify-only
$LASTEXITCODE
macOS / Linux: validate a simulated interruption · sh
python3 run_summary.py recovery-running --verify-only
echo $?

Make another copy of original run-01 as recovery-timeout, change only status.json to {"state":"timeout"}, and validate that folder; it must also fail. Both cases show that a plausible answer without completion evidence is unacceptable. Verify unchanged run-01 again; it should pass. Record all three outcomes as copy-based drills. They do not prove an actual host outage or model timeout occurred.

Step 3: Handle a real timeout or disconnection

For a desktop schedule, pause that same entry in Scheduled to prevent new triggers, then open its existing run to check activity. Restore power, network, app and folder after a host outage. Revisit the original task's last actions and results before creating any duplicate schedule. On every OS, inspect the actual host rather than assuming visible phone history means the computer is online. Determine missed-run behavior from real records and product behavior, not an assumption that every missed time is replayed.

For a CLI script, retain status, stderr and events, then check the original terminal for process completion. Cancel only the identified run using Ctrl+C or its tool's cancellation control, not every node, python or Codex process. In GitHub Actions, cancel the specific run and confirm its terminal state; disabling the workflow governs future triggers. A missing approval, renamed input path or exhausted API quota is not fixed by increasing timeout. Consult , paths or account lessons for the actual cause.

The supplied run_summary.py uses subprocess.run(timeout=120). Python documents that a timeout kills and waits for its direct child before raising TimeoutExpired; initial process creation can extend elapsed time. This does not prove descendants, remote work or already-sent requests stopped. Record the direct process separately from unverified work and preserve the timeout record.

Step 4: Retry one objective while keeping attempts distinct

Retry only after the original has stopped or duplicate side effects have been ruled out, and the cause is fixed. Even this fictional read-only task uses a new run-02 output folder. Record the same input revision with a new attempt; when input content changes, change the revision instead of mixing different data under one version. The record below is a template: fill in observed states, not an assumed accepted result.

recovery-notes.md (fill with observations) · markdown
# Recovery record
- Objective: summarize the fictional checklist
- Input revision: exec-practice-1
- Host and folder: fill in locally
- Original run and state: fill in
- Last observed action: fill in
- Cause and correction: fill in
- Retry output folder: run-02
- Exit status and event validation: fill in
- Manual counts: total 3, completed 1, pending 2
- Accepted result path: fill in only after verification
- Schedule state after practice: fill in

Before downstream use, check the new run's exit status, events, schema and manual counts, then record one accepted result path. A retry is not a reason to send, upload or charge again: external actions need service-supported idempotency keys or an explicit deduplication strategy, not a model promise. Refusing an existing output folder protects local files, not a distributed lock or exactly-once external execution. For real write operations, inspect existing side effects before completing missing work.

Stopping and acceptance

Keep the successful original and labeled failed copies, and revalidate the original once instead of repeatedly running the model. If a real schedule is enabled, pause it in Scheduled, confirm the saved state and separately check active runs. To resume, edit the same entry and verify next time, timezone, host and notification conditions. Notify on changes, completion or actionable failures rather than repeating identical empty checks. You pass when you can explain original state, cause, correction, retry and accepted result. Record official feature research, local copy drills, real host-outage tests and cloud CI runs separately.

Original workflow illustration, not a product screenshot.
Original workflow illustration, not a product screenshot. · Image: Mokaair (© Mokaair)
Read the full description

Three numbered stages: identify the starting point, perform the exercise, and verify the result. Original illustration, not a product screenshot.

Back to directory

  • Lifestyle

    Codex learning hub: tutorial directory

    A planned 60-lesson, ten-unit Codex curriculum, from setup and your first task to MD instructions and advanced integrations. Find your next lesson by experience, platform, goal or command; unpublished entries show their status.

  • Lifestyle

    Worktrees and isolated tasks

    A Git worktree gives one repository multiple working directories on different branches. It isolates file edits, but databases, ports and external services may still be shared. File isolation is not full resource isolation.

  • Lifestyle

    Usage and efficiency: reducing rework

    Record task conditions, model options, time and outcomes to reduce unnecessary retries and excess context.

Latest travel guides

Sources

Lifestyle