Lifestyle

JSON output and result validation

Separate event streams from final results, parse structured output and reject missing or invalid fields.

About 20 min read · Practice 25 min

Original workflow illustration, not a product screenshot.
Image: Mokaair (© Mokaair)
On this page
  1. Goal and materials
  2. Step 1: Separate events from the final answer
  3. Step 2: Prepare a fixed input and contract
  4. Step 3: Create a single-run, layered validator
  5. Step 4: Run and check the content
  6. Step 5: Three failure drills without model usage
  7. Recovery, limits and next steps

Goal and materials

Lessons and resources mentioned here:

Step 1: Separate events from the final answer

--json changes stdout to JSONL: each line is a separate JSON object. It is neither a single JSON array nor just the final answer, so do not call json.loads once on the whole events.jsonl. --output-schema specifies the final response shape, and -o saves that response as final.json. One run can retain both events and a structured answer, while stderr retains diagnostics. Extensions are names; flags and content determine the actual format.

FileContentAcceptance condition
status.jsonWrapper's exit recordexited with exit_code 0
events.jsonlLine-oriented CLI eventsParseable lines, a completion and no failure event
final.jsonFinal structured responseCheck fields, types, revision and values
stderr.logProgress and diagnosticsInvestigate it; do not treat it as the answer
schema.jsonFormat requested for this runKeep it with the expected contract

These filenames are exercise conventions, not a list of files Codex always creates automatically.

Step 2: Prepare a fixed input and contract

Run git init in your own empty exec-json-lab folder and create UTF-8 tasks.md below. If reusing the previous lesson's fixture, verify it is identical. This validator deliberately accepts only exec-practice-1 to prevent mixing revisions. Count total 3, completed 1 and pending 2 by hand. For real data, define a new revision and acceptance contract instead of removing checks until any output passes.

tasks.md · markdown
# Practice tasks
Revision: exec-practice-1

- [x] Read the guide
- [ ] Create a practice file
- [ ] Verify the result

Step 3: Create a single-run, layered validator

Save the complete program below as run_summary.py in that folder. It uses only Python's standard library, with no pip install. It creates a new output directory, records the schema and running state, then starts Codex with an argument array, never a shell command assembled from model text. The 120-second timeout is an exercise choice, not a fixed Codex limit. Timeout, launch failure or nonzero exit retains evidence and rejects the answer without automatic retry.

On Windows this lesson uses the standalone codex.exe. Run Get-Command codex -All in PowerShell and confirm an installed .exe. If it lists only npm .cmd/.ps1 wrappers or an alias, Python does not use PowerShell command resolution. Follow the official standalone option in , reopen the terminal and check again. Do not switch to shell=True or concatenate shell commands to fix this. On macOS/Linux ensure the terminal and Python process use the same PATH.

run_summary.py (complete program) · python
"""Run once, retain evidence, and accept only a validated practice summary."""
import argparse
import json
from pathlib import Path
import subprocess
import sys

SCHEMA = {
    "type": "object",
    "properties": {
        "revision": {"type": "string"},
        "total": {"type": "integer", "minimum": 0},
        "completed": {"type": "integer", "minimum": 0},
        "pending": {"type": "integer", "minimum": 0},
    },
    "required": ["revision", "total", "completed", "pending"],
    "additionalProperties": False,
}
PROMPT = (
    "Read only tasks.md. Return its Revision marker and checkbox counts "
    "as revision, total, completed and pending. Do not edit input files "
    "or use external services."
)


def verify(run):
    status = json.loads((run / "status.json").read_text(encoding="utf-8"))
    if (not isinstance(status, dict)
            or status != {"state": "exited", "exit_code": 0}
            or type(status.get("exit_code")) is not int):
        raise ValueError("Process did not exit successfully")
    events = []
    for number, line in enumerate((run / "events.jsonl").read_text(encoding="utf-8-sig").splitlines(), 1):
        if not line.strip():
            continue
        event = json.loads(line)
        if not isinstance(event, dict) or not isinstance(event.get("type"), str):
            raise ValueError(f"Invalid event on line {number}")
        events.append(event["type"])
    if (events.count("turn.completed") != 1 or events.count("turn.started") != 1
            or events.index("turn.started") > events.index("turn.completed")):
        raise ValueError("Missing or ambiguous completed turn")
    if "turn.failed" in events or "error" in events:
        raise ValueError("Failure event requires investigation")
    result = json.loads((run / "final.json").read_text(encoding="utf-8-sig"))
    if not isinstance(result, dict) or set(result) != set(SCHEMA["required"]):
        raise ValueError("Unexpected or missing final fields")
    if result["revision"] != "exec-practice-1":
        raise ValueError("Unexpected input revision")
    if any(type(result[key]) is not int or result[key] < 0 for key in ("total", "completed", "pending")):
        raise ValueError("Counts must be nonnegative integers, not booleans")
    if result["total"] != result["completed"] + result["pending"]:
        raise ValueError("Inconsistent count arithmetic")
    return result


def main():
    parser = argparse.ArgumentParser()
    parser.add_argument("run_name")
    parser.add_argument("--verify-only", action="store_true")
    args = parser.parse_args()
    root = Path.cwd().resolve()
    if Path(args.run_name).name != args.run_name or args.run_name in {".", ".."}:
        parser.error("Use a folder name, not a path")
    run = root / args.run_name
    if not args.verify_only:
        if not (root / "tasks.md").is_file():
            parser.error("Missing tasks.md in the current folder")
        run.mkdir()  # Refuse to overwrite an earlier run.
        schema_path = run / "schema.json"
        schema_path.write_text(json.dumps(SCHEMA, indent=2) + "\n", encoding="utf-8")
        status_path = run / "status.json"
        status_path.write_text(json.dumps({"state": "running"}) + "\n", encoding="utf-8")
        command = ["codex", "exec", "--sandbox", "read-only", "--ephemeral", "--json",
                   "--output-schema", str(schema_path), "-o", str(run / "final.json"), PROMPT]
        try:
            with (run / "events.jsonl").open("xb") as stdout, (run / "stderr.log").open("xb") as stderr:
                process = subprocess.run(command, cwd=root, stdout=stdout, stderr=stderr, timeout=120, check=False)
            status = {"state": "exited", "exit_code": process.returncode}
        except subprocess.TimeoutExpired:
            status = {"state": "timeout"}
        except OSError:
            status = {"state": "launch_failed"}
        status_path.write_text(json.dumps(status) + "\n", encoding="utf-8")
    result = verify(run)
    # Output data only: never execute text returned by the model.
    print(json.dumps(result, ensure_ascii=False))


if __name__ == "__main__":
    try:
        main()
    except (OSError, ValueError) as error:
        print(f"Not accepted: {error}", file=sys.stderr)
        sys.exit(1)

Step 4: Run and check the content

Windows PowerShell, in the exercise folder · powershell
py -3 --version
py -3 run_summary.py run-01
$practiceExit = $LASTEXITCODE
$practiceExit
macOS / Linux, in the exercise folder · sh
python3 --version
python3 run_summary.py run-01
practice_exit=$?
printf '%s\n' "$practice_exit"

If Windows lacks py but python --version reports Python 3, replace py -3 with python. Expect one object with revision exec-practice-1 and total 3, completed 1, pending 2, followed by exit 0. Inspect the five files in run-01. The validator rejects missing/extra fields, negatives, booleans masquerading as integers and inconsistent arithmetic. But 3/2/1 has valid arithmetic and is still wrong, so compare with the list. Schema compliance does not prove factual accuracy.

Step 5: Three failure drills without model usage

Copy the entire run-01 folder to run-bad-fields in the file manager and remove only pending from the copy's final.json. Run verify-only below; expect Not accepted on stderr and a nonzero exit. Make a separate copy of original run-01 as run-bad-events, remove the final closing brace from one event object, and verify that folder; it must fail too. In a third copy, run-bad-status, change only status.json exit_code to 7: even a correct final.json must be rejected. Start each case from the unchanged original rather than stacking errors.

Windows: validate the copy without calling Codex · powershell
py -3 run_summary.py run-bad-fields --verify-only
$LASTEXITCODE
macOS / Linux: validate the copy only · sh
python3 run_summary.py run-bad-fields --verify-only
echo $?

Extension: valid structure, incorrect content

Copy original run-01 to an unused run-wrong-counts folder and replace only final.json with the object below. Keep status and events; run verify-only with this folder name. Expect exit 0: types, revision and arithmetic satisfy the validator. Comparing tasks.md must still reject the answer. The verifier neither rereads the input nor checks that final.json matches the answer in events; it is not a general factual verifier.

run-wrong-counts/final.json: deliberately incorrect counts · json
{
  "revision": "exec-practice-1",
  "total": 3,
  "completed": 2,
  "pending": 1
}

Record structure passed/content rejected without altering the original result or using this answer downstream. Recheck untouched run-01 and its 3/1/2 counts. The next CI lesson adds exact fixture comparison to make this acceptance condition executable.

Recovery, limits and next steps

Do not ask Codex to repair JSON for recovery. Keep the failed copies as evidence, then verify original run-01 again; it should pass. Use run-02 for the next real execution: reusing run-01 refuses to overwrite an existing directory. For launch_failed, check codex --version in the same terminal. For timeout, consult before retrying. This wrapper manages one invocation, not every descendant process on the host or a distributed execution lock.

New CLI event kinds may appear, so the program retains unknown events and does not require a fixed item count. It conservatively rejects incomplete or failed turns for investigation. It validates one turn, not an arbitrary resumed multi-turn transcript. Windows, macOS and Linux use the same Python source; a phone does not run this local script. Record Python/CLI versions, input revision, files and manual counts before moving to . Reference checks distinguish synthetic-event tests from actual model output; authored demonstration JSON is not presented as an API response.

Original workflow illustration, not a product screenshot.
Original workflow illustration, not a product screenshot. · Image: Mokaair (© Mokaair)
Read the full description

Three numbered stages: identify the starting point, perform the exercise, and verify the result. Original illustration, not a product screenshot.

Back to directory

  • Lifestyle

    Codex learning hub: tutorial directory

    A planned 60-lesson, ten-unit Codex curriculum, from setup and your first task to MD instructions and advanced integrations. Find your next lesson by experience, platform, goal or command; unpublished entries show their status.

  • Lifestyle

    Worktrees and isolated tasks

    A Git worktree gives one repository multiple working directories on different branches. It isolates file edits, but databases, ports and external services may still be shared. File isolation is not full resource isolation.

  • Lifestyle

    Workshop: build a small website

    Plan and build the Small Steps task website from brief.md, with adding, completing, deleting, filtering and local persistence. Separate HTML, CSS, data functions, UI events and tests, verify with Node and browser checks, and document restart and recovery steps.

  • Lifestyle

    Usage and efficiency: reducing rework

    Record task conditions, model options, time and outcomes to reduce unnecessary retries and excess context.

Latest travel guides

Sources

Lifestyle