生活分享

JSON 輸出與結果驗證

區分事件串流與最終結果,解析結構化輸出並拒絕缺欄位或錯誤格式。

閱讀時間約 20 分鐘 · 操作 25 分鐘

原創流程示意圖,非產品介面截圖。
圖片:Mokaair (© Mokaair)
本篇目錄
  1. 目標與材料
  2. 步驟 1:分開事件與最後答案
  3. 步驟 2:準備固定的輸入與契約
  4. 步驟 3:建立一次執行、逐層驗證的程式
  5. 步驟 4:執行並核對內容
  6. 步驟 5:三個不花模型額度的失敗練習
  7. 還原、限制與下一步

目標與材料

本段提到的教學與資源:

步驟 1:分開事件與最後答案

--json 會把 stdout 改為 JSONL,也就是每一行各自是一個 JSON 物件。它不是單一 JSON 陣列,也不是只包含最後答案;不能直接對整份 events.jsonl 呼叫一次 json.loads。--output-schema 指定最後回答的結構,-o 把最後回答另存為 final.json。因此同一工作可以同時保留事件與結構化答案,stderr 則繼續保存診斷。副檔名只是名稱,真正格式由旗標與內容決定。

檔案內容接受條件
status.json包裝程式記錄的退出狀態exited 且 exit_code 為 0
events.jsonlCLI 的逐行事件可逐行解析,完成事件存在且無失敗事件
final.json最後的結構化回答欄位、型別、版本與數值可驗證
stderr.log過程與錯誤說明供診斷,不當成答案
schema.json本輪要求的格式與預期契約一起保存

表中的檔名是本教材約定,並非 Codex 固定自動產生的全部檔案。

步驟 2:準備固定的輸入與契約

在自己的空白 exec-json-lab 資料夾執行 git init,以 UTF-8 建立下方 tasks.md。若你沿用前篇資料,先確認內容完全一致;本篇驗證器刻意只接受 exec-practice-1,避免拿其他版本的回答混用。先人工計算總數 3、完成 1、未完成 2。需要改用真實資料時,應另訂版本與驗收契約,不是刪掉驗證直到任何輸出都能通過。

tasks.md · markdown
# Practice tasks
Revision: exec-practice-1

- [x] Read the guide
- [ ] Create a practice file
- [ ] Verify the result

步驟 3:建立一次執行、逐層驗證的程式

將下方完整程式存為同一資料夾的 run_summary.py。它只用 Python 標準函式庫,不需 pip install。程式先建立全新的輸出資料夾,保存 schema 與 running 狀態,再用參數陣列啟動 Codex;不把模型文字拼成 shell 指令。120 秒是教材自行設定的等待上限,不是 Codex 的固定限制。遇到逾時、啟動失敗或非零退出,保留產物並拒絕接受答案,不自動重試。

Windows 本篇使用獨立安裝版 codex.exe。在 PowerShell 先執行 Get-Command codex -All,確認來源是已安裝的 .exe;若只有 npm 的 .cmd/.ps1 包裝檔或別名,Python 不會套用 PowerShell 的命令解析。請依 使用官方獨立安裝管道,開新終端機後重查;不要為此改成 shell=True 或拼接 shell 指令。macOS/Linux 則確認終端機和 Python 程序使用相同 PATH。

run_summary.py(完整程式) · python
"""Run once, retain evidence, and accept only a validated practice summary."""
import argparse
import json
from pathlib import Path
import subprocess
import sys

SCHEMA = {
    "type": "object",
    "properties": {
        "revision": {"type": "string"},
        "total": {"type": "integer", "minimum": 0},
        "completed": {"type": "integer", "minimum": 0},
        "pending": {"type": "integer", "minimum": 0},
    },
    "required": ["revision", "total", "completed", "pending"],
    "additionalProperties": False,
}
PROMPT = (
    "Read only tasks.md. Return its Revision marker and checkbox counts "
    "as revision, total, completed and pending. Do not edit input files "
    "or use external services."
)


def verify(run):
    status = json.loads((run / "status.json").read_text(encoding="utf-8"))
    if (not isinstance(status, dict)
            or status != {"state": "exited", "exit_code": 0}
            or type(status.get("exit_code")) is not int):
        raise ValueError("Process did not exit successfully")
    events = []
    for number, line in enumerate((run / "events.jsonl").read_text(encoding="utf-8-sig").splitlines(), 1):
        if not line.strip():
            continue
        event = json.loads(line)
        if not isinstance(event, dict) or not isinstance(event.get("type"), str):
            raise ValueError(f"Invalid event on line {number}")
        events.append(event["type"])
    if (events.count("turn.completed") != 1 or events.count("turn.started") != 1
            or events.index("turn.started") > events.index("turn.completed")):
        raise ValueError("Missing or ambiguous completed turn")
    if "turn.failed" in events or "error" in events:
        raise ValueError("Failure event requires investigation")
    result = json.loads((run / "final.json").read_text(encoding="utf-8-sig"))
    if not isinstance(result, dict) or set(result) != set(SCHEMA["required"]):
        raise ValueError("Unexpected or missing final fields")
    if result["revision"] != "exec-practice-1":
        raise ValueError("Unexpected input revision")
    if any(type(result[key]) is not int or result[key] < 0 for key in ("total", "completed", "pending")):
        raise ValueError("Counts must be nonnegative integers, not booleans")
    if result["total"] != result["completed"] + result["pending"]:
        raise ValueError("Inconsistent count arithmetic")
    return result


def main():
    parser = argparse.ArgumentParser()
    parser.add_argument("run_name")
    parser.add_argument("--verify-only", action="store_true")
    args = parser.parse_args()
    root = Path.cwd().resolve()
    if Path(args.run_name).name != args.run_name or args.run_name in {".", ".."}:
        parser.error("Use a folder name, not a path")
    run = root / args.run_name
    if not args.verify_only:
        if not (root / "tasks.md").is_file():
            parser.error("Missing tasks.md in the current folder")
        run.mkdir()  # Refuse to overwrite an earlier run.
        schema_path = run / "schema.json"
        schema_path.write_text(json.dumps(SCHEMA, indent=2) + "\n", encoding="utf-8")
        status_path = run / "status.json"
        status_path.write_text(json.dumps({"state": "running"}) + "\n", encoding="utf-8")
        command = ["codex", "exec", "--sandbox", "read-only", "--ephemeral", "--json",
                   "--output-schema", str(schema_path), "-o", str(run / "final.json"), PROMPT]
        try:
            with (run / "events.jsonl").open("xb") as stdout, (run / "stderr.log").open("xb") as stderr:
                process = subprocess.run(command, cwd=root, stdout=stdout, stderr=stderr, timeout=120, check=False)
            status = {"state": "exited", "exit_code": process.returncode}
        except subprocess.TimeoutExpired:
            status = {"state": "timeout"}
        except OSError:
            status = {"state": "launch_failed"}
        status_path.write_text(json.dumps(status) + "\n", encoding="utf-8")
    result = verify(run)
    # Output data only: never execute text returned by the model.
    print(json.dumps(result, ensure_ascii=False))


if __name__ == "__main__":
    try:
        main()
    except (OSError, ValueError) as error:
        print(f"Not accepted: {error}", file=sys.stderr)
        sys.exit(1)

步驟 4:執行並核對內容

Windows PowerShell,在教材資料夾內 · powershell
py -3 --version
py -3 run_summary.py run-01
$practiceExit = $LASTEXITCODE
$practiceExit
macOS/Linux,在教材資料夾內 · sh
python3 --version
python3 run_summary.py run-01
practice_exit=$?
printf '%s\n' "$practice_exit"

若 Windows 沒有 py,但 python --version 顯示 Python 3,可把 py -3 換成 python。預期終端機輸出一個物件,revision 為 exec-practice-1,數值為 total 3、completed 1、pending 2,退出 0。開啟 run-01,核對五份檔案。驗證器會拒絕缺欄、多欄、負數、以 true 冒充整數及數學不一致;但 3/2/1 雖然總和正確仍是錯誤內容,因此最後還要和清單核對。結構符合 schema 不等於事實正確。

步驟 5:三個不花模型額度的失敗練習

用檔案管理員複製整個 run-01 成 run-bad-fields,只在副本 final.json 刪除 pending 欄位;執行下方 verify-only。預期 stderr 出現 Not accepted 且退出非 0。接著另複製原始 run-01 成 run-bad-events,把 events.jsonl 任意一個有效物件的最後一個右大括號刪掉;換成該資料夾名稱再驗證,也應失敗。第三個副本 run-bad-status 只把 status.json 的 exit_code 改成 7,即使 final.json 正確仍應拒絕。每個案例都從未改動的原始產物複製,不互相疊加錯誤。

Windows:只驗證副本,不呼叫 Codex · powershell
py -3 run_summary.py run-bad-fields --verify-only
$LASTEXITCODE
macOS/Linux:只驗證副本 · sh
python3 run_summary.py run-bad-fields --verify-only
echo $?

延伸:格式通過,但內容不符

再從原始 run-01 複製出未使用的 run-wrong-counts,只把 final.json 換成下方完整物件。保留 status 與 events 原樣,使用相同 verify-only 命令但改為這個資料夾。預期程式退出 0,因為型別、版本和總和都符合;人工比對 tasks.md 則必須拒絕這份答案。它沒有重讀輸入,也沒有核對最終檔與事件中的回答是否相同,不能作為通用事實驗證器。

run-wrong-counts/final.json:刻意錯誤的計數 · json
{
  "revision": "exec-practice-1",
  "total": 3,
  "completed": 2,
  "pending": 1
}

把此例記成「結構通過/內容拒絕」,不改原始成功紀錄,也不自動提交或使用答案。回到未改動 run-01 驗證並人工核對,應仍為 3/1/2。下一篇 CI 會對固定教材加上精確數值檢查,展示如何把這個人工判斷寫成可執行條件。

還原、限制與下一步

還原時不用請 Codex 修 JSON;保留失敗副本當證據,再對原始 run-01 執行 verify-only,應重新通過。下一次真正執行用 run-02;重用 run-01 會因目錄已存在而拒絕覆寫。若出現 launch_failed,先在同一終端機查 codex --version;若是 timeout,查看 再決定是否重試。程式只負責這個單次工作,不追蹤整台主機所有子程序,也不提供跨機器的唯一執行鎖。

CLI 事件可能加入新種類,因此程式保留未知事件,不靠固定的 item 數量判定成功;遇到未完成或含失敗事件則保守拒絕,讓人調查。這裡只驗證一個回合,不直接套用到續接多回合紀錄。Windows、macOS、Linux 都能以相同 Python 原始碼操作;手機不在本機執行此腳本。正式交付前記錄 Python/CLI 版本、輸入版本、檔案及人工計數結果,再進入 。本教材的參考驗證會區分模擬事件測試與真正模型輸出,不把自行生成的示範 JSON 稱為實際 API 回覆。

原創流程示意圖,非產品介面截圖。
原創流程示意圖,非產品介面截圖。 · 圖片:Mokaair (© Mokaair)
閱讀完整文字說明

Three numbered stages: identify the starting point, perform the exercise, and verify the result. Original illustration, not a product screenshot.

回總目錄

  • 生活分享

    Codex 學習中心:完整教學目錄

    從安裝、第一個任務到 MD 規則與進階整合,規劃 60 篇 Codex 教學、十個單元。依程度、平台、需求或指令搜尋下一篇;尚未公開的教學會標示狀態,方便安排學習路線。

  • 生活分享

    Worktree 與多任務隔離

    Worktree 讓同一個 Git 程式庫有不同的工作目錄,各自承接不同分支。它適合讓兩項工作分開改檔,但資料庫、連接埠與外部服務仍可能共用,不能把檔案隔離當成所有資源隔離。

  • 生活分享

    實戰:製作小網站

    從 brief.md 規劃並製作 Small Steps 待辦網站,完成新增、完成、刪除、篩選與本機資料保存。將 HTML、CSS、資料函式、畫面事件與測試分開,以 Node 測試和瀏覽器操作驗收,並留下可重新啟動與還原的交接紀錄。

  • 生活分享

    用量與效率:減少重工

    記錄任務條件、模型選項、時間與成果,找出能減少無效重試和過多上下文的調整。

最新旅遊情報攻略

資料來源

生活分享