生活分享

JSON 输出与结果验证

区分事件串流与最终结果,解析结构化输出并拒绝缺栏位或错误格式。

阅读时间约 20 分钟 · 操作 25 分钟

原创流程示意图,非产品界面截图。
图片:Mokaair (© Mokaair)
本篇目录
  1. 目标与材料
  2. 步骤 1:分开事件与最后答案
  3. 步骤 2:准备固定的输入与契约
  4. 步骤 3:建立一次执行、逐层验证的程式
  5. 步骤 4:执行并核对内容
  6. 步骤 5:三个不花模型额度的失败练习
  7. 还原、限制与下一步

目标与材料

本段提到的教学与资源:

步骤 1:分开事件与最后答案

--json 会把 stdout 改为 JSONL,也就是每一行各自是一个 JSON 物件。它不是单一 JSON 阵列,也不是只包含最后答案;不能直接对整份 events.jsonl 呼叫一次 json.loads。--output-schema 指定最后回答的结构,-o 把最后回答另存为 final.json。因此同一工作可以同时保留事件与结构化答案,stderr 则继续保存诊断。副档名只是名称,真正格式由旗标与内容决定。

档案内容接受条件
status.json包装程式记录的退出状态exited 且 exit_code 为 0
events.jsonlCLI 的逐行事件可逐行解析,完成事件存在且无失败事件
final.json最后的结构化回答栏位、型别、版本与数值可验证
stderr.log过程与错误说明供诊断,不当成答案
schema.json本轮要求的格式与预期契约一起保存

表中的档名是本教材约定,并非 Codex 固定自动产生的全部档案。

步骤 2:准备固定的输入与契约

在自己的空白 exec-json-lab 资料夹执行 git init,以 UTF-8 建立下方 tasks.md。若你沿用前篇资料,先确认内容完全一致;本篇验证器刻意只接受 exec-practice-1,避免拿其他版本的回答混用。先人工计算总数 3、完成 1、未完成 2。需要改用真实资料时,应另订版本与验收契约,不是删掉验证直到任何输出都能通过。

tasks.md · markdown
# Practice tasks
Revision: exec-practice-1

- [x] Read the guide
- [ ] Create a practice file
- [ ] Verify the result

步骤 3:建立一次执行、逐层验证的程式

将下方完整程式存为同一资料夹的 run_summary.py。它只用 Python 标准函式库,不需 pip install。程式先建立全新的输出资料夹,保存 schema 与 running 状态,再用参数阵列启动 Codex;不把模型文字拼成 shell 指令。120 秒是教材自行设定的等待上限,不是 Codex 的固定限制。遇到逾时、启动失败或非零退出,保留产物并拒绝接受答案,不自动重试。

Windows 本篇使用独立安装版 codex.exe。在 PowerShell 先执行 Get-Command codex -All,确认来源是已安装的 .exe;若只有 npm 的 .cmd/.ps1 包装档或别名,Python 不会套用 PowerShell 的命令解析。请依 使用官方独立安装管道,开新终端机后重查;不要为此改成 shell=True 或拼接 shell 指令。macOS/Linux 则确认终端机和 Python 程序使用相同 PATH。

run_summary.py(完整程式) · python
"""Run once, retain evidence, and accept only a validated practice summary."""
import argparse
import json
from pathlib import Path
import subprocess
import sys

SCHEMA = {
    "type": "object",
    "properties": {
        "revision": {"type": "string"},
        "total": {"type": "integer", "minimum": 0},
        "completed": {"type": "integer", "minimum": 0},
        "pending": {"type": "integer", "minimum": 0},
    },
    "required": ["revision", "total", "completed", "pending"],
    "additionalProperties": False,
}
PROMPT = (
    "Read only tasks.md. Return its Revision marker and checkbox counts "
    "as revision, total, completed and pending. Do not edit input files "
    "or use external services."
)


def verify(run):
    status = json.loads((run / "status.json").read_text(encoding="utf-8"))
    if (not isinstance(status, dict)
            or status != {"state": "exited", "exit_code": 0}
            or type(status.get("exit_code")) is not int):
        raise ValueError("Process did not exit successfully")
    events = []
    for number, line in enumerate((run / "events.jsonl").read_text(encoding="utf-8-sig").splitlines(), 1):
        if not line.strip():
            continue
        event = json.loads(line)
        if not isinstance(event, dict) or not isinstance(event.get("type"), str):
            raise ValueError(f"Invalid event on line {number}")
        events.append(event["type"])
    if (events.count("turn.completed") != 1 or events.count("turn.started") != 1
            or events.index("turn.started") > events.index("turn.completed")):
        raise ValueError("Missing or ambiguous completed turn")
    if "turn.failed" in events or "error" in events:
        raise ValueError("Failure event requires investigation")
    result = json.loads((run / "final.json").read_text(encoding="utf-8-sig"))
    if not isinstance(result, dict) or set(result) != set(SCHEMA["required"]):
        raise ValueError("Unexpected or missing final fields")
    if result["revision"] != "exec-practice-1":
        raise ValueError("Unexpected input revision")
    if any(type(result[key]) is not int or result[key] < 0 for key in ("total", "completed", "pending")):
        raise ValueError("Counts must be nonnegative integers, not booleans")
    if result["total"] != result["completed"] + result["pending"]:
        raise ValueError("Inconsistent count arithmetic")
    return result


def main():
    parser = argparse.ArgumentParser()
    parser.add_argument("run_name")
    parser.add_argument("--verify-only", action="store_true")
    args = parser.parse_args()
    root = Path.cwd().resolve()
    if Path(args.run_name).name != args.run_name or args.run_name in {".", ".."}:
        parser.error("Use a folder name, not a path")
    run = root / args.run_name
    if not args.verify_only:
        if not (root / "tasks.md").is_file():
            parser.error("Missing tasks.md in the current folder")
        run.mkdir()  # Refuse to overwrite an earlier run.
        schema_path = run / "schema.json"
        schema_path.write_text(json.dumps(SCHEMA, indent=2) + "\n", encoding="utf-8")
        status_path = run / "status.json"
        status_path.write_text(json.dumps({"state": "running"}) + "\n", encoding="utf-8")
        command = ["codex", "exec", "--sandbox", "read-only", "--ephemeral", "--json",
                   "--output-schema", str(schema_path), "-o", str(run / "final.json"), PROMPT]
        try:
            with (run / "events.jsonl").open("xb") as stdout, (run / "stderr.log").open("xb") as stderr:
                process = subprocess.run(command, cwd=root, stdout=stdout, stderr=stderr, timeout=120, check=False)
            status = {"state": "exited", "exit_code": process.returncode}
        except subprocess.TimeoutExpired:
            status = {"state": "timeout"}
        except OSError:
            status = {"state": "launch_failed"}
        status_path.write_text(json.dumps(status) + "\n", encoding="utf-8")
    result = verify(run)
    # Output data only: never execute text returned by the model.
    print(json.dumps(result, ensure_ascii=False))


if __name__ == "__main__":
    try:
        main()
    except (OSError, ValueError) as error:
        print(f"Not accepted: {error}", file=sys.stderr)
        sys.exit(1)

步骤 4:执行并核对内容

Windows PowerShell,在教材资料夹内 · powershell
py -3 --version
py -3 run_summary.py run-01
$practiceExit = $LASTEXITCODE
$practiceExit
macOS/Linux,在教材资料夹内 · sh
python3 --version
python3 run_summary.py run-01
practice_exit=$?
printf '%s\n' "$practice_exit"

若 Windows 没有 py,但 python --version 显示 Python 3,可把 py -3 换成 python。预期终端机输出一个物件,revision 为 exec-practice-1,数值为 total 3、completed 1、pending 2,退出 0。开启 run-01,核对五份档案。验证器会拒绝缺栏、多栏、负数、以 true 冒充整数及数学不一致;但 3/2/1 虽然总和正确仍是错误内容,因此最后还要和清单核对。结构符合 schema 不等于事实正确。

步骤 5:三个不花模型额度的失败练习

用档案管理员复制整个 run-01 成 run-bad-fields,只在副本 final.json 删除 pending 栏位;执行下方 verify-only。预期 stderr 出现 Not accepted 且退出非 0。接著另复制原始 run-01 成 run-bad-events,把 events.jsonl 任意一个有效物件的最后一个右大括号删掉;换成该资料夹名称再验证,也应失败。第三个副本 run-bad-status 只把 status.json 的 exit_code 改成 7,即使 final.json 正确仍应拒绝。每个案例都从未改动的原始产物复制,不互相叠加错误。

Windows:只验证副本,不呼叫 Codex · powershell
py -3 run_summary.py run-bad-fields --verify-only
$LASTEXITCODE
macOS/Linux:只验证副本 · sh
python3 run_summary.py run-bad-fields --verify-only
echo $?

延伸:格式通过,但内容不符

再从原始 run-01 复制出未使用的 run-wrong-counts,只把 final.json 换成下方完整物件。保留 status 与 events 原样,使用相同 verify-only 命令但改为这个资料夹。预期程式退出 0,因为型别、版本和总和都符合;人工比对 tasks.md 则必须拒绝这份答案。它没有重读输入,也没有核对最终档与事件中的回答是否相同,不能作为通用事实验证器。

run-wrong-counts/final.json:刻意错误的计数 · json
{
  "revision": "exec-practice-1",
  "total": 3,
  "completed": 2,
  "pending": 1
}

把此例记成「结构通过/内容拒绝」,不改原始成功纪录,也不自动提交或使用答案。回到未改动 run-01 验证并人工核对,应仍为 3/1/2。下一篇 CI 会对固定教材加上精确数值检查,展示如何把这个人工判断写成可执行条件。

还原、限制与下一步

还原时不用请 Codex 修 JSON;保留失败副本当证据,再对原始 run-01 执行 verify-only,应重新通过。下一次真正执行用 run-02;重用 run-01 会因目录已存在而拒绝覆写。若出现 launch_failed,先在同一终端机查 codex --version;若是 timeout,查看 再决定是否重试。程式只负责这个单次工作,不追踪整台主机所有子程序,也不提供跨机器的唯一执行锁。

CLI 事件可能加入新种类,因此程式保留未知事件,不靠固定的 item 数量判定成功;遇到未完成或含失败事件则保守拒绝,让人调查。这里只验证一个回合,不直接套用到续接多回合纪录。Windows、macOS、Linux 都能以相同 Python 原始码操作;手机不在本机执行此脚本。正式交付前记录 Python/CLI 版本、输入版本、档案及人工计数结果,再进入 。本教材的参考验证会区分模拟事件测试与真正模型输出,不把自行生成的示范 JSON 称为实际 API 回复。

原创流程示意图,非产品界面截图。
原创流程示意图,非产品界面截图。 · 图片:Mokaair (© Mokaair)
阅读完整文字说明

Three numbered stages: identify the starting point, perform the exercise, and verify the result. Original illustration, not a product screenshot.

回总目录

  • 生活分享

    Codex 学习中心:完整教程目录

    从安装、第一个任务到 MD 规则与进阶集成,规划 60 篇 Codex 教程、十个单元。按程度、平台、需求或命令搜索下一篇;尚未公开的教程会标示状态,方便安排学习路线。

  • 生活分享

    Worktree 与多任务隔离

    Worktree 让同一个 Git 程式库有不同的工作目录,各自承接不同分支。它适合让两项工作分开改档,但资料库、连接埠与外部服务仍可能共用,不能把档案隔离当成所有资源隔离。

  • 生活分享

    实战:制作小网站

    从 brief.md 规划并制作 Small Steps 待办网站,完成新增、完成、删除、筛选与本机资料保存。将 HTML、CSS、资料函式、画面事件与测试分开,以 Node 测试和浏览器操作验收,并留下可重新启动与还原的交接纪录。

  • 生活分享

    用量与效率:减少重工

    记录任务条件、模型选项、时间与成果,找出能减少无效重试和过多上下文的调整。

最新旅游情报攻略

资料来源

生活分享