Lifestyle

Testing skill selection and results

Test format, selection and outcomes separately, using matching and non-matching requests to refine descriptions.

About 15 min read · Practice 30 min

Original workflow illustration, not a product screenshot.
Image: Mokaair (© Mokaair)
Back to directory:Codex learning hub: tutorial directory

Advanced · Desktop / CLI / VS Code / JetBrains

On this page
  1. Goal and preparation
  2. Step 1: Define results and failure conditions
  3. Step 2: Add eight repeatable tests
  4. Step 3: Prove the tests can catch a defect
  5. Step 4: Design skill behavior checks
  6. Step 5: Record, repair and repeat

Goal and preparation

Lessons and resources mentioned here:

Eight automated tests check the program's specified inputs. Skill acceptance separately checks selection, file reads, tool records and replies; Node.js success does not prove activation. No specific model is required: record your actual surface and available settings.

Step 1: Define results and failure conditions

Read the matrix before adding tests. The normal case keeps both Read tasks; an empty array is valid. Duplicate IDs, invalid types and malformed JSON must fail explicitly instead of dropping troublesome rows and reporting plausible counts. Invalid input must not produce partial success JSON on stdout, which downstream scripts could mistake for a result.

CaseChecker exitAccepted result
Three tasks, repeated title0total 3, active 1, completed 2
Empty array0All three counts are 0
Duplicate ID1Duplicate id; no success JSON
String completed value1Invalid task; no coercion
Non-array, malformed JSON, absent file or argument1Explicit error; no success counts

Step 2: Add eight repeatable tests

Create skill.test.mjs at the codex-skill-lab root and paste the complete program below. It uses Node.js's built-in test runner, requiring no extra package. Each case creates its own temporary input, executes the real checker and compares exit status, output and original bytes. Cleanup removes only its own newly created temporary folder, preserving data/tasks.json.

skill.test.mjs · javascript
import test from 'node:test';
import assert from 'node:assert/strict';
import { mkdtempSync, writeFileSync, readFileSync, rmSync } from 'node:fs';
import { tmpdir } from 'node:os';
import { resolve, join, dirname } from 'node:path';
import { spawnSync } from 'node:child_process';

const script = resolve('.agents/skills/todo-summary/scripts/count-tasks.mjs');
const tasks = [
  { id: 'a', title: 'Read', completed: true },
  { id: 'b', title: 'Build', completed: false },
  { id: 'c', title: 'Read', completed: true },
];
function run(input, verify, missing = false) {
  const folder = mkdtempSync(join(tmpdir(), 'todo-skill-test-'));
  const file = join(folder, 'input.json');
  try {
    if (!missing) writeFileSync(file, input, 'utf8');
    const before = missing ? null : readFileSync(file);
    const result = spawnSync(process.execPath, [script, file], { encoding: 'utf8' });
    assert.equal(result.error, undefined);
    verify(result);
    if (!missing) assert.deepEqual(readFileSync(file), before, 'Input changed');
  } finally {
    assert.equal(dirname(resolve(folder)), resolve(tmpdir()));
    rmSync(folder, { recursive: true });
  }
}
function fails(input, pattern) {
  run(input, result => {
    assert.equal(result.status, 1);
    assert.equal(result.stdout, '');
    assert.match(result.stderr, pattern);
  });
}
test('keeps repeated titles as three records', () => run(JSON.stringify(tasks), result => {
  assert.equal(result.status, 0);
  assert.equal(result.stderr, '');
  assert.deepEqual(JSON.parse(result.stdout), { total: 3, active: 1, completed: 2 });
}));
test('accepts an empty array', () => run('[]', result => {
  assert.equal(result.status, 0);
  assert.deepEqual(JSON.parse(result.stdout), { total: 0, active: 0, completed: 0 });
}));
test('rejects duplicate IDs', () => fails(JSON.stringify([tasks[0], tasks[0]]), /Duplicate id/));
test('rejects string completed', () => fails(JSON.stringify([{ ...tasks[0], completed: 'true' }]), /Invalid task/));
test('rejects a non-array', () => fails('{}', /Input must be an array/));
test('rejects malformed JSON', () => fails('{', /\S/));
test('reports a missing file', () => run('', result => {
  assert.equal(result.status, 1);
  assert.equal(result.stdout, '');
  assert.match(result.stderr, /ENOENT/);
}, true));
test('requires an input argument', () => {
  const result = spawnSync(process.execPath, [script], { encoding: 'utf8' });
  assert.equal(result.status, 1);
  assert.equal(result.stdout, '');
  assert.match(result.stderr, /Usage:/);
});

From the practice root, run the following command in PowerShell, macOS or Linux. The correct checker should produce eight passes, zero failures and test-runner exit 0. Inputs expected to fail make their tests pass when correctly rejected. Distinguish a checker exit of 1 from failure of the test suite itself.

Run at the practice root · sh
node --test skill.test.mjs

If the first case cannot find the script, verify the root and the exact script location. If every case prints Usage, inspect whether the test still passes file. Malformed-JSON wording can vary by Node version; that test requires a nonempty error rather than version-specific punctuation. Preserve the failing output and compare versions instead of deleting negative cases to get a green result.

Step 3: Prove the tests can catch a defect

Copy count-tasks.mjs to count-tasks.saved.mjs and confirm the backup exists. In the original result line, temporarily replace only total: tasks.length with the fragment below. This simulates counting titles: two Read records become one while active and completed retain the original counts. It is a deliberate local fault, not the final version.

Replace only total inside result · javascript
total: new Set(tasks.map(task => task.title)).size

Run the same command: expect seven passes and one failure named keeps repeated titles as three records. Its actual total is 2 and expected total is 3, demonstrating detection of a contract violation. Restore only the original checker from your backup and rerun for eight passes. Keep summaries of baseline, fault and restoration rather than only the last green result.

Step 4: Design skill behavior checks

After restoring the program, test Codex. Use a fresh task for each row, verify the working folder and send the specified request. Prevent previously read contracts, finished reports or manually supplied answers from contaminating later cases. Record the exact request, manual selection, files read, command, exit status and result to distinguish selection failures from execution failures.

SituationNatural-language requestWhat to check
ExplicitSelect todo-summary; summarize data/tasks.json in the reply onlyRead resources, execute, 3/1/2, preserve input
Matching contextCount total, active and completed tasks in this local JSONRecord selection; no false claim of using a skill
Outside scopePlan a blue website background; do not edit filesDo not run the task counter for this request
Required resource missingRename the contract, then explicitly request a verified skill reportReport missing resource and stop; invent nothing

Use the desktop Skills entry or @ picker; in CLI/IDE use /skills or $. Implicit selection depends on description and context, so non-selection does not by itself establish an installation failure. First verify explicit selection, then inspect whether description names the use case and exclusions. Do not broaden it to all tasks, which would load this workflow for unrelated work.

Compare explicit-only invocation

Create agents/openai.yaml in this isolated todo-summary skill with the setting below. If it exists, save a copy and merge only the policy field, preserving interface and dependencies. The official skill documentation says false disables implicit invocation while explicit mentions still work. This controls invocation, not sandboxing or tool permissions.

agents/openai.yaml: merge fragment · yaml
policy:
  allow_implicit_invocation: false

Use two fresh tasks with data/tasks-next.json: one requests counts without selecting a skill; the other explicitly selects todo-summary. Expect no implicit invocation in the first and explicit use in the second. Record reads and tool activity: correct counts from ordinary file tools do not prove skill use. If unchanged, restart Codex and retest; retain unconfirmed fields.

Verify this policy on a supported surface; parsed YAML or eight passing Node tests cannot prove it. Afterwards remove only this exercise's new openai.yaml, or restore its backup, then check a fresh task. To retain explicit-only use, keep false and record the decision and path.

Step 5: Record, repair and repeat

Use this acceptance record and mark unmeasured fields not run instead of filling expected values. Fix one issue at a time: resource paths, selection description or checker logic. Preserve the previous version and failing case, then rerun affected program tests and the corresponding fresh task. Editing SKILL.md does not establish that an old task has reloaded it.

Acceptance record template · markdown
# Skill acceptance record

Date and surface: <observed>
Node / Codex versions: <observed>
Model and effort, if shown: <observed or unavailable>
Skill path and revision: <exact local path and saved version>
Program tests: <command, exit, passes, failures>
Behavior case: <explicit / matching / outside-scope / missing-resource>
Request: <exact text>
Selected skill: <observed / not selected / unconfirmed>
Resources read and command executed: <evidence or not run>
Actual result and preserved input: <evidence>
Failure and one change: <description>
Fresh-task retest: <result or not run>
Remaining checks: <not run>

If the agent claims execution but supplies only expectations, ask for the actual command and tool output. Keep the status unconfirmed if evidence remains unavailable. If duplicate skill names appear, record the selected path and disable or move only the duplicate you created before retesting, preserving others' skills. Avoid global-setting changes that merely hide a local fixture-path problem.

Before finishing, restore input-format.md, restore the correct checker, confirm data/tasks.json is unchanged and obtain eight passing tests. Report program and skill checks separately. Mark macOS/Linux as documentation and cross-platform Node.js checks if you did not operate those systems. Reference tests are not claims of model execution in every surface.

Apply the same method to future skills: fix the input, define failure conditions and retain repeatable checks. For reusable requests or document outlines, continue to ; for distributing multiple skills, read . Diagram 1 fixes cases, 2 executes and observes, and 3 rechecks a repair.

Original workflow illustration, not a product screenshot.
Original workflow illustration, not a product screenshot. · Image: Mokaair (© Mokaair)
Read the full description

Three numbered stages: identify the starting point, perform the exercise, and verify the result. Original illustration, not a product screenshot.

Back to directory

  • Lifestyle

    Codex learning hub: tutorial directory

    A planned 60-lesson, ten-unit Codex curriculum, from setup and your first task to MD instructions and advanced integrations. Find your next lesson by experience, platform, goal or command; unpublished entries show their status.

  • Lifestyle

    Worktrees and isolated tasks

    A Git worktree gives one repository multiple working directories on different branches. It isolates file edits, but databases, ports and external services may still be shared. File isolation is not full resource isolation.

  • Lifestyle

    Workshop: build a small website

    Plan and build the Small Steps task website from brief.md, with adding, completing, deleting, filtering and local persistence. Separate HTML, CSS, data functions, UI events and tests, verify with Node and browser checks, and document restart and recovery steps.

  • Lifestyle

    Usage and efficiency: reducing rework

    Record task conditions, model options, time and outcomes to reduce unnecessary retries and excess context.

Latest travel guides

Sources

Lifestyle