Lifestyle
Testing skill selection and results
Test format, selection and outcomes separately, using matching and non-matching requests to refine descriptions.
About 15 min read · Practice 30 min

Advanced · Desktop / CLI / VS Code / JetBrains
Before you start
On this page
Back to the Codex learning hubCodex learning hub: tutorial directoryA planned 60-lesson, ten-unit Codex curriculum, from setup and your first task to MD instructions and advanced integrations. Find your next lesson by experience, platform, goal or command; unpublished entries show their status.Read the full article
Goal and preparation
Lessons and resources mentioned here: Skill resourcesSkill scripts, references and assetsMove repeated logic and large references into useful resources, then verify conditional loading and relative paths.Read the full article
Eight automated tests check the program's specified inputs. Skill acceptance separately checks selection, file reads, tool records and replies; Node.js success does not prove activation. No specific model is required: record your actual surface and available settings.
Step 1: Define results and failure conditions
Read the matrix before adding tests. The normal case keeps both Read tasks; an empty array is valid. Duplicate IDs, invalid types and malformed JSON must fail explicitly instead of dropping troublesome rows and reporting plausible counts. Invalid input must not produce partial success JSON on stdout, which downstream scripts could mistake for a result.
| Case | Checker exit | Accepted result |
|---|---|---|
| Three tasks, repeated title | 0 | total 3, active 1, completed 2 |
| Empty array | 0 | All three counts are 0 |
| Duplicate ID | 1 | Duplicate id; no success JSON |
| String completed value | 1 | Invalid task; no coercion |
| Non-array, malformed JSON, absent file or argument | 1 | Explicit error; no success counts |
Step 2: Add eight repeatable tests
Create skill.test.mjs at the codex-skill-lab root and paste the complete program below. It uses Node.js's built-in test runner, requiring no extra package. Each case creates its own temporary input, executes the real checker and compares exit status, output and original bytes. Cleanup removes only its own newly created temporary folder, preserving data/tasks.json.
import test from 'node:test';
import assert from 'node:assert/strict';
import { mkdtempSync, writeFileSync, readFileSync, rmSync } from 'node:fs';
import { tmpdir } from 'node:os';
import { resolve, join, dirname } from 'node:path';
import { spawnSync } from 'node:child_process';
const script = resolve('.agents/skills/todo-summary/scripts/count-tasks.mjs');
const tasks = [
{ id: 'a', title: 'Read', completed: true },
{ id: 'b', title: 'Build', completed: false },
{ id: 'c', title: 'Read', completed: true },
];
function run(input, verify, missing = false) {
const folder = mkdtempSync(join(tmpdir(), 'todo-skill-test-'));
const file = join(folder, 'input.json');
try {
if (!missing) writeFileSync(file, input, 'utf8');
const before = missing ? null : readFileSync(file);
const result = spawnSync(process.execPath, [script, file], { encoding: 'utf8' });
assert.equal(result.error, undefined);
verify(result);
if (!missing) assert.deepEqual(readFileSync(file), before, 'Input changed');
} finally {
assert.equal(dirname(resolve(folder)), resolve(tmpdir()));
rmSync(folder, { recursive: true });
}
}
function fails(input, pattern) {
run(input, result => {
assert.equal(result.status, 1);
assert.equal(result.stdout, '');
assert.match(result.stderr, pattern);
});
}
test('keeps repeated titles as three records', () => run(JSON.stringify(tasks), result => {
assert.equal(result.status, 0);
assert.equal(result.stderr, '');
assert.deepEqual(JSON.parse(result.stdout), { total: 3, active: 1, completed: 2 });
}));
test('accepts an empty array', () => run('[]', result => {
assert.equal(result.status, 0);
assert.deepEqual(JSON.parse(result.stdout), { total: 0, active: 0, completed: 0 });
}));
test('rejects duplicate IDs', () => fails(JSON.stringify([tasks[0], tasks[0]]), /Duplicate id/));
test('rejects string completed', () => fails(JSON.stringify([{ ...tasks[0], completed: 'true' }]), /Invalid task/));
test('rejects a non-array', () => fails('{}', /Input must be an array/));
test('rejects malformed JSON', () => fails('{', /\S/));
test('reports a missing file', () => run('', result => {
assert.equal(result.status, 1);
assert.equal(result.stdout, '');
assert.match(result.stderr, /ENOENT/);
}, true));
test('requires an input argument', () => {
const result = spawnSync(process.execPath, [script], { encoding: 'utf8' });
assert.equal(result.status, 1);
assert.equal(result.stdout, '');
assert.match(result.stderr, /Usage:/);
});
From the practice root, run the following command in PowerShell, macOS or Linux. The correct checker should produce eight passes, zero failures and test-runner exit 0. Inputs expected to fail make their tests pass when correctly rejected. Distinguish a checker exit of 1 from failure of the test suite itself.
node --test skill.test.mjs
If the first case cannot find the script, verify the root and the exact script location. If every case prints Usage, inspect whether the test still passes file. Malformed-JSON wording can vary by Node version; that test requires a nonempty error rather than version-specific punctuation. Preserve the failing output and compare versions instead of deleting negative cases to get a green result.
Step 3: Prove the tests can catch a defect
Copy count-tasks.mjs to count-tasks.saved.mjs and confirm the backup exists. In the original result line, temporarily replace only total: tasks.length with the fragment below. This simulates counting titles: two Read records become one while active and completed retain the original counts. It is a deliberate local fault, not the final version.
total: new Set(tasks.map(task => task.title)).size
Run the same command: expect seven passes and one failure named keeps repeated titles as three records. Its actual total is 2 and expected total is 3, demonstrating detection of a contract violation. Restore only the original checker from your backup and rerun for eight passes. Keep summaries of baseline, fault and restoration rather than only the last green result.
Step 4: Design skill behavior checks
After restoring the program, test Codex. Use a fresh task for each row, verify the working folder and send the specified request. Prevent previously read contracts, finished reports or manually supplied answers from contaminating later cases. Record the exact request, manual selection, files read, command, exit status and result to distinguish selection failures from execution failures.
| Situation | Natural-language request | What to check |
|---|---|---|
| Explicit | Select todo-summary; summarize data/tasks.json in the reply only | Read resources, execute, 3/1/2, preserve input |
| Matching context | Count total, active and completed tasks in this local JSON | Record selection; no false claim of using a skill |
| Outside scope | Plan a blue website background; do not edit files | Do not run the task counter for this request |
| Required resource missing | Rename the contract, then explicitly request a verified skill report | Report missing resource and stop; invent nothing |
Use the desktop Skills entry or @ picker; in CLI/IDE use /skills or $. Implicit selection depends on description and context, so non-selection does not by itself establish an installation failure. First verify explicit selection, then inspect whether description names the use case and exclusions. Do not broaden it to all tasks, which would load this workflow for unrelated work.
Compare explicit-only invocation
Create agents/openai.yaml in this isolated todo-summary skill with the setting below. If it exists, save a copy and merge only the policy field, preserving interface and dependencies. The official skill documentation says false disables implicit invocation while explicit mentions still work. This controls invocation, not sandboxing or tool permissions.
policy:
allow_implicit_invocation: false
Use two fresh tasks with data/tasks-next.json: one requests counts without selecting a skill; the other explicitly selects todo-summary. Expect no implicit invocation in the first and explicit use in the second. Record reads and tool activity: correct counts from ordinary file tools do not prove skill use. If unchanged, restart Codex and retest; retain unconfirmed fields.
Verify this policy on a supported surface; parsed YAML or eight passing Node tests cannot prove it. Afterwards remove only this exercise's new openai.yaml, or restore its backup, then check a fresh task. To retain explicit-only use, keep false and record the decision and path.
Step 5: Record, repair and repeat
Use this acceptance record and mark unmeasured fields not run instead of filling expected values. Fix one issue at a time: resource paths, selection description or checker logic. Preserve the previous version and failing case, then rerun affected program tests and the corresponding fresh task. Editing SKILL.md does not establish that an old task has reloaded it.
# Skill acceptance record
Date and surface: <observed>
Node / Codex versions: <observed>
Model and effort, if shown: <observed or unavailable>
Skill path and revision: <exact local path and saved version>
Program tests: <command, exit, passes, failures>
Behavior case: <explicit / matching / outside-scope / missing-resource>
Request: <exact text>
Selected skill: <observed / not selected / unconfirmed>
Resources read and command executed: <evidence or not run>
Actual result and preserved input: <evidence>
Failure and one change: <description>
Fresh-task retest: <result or not run>
Remaining checks: <not run>
If the agent claims execution but supplies only expectations, ask for the actual command and tool output. Keep the status unconfirmed if evidence remains unavailable. If duplicate skill names appear, record the selected path and disable or move only the duplicate you created before retesting, preserving others' skills. Avoid global-setting changes that merely hide a local fixture-path problem.
Before finishing, restore input-format.md, restore the correct checker, confirm data/tasks.json is unchanged and obtain eight passing tests. Report program and skill checks separately. Mark macOS/Linux as documentation and cross-platform Node.js checks if you did not operate those systems. Reference tests are not claims of model execution in every surface.
Apply the same method to future skills: fix the input, define failure conditions and retain repeatable checks. For reusable requests or document outlines, continue to Templates and cheatsheetsPrompt, rules and handoff templatesChoose a prompt, rule or handoff template, fill its required fields and use the correct input location.Read the full article; for distributing multiple skills, read PluginsPlugins and connected servicesA plugin can distribute skills and MCP tools as an installable capability. Installing it, connecting an external account and successfully calling a tool are distinct stages. A listing does not establish access to your data.Read the full article. Diagram 1 fixes cases, 2 executes and observes, and 3 rechecks a repair.
Back to the Codex learning hubCodex learning hub: tutorial directoryA planned 60-lesson, ten-unit Codex curriculum, from setup and your first task to MD instructions and advanced integrations. Find your next lesson by experience, platform, goal or command; unpublished entries show their status.Read the full article
Read the full description
Three numbered stages: identify the starting point, perform the exercise, and verify the result. Original illustration, not a product screenshot.
Lifestyle
Codex learning hub: tutorial directory
A planned 60-lesson, ten-unit Codex curriculum, from setup and your first task to MD instructions and advanced integrations. Find your next lesson by experience, platform, goal or command; unpublished entries show their status.
Lifestyle
Worktrees and isolated tasks
A Git worktree gives one repository multiple working directories on different branches. It isolates file edits, but databases, ports and external services may still be shared. File isolation is not full resource isolation.
Lifestyle
Workshop: build a small website
Plan and build the Small Steps task website from brief.md, with adding, completing, deleting, filtering and local persistence. Separate HTML, CSS, data functions, UI events and tests, verify with Node and browser checks, and document restart and recovery steps.
Lifestyle
Usage and efficiency: reducing rework
Record task conditions, model options, time and outcomes to reduce unnecessary retries and excess context.
Articles that cite this one
Latest travel guides

GuideTokyo
Where to Stay in Tokyo: Comparing Shinjuku, Ueno, Tokyo Station, Shibuya, Asakusa, Ikebukuro, and Ginza, Plus Airport Access, Accommodation Tax, and Luggage Delivery
Where should you stay in Tokyo? Compare Shinjuku, Ueno, Tokyo Station, Shibuya, Asakusa, Ikebukuro, and Ginza by the same criteria: access from Narita and Haneda, transit routes, nearby attractions, neighborhood character, and who each area suits. Includes a comparison table, a Yamanote Line diagram, Tokyo’s accommodation tax as verified in 2026/9 (changing to 3% in 2027/4), and Airport TA-Q-BIN luggage shipping rules.
- Budget
- Hotels

GuideTokyo
How to Choose Tokyo Transit Passes: Are Suica, Welcome Suica, the Tokyo Subway Ticket, and the JR Pass Worth It?
On a first Tokyo trip, start with an IC card and pay per ride (Welcome Suica has no deposit and is valid for 28 days). If you take four or more subway rides in a day, add a 72-hour Tokyo Subway Ticket for 2,000 yen; a JR Pass is never worthwhile if you stay in Tokyo and do not go to Kansai. See what TOURIST PASMO, Suica on iPhone, and the Tokyo Metro day pass do and do not cover, with a decision chart. Prices verified in September 2026.
- Transport
- Budget

GuideTokyo
Tokyo Disneyland and DisneySea Guide: Ticket Prices, Fantasy Springs, Disney Premier Access (DPA), Standby Pass, and Which Park to Choose for Your First Visit
Tokyo Disney one-day Passport prices vary: most weekdays in 9/2026 cost ¥9,900 and weekends ¥10,900. At 14:00 daily, tickets go on sale for the same date two months later. Free Priority Pass is no longer on the official service list; only paid Disney Premier Access (¥1,000–3,500 per person per use) shortens waits. Covers hours, the 25th anniversary, Standby Pass, Entry Request, Fantasy Springs access and first-visit park choice; checked on the official site in 9/2026.
- Itineraries
- Family
Sources
- Build skills: resources and invocation · Checked: