Lifestyle

Usage and efficiency: reducing rework

Record task conditions, model options, time and outcomes to reduce unnecessary retries and excess context.

About 15 min read · Practice 25 min

Original workflow illustration, not a product screenshot.
Image: Mokaair (© Mokaair)
Back to directory:Codex learning hub: tutorial directory

Advanced · Desktop / mobile / CLI / VS Code / JetBrains / cloud

On this page
  1. Goal and preparation
  2. Step 1: Identify what you are measuring
  3. Step 2: Fix comparison conditions and ground truth
  4. Step 3: Change only prompt specificity
  5. Step 4: Evaluate both results with the same criteria
  6. Step 5: Turn observations into a repeatable habit
  7. Completion and common misinterpretations

Goal and preparation

Lessons and resources mentioned here: exercise pack ·

Step 1: Identify what you are measuring

Plan quota in ChatGPT/Codex and API charges are different metrics. Available percentages can describe an account-wide shared window affected by other tasks, resets and plans; a one-point change is not this task's exact price. API tokens/cost require the actual model, pricing and usage record, not multiplying every CLI token by one rate. This lesson links to centrally maintained account/model guidance instead of repeating changing price figures.

MetricRecord it asDoes not directly establish
Human timePrompt preparation through verificationModel processing time
Waiting timeStart to final answerPure reasoning speed
Rework countClarifications, corrections and repeatsFewer replies do not guarantee correctness
Acceptance coverageCheck four predefined criteriaConfidence of wording is not evidence
Quota/usageActual visible informationAccount percentages are not per-run cost

Use unavailable for missing data and estimated with a method for estimates. Do not substitute zero for unknown information.

Step 2: Fix comparison conditions and ground truth

Use identical untouched broken copies. Run node --test core.test.mjs once in each and confirm the same two passes and one failure; this Node-only check consumes no model quota. The defect is visibleTasks selecting !task.completed in the completed branch. Record that ground truth without inserting the solution into comparison prompts. Both A and B propose a fix plan without editing files or starting services.

Create a fresh task for each folder on the same surface and host, with the same model and reasoning setting. If a choice is unavailable, record the current default without guessing its name. Run sequentially to reduce contention. Fresh tasks reduce answer leakage between conversations, but caching, networking and service load still differ. This is a personal workflow exercise, not a publishable model-performance ranking.

Step 3: Change only prompt specificity

A provides less context but still has an objective; B adds reproducible input, expected/actual behavior and file scope. Both target the same outcome. The variable is how much reliable context you provide in advance. Record prompt-writing time and submit each once. Answer clarification questions normally and count them; do not withhold needed information from A to make B win.

A: less contextual detail · text
Find why Completed shows unfinished tasks in this project.
Explain the cause, propose the smallest fix and give verification steps.
Read only; do not edit files, run tests or start services.
B: reproduction and scope · text
Investigate the Completed filter in this Small Steps practice copy.
Reproduction: add Read and Build, complete Read, then choose Completed.
Expected: only Read. Actual: only Build.
Inspect core.mjs, app.js and core.test.mjs as needed.
Explain the cause, propose the smallest fix while preserving Active/All behavior,
and give exact Node and browser verification steps.
Read only; do not edit files, run tests or start services.

Step 4: Evaluate both results with the same criteria

Manually check four criteria: locate the completed condition in visibleTasks; propose selecting task.completed only in that branch; preserve Active/All behavior; and specify node --test core.test.mjs plus the same Read/Build UI reproduction. Also verify it neither claims to have run prohibited tests nor edits files. A long generic risk list earns no extra credit, while a short answer missing verification is incomplete.

efficiency-notes.md (observed values only) · markdown
# Workflow comparison
Date / host / surface / CLI or app version:
Model and reasoning setting:
Fixture: unchanged broken copy

| Metric | A | B |
| --- | --- | --- |
| Prompt preparation time | | |
| Wait until final answer | | |
| Human verification time | | |
| Clarification/correction turns | | |
| Accepted criteria out of 4 | | |
| Actual visible usage, or unavailable | | |
| Files unchanged | | |
| Remaining uncertainty | | |

Decision and evidence:
One adjustment to try next:

Step 5: Turn observations into a repeatable habit

Compare preparation plus verification time, waiting and rework, not just the fastest answer. If B takes two more minutes to prepare but saves several clarifications, it may suit daily work. If A also passes completely in one turn, there is no evidence every small task needs a long brief. Save an effective reproduction format in , retaining objective, source, scope and acceptance instead of pasting entire chat histories. This reduces stale and contradictory context.

Use a clearly fictional interpretation exercise: A takes 1 minute to prepare, 3 to wait and 4 to verify, meeting three of four criteria. B takes 3, 4 and 1 minutes, meeting all four. Both recorded totals are 8 minutes. B waits longer initially but delivers more; A's remediation time is unknown, so equal end-to-end completion time is unproven. These are not model performance measurements.

Check that criteria were set beforehand, quality was verified and missing usage stays unavailable. Compare complete time and rework once both meet all four criteria. Keep these fictional numbers separate from your observations. Reuse the criteria and change one factor per follow-up so differences remain interpretable.

To compare models or reasoning next, hold the accepted B prompt and input fixed and change one available setting using . Do not change model, delegation and prompt together and attribute everything to speed. Parallel tasks and prolonged retries also use quota. For a small problem, narrowing the objective is usually easier to verify than adding tools. For long work, retain verified state in a to avoid repeating all exploration.

Completion and common misinterpretations

Both copies should remain unchanged. Save observations and both answers; if an agent edited files, record a scope failure, preserve the diff and restore your own copy. One comparison supports a choice only under those conditions, not permanent model economy or a fixed number of tasks per plan. Mark usage incomparable if the display is stale, a reset occurred or other tasks used the account. Completion requires evidenced quality judgments, a timing method, no invented usage and one next adjustment, not a table where every metric improves.

Original workflow illustration, not a product screenshot.
Original workflow illustration, not a product screenshot. · Image: Mokaair (© Mokaair)
Read the full description

Three numbered stages: identify the starting point, perform the exercise, and verify the result. Original illustration, not a product screenshot.

Back to directory

  • Lifestyle

    Codex learning hub: tutorial directory

    A planned 60-lesson, ten-unit Codex curriculum, from setup and your first task to MD instructions and advanced integrations. Find your next lesson by experience, platform, goal or command; unpublished entries show their status.

  • Lifestyle

    Worktrees and isolated tasks

    A Git worktree gives one repository multiple working directories on different branches. It isolates file edits, but databases, ports and external services may still be shared. File isolation is not full resource isolation.

  • Lifestyle

    Workshop: build a small website

    Plan and build the Small Steps task website from brief.md, with adding, completing, deleting, filtering and local persistence. Separate HTML, CSS, data functions, UI events and tests, verify with Node and browser checks, and document restart and recovery steps.

  • Lifestyle

    Understanding an existing codebase

    Use a read-only workflow to locate entry points, data flow and tests, with file-backed explanations.

Latest travel guides

Sources

Lifestyle