Skip to content

Assess, run, and write tests

Goal: know which tests are missing, run the ones you trust, and add the ones you need, without hidden side effects or invented results. Skill: Test · Walkthrough: WT-08 (unavailable environment)

You want Mode Request
Find gaps without touching anything assess Assess authorization test gaps in this module; no writes or execution.
Run an existing suite run Run the existing approved unit suite after inspecting its script and environment.
Add tests for agreed cases write Add tests for these agreed cases in the existing framework.
Add and run write + run Write tests for AC-1 to AC-3 and run them with the existing unit runner.

“Check the tests” is ambiguous. The agent starts with assess or asks which you mean.

The agent reads requirements, the change, and test source, and proposes cases for:

  • happy paths;
  • errors and boundaries;
  • authorization and validation, where relevant.

It does not run discovery or collection commands to count tests, because those can execute project code.

Before running, the agent reports the preflight:

  1. the exact command, filters, working directory, and revision;
  2. scripts, hooks, fixtures, and setup or teardown that affect effects;
  3. the target environment and why it is non-production;
  4. data and destinations, with no secrets read;
  5. expected artifacts, migrations, services, or network calls;
  6. authority: whether your request and policy cover those effects.

If the command needs something missing (a database, browser, container, or dependency), the check is BLOCKED. The agent will not install tools or start infrastructure you did not authorize.

Illustrative blocked check (walkthrough WT-08)
**Execution status:** BLOCKED
**Observed result:** integration suite requires a local PostgreSQL service; none reachable on the configured port. No tests executed.
**Baseline relation:** Unknown — no comparable prior run available

Only the requested test files and test-only fixtures change. Production code, new dependencies, and config changes are out of scope unless you authorize them separately. Tests assert the sourced intended behavior, not whatever the current implementation does. Writing tests is not running them; results are NOT_RUN until executed.

Report says It means
PASS for a named check That exact check ran and met its criterion on the checked state
Counts: 42 passed, 1 skipped Observed runner output, not an estimate
Coverage: not measured No coverage tool ran; no percentage is implied
Baseline failure retained The failure existed before the change and is still reported

A run-results report can be DONE with failing tests, because delivering actual results was the task.