Assess, run, and write tests
Goal: know which tests are missing, run the ones you trust, and add the ones you need, without hidden side effects or invented results. Skill: Test · Walkthrough: WT-08 (unavailable environment)
Pick the mode with your wording
Section titled “Pick the mode with your wording”| You want | Mode | Request |
|---|---|---|
| Find gaps without touching anything | assess | Assess authorization test gaps in this module; no writes or execution. |
| Run an existing suite | run | Run the existing approved unit suite after inspecting its script and environment. |
| Add tests for agreed cases | write | Add tests for these agreed cases in the existing framework. |
| Add and run | write + run | Write tests for AC-1 to AC-3 and run them with the existing unit runner. |
“Check the tests” is ambiguous. The agent starts with assess or asks which you mean.
assess: read-only strategy
Section titled “assess: read-only strategy”The agent reads requirements, the change, and test source, and proposes cases for:
- happy paths;
- errors and boundaries;
- authorization and validation, where relevant.
It does not run discovery or collection commands to count tests, because those can execute project code.
run: the preflight matters
Section titled “run: the preflight matters”Before running, the agent reports the preflight:
- the exact command, filters, working directory, and revision;
- scripts, hooks, fixtures, and setup or teardown that affect effects;
- the target environment and why it is non-production;
- data and destinations, with no secrets read;
- expected artifacts, migrations, services, or network calls;
- authority: whether your request and policy cover those effects.
If the command needs something missing (a database, browser, container, or dependency), the check is BLOCKED. The agent will not install tools or start infrastructure you did not authorize.
**Execution status:** BLOCKED**Observed result:** integration suite requires a local PostgreSQL service; none reachable on the configured port. No tests executed.**Baseline relation:** Unknown — no comparable prior run availablewrite: tests only
Section titled “write: tests only”Only the requested test files and test-only fixtures change. Production code, new dependencies, and config changes are out of scope unless you authorize them separately. Tests assert the sourced intended behavior, not whatever the current implementation does. Writing tests is not running them; results are NOT_RUN until executed.
Read the results honestly
Section titled “Read the results honestly”| Report says | It means |
|---|---|
PASS for a named check |
That exact check ran and met its criterion on the checked state |
| Counts: 42 passed, 1 skipped | Observed runner output, not an estimate |
| Coverage: not measured | No coverage tool ran; no percentage is implied |
| Baseline failure retained | The failure existed before the change and is still reported |
A run-results report can be DONE with failing tests, because delivering actual results was the task.