Skip to content

Test

Logical ID: kiyo.test · Entry: src/kiyo/skills/test/SKILL.md · Procedure: workflows/test.md (KIYO-TEST-001) · Safety matrix: framework/test-mode-safety.md

Test covers three different kinds of test work:

  • assessing gaps;
  • running authorized checks;
  • writing scoped tests and test-only fixtures.

It reports observed results, baseline failures, and environment limits. It never invents coverage and never repairs production code.

assess, run, and write are logical modes chosen from your intent. They are not command arguments.

Mode Permitted effects Not granted Completion evidence
assess Analyze requirements, changes, relevant Memory, and test source; return gaps and strategy in chat File, report, or Memory writes; test execution; discovery or collection scripts; module imports; installs Inspected scope, behavior-to-case gaps, proposals; execution NOT_RUN
run Named checks on a known non-production target, plus expected artifact writes Source, test, fixture, or config repair; snapshot acceptance; production resources; tool installs; Memory sync Preflight, actual command and result, counts, skips, blockers, baseline relation
write Create or update the specified tests and test-only fixtures Production source, unrelated tests, new dependencies, config changes, execution unless separately covered Minimal diff, sourced assertions, test review, real run results or explicit NOT_RUN

Unclear intent, such as “check tests”, starts with assess or a clarifying question. “Write and run these tests” can authorize both phases without asking twice.

Before each materially different command, the agent identifies:

  1. Command and checked state: the exact command, arguments, working directory, filters, and revision, plus scripts, hooks, fixtures, and setup or teardown.
  2. Environment and target: tool availability and non-production identity. A variable named TEST_DB proves no isolation.
  3. Data and destinations: synthetic data, and file, database, or network targets. No connection secrets or bulk environment reads.
  4. Effects and recovery: caches, artifacts, migrations, services, and containers. A migration file is not permission to run it.
  5. Authority and gaps: effects compared with the request, policy, approvals, and host permission.
Observed condition Treatment
Isolated local unit runner, effects within the request Run the exact scoped check; report results and artifacts
Missing browser, driver, container, service, or dependency Don’t install; the check is BLOCKED, and no results are invented
Script touches an unknown, shared, or production database Hold until a non-production target is established; production stays excluded
Command denied by host or organization BLOCKED; no bypass or recycled approval
Runner exits before setup completes Report the failure; tests not reached get no PASS and no guessed zero
Unrelated baseline failure or suspected flakiness Keep the failure; don’t skip, delete, or weaken assertions

The testing standard (KIYO-ENG-005) selects:

  • unit tests for isolated logic;
  • integration or API tests for boundaries and contracts;
  • E2E tests for critical flows the project supports;
  • regression cases for bugs.

The existing test framework and patterns are preserved. Missing business rules remain open decisions; tests do not invent permissions, HTTP behavior, or validation rules.

The test report records check records plus counts and coverage: runner outcomes, skipped or blocked, coverage, and baseline and freshness. Unknown counts or coverage are reported as Unknown or not measured. Coverage requires actual measurement.