Skip to content

Testing strategy

The framework separates three evidence layers that never upgrade each other:

  • static inspection and validation;
  • behavioral evaluation of an agent following the Skills;
  • live host tests on each of the six targets.

A PASS in one layer says nothing about another.

Suite Command What it verifies
Static contracts python -B tests/static/test_contracts.py --report <fresh.json> (see the note below) 16 contract groups plus 21 negative fixtures: 37 tests
Packaging python -B tests/packaging/test_distributions.py --report <fresh.json> PKG-01 to PKG-09: reproducible builds, extraction, rejection of bad payloads, line endings, ZIP safety, exclusions, no-op and refusal, symlinks, fixed metadata
Release python -B tests/release/test_release.py --report <fresh.json> 10 regression tests for version consistency, identity, overwrite refusal, and “never plain VALIDATED on failure”
Audit python -B tests/audit/test_gap_audit.py --report <fresh.json> 8 tests over the requirement ledger (no omission, duplicate, fake pass, or release claim)
Live records python -B tests/live/check_records.py --report <fresh.json> Offline audit of recorded host checks; pinned to a historical state
Behavioral helper python -B tests/behavioral/evaluation/test_suite.py --report <fresh.json> 13 offline tests of the evaluation catalog, fixtures, and prepare/capture helper; never runs an agent

All suites are standalone scripts using unittest, and each refuses an existing report path. There is no CI: the repository has no .github/workflows, and the release runbook says any future workflow should be manual (workflow_dispatch) with least privilege and no implicit publishing.

Group Checks
G01–G02 Exactly eight Skills; frontmatter is name and description only, within length limits
G03–G04 37 mandatory resources exist; every Skill reaches Core, its workflow, and its templates; links are contained, exact-case, and anchored
G05–G06 Manifest fields and parity; 68 unique controls; logical IDs; byte-identical shared copies; no version, author, or publisher fields
G07–G08 Pinned read-only and approval clauses; status, governance, risk, and data enums; example roles
G09–G10 AST01–AST10 coverage and owners; Memory templates with 13 fields and UNKNOWN defaults
G11–G12 Neutral templates (no private keys, home paths, or credentials); input and payload allowlists; hash agreement
G13–G14 Safe extraction plus an isolated verifier run; context budgets
G15–G16 No TODO-only files; required sections have substance; REQ-001 to REQ-080 coverage and traceability

The 21 fixtures in tests/static/fixtures/ are synthetic mutations, for example a broken resource, a duplicate Skill, a fake PASS, or a private file in a package. Each must be rejected with the exact expected violation code.

As recorded by the framework (all 2026-09-29, Python 3.11.9, Windows):

Layer Result Record
Static contracts 37 run, 0 failures (16 positive, 21 negative) test-results-final.json
Packaging 10 PASS, 1 BLOCKED (PKG-08: OS denied symlink creation, WinError 1314) test-results-final.json
Release rehearsal All six stages pass; PACKAGE_VALIDATED_WITH_LIMITATIONS validation-report.md
Behavioral helper 13 tests, exit 0 docs/evidence/behavioral/harness-results-01.json
Behavioral host cases 48 cases NOT_RUN catalog.json
Live host checks 72 records: 2 complete checks PASS, 70 NOT_RUN live-test-matrix.md

These records were produced before the rename to kiyo-axiom-framework.

The static contract suite was run against the framework at revision 897516e, with the report written outside the repository. The framework’s working tree was unchanged afterwards.

Command Result
python -B tests/static/test_contracts.py --report <path> (defaults) Exit 1. The default inventory still references kiyo-compass-*-development.zip, which no longer exists in dist/archives/
python -B tests/static/test_contracts.py --report <path> --inventory docs/evidence/packaging/artifact-inventory-axiom-rename.json --archives dist/archives Exit 0. 37 tests run, 0 failures, 0 errors, 0 skipped (Python 3.11.9, Windows)

Until the default inventory is updated, pass the post-rename inventory explicitly. --inventory and --archives must be given together.

tests/behavioral/evaluation/ defines 48 cases: 4 per Skill (BEH-INIT-01 … BEH-MEM-04) and 16 cross-cutting cases (BEH-X-01..16), with 31 fixture bundles.

  • runner.py prepare --case <id> --output <dir> creates a fresh run directory with a workspace, a copy of the 105 canonical files, and operator inputs. It records host_execution: "NOT_RUN by preparation".
  • runner.py capture --run <dir> --label <label> snapshots files and classifies each change as added, deleted, timestamp-only, or content, flagging writes outside the allowed paths.
  • The grading protocol has three layers: objective evidence, a human rubric, and optional model grading. It defines eight future metrics: routing accuracy, unauthorized writes, fact hallucination, approval handling, memory conflict handling, evidence honesty, user interruptions, and context overhead. All are null today. Critical failures such as destructive actions, fabricated results, or approval bypass block release.

Running these cases needs an authorized host, account, and quota. That is open owner action OA-05.

For your own application, Kiyo does not replace your test suite. It governs how the agent uses it:

  1. Keep your existing test framework. The Test Skill preserves it and never swaps runners.
  2. Ask for assess first on unfamiliar code, then run with an explicit, non-production target.
  3. Expect a nine-field record per check, and treat NOT_RUN and BLOCKED as real gaps.
  4. Keep CI as the authority for merges. Local agent checks cover only their stated scope.
  5. For read-only Skills, verify with git status that nothing changed.

To check that Kiyo itself is available in a session, use the Security Skill’s self-check. It reports exposed identity, readable resources, and activation evidence, and it does not certify integrity or isolation.