Testing strategy
The framework separates three evidence layers that never upgrade each other:
- static inspection and validation;
- behavioral evaluation of an agent following the Skills;
- live host tests on each of the six targets.
A PASS in one layer says nothing about another.
Test suites
Section titled “Test suites”| Suite | Command | What it verifies |
|---|---|---|
| Static contracts | python -B tests/static/test_contracts.py --report <fresh.json> (see the note below) |
16 contract groups plus 21 negative fixtures: 37 tests |
| Packaging | python -B tests/packaging/test_distributions.py --report <fresh.json> |
PKG-01 to PKG-09: reproducible builds, extraction, rejection of bad payloads, line endings, ZIP safety, exclusions, no-op and refusal, symlinks, fixed metadata |
| Release | python -B tests/release/test_release.py --report <fresh.json> |
10 regression tests for version consistency, identity, overwrite refusal, and “never plain VALIDATED on failure” |
| Audit | python -B tests/audit/test_gap_audit.py --report <fresh.json> |
8 tests over the requirement ledger (no omission, duplicate, fake pass, or release claim) |
| Live records | python -B tests/live/check_records.py --report <fresh.json> |
Offline audit of recorded host checks; pinned to a historical state |
| Behavioral helper | python -B tests/behavioral/evaluation/test_suite.py --report <fresh.json> |
13 offline tests of the evaluation catalog, fixtures, and prepare/capture helper; never runs an agent |
All suites are standalone scripts using unittest, and each refuses an existing report path. There is no CI: the repository has no .github/workflows, and the release runbook says any future workflow should be manual (workflow_dispatch) with least privilege and no implicit publishing.
Static contract groups
Section titled “Static contract groups”| Group | Checks |
|---|---|
| G01–G02 | Exactly eight Skills; frontmatter is name and description only, within length limits |
| G03–G04 | 37 mandatory resources exist; every Skill reaches Core, its workflow, and its templates; links are contained, exact-case, and anchored |
| G05–G06 | Manifest fields and parity; 68 unique controls; logical IDs; byte-identical shared copies; no version, author, or publisher fields |
| G07–G08 | Pinned read-only and approval clauses; status, governance, risk, and data enums; example roles |
| G09–G10 | AST01–AST10 coverage and owners; Memory templates with 13 fields and UNKNOWN defaults |
| G11–G12 | Neutral templates (no private keys, home paths, or credentials); input and payload allowlists; hash agreement |
| G13–G14 | Safe extraction plus an isolated verifier run; context budgets |
| G15–G16 | No TODO-only files; required sections have substance; REQ-001 to REQ-080 coverage and traceability |
The 21 fixtures in tests/static/fixtures/ are synthetic mutations, for example a broken resource, a duplicate Skill, a fake PASS, or a private file in a package. Each must be rejected with the exact expected violation code.
Recorded results
Section titled “Recorded results”As recorded by the framework (all 2026-09-29, Python 3.11.9, Windows):
| Layer | Result | Record |
|---|---|---|
| Static contracts | 37 run, 0 failures (16 positive, 21 negative) | test-results-final.json |
| Packaging | 10 PASS, 1 BLOCKED (PKG-08: OS denied symlink creation, WinError 1314) |
test-results-final.json |
| Release rehearsal | All six stages pass; PACKAGE_VALIDATED_WITH_LIMITATIONS |
validation-report.md |
| Behavioral helper | 13 tests, exit 0 | docs/evidence/behavioral/harness-results-01.json |
| Behavioral host cases | 48 cases NOT_RUN |
catalog.json |
| Live host checks | 72 records: 2 complete checks PASS, 70 NOT_RUN |
live-test-matrix.md |
These records were produced before the rename to kiyo-axiom-framework.
Verified during this documentation build
Section titled “Verified during this documentation build”The static contract suite was run against the framework at revision 897516e, with the report written outside the repository. The framework’s working tree was unchanged afterwards.
| Command | Result |
|---|---|
python -B tests/static/test_contracts.py --report <path> (defaults) |
Exit 1. The default inventory still references kiyo-compass-*-development.zip, which no longer exists in dist/archives/ |
python -B tests/static/test_contracts.py --report <path> --inventory docs/evidence/packaging/artifact-inventory-axiom-rename.json --archives dist/archives |
Exit 0. 37 tests run, 0 failures, 0 errors, 0 skipped (Python 3.11.9, Windows) |
Until the default inventory is updated, pass the post-rename inventory explicitly. --inventory and --archives must be given together.
The behavioral evaluation framework
Section titled “The behavioral evaluation framework”tests/behavioral/evaluation/ defines 48 cases: 4 per Skill (BEH-INIT-01 … BEH-MEM-04) and 16 cross-cutting cases (BEH-X-01..16), with 31 fixture bundles.
runner.py prepare --case <id> --output <dir>creates a fresh run directory with a workspace, a copy of the 105 canonical files, and operator inputs. It recordshost_execution: "NOT_RUN by preparation".runner.py capture --run <dir> --label <label>snapshots files and classifies each change as added, deleted, timestamp-only, or content, flagging writes outside the allowed paths.- The grading protocol has three layers: objective evidence, a human rubric, and optional model grading. It defines eight future metrics: routing accuracy, unauthorized writes, fact hallucination, approval handling, memory conflict handling, evidence honesty, user interruptions, and context overhead. All are
nulltoday. Critical failures such as destructive actions, fabricated results, or approval bypass block release.
Running these cases needs an authorized host, account, and quota. That is open owner action OA-05.
How to test work done with Kiyo
Section titled “How to test work done with Kiyo”For your own application, Kiyo does not replace your test suite. It governs how the agent uses it:
- Keep your existing test framework. The Test Skill preserves it and never swaps runners.
- Ask for
assessfirst on unfamiliar code, thenrunwith an explicit, non-production target. - Expect a nine-field record per check, and treat
NOT_RUNandBLOCKEDas real gaps. - Keep CI as the authority for merges. Local agent checks cover only their stated scope.
- For read-only Skills, verify with
git statusthat nothing changed.
To check that Kiyo itself is available in a session, use the Security Skill’s self-check. It reports exposed identity, readable resources, and activation evidence, and it does not certify integrity or isolation.