improving-test-suites
Improve existing test suites into minimal, high-signal behavior-focused harnesses. Use this skill when the user asks to improve, trim, rewrite, delete, review, or harden tests around public contracts, critical business logic, schema validation, security-sensitive behavior, meaningful failures, realistic edge cases, readability, or maintainability. Delegates inspection, reference lookup, editing, and validation to co-located subagents and fetches external testing guidance only when it changes a concrete decision.
What this skill does
# Improving Test Suites You are a test-suite improvement orchestrator. Your job is to turn an existing test suite into the smallest useful harness that protects behavior the users, callers, and operators of the system depend on. The orchestrator does three things: **think** from compact reports, **decide** the minimal target harness, and **dispatch** focused subagents. Subagents inspect raw files, fetch external URLs when needed, edit tests, run commands, and return structured summaries. ## Inputs | Input | Required | Example | | ----- | -------- | ------- | | `TARGET_TEST_FILES` | Yes | `tests/test_billing.py` | | `USER_GOAL` | No | `"reduce brittle implementation-coupled tests"` | | `TEST_COMMAND` | No | `pytest tests/test_billing.py -q` | | `SCOPE_LIMITS` | No | `"test files only"` | | `REFERENCE_NEED` | No | `"pytest parametrization"` | `TARGET_TEST_FILES` may be one path, multiple explicit paths, a directory, or a glob. Ask one focused question for the target only when it is missing and cannot be inferred safely. ## Pipeline Overview | Phase | Mode | Goal | Output | | ----- | ---- | ---- | ------ | | Intake | Inline | Normalize target, goal, scope, and validation inputs | Dispatch packet | | Test value review | Subagent | Identify low-value tests, missing high-value coverage, and routed reviews | `TEST_VALUE_REVIEW` | | API/security review | Subagent when routed | Check public contract, schema, authorization, validation, and unsafe-input coverage | `API_SECURITY_REVIEW` | | Maintainability review | Subagent when routed | Check readability, mocking, duplication, fixtures, and parametrization | `MAINTAINABILITY_REVIEW` | | Synthesis | Inline | Choose the smallest target harness from compact reports | `MINIMAL_HARNESS_DECISION` | | Refactor | Subagent | Apply approved test edits | `TEST_REFACTOR` | | Validate | Subagent | Run the narrow relevant command and classify failures | `TEST_VALIDATION` | | Repair or handoff | Inline dispatch | Route targeted repair, escalate blockers, or summarize result | `CHANGED_PASS`, `COMPLETE_NO_SAFE_CHANGE`, `COMPLETE_PRODUCTION_BUG_EXPOSED`, `VALIDATION_FAILED_AFTER_REPAIR`, `COMPLETE_ERROR`, or `COMPLETE_BLOCKED` | Inline phases exist only where the orchestrator needs the output for routing or trade-off decisions. File inspection, code editing, reference lookup, and command execution are delegated. After every subagent dispatch, route the returned status before doing the next phase. Use `VALUE_STATUS`, `API_STATUS`, `MAINT_STATUS`, `REFACTOR_STATUS`, and `VALIDATION_STATUS` as the status decision names. Use `API_ROUTE` and `MAINT_ROUTE` for routed coverage reviews. Required reviewer blockers stop or ask; optional reviewer blockers continue only when the value review gives enough evidence for a safe decision, and are recorded as remaining risk. ## Subagent Registry | Subagent | Path | Purpose | | -------- | ---- | ------- | | `test-value-reviewer` | `./subagents/test-value-reviewer.md` | Reviews behavior value, deletion candidates, missing high-signal coverage, and follow-up review routing | | `api-security-reviewer` | `./subagents/api-security-reviewer.md` | Reviews API, schema, authorization, validation, and security-sensitive coverage | | `test-maintainability-reviewer` | `./subagents/test-maintainability-reviewer.md` | Reviews fixture design, mocking, duplication, readability, parametrization, and cognitive cost | | `test-refactorer` | `./subagents/test-refactorer.md` | Applies approved minimal harness edits to tests and directly related test helpers | | `test-validator` | `./subagents/test-validator.md` | Runs the relevant test command after refactoring or a no-op decision and returns a compact pass/fail/error verdict | Read a subagent definition only when dispatching that subagent. Retain only its structured report, fetched URLs, changed file paths, blockers, and concise decision summaries. ## Progressive Disclosure | Need | Load | When | | ---- | ---- | ---- | | Detailed phase routing and status handling | `./references/orchestration-protocol.md` | After intake, before dispatching the first reviewer | | Trade-off priority, low/high-value test categories, minimal harness rules | `./references/test-quality-heuristics.md` | Before synthesizing `MINIMAL_HARNESS_DECISION`, or whenever a reviewer needs operational categories | | External testing, framework, and security URLs | `./references/external-sources.md` | Only when a concrete decision needs source-backed support beyond local code and bundled heuristics | | Targeted validation repair rules | `./references/repair-protocol.md` | Only after changed-file validation fails, or after `BLOCKED`/repeated `ERROR` while already in a repair cycle | | Report examples | `./references/report-examples.md` | Only when a template needs an example to resolve formatting ambiguity | | Final user handoff format | `./references/final-handoff-template.md` | Immediately before the final response | | Subagent report format | Template path listed in the dispatch packet | Immediately before the subagent returns its report | Bundled paths are relative to the file that names them and must stay inside this skill folder. When dispatching a subagent, pass template and reference paths exactly as listed in that subagent's input contract. This skill is standalone: use only co-located files under this skill folder, public web URLs from `./references/external-sources.md`, or an official documentation URL supplied by the user. If a public source cannot be fetched, make the local-code decision when safe and record the unavailable source as a remaining risk; block only when freshness or framework behavior is essential. ## Mental Model Treat tests as executable contracts, not coverage inventory. A test earns its place when it would fail for a real break in public behavior, validation, security behavior, meaningful failure handling, or production-relevant edge cases. Prefer deleting, rewriting, or consolidating tests that mainly protect internal structure, mock call order, trivial construction, incidental fixture shape, or the current implementation layout. For the trade-off priority, classification categories, and minimal harness rules used during synthesis, load `./references/test-quality-heuristics.md`. For source-backed rationale, fetch the smallest relevant URL from `./references/external-sources.md`. ## Execution 1. Normalize the dispatch packet from the inputs. Ask the smallest clarifying question only when `TARGET_TEST_FILES` is missing and cannot be inferred safely. 2. Load `./references/orchestration-protocol.md` and follow its phase routing and status handling. 3. Dispatch subagents with explicit inputs only. Include `HEURISTICS_PATH` and the relevant one-hop report template path using the path values listed in the receiving subagent's input contract. Include `EXTERNAL_SOURCES_PATH` only when the user requested a source-backed decision or the subagent reaches a concrete source need. Ask before using unsupported external sources. 4. Synthesize `MINIMAL_HARNESS_DECISION` from concise reports using the priorities and rules in `./references/test-quality-heuristics.md`. Record no-op rationale when no safe edit is justified. 5. Dispatch `test-validator` with the supplied command, the refactorer's suggested command, or an inferable narrow command. For no-op decisions, dispatch it with `CHANGED_FILES=none`. Ask for a command or prerequisite only when `TEST_VALIDATION: BLOCKED` returns that decision. 6. When changed-file validation fails, load `./references/repair-protocol.md` and use targeted repair cycles instead of rerunning the whole workflow. 7. Load `./references/final-handoff-template.md` and return the final handoff with exactly one named handoff status. ## Output Contract Return the final answer using `./references/final-handoff-template.md`. Match the result language to the selected status: changed results explain why
Related in Security
mac-ops
IncludedComprehensive macOS workstation operations — diagnose kernel panics, identify failing drives, audit launchd startup items, decode wake reasons, triage TCC permission denials, manage APFS snapshots, recover from no-boot. Use for: Mac is slow, slow bootup, won't boot, kernel panic, kernel_task hot, mds_stores CPU, photoanalysisd, cloudd, login loop, gray screen, sleep wake failure, drive failing, IO errors, APFS snapshots eating space, Time Machine local snapshots, Spotlight indexing, launchd, LaunchAgent, LaunchDaemon, login items, TCC permissions, Full Disk Access, Screen Recording denied, Gatekeeper, quarantine, com.apple.quarantine, app is damaged, helper tool, /Library/PrivilegedHelperTools, pmset, wake reasons, dark wake, sysdiagnose, panic.ips, DiagnosticReports, configuration profile, MDM profile, remote diagnostics over SSH.
a11y-audit
IncludedRun accessibility audits on web projects combining automated scanning (axe-core, Lighthouse) with WCAG 2.1 AA compliance mapping, manual check guidance, and structured reporting. Output is configurable: markdown report only, markdown plus machine-readable JSON, or markdown plus issue tracker integration. Use this skill whenever the user mentions "accessibility audit", "a11y audit", "WCAG audit", "accessibility check", "compliance scan", or asks to check a web project for accessibility issues. Also trigger when the user wants to verify WCAG conformance or map findings to a specific standard (CAN-ASC-6.2, EN 301 549, ADA/AODA).
erpclaw
IncludedAI-native ERP system with self-extending OS. Full accounting, invoicing, inventory, purchasing, tax, billing, HR, payroll, advanced accounting (ASC 606/842, intercompany, consolidation), and financial reporting. 413 actions across 14 domains, 43 expansion modules. Constitutional guardrails, adversarial audit, schema migration. Double-entry GL, immutable audit trail, US GAAP.
assess
IncludedAssesses and rates quality 0-10 across multiple dimensions (correctness, maintainability, security, performance, testability, simplicity) with pros/cons analysis. Compares against project conventions and prior decisions from memory. Produces structured evaluation reports with actionable improvement suggestions. Use when evaluating code, designs, architectures, or comparing alternative approaches.
spring-boot-security-jwt
IncludedProvides JWT authentication and authorization patterns for Spring Boot 3.5.x covering token generation with JJWT, Bearer/cookie authentication, database/OAuth2 integration, and RBAC/permission-based access control using Spring Security 6.x. Use when implementing authentication or authorization in Spring Boot applications.
code-hardcode-audit
IncludedDetect hardcoded values, magic numbers, and leaked secrets. TRIGGERS - hardcode audit, magic numbers, PLR2004, secret scanning.