CF qa-methodology
Design and apply QA methodology for software teams: test strategy, regression testing, CI failure triage, test automation, quality gates and metrics, risk-based testing, exploratory testing, test design techniques, AI code quality gates (independent verification, artifact provenance, spec-first oracles, AI-review-comment triage, acceptance-criteria testability for Spec-Driven Development), mutation-guided test hardening and review evidence (surviving mutants, weak assertions, diff-aware mutation testing), agentic eval design (dataset test design, judge-as-system-under-test, flaky-eval discipline), QA career levels (Senior/Staff/Principal), and SDET engineering (test infrastructure, gTAA, CI/CD integration). Do not use for root-cause debugging of production incidents, security implementation or threat modeling, or evaluation framework governance and statistical analysis — route those to systematic-debugging, secure-software-engineering, and agent-evals-and-observability respectively.
Design and apply QA methodology for software teams: test strategy, regression testing, CI failure triage, test automation, quality gates and metrics…
As a process F 47/100 · Will not run — References files that are not bundled: ../systematic-debugging/SKILL.md, ../secure-software-engineering/SKILL.md, ../playwright/SKILL.md
How to improve
- The text references files that are not there: add them or drop the references.
- A spec.yaml with trigger phrases and assertions — a behaviour contract for CI; `skilltest init` writes a template.
Guard findings · 4
✓ No critical or high findings
Medium and low: 4
-
low Risky intent
intent-offensive-securityreferences/security-testing.md:13Offensive-security / dual-use content (legitimate for authorised testing; review intended use) (test fixture / example file)| Pre-release | Pen test (manual) | External firm or red team | Quarterly / major release |
fixture -
low Risky intent
intent-offensive-securityreferences/security-testing.md:21Offensive-security / dual-use content (legitimate for authorised testing; review intended use) (test fixture / example file)| A01 | Broken Access Control | IDOR, privilege escalation, path traversal | Access other users' resources by ID; test admin endpoints as regular user; fuzz path parameters with `../` |
fixture -
low Secrets in code
secret-aws-keyreferences/security-testing.md:128AWS access key ID (placeholder value)- Credential testing uses obviously-fake values: `AKIA…PLE`
placeholder -
low Secrets in code
secret-aws-keyreferences/test-data-management.md:138AWS access key ID (placeholder value)| Use obviously-fake credentials for auth tests | `AKIA…PLE` (AWS example key) |
placeholder
Files scanned: 32. Evidence is masked. Grey chips explain why severity was lowered.
Against the Agent Skills spec
- warning
missing-refreference to a missing file: ../systematic-debugging/SKILL.md - warning
missing-refreference to a missing file: ../secure-software-engineering/SKILL.md - warning
missing-refreference to a missing file: ../playwright/SKILL.md - warning
missing-refreference to a missing file: ../spec-driven-development/SKILL.md - warning
missing-refreference to a missing file: ../agent-evals-and-observability/SKILL.md - warning
missing-refreference to a missing file: ../verification-methodology/SKILL.md - warning
missing-refreference to a missing file: ../release-engineering/SKILL.md
Process rating: all ten parameters 47/100
- 0Tools and files. 7 referenced file(s) missing: ../systematic-debugging/SKILL.md, ../secure-software-engineering/SKILL.md, ../playwright/SKILL.md
- 0Inputs and preconditions. Does not say what the process needs to start
- 0Progress reporting. Says nothing while it works
- 30Running it twice. 1 mutating operations with no state check
- 40Result and completion. Does not say what the result is
- 50When it triggers. No condition that starts the skill
- 50Failures and branches. 0 branches, has a failure section
- 100Steps. 23 steps
- 100Consistency. Name and required fields are in place
- 100Execution cost. Instruction body is 2620 tokens
- low No test case covers injection arriving through data
Everything here is measured from the skill text rather than judged by a model, so the numbers are checkable. A parameter weighs more when it is a more common reason for the process to stall.
Quality signals
- +5Description has no quoted example phrases that should trigger the skill
- +3Description length 997: 120–800 characters recommended
- +3Output format is not stated: the model decides each time
- +4No input/output examples
- +2Single-language instructions
- +4Description says when NOT to use the skill
- +4Structure: 8 headings
- +3Step-by-step instructions: 23 items
- +4Reference files are cited in the instructions (17 of 17)
- +3All 2 scripts are documented
- +1License stated
Quality base 70; lint remarks subtract, signals add up to 100. Result: 77.