This is a copy of a skill from another catalog; the rating counts the canonical one: skillboss(ClawHub)
What is at stake
Medium-severity findings: the skill is probably honest, but read what alarmed the scanner.
Broad scope medium severity
Below is the worst case for this category. The finding here is medium: the guard saw a sign, not a proof.
If you install
The skill asks for more than the task needs: broad tool access, credential environment variables, binaries. Every extra permission widens the damage from a mistake or a compromise.
For the author
Narrow allowed-tools and the variable list to the minimum; replace binaries with readable sources or scripts.
How to improve
For the model run — optional
Your own cases (evals/evals.json, 4–6 real requests with expected answers): the full check would then run those instead of a model-drafted suite.
A spec.yaml with trigger phrases and assertions — a behaviour contract for CI; `skilltest init` writes a template.
Files scanned: 80. Evidence is masked. Grey chips explain why severity was lowered.
Against the Agent Skills spec
✓ No remarks against the Agent Skills spec
Process rating: all ten parameters 64/100
0Result and completion. Does not say what the result is
0Inputs and preconditions. Does not say what the process needs to start
30Running it twice. 8 mutating operations with no state check
40Consistency. Frontmatter name (skillboss) differs from the folder (skillboss-2)
70When it triggers. States when to use, but not when not to
100Tools and files. Tools declared in frontmatter
100Steps. 28 steps
100Failures and branches. 7 branches, has a failure section
100Execution cost. Instruction body is 2153 tokens
100Progress reporting. Reports progress
Everything here is measured from the skill text rather than judged by a model, so the numbers are checkable. A parameter weighs more when it is a more common reason for the process to stall.
Quality signals
+5Description has no quoted example phrases that should trigger the skill
+4Description does not say when NOT to use the skill (false activations)
+3Output format is not stated: the model decides each time
-33 of 4 scripts are never mentioned in SKILL.md
+1No license
+2Single-language instructions
+3Description length 379: enough signal without eating the budget
+4Structure: 21 headings
+3Step-by-step instructions: 28 items
+4Has examples (13 code blocks)
Quality base 70; lint remarks subtract, signals add up to 100. Result: 81.
External checks
ClawHub: suspicious
The skill broadly matches its advertised AI/deployment gateway purpose, but it needs review because it can handle credentials, deploy code, send messages, upload environment data, and attempt update behavior with limited confirmation.
LLM: suspicious (high) · VirusTotal: · 29 May 2026