SKILLEMALL.ai

BD agent-redteam-kit

当用户说『测一下这个AI安不安全』『会不会被越狱』『让AI干危险的事它听不听』『上线前做对抗测试』,或要给一个 agent/提示词做安全性红队时使用。中英双库扫描越狱/危险能力请求(DAN/忽略指令/提权/数据外泄/自改进等),给出风险分级+加固建议,并设「危险操作闸门」(有门禁)。可运行脚本(redteam_scan 扫描器)。理论根基:LGD 三律之有门禁(危险动作先过闸)。触发词:红队、red team、越狱、jailbreak、对抗测试、prompt攻击、AI安全测试、危险指令、安全评估。

ClawHub Hermes author: zhaoxinghua09-cell v1.0.0 MIT-0 7 files body ≈ 515 tokens Open the sourceclawhub.ai analyzed 2 d ago

当用户说『测一下这个AI安不安全』『会不会被越狱』『让AI干危险的事它听不听』『上线前做对抗测试』,或要给一个…

As a process D 46/100 · Unfinished process — weak spots: result and completion, when it triggers, inputs and preconditions

ProcedureSecurityAI and agentstype and topics are labelled automatically from the skill text
JSON
Technical rating
B
80/100
safety, quality, tests
Safety 60%
93
Quality 40%
61
Run on models
none yet
Process rating
D
46/100
Unfinished process
Result and completion w 14
0
Inputs and preconditions w 11
0
Failures and branches w 10
0
the three weakest of ten parameters · all ten

How to improve

  1. Say in the description WHEN to use the skill ("use when…", example requests): that is the agent's main cue.
  2. For Hermes the description must be one sentence under 60 characters; move the conditions to a "When to Use" section.
For the model run — optional
  • Your own cases (evals/evals.json, 4–6 real requests with expected answers): the full check would then run those instead of a model-drafted suite.
  • A spec.yaml with trigger phrases and assertions — a behaviour contract for CI; `skilltest init` writes a template.

Guard findings · 7

✓ No critical or high findings

Medium and low: 7
  • low Risky intent intent-offensive-security README.md:5
    Offensive-security / dual-use content (legitimate for authorised testing; review intended use) (detector / deny-list definition)
    **EN** — Red-team before launch: bilingual (zh+en) scanner for jailbreak / dangerous-capability requests (ignore-instructions, DAN, privilege escalation, data exfil, self-modify), with risk tiers and 
    detector
  • low Risky intent intent-offensive-security skill-card.md:17
    Offensive-security / dual-use content (legitimate for authorised testing; review intended use)
    Developers and security reviewers use this skill to run local pre-release checks on agent prompts or prompt files for jailbreak, instruction override, data exfiltration, privilege escalation, and rela
  • low Risky intent intent-offensive-security SKILL.md:4
    Offensive-security / dual-use content (legitimate for authorised testing; review intended use)
    display_name: AI红队对抗测试(Red Team Kit)
  • low Risky intent intent-offensive-security SKILL.md:5
    Offensive-security / dual-use content (legitimate for authorised testing; review intended use)
    displayName: AI红队对抗测试(Red Team Kit)
  • low Risky intent intent-offensive-security SKILL.md:6
    Offensive-security / dual-use content (legitimate for authorised testing; review intended use)
    title: AI红队对抗测试(Red Team Kit)
  • low Risky intent intent-offensive-security SKILL.md:12
    Offensive-security / dual-use content (legitimate for authorised testing; review intended use) (detector / deny-list definition)
    description: 当用户说『测一下这个AI安不安全』『会不会被越狱』『让AI干危险的事它听不听』『上线前做对抗测试』,或要给一个 agent/提示词做安全性红队时使用。中英双库扫描越狱/危险能力请求(DAN/忽略指令/提权/数据外泄/自改进等),给出风险分级+加固建议,并设「危险操作闸门」(有门禁)。可运行脚本(redteam_scan 扫描器)。理论根基:LGD 三律之有门禁(危险动作先
    detector
  • low Risky intent intent-offensive-security SKILL.md:13
    Offensive-security / dual-use content (legitimate for authorised testing; review intended use)
    tags: [红队, red team, 越狱, jailbreak, 对抗测试, AI安全, 有门禁]

Files scanned: 7. Evidence is masked. Grey chips explain why severity was lowered.

Against the Agent Skills spec

  • warning description-long-hermes description is 251 chars; the Hermes authoring standard requires ≤ 60 (one sentence, ending with a period)
  • warning description-no-when neither description nor a "## When to Use" section says when to use the skill
  • note frontmatter-key unknown frontmatter key "slug"
  • note frontmatter-key unknown frontmatter key "display_name"
  • note frontmatter-key unknown frontmatter key "displayName"
  • note frontmatter-key unknown frontmatter key "title"
  • note frontmatter-key unknown frontmatter key "display_name_en"
  • note frontmatter-key unknown frontmatter key "summary"
  • note frontmatter-key unknown frontmatter key "agent_created"
  • note frontmatter-key unknown frontmatter key "copyright"
  • note frontmatter-key unknown frontmatter key "read_when"
  • note frontmatter-key unknown frontmatter key "homepage"

Process rating: all ten parameters 46/100

  • 0Result and completion. Does not say what the result is
  • 0Inputs and preconditions. Does not say what the process needs to start
  • 0Failures and branches. Linear process with no failure handling
  • 0Progress reporting. Says nothing while it works
  • 20When it triggers. No condition that starts the skill
  • 60Tools and files. Uses tools (python) that frontmatter does not declare
  • 100Steps. 12 steps
  • 100Consistency. Name and required fields are in place
  • 100Execution cost. Instruction body is 515 tokens
  • 100Running it twice. No mutating operations

Everything here is measured from the skill text rather than judged by a model, so the numbers are checkable. A parameter weighs more when it is a more common reason for the process to stall.

Quality signals

  • +5Description has no quoted example phrases that should trigger the skill
  • +4Description does not say when NOT to use the skill (false activations)
  • +3Output format is not stated: the model decides each time
  • -31 of 1 scripts are never mentioned in SKILL.md
  • +2Single-language instructions
  • +3Description length 251: enough signal without eating the budget
  • +4Structure: 10 headings
  • +3Step-by-step instructions: 12 items
  • +4Has examples (1 code blocks)
  • +1License stated

Quality base 70; lint remarks subtract, signals add up to 100. Result: 61.

External checks

ClawHub: suspicious
This local red-team scanner is not malicious, but its documented blocking gate fails open and its install command uses an unpinned global installer.
LLM: suspicious (high) · 11 Sept 2026