SKILLEMALL.ai

BC war-room

Convenes a multi-LLM expert panel to pressure-test hard-to-reverse decisions

ClawHub Agent Skills author: athola v1.9.19 MIT-0 8 files body ≈ 3 741 tokens Open the sourceclawhub.ai analyzed 4 d ago

As a process C 60/100 · Has gaps — weak spots: when it triggers, inputs and preconditions, consistency

ProcedureAI and agentsInfrastructuretype and topics are labelled automatically from the skill text
JSON
Technical rating
B
81/100
safety, quality, tests
Safety 60%
81
Quality 40%
81
Run on models
none yet
Process rating
C
60/100
Has gaps
Inputs and preconditions w 11
0
Progress reporting w 2
0
When it triggers w 12
20
the three weakest of ten parameters · all ten

How to improve

    For the model run — optional
    • Your own cases (evals/evals.json, 4–6 real requests with expected answers): the full check would then run those instead of a model-drafted suite.
    • A spec.yaml with trigger phrases and assertions — a behaviour contract for CI; `skilltest init` writes a template.

    Guard findings · 19

    ✓ No critical or high findings

    Medium and low: 19
    • low Risky intent intent-offensive-security modules/deliberation-protocol.md:49
      Offensive-security / dual-use content (legitimate for authorised testing; review intended use)
      | Phase 4: Red Team |  <-- Red Team Commander
    • low Risky intent intent-offensive-security modules/deliberation-protocol.md:222
      Offensive-security / dual-use content (legitimate for authorised testing; review intended use)
      ### Phase 4: Red Team and Wargaming
    • low Risky intent intent-offensive-security modules/deliberation-protocol.md:226
      Offensive-security / dual-use content (legitimate for authorised testing; review intended use)
      **Expert**: Red Team Commander (Gemini Flash)
    • low Risky intent intent-offensive-security modules/deliberation-protocol.md:251
      Offensive-security / dual-use content (legitimate for authorised testing; review intended use)
      - All COAs with Red Team challenges
    • low Risky intent intent-offensive-security modules/deliberation-protocol.md:308
      Offensive-security / dual-use content (legitimate for authorised testing; review intended use)
      - Experts revise positions based on Red Team feedback
    • low Risky intent intent-offensive-security modules/discussion-publishing.md:114
      Offensive-security / dual-use content (legitimate for authorised testing; review intended use)
      | Red Team | ... | ... |
    • low Risky intent intent-offensive-security modules/discussion-publishing.md:128
      Offensive-security / dual-use content (legitimate for authorised testing; review intended use)
      4. **Red Team Challenges** — key challenges and responses (Phase 4)
    • low Risky intent intent-offensive-security modules/discussion-publishing.md:131
      Offensive-security / dual-use content (legitimate for authorised testing; review intended use) (quoted — discussed, not commanded)
      > Phases 5-6 (War Game execution and Decision Briefing) are internal deliberation artifacts that don't warrant separate comments. Their outputs are incorporated into the Red Team Challenges and Suprem
      quoted
    • low Risky intent intent-offensive-security modules/expert-roles.md:70
      Offensive-security / dual-use content (legitimate for authorised testing; review intended use) (quoted — discussed, not commanded)
      "role": "Red Team Commander",
      quoted
    • low Risky intent intent-offensive-security modules/expert-roles.md:205
      Offensive-security / dual-use content (legitimate for authorised testing; review intended use)
      | Red Team | Red Team Commander | - |
    • low Risky intent intent-offensive-security modules/expert-roles.md:257
      Offensive-security / dual-use content (legitimate for authorised testing; review intended use)
      - **Sonnet** for Strategist, Intel, Tactician, Red Team — strong reasoning at moderate cost
    • low Risky intent intent-offensive-security modules/merkle-dag.md:44
      Offensive-security / dual-use content (legitimate for authorised testing; review intended use) (quoted — discussed, not commanded)
      expert_role: str       # "Intelligence Officer", "Red Team", etc.
      quoted
    • low Risky intent intent-offensive-security modules/merkle-dag.md:230
      Offensive-security / dual-use content (legitimate for authorised testing; review intended use)
      # Present to Red Team anonymized
    • low Risky intent intent-offensive-security modules/reversibility-assessment.md:142
      Offensive-security / dual-use content (legitimate for authorised testing; review intended use)
      *Recommendation*: Convene full council. Extensive Red Team review required.
    • low Risky intent intent-offensive-security SKILL.md:76
      Offensive-security / dual-use content (legitimate for authorised testing; review intended use)
      | Red Team | Gemini Flash | Adversarial challenge, failure modes |
    • low Risky intent intent-offensive-security SKILL.md:87
      Offensive-security / dual-use content (legitimate for authorised testing; review intended use)
      | Red Team Commander | Gemini Flash | Adversarial challenge |
    • low Risky intent intent-offensive-security SKILL.md:102
      Offensive-security / dual-use content (legitimate for authorised testing; review intended use)
      - Phase 4: Red Team Review (all COAs)
    • low Risky intent intent-offensive-security SKILL.md:278
      Offensive-security / dual-use content (legitimate for authorised testing; review intended use)
      - Red Team challenges
    • low Risky intent intent-offensive-security SKILL.md:385
      Offensive-security / dual-use content (legitimate for authorised testing; review intended use)
      4. **Phase 4 (Red Team)**: Red-team teammate receives all COAs, posts challenges; other teammates can **respond to challenges in real-time**

    Files scanned: 8. Evidence is masked. Grey chips explain why severity was lowered.

    Against the Agent Skills spec

    • note frontmatter-key unknown frontmatter key "triggers"
    • note frontmatter-key unknown frontmatter key "source"
    • note frontmatter-key unknown frontmatter key "source_plugin"

    Process rating: all ten parameters 60/100

    • 0Inputs and preconditions. Does not say what the process needs to start
    • 0Progress reporting. Says nothing while it works
    • 20When it triggers. No condition that starts the skill
    • 30Running it twice. 13 mutating operations with no state check
    • 40Consistency. Frontmatter name (war-room) differs from the folder (nm-attune-war-room)
    • 60Result and completion. Output format stated, no completion criterion
    • 60Failures and branches. 2 branches
    • 100Tools and files. No external tools needed
    • 100Steps. 77 steps
    • 100Execution cost. Instruction body is 3741 tokens
    • low 16 top-level sections: this looks like several domains in one skill
    • high The skill tells the model to perform an irreversible action with no human approval

    Everything here is measured from the skill text rather than judged by a model, so the numbers are checkable. A parameter weighs more when it is a more common reason for the process to stall.

    Quality signals

    • +5Description has no quoted example phrases that should trigger the skill
    • +4Description does not say when NOT to use the skill (false activations)
    • +3Description length 76: 120–800 characters recommended
    • +1No license
    • +2Single-language instructions
    • +4Structure: 47 headings
    • +3Step-by-step instructions: 77 items
    • +3Output format is stated explicitly
    • +4Has examples (15 code blocks)

    Quality base 70; lint remarks subtract, signals add up to 100. Result: 81.

    External checks

    ClawHub: suspicious
    This decision-review skill is mostly transparent, but it defaults to publishing sensitive deliberations to GitHub and recommends permission-bypassing execution for one external model.
    LLM: suspicious (high) · 26 Aug 2026