AC zmm-benchmark
📐 詹明明·找对标 ——找对标。给一个方向就去抖音/小红书/视频号搜人,给一个名字就去认人;三筛过滤(赚钱 / 看懂 / 能仿)挑出真正值得抄的那个,然后把他的**三批内容**全扒下来——最早 10 条(他怎么起的号)、数据最好 10 条(什么能爆)、最新 10 条(他现在在哪)——封面、大字、逐字稿、互动数据一条不落,最后出完整拆解和抄袭路线图。 触发方式:/zmm-benchmark、/找对标、/对标、「我该学谁」「帮我找个对标」「这个号值不值得学」「把这个博主拆一下」「他是怎么起号的」「扒一下这个账号」 Find and dissect a benchmark creator: search by direction or by name, filter on money/understandable/copyable, then pull the earliest 10, best-performing 10, and latest 10 posts with covers, cover text, transcripts and engagement data, and produce a full teardown plus a copy roadmap. Trigger: /zmm-benchmark, "find me a benchmark account", "is this creator worth learning from", "tear down this account" —— 📐 詹明明 · 不给公式,给判据。每条规则都标了实测代价。
📐 詹明明·找对标 ——找对标。给一个方向就去抖音/小红书/视频号搜人,给一个名字就去认人;三筛过滤(赚钱 / 看懂 / 能仿)挑出真正值得抄的那个,然后把他的三批内容全扒下来——最早 10 条(他怎么起的号)、数据最好 10 条(什么能爆)、最新 10…
As a process C 53/100 · Has gaps — weak spots: result and completion, when it triggers, inputs and preconditions
How to improve
- Your own cases (evals/evals.json, 4–6 real requests with expected answers): the full check would then run those instead of a model-drafted suite.
- A spec.yaml with trigger phrases and assertions — a behaviour contract for CI; `skilltest init` writes a template.
Guard findings · 0
✓ No critical or high findings
Files scanned: 5. Evidence is masked. Grey chips explain why severity was lowered.
Against the Agent Skills spec
- note
frontmatter-keyunknown frontmatter key "slug" - note
frontmatter-keyunknown frontmatter key "displayName"
Process rating: all ten parameters 53/100
- 0Result and completion. Does not say what the result is
- 0Inputs and preconditions. Does not say what the process needs to start
- 0Failures and branches. Linear process with no failure handling
- 0Progress reporting. Says nothing while it works
- 20When it triggers. No condition that starts the skill
- 100Tools and files. No external tools needed
- 100Steps. 22 steps
- 100Consistency. Name and required fields are in place
- 100Execution cost. Instruction body is 1743 tokens
- 100Running it twice. No mutating operations
Everything here is measured from the skill text rather than judged by a model, so the numbers are checkable. A parameter weighs more when it is a more common reason for the process to stall.
Quality signals
- +4Description does not say when NOT to use the skill (false activations)
- +3Output format is not stated: the model decides each time
- +1No license
- +2Single-language instructions
- +5Description quotes 3 example trigger phrases
- +3Description length 697: enough signal without eating the budget
- +4Structure: 26 headings
- +3Step-by-step instructions: 22 items
- +4Has examples (4 code blocks)
- +4Reference files are cited in the instructions (3 of 3)
Quality base 70; lint remarks subtract, signals add up to 100. Result: 91.