SKILLEMALL.ai

CC launch

Launch Code OSS (VS Code from sources) into an isolated throwaway profile with unique debug ports so you can drive it with @playwright/cli AND attach a Node debugger via dap-cli in the same session. Use when working on VS Code itself and you want to interact with the running workbench, automate chat or UI flows, test UI features, take screenshots, set breakpoints in the renderer / extension host / main process, or combine UI driving with debugging.

The skillemall take

The skill launches VS Code from sources into an isolated profile with unique debug ports. It promises to drive the UI via Playwright and attach a Node debugger through dap-cli simultaneously — for working on the editor itself, automating chat flows, testing UI features, taking screenshots, and setting breakpoints across renderer, extension host, and main process.

The package contains 13 files and 12 scripts. Quality score sits at 69, process score at 57 — mediocre. No critical issues, but five medium and low-severity findings. Sandbox tests showed quiet behavior: launch.sh exited with code 133, others stayed within their folder. Platform support is broad.

Install if you're actually hacking on VS Code and need UI automation combined with debugging. For regular use, it's overkill.

microsoft/vscode Agent Skills author: microsoft MIT 13 files · 12 scripts body ≈ 8 576 tokens Open the sourcegithub.com↗ analyzed 2 d ago

Launch Code OSS (VS Code from sources) into an isolated throwaway profile with unique debug ports so you can drive it with @playwright/cli AND attach a Node…

As a process C 57/100 · Has gaps — weak spots: result and completion, when it triggers, execution cost

ProcedureVS CodePlaywrightGitHubSoftware developmentAI and agentstype and topics are labelled automatically from the skill text
JSON
Technical rating
C
85/100
safety, quality, tests
Safety 60%
95
Quality 40%
69
Run on models
none yet
Process rating
C
57/100
Has gaps
Result and completion w 14
0
When it triggers w 12
20
Running it twice w 4
30
the three weakest of ten parameters · all ten

How to improve

  1. The SKILL.md body is over 5,000 tokens: move reference detail into references/ and load it when needed.
For the model run — optional
  • Your own cases (evals/evals.json, 4–6 real requests with expected answers): the full check would then run those instead of a model-drafted suite.
  • A spec.yaml with trigger phrases and assertions — a behaviour contract for CI; `skilltest init` writes a template.

Guard findings · 5

✓ No critical or high findings

Medium and low: 5
  • low Dangerous commands cmd-execpolicy-bypass scripts/bootstrap/bootstrap-profile.ps1:74
    Runs PowerShell with execution policy bypassed (detector / deny-list definition; string literal in code, not executed)
    $monitorArguments = "-NoProfile -NonInteractive -ExecutionPolicy Bypass -File `"$monitorScript`" -RootProcessId $($process.Id) -ExitMarker `"$exitMarker`" -StopMarker `"$stopMarker`""
    detectorcode literal
  • low Dangerous commands cmd-background-process scripts/bootstrap/bootstrap-profile.sh:113
    Starts a background / autostarted process
    disown $! 2>/dev/null || true
  • low Dangerous commands cmd-background-process scripts/launch.sh:261
    Starts a background / autostarted process (string literal in code, not executed; code comment)
    # Launch code.sh in the background. Detaching with `nohup ... & disown` is
    code literalcomment
  • low Dangerous commands cmd-background-process scripts/launch.sh:269
    Starts a background / autostarted process
    disown $PID 2>/dev/null || true
  • low Dangerous commands cmd-execpolicy-bypass SKILL.md:78
    Runs PowerShell with execution policy bypassed (detector / deny-list definition)
    If the local execution policy blocks scripts, invoke it with `powershell -ExecutionPolicy Bypass -File <path…ps1>`. The Windows implementation has the same profile isolation, slim-copy exclu
    detector

Files scanned: 13. Evidence is masked. Grey chips explain why severity was lowered.

Against the Agent Skills spec

  • warning body-long SKILL.md body ≈ 8576 tokens (recommended < 5000); move details to references/
  • note edit-residue the text marks something as outdated (lines 149): check that old rules are not kept next to new ones — the full check reads the text for contradictions

Process rating: all ten parameters 57/100

  • 0Result and completion. Does not say what the result is
  • 20When it triggers. No condition that starts the skill
  • 30Running it twice. 2 mutating operations with no state check
  • 40Execution cost. Instruction body is 8576 tokens: crowds the task out of the window
  • 60Tools and files. Uses tools (bash, node) that frontmatter does not declare
  • 70Inputs and preconditions. Inputs and preconditions are listed
  • 85Steps. 44 steps, 2 vague phrases
  • 100Failures and branches. 5 branches, has a failure section
  • 100Consistency. Name and required fields are in place
  • 100Progress reporting. Reports progress
  • medium Safety rules and hard prohibitions inside a skill: they belong in the system prompt, here they protect nothing
  • low The response is described with custom markup (9 tags): a typed call is more reliable

Everything here is measured from the skill text rather than judged by a model, so the numbers are checkable. A parameter weighs more when it is a more common reason for the process to stall.

Quality signals

  • +5Description has no quoted example phrases that should trigger the skill
  • +4Description does not say when NOT to use the skill (false activations)
  • +3Output format is not stated: the model decides each time
  • -2localhost URLs: will not work for another user
  • -32 of 6 scripts are never mentioned in SKILL.md
  • +1No license
  • +2Single-language instructions
  • +3Description length 452: enough signal without eating the budget
  • +4Structure: 20 headings
  • +3Step-by-step instructions: 44 items
  • +4Has examples (27 code blocks)

Quality base 70; lint remarks subtract, signals add up to 100. Result: 69.

In the sandbox The scripts kept to themselves

The skill's scripts were run in a throwaway machine: no network, fake keys in the home directory, a tracer watching. We wrote down what they did. Reaching for the network or for secrets caps the technical grade at C; a quiet run adds no points.

Запущено 7 скриптов; каждому дали двадцать секунд, поддельный домашний каталог с ключами и сеть, в которой ничего нет.

Из них 1 не дошёл до работы, и об их поведении мы ничего не узнали.

playwrightScripts/focus-chat-input.tsничего за пределами своей папки
scripts/launch.shзавершился с кодом 133
scripts/monaco-paste.shничего за пределами своей папки
scripts/monaco-paste.tsничего за пределами своей папки
scripts/updateSettings.tsничего за пределами своей папки
scripts/waitForCdp.tsничего за пределами своей папки
scripts/bootstrap/bootstrap-profile.shне запустился: Could not find an executable Code OSS launcher at //scripts/code.sh.

3 Oct 2026