SKILLEMALL.ai

Skills in a company: a process, not a folder of files

Skills appear in a company by themselves. An analyst writes one so the agent can answer questions about reporting; support copies something from an open catalog and edits it; a developer commits a folder with instructions and forgets it for six months. A year later there are a hundred of them, and nobody can answer the simple question — which ones still work, which duplicate each other, which quietly send data outside. We put a process in place that answers it on any given day, and leave it with you along with the people who can keep it running.

A month into the pilot, one team has a list of its skills with grades, a gate that keeps the dangerous ones out of the shared catalog, and people who can keep both running.

What we do

1. Authoring: so skills are written the same way

A company standard: what a SKILL.md must contain, how triggers are described, which tests have to ship with a skill for it to be accepted at all. Templates for your recurring tasks, and naming rules so two skills stop intercepting the same user request — the most common reason behind "the agent somehow doesn't do what we asked".

2. Verification: a gate before publishing, and a run every month

The engine behind the public rating moves inside your process. Before publication a skill passes a gate: a scan for secrets, auto-run commands, injections and instruction overrides — in Russian as well as English, because English-only scanners do not see "игнорируй предыдущие инструкции". The gate sits on pre-commit and in the build, so a dangerous skill never reaches the shared catalog.

Then the regular model run: every case runs twice, with the skill and without, which shows whether the skill still adds anything. It works on the models you already have, including GigaChat and YandexGPT.

Models change, and a skill that worked in March can quietly stop firing by autumn. The run catches that before your users do.

The result is an internal catalog with the same two grades as the public one: technical (safety and the quality of the text) and process maturity (does the skill have tests, a version, a license, signs of upkeep). Private, yours only.

3. Training: so it lives without us

A workshop for teams: how a skill actually works, why it fires when you did not expect it, how to write checkable assertions instead of wishes. We work through your own skills rather than abstract examples — usually the most useful part. Plus a short session for whoever will hold the gate: how to read a report, when a finding is real and when the scanner is quibbling with a line of documentation.

How the work goes

Audit. We take what is already written, run it through the checks and show you the map: what is alive, what is duplicated, what is dangerous, where keys sit in plain text. This is usually where it turns out there are twice as many skills as anyone thought. The audit is what the estimate for everything else is built on.

Pilot. One team: the standard, the gate before publishing, the internal catalog. A month, to find out whether it takes root before it touches everyone.

Rollout. The remaining teams, the hook into your build, training, handover of the rules.

Upkeep. A monthly run and report, rules updated as new models appear. Or, if you prefer, we simply send the report and you hold the process yourself.

What stays with you

What it costs

Start with the audit: a couple of weeks, inexpensive — and an honest estimate for the rest, along with an answer to whether you need a rollout at all.

After that the price depends on the number of teams and how many skills are already written. The public service stays free meanwhile: the rating and the quick scan need neither payment nor an account, and you can use them in parallel.

Write to us about a rollout

We reply to that address. Nothing but a reply address is required.