The gist
Imagine you have a mentor who asks the right questions, helps you put a process into words, and then tests the result on their own, showing you what works and what doesn't. That's exactly how Skill Creator works: you describe what you want, it structures it, builds it and evaluates it. You iterate.
Key concepts
- Skill Creator: the
skill-creatorplugin for creating and checking skills (from Anthropic's official catalog) - Eval Framework: automatic testing of a skill with a pass rate and metrics
- Trigger tuning: adjusting the description (
description) the agent (an autonomous task performer) uses to recognize when to use the skill - A live example: YouTube Weekly Roundup with 3 parallel agents
- Iteration: quality grows from version to version
Theory
Skill Creator: the skill that creates skills
It's like hiring a specialist in writing instructions. You say "I want to automate this," and it asks the right questions, structures the process, generates the frontmatter (metadata at the top of the file) and the workflow, and then checks the quality on its own.
Installation (the command from the Claude Code docs as of October 2026; the /plugin menu will show the exact catalog name):
/plugin install skill-creator@claude-plugins-official
Running it:
Create a skill for a weekly analysis of a YouTube channel
Skill Creator asks follow-up questions:
- "Which metrics should we track?"
- "What format do you need the report in?"
- "Do we need to compare against competitors?"
- "Which model do you prefer?"
You answer, and it builds.
A live example: YouTube Weekly Roundup
We're building a skill that analyzes a YouTube channel every week and produces a structured report.
What the skill does:
- Pulls channel data for the last 7 days (views, subscribers, ER)
- Analyzes each video: CTR, retention, comment sentiment
- A SWOT analysis of the current period
- A comparison with competitors (the top 3 channels in the niche)
- Trending content in the niche this week
- Recommendations for next week
- A branded PDF report
Architecture: 3 parallel agents
Instead of one agent doing everything in sequence, three work at the same time:
Coordinating agent
├── Agent 1: Data collection
│ ├── YouTube Analytics API
│ ├── Data on all videos for 7 days
│ └── Returns: JSON with metrics
│
├── Agent 2: Competitor analysis
│ ├── Data on the top 3 competitors
│ ├── Benchmarking vs. our channel
│ └── Returns: a comparison table
│
└── Agent 3: Trending content
├── Searching for trends in the niche
├── Analyzing viral videos
└── Returns: a list of ideas
→ The coordinating agent puts it all together → generates the report → renders the PDFWhy in parallel? Collecting YouTube data, analyzing competitors and searching for trends are independent tasks. Roughly: 15 minutes in sequence, 5 in parallel (the numbers are an illustration). But each agent has its own context, so you don't use fewer tokens, usually more: parallelism saves time, not money.
The Eval Framework: automatic quality checks
After Skill Creator has generated a skill, you can ask it to check the skill (for example: "evaluate my youtube-weekly-roundup skill with skill-creator"). What it does:
How it works:
- Test cases are stored in the
evals/evals.jsonfile inside the skill folder - Each test runs in isolation, in its own subagent
- The result is checked against assertions, and the scores are written to
grading.json - You can compare the work with the skill and without it (
benchmark.json) and compare versions of the skill against each other - The description is tuned separately: it picks requests the skill should and shouldn't fire on
What gets tested:
- 10-20 different prompts that should activate the skill
- Whether the skill activates correctly (trigger precision)
- Output quality (against set criteria)
- Run time
- Token usage
Eval metrics (the format is illustrative, the numbers are an illustration):
Eval Results: YouTube Weekly Roundup v1.0
─────────────────────────────────────────
Trigger accuracy: 78% (7/9 correct activations)
Output quality: 6.2/10
Avg tokens (tokens are units of text for AI): 3,847
Avg time: 47 seconds
Pass rate: 44% (4/9 tests passed)
─────────────────────────────────────────
Problems:
- The "quick channel overview" test didn't activate the skill (78% trigger)
- The SWOT section is too generic (6.2 quality)
- The competitor analysis ignores the region (44% pass rate)It's not just "works / doesn't work." These are concrete numbers showing where it fails.
Trigger tuning: how the agent "hears" a request
The problem: a user writes "give me an overview of the channel for the week" and the skill doesn't activate. They write "youtube weekly roundup" and it activates. Why?
Because the conditions for firing are spelled out too narrowly in the description. A skill has no separate list of triggers: the agent decides based on the text of description (and the optional when_to_use). You need to broaden the description:
# Before tuning
description: >
Weekly analysis of a YouTube channel.
Use for "youtube weekly roundup" and "weekly YouTube report".
# After tuning
description: >
Weekly analysis of a YouTube channel: metrics, competitors, recommendations.
Use for requests like "youtube weekly roundup",
"weekly YouTube report", "channel overview for the week",
"how did the channel do this week", "YouTube stats for 7 days",
"channel analytics". Do not use for reviewing a single video.Based on the eval results, Skill Creator suggests refining the description. Keep the limit in mind: in the skills list, the description is cut off at about 1,500 characters, so put the most important part first.
The iteration process
Iteration 1 (pass rate 44%):
- You see the trigger misses requests
- You see the SWOT is generic
- You see the region isn't taken into account
What you do:
The skill showed a pass rate of 44%. Problems: 1. Broaden the description with variations of the "channel overview" request 2. In the SWOT, add concrete numbers, not generic words 3. Add a region parameter: "global" by default, a country can be specified
Skill Creator updates the skill and runs the eval again.
Below is an illustrative trajectory (the numbers are an illustration, not a promised result: for you, everything will depend on the task).
Iteration 5 (pass rate 71%): Better, but quality is still 7.1/10. The report is dull, no visualizations.
Iteration 10 (pass rate 89%, quality 8.4/10): The skill works reliably. A report with charts, the right region, accurate triggers.
By iterations 20-30: the skill becomes your own personally tuned tool.
The frontmatter and skill folder after Skill Creator
---
name: youtube-weekly-roundup
description: >
Weekly analysis of a YouTube channel with competitor benchmarking
and a PDF report. Use for requests like "channel overview for the week",
"youtube weekly roundup", "channel analytics".
argument-hint: "[channel_id] [region]"
---And the folder structure:
youtube-weekly-roundup/
SKILL.md # instructions: steps, rules, report format
brand/guidelines.md # branding for the layout
reports/weekly-template.html # report template
data/competitors.json # list of competitors
evals/evals.json # test cases for checking the skillThe parameters (channel_id, region, competitors), the set of subagents, the report format and links to files are described in the text of SKILL.md in plain words. Claude Code has no separate parameters, subagents, output, references, eval, model_invocation, tags or version fields in the frontmatter: they appeared in older versions of the course, but they don't work.
Practice
Assignment: Create your own skill with Skill Creator
- Pick a task you do regularly (at least once a week):
- Examples: "weekly sales analysis," "preparing a content plan," "monitoring brand mentions"
- Install Skill Creator:
/plugin install skill-creator@claude-plugins-official(or find theskill-creatorplugin through the/pluginmenu) - Run it: describe your task
- Answer all the follow-up questions; the more detail, the better
- Get the first version of the skill
- Ask Skill Creator to check the skill ("evaluate my skill with skill-creator") and look at the pass rate and the test results
- Do at least 3 rounds of improvements
- Write down: what changed between version 1 and version 3?
Tools and resources
skill-creator: the plugin for creating and checking skills (/plugin install skill-creator@claude-plugins-official)- Eval Framework: part of Skill Creator, runs on request (tests in
evals/evals.json) - Claude Code Skills documentation: the official guide to skills
.claude/skills/<name>/SKILL.md: the folder where your skills are stored- YouTube Data API v3 (API: an application programming interface): for getting real channel data (requires an API key)
Key takeaways
Skill Creator asks the right questions, better than if you came up with the structure from scratch yourself. Trust the process.
A low pass rate on the first version is normal. A skill is built through iterations, not on the first try. In the example above, the pass rate rose to ~90% by the tenth iteration (illustrative numbers).
3 parallel agents instead of 1 sequential one = noticeably faster. But tokens aren't free: each agent has its own context, so usage is usually higher. Parallelism speeds things up, but doesn't make them cheaper.
Related lessons
- ← What Skills are: the basic concepts of skills, frontmatter, progressive loading
- → Skill architecture: Capability Uplift vs. Encoded Preference, and when to use which
- → Evals: self-improving skills: a deep dive into testing skills
Next lesson
→ Skill architecture: 2 archetypes
The mark stays in this browser only and is never sent anywhere. My progress