Library · Skills: teach the agent your way of working

Building a skill (a reusable instruction for Claude) from scratch, LIVE: Skill Creator + the Eval (automatic quality check) Framework

Builder80 minUpdated: October 2026
30 of 105 in the library

Module: Skills: reusable expertise | Time: ~20 min theory + 60 min practice


The gist

Imagine you have a mentor who asks the right questions, helps you put a process into words, and then tests the result on their own, showing you what works and what doesn't. That's exactly how Skill Creator works: you describe what you want, it structures it, builds it and evaluates it. You iterate.


Key concepts

  • Skill Creator: the skill-creator plugin for creating and checking skills (from Anthropic's official catalog)
  • Eval Framework: automatic testing of a skill with a pass rate and metrics
  • Trigger tuning: adjusting the description (description) the agent (an autonomous task performer) uses to recognize when to use the skill
  • A live example: YouTube Weekly Roundup with 3 parallel agents
  • Iteration: quality grows from version to version

Theory

Skill Creator: the skill that creates skills

It's like hiring a specialist in writing instructions. You say "I want to automate this," and it asks the right questions, structures the process, generates the frontmatter (metadata at the top of the file) and the workflow, and then checks the quality on its own.

🎨 Picture this: Skill Creator is like an experienced architect at the first meeting with a client. You say "I want a house." He doesn't start drawing right away. He asks: how many rooms, do you have kids, what style, what budget. Only after the right questions does he draw a plan.

Installation (the command from the Claude Code docs as of October 2026; the /plugin menu will show the exact catalog name):

Type this into the chat
/plugin install skill-creator@claude-plugins-official

Running it:

Type this into the chat
Create a skill for a weekly analysis of a YouTube channel

Skill Creator asks follow-up questions:

  • "Which metrics should we track?"
  • "What format do you need the report in?"
  • "Do we need to compare against competitors?"
  • "Which model do you prefer?"

You answer, and it builds.


A live example: YouTube Weekly Roundup

We're building a skill that analyzes a YouTube channel every week and produces a structured report.

What the skill does:

  1. Pulls channel data for the last 7 days (views, subscribers, ER)
  2. Analyzes each video: CTR, retention, comment sentiment
  3. A SWOT analysis of the current period
  4. A comparison with competitors (the top 3 channels in the niche)
  5. Trending content in the niche this week
  6. Recommendations for next week
  7. A branded PDF report

Architecture: 3 parallel agents

Instead of one agent doing everything in sequence, three work at the same time:

Code
Coordinating agent
├── Agent 1: Data collection
│   ├── YouTube Analytics API
│   ├── Data on all videos for 7 days
│   └── Returns: JSON with metrics
│
├── Agent 2: Competitor analysis
│   ├── Data on the top 3 competitors
│   ├── Benchmarking vs. our channel
│   └── Returns: a comparison table
│
└── Agent 3: Trending content
    ├── Searching for trends in the niche
    ├── Analyzing viral videos
    └── Returns: a list of ideas

→ The coordinating agent puts it all together → generates the report → renders the PDF

🎨 Picture this: three parallel agents are like three cooks in one kitchen. One chops vegetables, the second makes the broth, the third grills the meat. All at once. One cook makes the same dish, but it takes three times as long.

Why in parallel? Collecting YouTube data, analyzing competitors and searching for trends are independent tasks. Roughly: 15 minutes in sequence, 5 in parallel (the numbers are an illustration). But each agent has its own context, so you don't use fewer tokens, usually more: parallelism saves time, not money.


The Eval Framework: automatic quality checks

After Skill Creator has generated a skill, you can ask it to check the skill (for example: "evaluate my youtube-weekly-roundup skill with skill-creator"). What it does:

🎨 Picture this: an eval is a test drive of a car after it's assembled. The engineer doesn't just look it over. He drives it up and down hills, brakes, accelerates. He finds what squeaks in the turns. Fixes it. Drives again.

How it works:

  • Test cases are stored in the evals/evals.json file inside the skill folder
  • Each test runs in isolation, in its own subagent
  • The result is checked against assertions, and the scores are written to grading.json
  • You can compare the work with the skill and without it (benchmark.json) and compare versions of the skill against each other
  • The description is tuned separately: it picks requests the skill should and shouldn't fire on

What gets tested:

  • 10-20 different prompts that should activate the skill
  • Whether the skill activates correctly (trigger precision)
  • Output quality (against set criteria)
  • Run time
  • Token usage

Eval metrics (the format is illustrative, the numbers are an illustration):

Code
Eval Results: YouTube Weekly Roundup v1.0
─────────────────────────────────────────
Trigger accuracy:    78% (7/9 correct activations)
Output quality:      6.2/10
Avg tokens (tokens are units of text for AI):  3,847
Avg time:            47 seconds
Pass rate:           44% (4/9 tests passed)
─────────────────────────────────────────
Problems:
- The "quick channel overview" test didn't activate the skill (78% trigger)
- The SWOT section is too generic (6.2 quality)
- The competitor analysis ignores the region (44% pass rate)

It's not just "works / doesn't work." These are concrete numbers showing where it fails.


Trigger tuning: how the agent "hears" a request

🎨 Picture this: trigger tuning is like adjusting a door lock. Too tight, and the right key won't go in. Too loose, and any key goes in. Ideally: only your key, from either side, with either hand.

The problem: a user writes "give me an overview of the channel for the week" and the skill doesn't activate. They write "youtube weekly roundup" and it activates. Why?

Because the conditions for firing are spelled out too narrowly in the description. A skill has no separate list of triggers: the agent decides based on the text of description (and the optional when_to_use). You need to broaden the description:

yaml
# Before tuning
description: >
  Weekly analysis of a YouTube channel.
  Use for "youtube weekly roundup" and "weekly YouTube report".

# After tuning
description: >
  Weekly analysis of a YouTube channel: metrics, competitors, recommendations.
  Use for requests like "youtube weekly roundup",
  "weekly YouTube report", "channel overview for the week",
  "how did the channel do this week", "YouTube stats for 7 days",
  "channel analytics". Do not use for reviewing a single video.

Based on the eval results, Skill Creator suggests refining the description. Keep the limit in mind: in the skills list, the description is cut off at about 1,500 characters, so put the most important part first.


The iteration process

🎨 Picture this: iterating on a skill is like a sniper sighting in a rifle. The first shot rarely hits the center. You look where it landed. Adjust the scope. Shoot again. By the tenth shot, bull's-eye.

Iteration 1 (pass rate 44%):

  • You see the trigger misses requests
  • You see the SWOT is generic
  • You see the region isn't taken into account

What you do:

Type this into the chat
The skill showed a pass rate of 44%. Problems:
1. Broaden the description with variations of the "channel overview" request
2. In the SWOT, add concrete numbers, not generic words
3. Add a region parameter: "global" by default, a country can be specified

Skill Creator updates the skill and runs the eval again.

Below is an illustrative trajectory (the numbers are an illustration, not a promised result: for you, everything will depend on the task).

Iteration 5 (pass rate 71%): Better, but quality is still 7.1/10. The report is dull, no visualizations.

Iteration 10 (pass rate 89%, quality 8.4/10): The skill works reliably. A report with charts, the right region, accurate triggers.

By iterations 20-30: the skill becomes your own personally tuned tool.


The frontmatter and skill folder after Skill Creator

yaml
---
name: youtube-weekly-roundup
description: >
  Weekly analysis of a YouTube channel with competitor benchmarking
  and a PDF report. Use for requests like "channel overview for the week",
  "youtube weekly roundup", "channel analytics".
argument-hint: "[channel_id] [region]"
---

And the folder structure:

Code
youtube-weekly-roundup/
  SKILL.md                      # instructions: steps, rules, report format
  brand/guidelines.md           # branding for the layout
  reports/weekly-template.html  # report template
  data/competitors.json         # list of competitors
  evals/evals.json              # test cases for checking the skill

The parameters (channel_id, region, competitors), the set of subagents, the report format and links to files are described in the text of SKILL.md in plain words. Claude Code has no separate parameters, subagents, output, references, eval, model_invocation, tags or version fields in the frontmatter: they appeared in older versions of the course, but they don't work.


Practice

Assignment: Create your own skill with Skill Creator

  1. Pick a task you do regularly (at least once a week):
    • Examples: "weekly sales analysis," "preparing a content plan," "monitoring brand mentions"
  2. Install Skill Creator: /plugin install skill-creator@claude-plugins-official (or find the skill-creator plugin through the /plugin menu)
  3. Run it: describe your task
  4. Answer all the follow-up questions; the more detail, the better
  5. Get the first version of the skill
  6. Ask Skill Creator to check the skill ("evaluate my skill with skill-creator") and look at the pass rate and the test results
  7. Do at least 3 rounds of improvements
  8. Write down: what changed between version 1 and version 3?

Tools and resources

  • skill-creator: the plugin for creating and checking skills (/plugin install skill-creator@claude-plugins-official)
  • Eval Framework: part of Skill Creator, runs on request (tests in evals/evals.json)
  • Claude Code Skills documentation: the official guide to skills
  • .claude/skills/<name>/SKILL.md: the folder where your skills are stored
  • YouTube Data API v3 (API: an application programming interface): for getting real channel data (requires an API key)

Key takeaways

Skill Creator asks the right questions, better than if you came up with the structure from scratch yourself. Trust the process.

A low pass rate on the first version is normal. A skill is built through iterations, not on the first try. In the example above, the pass rate rose to ~90% by the tenth iteration (illustrative numbers).

3 parallel agents instead of 1 sequential one = noticeably faster. But tokens aren't free: each agent has its own context, so usage is usually higher. Parallelism speeds things up, but doesn't make them cheaper.



Next lesson

→ Skill architecture: 2 archetypes

The mark stays in this browser only and is never sent anywhere. My progress