Writing Guide

Overview introduces where Agent Skills came from, how they are formatted, and how progressive disclosure works. This page starts from the official specification and Anthropic’s authoring best practices, then moves into the development process: deciding whether a capability belongs in a Skill, writing SKILL.md, testing trigger behavior and output quality, and iterating without overfitting to a few examples.

Writing a Skill is not the same as moving a manual into a folder. A good Skill lets an agent discover, load, execute, and reuse a class of capability. Effective Skills usually have three properties: trustworthy sources, a compact structure, and validation against realistic tasks.


Official Authoring Principles

The center of the official best practices is not “write more rules.” It is to make the Skill light enough, precise enough, and verifiable enough to work inside the agent’s context.

PrincipleAuthoring implication
Concise is keyOnce triggered, SKILL.md enters context; every paragraph should justify its token cost.
Assume the model is already smartAdd project conventions, operation paths, boundaries, and validation methods, not general knowledge.
Set the right degree of freedomNarrow fragile workflows; leave room for judgment in open tasks.
Test target modelsSkill behavior depends on the underlying model; test important Skills on the models that will use them.
Description drives discoveryInclude what the Skill does, when to use it, and concrete trigger terms.
Progressive disclosureTreat SKILL.md as an entry point; move depth into shallow references, assets, or scripts.
Build validation loopsAutomate verification where possible; give clear review criteria where automation does not fit.
Avoid stale knowledgePoint time-sensitive or version-specific information to a source of truth instead of freezing it as a long-term rule.

The authoring loop below is a practical workflow built on top of these principles.


Common Structure Choices

The official specification does not require Skills to follow named design patterns. In practice, authors often choose among a few lightweight structures:

StructureBest fitCommon resources
Tool WrapperAdd tool, framework, or team conventions the agent needsreferences/
GeneratorProduce fixed-structure reports, configs, pages, emails, or scaffoldsassets/ + references/
ReviewerReview code, documents, designs, or compliance risks against a checklist or rubricreferences/
Ask FirstMissing goals, constraints, inputs, or acceptance criteria would lead to wrong workreferences/ or assets/
PipelineMulti-step tasks need order, intermediate artifacts, and validationreferences/ + assets/ + scripts/

These structures are useful only when they implement the official principles: keep SKILL.md concise, make the description trigger correctly, load resources on demand, and validate against real tasks. If a Skill only needs one page of instructions, do not add templates or scripts just to fit a structure.


The Authoring Loop

The first version of a Skill is usually a hypothesis. Quality comes from iteration: capture intent, draft, test, observe, and improve.

Authoring Loop for Skills Draft → Test → Review → Improve → Repeat ① Capture Intent once per skill ② Draft pick a pattern ③ Test with-skill runs ④ Review qualitative + metrics ⑤ Improve generalize, don't overfit SKILL.md the artifact you shape run grade feedback revise

  1. Capture intent: Define the capability boundary, trigger cases, and non-trigger cases.
  2. Check sources: Confirm that rules, workflows, tool behavior, terminology, and examples are reliable and versioned when needed.
  3. Draft the smallest structure: Start with the smallest useful SKILL.md; add references/, assets/, and scripts/ only when they earn their place.
  4. Test on realistic tasks: Compare runs with and without the Skill. For important Skills, test the target models.
  5. Read transcripts: Did the agent trigger the Skill, load the right resources, wander, or misunderstand constraints?
  6. Improve failure modes: Fix classes of problems instead of patching individual prompts.

Keep each loop light. Complexity should come from test evidence, not from authoring-time imagination.


1. Capture Intent

Answer these questions before writing:

QuestionWhy it matters
What should this Skill help the agent do?The description and body should describe a capability, not a folder layout.
When should it trigger?The agent sees metadata first. A weak trigger can hide an otherwise good Skill.
When should it not trigger?Near-miss cases prevent the Skill from stealing adjacent tasks.
What is the output?A file, patch, review, table, report, or command result determines how you test it.
What are the sources of truth?Technical rules, product behavior, research claims, and policies need verifiable sources.
Does it depend on tools or permissions?A Skill can guide the agent to use available tools. It cannot grant missing permissions.
Which steps should be scripted?Parsing, conversion, validation, and other deterministic work often belong in scripts/.

If the user is already mid-conversation and says “turn this into a Skill,” extract the answers from the conversation: the goal, corrections, tools used, accepted output, and failed paths. Ask only when a missing answer would change the design.


2. Check Sources

Skills are reused, so wrong knowledge gets reused too. Verify the content before baking it into instructions.

Check especially:

  • Product or platform behavior: prefer official docs, specifications, and changelogs.
  • APIs, CLIs, and configuration: confirm versions, parameter names, defaults, and deprecations.
  • Papers or method claims: distinguish the paper’s claim, the experiment scope, and secondary commentary.
  • Safety, compliance, legal, finance, or medical guidance: state the scope and avoid turning general advice into a universal rule.
  • Time-sensitive facts: do not encode prices, quotas, rosters, or release dates as if they were long-term rules.

The best Skill content is not “the fact I found today.” It is stable procedure, judgment structure, team convention, template, or verification workflow. If content must change, point to the source of truth and tell the agent when to re-check it.


3. Draft SKILL.md

Frontmatter

A Skill needs at least name and description.

---
name: pdf-processing
description: Extract text and tables from PDFs, fill PDF forms, and merge or split PDF files. Use when the user asks to inspect, transform, validate, or generate PDF artifacts.
---

Frontmatter guidance:

  • name should match the directory name and use lowercase letters, numbers, and hyphens.
  • description should include both what the Skill does and when to use it.
  • Do not put the whole workflow in description. Metadata helps the agent decide whether to load the Skill; the body carries the process.
  • Add license, compatibility, metadata, or allowed-tool fields only when they are useful.

Body

The body should answer three questions:

  1. What should the agent do first after the Skill triggers?
  2. Which resources should it read, and when?
  3. How should it verify completion?

Recommendations:

  • Keep SKILL.md under 500 lines.
  • Move details into references/, templates into assets/, and deterministic operations into scripts/.
  • Keep reference links shallow. The agent should know from SKILL.md which file to read next.
  • Give large reference files a table of contents or clear section names.
  • Do not describe tool capabilities that the runtime does not actually provide.

Multi-Variant Skills

When a Skill covers variants such as AWS/GCP/Azure or Python/TypeScript/Rust, let SKILL.md route to the right reference:

cloud-deploy/
  SKILL.md
  references/
    aws.md
    gcp.md
    azure.md

Write the selection rules in the body: when to read which reference. Do not put every variant into one large body.


4. Set The Right Degree Of Freedom

Reliability does not come from making every instruction forceful. It comes from matching the amount of freedom to the task.

Low freedom fits:

  • Fragile workflows such as releases, migrations, form filling, and data conversion.
  • High-consistency outputs such as compliance reports, configuration files, and fixed templates.
  • Clear safety boundaries such as not committing credentials or bypassing approval.

Use explicit steps, inputs, outputs, validation conditions, and stop conditions.

High freedom fits:

  • Writing, explanation, research, and design review.
  • Open-ended tasks that need synthesis.
  • Tasks where user preference materially changes the result.

Use principles, examples, and evaluation criteria rather than locking down every sentence.

If you are reaching for many all-caps ALWAYS, NEVER, or MUST commands, pause. Often the better move is to explain the reason behind the rule so the model can generalize in cases the Skill did not explicitly cover.


5. Tune The Description

description is the most important signal for whether the agent loads the Skill. It should read like a trigger, not an article summary.

Weak:

How to build a dashboard to display internal data.

Better:

Build dashboards for internal metrics and company data. Use when the user mentions dashboards,
charts, internal reporting, KPI views, operational metrics, or displaying business data.

A good description usually includes:

  • Trigger phrases users actually type.
  • File types, tools, domains, or output forms.
  • Scope boundaries: what is in and what is out.
  • Synonyms and common phrasing, not just the Skill name.

When testing triggers, do not run only positive examples. Include near-miss negatives: tasks that share keywords but should not trigger the Skill. For a PDF Skill, “write a Fibonacci function” is a weak negative. Better negatives are “explain the history of the PDF format” or “build a web page with a PDF download button.”


6. Test And Iterate

Prepare several kinds of tests:

Test typeWhat to check
Positive triggerDoes the Skill load when it should?
Near-miss negativeDoes the Skill stay out of adjacent tasks?
With/without comparisonDoes the Skill improve the result instead of only adding tokens?
Target model testDo the target models follow the Skill?
Transcript reviewDid the agent read the right resources, call the right scripts, and avoid unnecessary loops?
Objective assertionsFile structure, field completeness, builds, unit tests, schemas, validation scripts.

Subjective tasks still need testing, but they may not need automated assertions. Writing quality, explanation clarity, review value, and design judgment are often better evaluated with human review, sample comparison, and user feedback.

Change only one or two things per iteration:

  1. Record the failure mode.
  2. Decide whether it is a trigger issue, resource organization issue, instruction issue, script issue, or a sign that the task does not belong in a Skill.
  3. Change the smallest necessary part.
  4. Re-run positives and near-misses to make sure the fix did not break another behavior.

Do not copy test prompts into the Skill as rules. If the user asks Q4, output X patches one example and weakens generalization.


7. When To Add scripts/

Scripts are useful for stable, repeatable, verifiable work:

  • Parsing files, extracting structured data, converting formats.
  • Running lint, tests, schema validation, PDF rendering, or screenshot checks.
  • Generating fixed assets or batch-processing files.
  • Wrapping complex API calls or command sequences that are easy to get wrong.

Script guidance:

  • Have the script solve the problem directly, not print instructions for the model to continue.
  • Prefer structured inputs and outputs: JSON, CSV, or explicit file paths.
  • Make errors actionable: say what failed and what to try next.
  • Document dependencies and compatibility.
  • Avoid hard-coded local paths or shell-specific assumptions if the Skill should be portable.

Do not add scripts just to make the Skill look advanced. If SKILL.md plus references works reliably, keep it light.


8. Pattern Evolution

A Skill’s structure should evolve with evidence.

Pattern Evolution Through Iteration Let testing reveal when to upgrade the structure Tool Wrapper SKILL.md + references/ starting point 3/3 tests generate similar templates → extract to assets/ Generator + assets/template adds structure steps need enforced order + validation → add scripts/ + gates Pipeline + scripts/ + sequential gates full orchestration Patterns can also shrink — simplify when structure isn't earning its cost.

Common signals:

  • The agent repeatedly generates the same structure: add assets/template.* and evolve from Tool Wrapper to Generator.
  • It repeatedly reviews against the same checklist: move standards into references/checklist.md and make it a Reviewer.
  • It repeatedly skips order or validation: add a stepwise workflow and scripts, evolving toward Pipeline.
  • It repeatedly starts wrong and asks clarifying questions too late: add an Ask First phase.
  • Scripts or references stay unused: delete or merge them to reduce maintenance cost.

The goal is not to make the Skill more complex. The goal is to make the structure just strong enough for the recurring task.


Anti-Patterns

  • Treating the Skill as a knowledge dump. Piling in material without triggers, workflow, and verification makes it hard for the agent to use reliably.
  • Unclear sources. Secondary claims, stale APIs, temporary prices, and old product behavior should not become long-term instructions.
  • Abstract descriptions. “Help with documents” will not trigger reliably.
  • Overlong body. Dense detail in SKILL.md increases context cost and hides the main path.
  • Deep reference chains. If the agent must chase several links to find the rule, the structure probably needs work.
  • Pretending to have tools. A Skill cannot create APIs, accounts, permissions, or network access that the runtime does not provide.
  • Decorative scripts. If a script does not reduce repeated reasoning or improve determinism, it probably should not exist.
  • Testing only positives. Without near-miss negatives, you will miss false triggers.

Was this page helpful?