Perplexity Case Study
Perplexity Research published Designing, refining, and maintaining agent skills at Perplexity on May 1, 2026. The article is not the Agent Skills specification. It is Perplexity’s production experience maintaining Skills for its agent product, Computer.
In this section, its value is not that it defines another standard. It shows what happens when a team operates a real Skill catalog: the problem expands from “how do we write SKILL.md?” to “how do we manage routing, context cost, evals, and long-term maintenance?”
Core Claim
Perplexity’s central claim is direct: writing a Skill is not traditional software engineering. It is constructing context for the model and its execution environment.
That reverses several engineering instincts:
- In code, explicitness is usually good. In Skills, activation depends on implicit matching, so
descriptionmatters more than long explanation. - In documentation, complete explanation is often good. In Skills, content the model already knows should be deleted.
- In ordinary systems, special cases are exceptions. In Skills, gotchas are often the highest-value content.
- In context, every token has a cost. Brevity is a runtime constraint, not a style preference.
This aligns with the official principles in the Writing Guide: a Skill should add task context the model lacks, not repeat general knowledge the model already has.
Three-Tier Cost Model
Perplexity Computer divides Skill cost into three tiers.
| Tier | What loads | Perplexity’s reported budget | Design implication |
|---|---|---|---|
| Index | name + description for every non-hidden Skill | About 100 tokens per Skill | Paid every session, so it must be short and high-signal |
| Load | Full SKILL.md body | Ideally under about 5,000 tokens | Once loaded, it occupies context until the next compaction boundary |
| Runtime | scripts/, references/, assets/, subskills, formatting files | No fixed upper bound | Truly on demand; put heavy material here |
These numbers are Perplexity Computer implementation experience, not specification limits. But they make a general point:
description is the most expensive text because it is always present; SKILL.md is next because it stays in context
after loading; runtime resources are closest to pay-as-needed.
That is why Perplexity applies a strict test to every sentence: would the agent get the task wrong without it? If not, delete it.
When A Skill Is Worth Writing
Perplexity’s criteria can be compressed into three groups.
Write a Skill when:
- The agent fails or behaves inconsistently without specialized context.
- The knowledge is durable but absent from training data, such as enterprise workflows, post-cutoff information, or team taste and judgment.
- The desired behavior cannot be changed reliably with one prompt sentence.
Do not write a Skill when:
- The model already knows it, such as common Git command sequences.
- The instruction is universal across most requests; that belongs in the system prompt or harness policy.
- The content changes faster than the team can maintain it. A stale Skill makes the agent confidently take the wrong path.
This is the point behind Perplexity’s “every Skill is a tax” framing. More Skills are not automatically better. Each new Skill adds index cost and may disturb the routing boundaries of existing Skills.
Build Flow
Perplexity’s build order is eval-first.
-
Write evals first
Build the test set from real production queries, known failures, and adjacent negative examples that should not trigger the Skill. Negative examples are especially important because false activation is one of the main failure modes. -
Tune the description first
descriptionis a routing trigger, not internal documentation. Perplexity recommends phrasing it close to user intent, often with “Load when …”, rather than summarizing what the Skill does. For PR monitoring, this means using real user language such as “watch CI” or “make sure this lands.” -
Then write the body
The body should not list command sequences the model already knows. It should give the model the goal, boundaries, failure handling, and gotchas. Conditional or heavy material should move into resource files. -
Use the hierarchy
Put deterministic logic inscripts/, conditional heavy documentation inreferences/, templates and schemas inassets/, and first-run setup in configuration when needed. -
Iterate on a branch and submit evals with the change
A small wording change indescriptioncan shift routing. Perplexity recommends reviewing the Skill and its eval set as one changeset.
The Tax Skill Lesson
The most useful case in the article is Perplexity’s U.S. tax Skill.
Perplexity first put all 1,945 sections of the Internal Revenue Code into a flat folder. The result was worse than not loading the Skill. The problem was not merely missing knowledge; the entry structure was too poor. Asking the model to locate the right material across thousands of items made routing itself unreliable.
The fix was a hierarchy: broad categories, then topic clusters, then specific sections, supported by quick references and search utilities.
The lesson is clear: when reference material is dense, the core Skill work is not “put all the documents in.” It is designing an index the model can navigate.
Maintenance
Perplexity maintains Skills through gotchas and evals.
Common maintenance actions:
- The agent fails in production: append a gotcha.
- The wrong Skill loads: tighten
descriptionand add a negative eval. - The right Skill fails to load: add trigger language and a positive eval.
- The system prompt changes: check for duplication and routing contention.
Perplexity also stresses that evals cannot remain skill-local. Adding a new Skill can cause an existing Skill to trigger incorrectly or stop triggering. This action at a distance is why catalog-wide evals matter.
The article mentions several eval types:
- Skill loading precision and recall, plus forbidden-load negative cases.
- Progressive-loading checks: when the body asks the agent to read an accessory file, does it actually read it?
- End-to-end domain task completion, graded against a rubric.
- Multi-model tests, because different model families may trigger and follow the same Skill differently.
Portable Lessons
-
A Skill catalog is a routing system, not only a knowledge base.
Description boundaries affect the entire catalog, not just the current Skill. -
Negative knowledge is often the highest-value content.
Gotchas, counterexamples, and boundary cases are usually more valuable than general tutorials. -
Evals should come before the Skill and cover catalog-level regressions.
It is not enough to prove that one Skill works. You also need to know whether it breaks another Skill.