Perplexity Case Study

Perplexity Research published Designing, refining, and maintaining agent skills at Perplexity on May 1, 2026. The article is not the Agent Skills specification. It is Perplexity’s production experience maintaining Skills for its agent product, Computer.

In this section, its value is not that it defines another standard. It shows what happens when a team operates a real Skill catalog: the problem expands from “how do we write SKILL.md?” to “how do we manage routing, context cost, evals, and long-term maintenance?”


Core Claim

Perplexity’s central claim is direct: writing a Skill is not traditional software engineering. It is constructing context for the model and its execution environment.

That reverses several engineering instincts:

  • In code, explicitness is usually good. In Skills, activation depends on implicit matching, so description matters more than long explanation.
  • In documentation, complete explanation is often good. In Skills, content the model already knows should be deleted.
  • In ordinary systems, special cases are exceptions. In Skills, gotchas are often the highest-value content.
  • In context, every token has a cost. Brevity is a runtime constraint, not a style preference.

This aligns with the official principles in the Writing Guide: a Skill should add task context the model lacks, not repeat general knowledge the model already has.


Three-Tier Cost Model

Perplexity Computer divides Skill cost into three tiers.

Three-tier context cost Width tracks how often a tier is paid; per-payment cost rises as width shrinks Index name + description for every visible skill ~100 tok / skill every session, every user, always Load full SKILL.md body ~5,000 tok per load_skill, sticky to compaction Runtime scripts / references / assets / sub-skills unbounded only on explicit file open Computer reports a steady state of 3–5 skills loaded per thread

TierWhat loadsPerplexity’s reported budgetDesign implication
Indexname + description for every non-hidden SkillAbout 100 tokens per SkillPaid every session, so it must be short and high-signal
LoadFull SKILL.md bodyIdeally under about 5,000 tokensOnce loaded, it occupies context until the next compaction boundary
Runtimescripts/, references/, assets/, subskills, formatting filesNo fixed upper boundTruly on demand; put heavy material here

These numbers are Perplexity Computer implementation experience, not specification limits. But they make a general point: description is the most expensive text because it is always present; SKILL.md is next because it stays in context after loading; runtime resources are closest to pay-as-needed.

That is why Perplexity applies a strict test to every sentence: would the agent get the task wrong without it? If not, delete it.


When A Skill Is Worth Writing

Perplexity’s criteria can be compressed into three groups.

Write a Skill when:

  • The agent fails or behaves inconsistently without specialized context.
  • The knowledge is durable but absent from training data, such as enterprise workflows, post-cutoff information, or team taste and judgment.
  • The desired behavior cannot be changed reliably with one prompt sentence.

Do not write a Skill when:

  • The model already knows it, such as common Git command sequences.
  • The instruction is universal across most requests; that belongs in the system prompt or harness policy.
  • The content changes faster than the team can maintain it. A stale Skill makes the agent confidently take the wrong path.

This is the point behind Perplexity’s “every Skill is a tax” framing. More Skills are not automatically better. Each new Skill adds index cost and may disturb the routing boundaries of existing Skills.


Build Flow

Perplexity’s build order is eval-first.

  1. Write evals first
    Build the test set from real production queries, known failures, and adjacent negative examples that should not trigger the Skill. Negative examples are especially important because false activation is one of the main failure modes.

  2. Tune the description first
    description is a routing trigger, not internal documentation. Perplexity recommends phrasing it close to user intent, often with “Load when …”, rather than summarizing what the Skill does. For PR monitoring, this means using real user language such as “watch CI” or “make sure this lands.”

  3. Then write the body
    The body should not list command sequences the model already knows. It should give the model the goal, boundaries, failure handling, and gotchas. Conditional or heavy material should move into resource files.

  4. Use the hierarchy
    Put deterministic logic in scripts/, conditional heavy documentation in references/, templates and schemas in assets/, and first-run setup in configuration when needed.

  5. Iterate on a branch and submit evals with the change
    A small wording change in description can shift routing. Perplexity recommends reviewing the Skill and its eval set as one changeset.


The Tax Skill Lesson

The most useful case in the article is Perplexity’s U.S. tax Skill.

Perplexity first put all 1,945 sections of the Internal Revenue Code into a flat folder. The result was worse than not loading the Skill. The problem was not merely missing knowledge; the entry structure was too poor. Asking the model to locate the right material across thousands of items made routing itself unreliable.

The fix was a hierarchy: broad categories, then topic clusters, then specific sections, supported by quick references and search utilities.

The lesson is clear: when reference material is dense, the core Skill work is not “put all the documents in.” It is designing an index the model can navigate.

1,945-section tax skill — flat vs hierarchical Perplexity Computer, before / after restructuring Flat — one directory 1,945 IRC sections, one folder no internal index, no search utility Result worse than loading no skill at all Hierarchical — three levels L1 — ~20 broad categories L2 — ~15 topic clusters per category L3 — specific section content Result routing reliability improves at each level


Maintenance

Perplexity maintains Skills through gotchas and evals.

Gotchas flywheel Production observations append to SKILL.md, which informs the next cycle Production observations ① Agent fails in production ↓ append a gotcha ② Wrong skill loads ↓ tighten description, add negative eval ③ Skill missed when needed ↓ add keywords, add positive eval ④ System prompt changes ↓ audit for duplication / contention SKILL.md append-only accumulator feeds next cycle Steady state — SKILL.md grows slowly, accumulating negative knowledge

Common maintenance actions:

  • The agent fails in production: append a gotcha.
  • The wrong Skill loads: tighten description and add a negative eval.
  • The right Skill fails to load: add trigger language and a positive eval.
  • The system prompt changes: check for duplication and routing contention.

Perplexity also stresses that evals cannot remain skill-local. Adding a new Skill can cause an existing Skill to trigger incorrectly or stop triggering. This action at a distance is why catalog-wide evals matter.

The article mentions several eval types:

  • Skill loading precision and recall, plus forbidden-load negative cases.
  • Progressive-loading checks: when the body asks the agent to read an accessory file, does it actually read it?
  • End-to-end domain task completion, graded against a rubric.
  • Multi-model tests, because different model families may trigger and follow the same Skill differently.

Portable Lessons

  1. A Skill catalog is a routing system, not only a knowledge base.
    Description boundaries affect the entire catalog, not just the current Skill.

  2. Negative knowledge is often the highest-value content.
    Gotchas, counterexamples, and boundary cases are usually more valuable than general tutorials.

  3. Evals should come before the Skill and cover catalog-level regressions.
    It is not enough to prove that one Skill works. You also need to know whether it breaks another Skill.


References

Was this page helpful?