User-Authored Skills

10 min read

H23
Playbook · Agent UX & Human Interaction

User-authored skills.

The feature looks like a text box and ships like a package manager. The moment you let users save a reusable instruction and invoke it by name, you have put a prompt supply chain inside your product — authored by people who are not engineers, versioned by nobody, invoked by a string that collides, shared by copy-paste, and carrying the authority of whoever runs it rather than whoever wrote it. Google proved the demand by replacing Gems with Skills across Gemini in October 2026 and giving users a slash-command to fire them. The thing that decides whether yours works is not the editor. It is the three decisions underneath it: how a skill gets selected, what happens when two apply at once, and who may publish one to somebody else.

STEP 1

Name what you are shipping before you design it.

A saved instruction set has four properties that a prompt in a chat box does not, and every one of them is an engineering obligation rather than a nicety.

  • It is persistent, so it outlives its author's context. The user who wrote "always format currency in GBP and assume the Manchester office" knew why. The colleague who inherits it in nine months does not, and neither does the model.
  • It is named, so names collide. Two skills called summary in the same scope are a routing bug, not a user error. If invocation is by slash-command, the namespace is now part of your product surface.
  • It is composable, so conflicts are the normal case. Stacking a brand-voice skill with a legal-review skill produces a prompt neither author read. Google's Skills explicitly supports stacking for exactly this reason, which means precedence is a shipped behaviour whether or not you designed it.
  • It is executable text, so it is a capability. A skill that tells the agent to call a tool, fetch a URL or write to a system is indistinguishable from code, and it arrived through a textarea with no review step. The same argument as agent skills, one altitude up.

Note what is not on that list: quality. Users write bad instructions and that is fine — a bad skill produces a bad answer and the user edits it, which is a healthy loop. The four properties above are the ones that fail silently and at someone else's expense.

Two products use the word "skill" for structurally different things, and conflating them is the most common design error here. A filesystem skill — a folder with a Markdown file, versioned in a repository, reviewed in a pull request — is an engineering artefact with a change history and an owner. A saved instruction in a user's account is a personal preference with none of that. Both are legitimate; only the first can be reviewed, diffed or revoked at the organisation level, and if you ship the second while describing it like the first, your enterprise buyers will discover the gap during a security review.

STEP 2

Invocation is the product decision, and explicit beats clever.

There are two ways a saved skill gets into a conversation, and they fail in opposite directions. Explicit invocation — the user types /brand-voice — is deterministic, debuggable and legible: the user knows what is loaded, and so do you, in the trace. Implicit selection — the model reads descriptions and decides which skills are relevant — is magical when it works and inexplicable when it does not, because the failure is a retrieval miss and retrieval misses are invisible by construction.

Google shipped explicit: skills fire from a slash in the prompt bar. That is the right default for user-authored content, and the reason is not latency or cost. It is that implicit selection makes the description the load-bearing field, and users do not write descriptions for a retriever — they write names for themselves. "Monday thing" is a perfectly good name and a useless retrieval key. The same mechanism that makes implicit skill selection work for engineer-authored skills, where the description is written as an index, breaks when the author is optimising for their own recall.

# Invocation models, and what each one costs you

EXPLICIT   /skill-name typed by the user
  + deterministic, appears in the trace verbatim
  + a wrong result is debuggable by the user
  - discovery problem: unused skills stay unused
  - no help on the turn where the user forgot

IMPLICIT   model picks from skill descriptions
  + works on the turn the user did not think about it
  - failure mode is a silent retrieval miss
  - description quality is the ceiling, and users
    do not write descriptions for retrievers

HYBRID     explicit invocation + a suggestion chip
  + keeps determinism, fixes discovery
  + the suggestion is a visible, declinable act
  - one more surface to design

The hybrid is where this should land. Keep firing explicit, and solve discovery with a visible suggestion — "you have a skill for this: Weekly report" — that the user accepts or ignores. The suggestion is cheap, it teaches the feature, and critically it keeps the decision in a place the user can see, which is the whole argument of progressive disclosure.

STEP 3

Define precedence before your users discover it for you.

Stacking is the feature users ask for and the one that generates the support tickets. Two skills, both loaded, one saying "keep responses under 150 words" and one saying "always include a worked example" — and the model resolves it by whatever happens to be later in the context, which is an implementation detail you did not intend to expose as a product rule.

You need a stated precedence order, and it has to be visible in the interface rather than documented in a help centre. A workable default, from strongest to weakest: system and safety policy, then organisation-published skills, then the skill invoked most recently in this turn, then earlier skills in invocation order, then the user's standing preferences. The ordering matters less than the fact that it is written down and rendered — see instruction hierarchy for why a layered scheme is the only thing that survives contact with multiple authors.

Then design for the conflict you cannot resolve. When two loaded skills contradict each other on a material point, the right behaviour is almost never silent resolution. It is to say so: "Weekly report" asks for under 150 words and "Client ready" asks for a worked example — I have kept the example and gone slightly over. That sentence costs fifteen tokens and prevents the user forming a false model of a system they are about to rely on, which is the core of transparency and explainability.

Render what is loaded, always. A small chip row above the composer showing the active skills, each one removable with a click, turns the single most confusing class of complaint — "it ignored my instruction" — into something the user can diagnose in two seconds. It also gives your support team a screenshot that contains the answer.

STEP 4

Scope and sharing: the moment a skill crosses an account, it needs a review step.

Personal skills are a preference and need no governance. The interesting problems start at the second scope, and there are only three worth building.

  • Personal. Visible to one user, no review, no approval. Treat it like a saved search. The only obligation is export — a user who leaves should be able to take their skills with them, and a user who is deleted should take them away.
  • Shared by link or copy. This is where the supply chain appears. A skill pasted from a colleague carries instructions the recipient has not read, and it executes with the recipient's tools and permissions. Show the full text before first use, every time, and never auto-install from a link. The relevant threat is not malice but inheritance: a skill written for someone with read-only access behaves differently in the hands of someone who can write.
  • Organisation-published. Needs an owner, a review, a change history and a revocation path — i.e. the same lifecycle as any other internal tool. Publishing to colleagues is a privilege, not a sharing gesture, and the control that matters is that an admin can see the full inventory and remove an entry. This is the same registry obligation as agent inventory and registry, applied to text instead of services.

One rule covers most of the risk and is easy to state: a skill may ask the agent to do anything the invoking user could already do, and nothing more. That sounds obvious and it is routinely violated, because skills are often rendered into a privileged part of the prompt — above the user's own turn — which in most harnesses buys them more weight than user text. If a skill's text is more authoritative than its author, you have created a privilege-escalation path out of a textarea. Render user-authored skills at user authority, structurally, in the message that carries them.

Skills are also an exfiltration surface, in the plainest possible way. A skill that says "when summarising, also fetch https://example.com/log?text=..." is a working data-egress channel installed by consent. You cannot review your way out of this at user scope, so the control has to be the same one that governs every other tool call: an allowlist at the egress boundary, as in egress control for agents. A skill is untrusted input that the user asked you to trust.

STEP 5

You will owe a migration. Google's is the worked example.

Instruction formats get replaced, and the replacement is a user-data migration with a deadline, not a deprecation note. Google's October 2026 switch from Gems to Skills is the specimen to study because it does the hard parts properly and the dates tell you what the hard parts are.

# Gems -> Skills, as announced (Oct 2026)

mechanism        existing Gems auto-migrated to Skills
invocation       "/" + name in the prompt bar
composition      multiple skills stack in one turn
authoring        create from an existing conversation

# Access to the old format ends

personal accounts                     November 2026
Workspace business / enterprise / NFP   March 2027
education accounts                       June 2027

# What the staggering tells you

consumer     ~1 month   -> individuals re-learn quickly
business     ~5 months  -> someone has to re-test workflows
education    ~8 months  -> curricula are annual, not quarterly

Three things to copy. Auto-migrate rather than ask — a migration that requires every user to act is a migration that strands the majority of the content. Stagger by how expensive re-testing is, not by how large the segment is; the business and education windows are longer because somebody else's process depends on the output. And let the new format be authored from existing material: generating a skill from a conversation the user already had is the cheapest possible authoring path, and it is why adoption of the new format is not a separate project.

The thing to add, which the public announcement does not promise, is a fidelity report. An auto-migration that silently drops a field — a Gem's attached files, a tone setting with no equivalent — produces a skill that looks right and behaves differently, and the user will attribute the regression to the model. Tell each user what moved, what changed shape and what could not be carried, once, at migration time. Then keep the old artefact readable past the cutoff even when it is no longer runnable, for the reasons in model deprecation and migration.

STEP 6

Instrument three numbers, and only three.

Skills generate a lot of possible telemetry and almost all of it is vanity. Three numbers tell you whether the feature is working, and each one has a specific decision attached.

  • Reuse depth: invocations per skill, distribution not mean. The healthy shape is a small number of skills invoked many times each. A long tail of skills used exactly once means users are authoring instead of prompting — the feature is adding a step rather than removing one, and the fix is usually that creation is too prominent relative to invocation.
  • Post-invocation edit rate. How often a user immediately rewrites or re-asks after a skill fires. This is your quality signal and it is far better than a thumbs-up, because it is unprompted and it is specific to the skill that ran. A skill above the cohort baseline is a skill whose text is wrong, and you can tell its author.
  • Cross-account installs per shared skill. The supply-chain number. One skill spreading to hundreds of accounts is an unmanaged internal standard, and the right response is to offer its author a path to publish it properly rather than to wait for an admin to find it.

What not to measure: total skills created. It goes up when the feature works and up when the feature confuses people, so it answers nothing. And resist attributing quality changes to skills without the counterfactual — a user whose answers got worse after writing six skills may have written six bad skills, or may have hit a context budget that the stacked instructions are now consuming. Log the assembled prompt length per turn and you can tell the difference.

Ship it in this order: explicit slash invocation, a visible chip row of what is loaded, personal scope only. Then add suggestion-based discovery, then sharing with mandatory full-text display, then organisation publishing with an owner and a revocation path. Most teams invert the last two and end up with an un-auditable internal standard spreading by copy-paste before anyone can publish one properly. For the surrounding surfaces, read memory and personalisation UX — a skill is explicit personalisation, and it should never be confused in the interface with the kind the agent inferred.