Safety, credentials and what rta does to your process
Safety is a claim, not a label
Every capability declares exactly one of plugin.Read, plugin.Write or plugin.Destructive — it is one value, not a set of flags.
It decides what an AI agent can reach without a human, so it is a statement about blast radius rather than about whether you touch the disk:
| Class | Meaning | What an agent needs |
|---|---|---|
Read | changes nothing, reveals nothing sensitive | nothing |
Write | changes something, or reveals a secret | a grant a person issued |
Destructive | removes something with no undo | a grant a person issued, naming the capability |
The rule that catches people: a capability that reveals a secret's plaintext is Write, even though it mutates nothing. kv get is the canonical case. If yours prints a token, signs with a private key, or dumps an environment, it is not Read.
A Write or Destructive capability honours a dry run. req.DryRun is true when the caller asked for one — --dry-run at a terminal — and the handler returns, before its first side effect, a view saying what it would have done (would add station beta), having touched nothing. sdktest drives every such capability that way and fails one that writes. The confirmation a destructive capability needs is the host's, and your handler is not run for it: on the command line a person at a terminal is shown the inputs the call will run with and asked, --yes skips the question, and anything that is not a terminal — a script, a pipe, an agent's shell — stops with exit code 3 until --yes is given, the TUI shows a confirmation screen first — for a plugin, the inputs the call will run with, because rta does not run a plugin's own claim about what it would do before anyone has said yes — and an agent's call has to clear a grant before it gets that far, so there is no prompt of yours to write. A plugin's handler is never run to preview a parked agent call either; its dry run is the one a person asks for.
Idempotent: true is a second claim, and a smaller one: running the capability again with the same inputs changes nothing more. It reaches an MCP client as the tool's idempotentHint and the TUI prints it beside the safety class, so say it only when it holds — a list, or a put that overwrites with the same value, never an add that appends.
Set NeedsGrant: true when the class understates it, and Scope: "city" to name the input a grant can be narrowed to — then a person can allow one record rather than the capability.
A capability that names a second record — the destination of a rename or a copy — lists that input in ScopeAlso as well, and a grant then has to cover both. Record names are what grants are scoped by, so a move checked only at its source carries a record from under one grant to under another: a grant to rename one key plus a grant to read another folder would add up to reading the key.
Set HumanOnly: true for a capability that answers to the person at the terminal and to nobody else — one that would hand an agent authority, lift a restriction, or read out the map of what every other agent may reach. It is never registered as an MCP tool, on any transport: absent from tools/list rather than advertised and refused, so there is no per-call surface check for your handler to forget. rta's own grant, agent, lock, operator and pkg namespaces declare it, and so do kv copy, kv edit, audit clients and the three keys verbs that move key material. It is stronger than NeedsGrant, and the two do not combine: a grant is consent for a call an agent may legitimately make, and these are calls it must not make at all.
Set Reveals: true on the capability whose answer is the value a masked view withholds: the one view.RevealKey points at, kv get against kv show. It is a statement about what the capability returns and lifts nothing: the mask on the other view stays, and the reveal is a capability of its own, so it is granted, narrowed to one record, limited to a few uses and written in the record like any other. Declaring it is coupled to that gate when the plugin loads: the capability is not Read, and it is either HumanOnly or has both NeedsGrant and a Scope, so the grant that lets a caller see the value names the record. A reveal returns the value unmarked — one that marks it Redacted would be masked for every reader whatever the grant, and sdktest fails it — and is not asked whether it marks nothing.
Set Keywords: []string{"todo", "ssl"} to the words a person would search for that the ID and summary do not contain: "todo" for a note list, "ssl" for a certificate check. Search on the human surfaces matches them as it matches the words of the ID, so a query in the person's vocabulary finds the capability written in the catalogue's. They are single lowercase words, a dozen at most, and never published to an agent; sdktest notes one that the ID or summary already starts with.
Declare Examples for the two or three calls a reader pastes first: plugin.Example{Title: "the five biggest processes", Inputs: map[string]any{"sort": "mem", "limit": 5}}. An example is the inputs it gives and not a command line, because the same example has a spelling for each surface — rta sys ps --sort mem --limit 5 at a terminal, the sys_ps tool and its arguments to an agent, a filled-in form in the TUI (Capability.ExampleCall(surface, example) is that spelling) — and one string cannot be all three. rta <capability> --help and rta explain print the terminal spelling, with the title beside it. Registration holds each to the capability it belongs to, so an example cannot go on naming a flag that was renamed: an input that does not exist, an operator-only (Local) one, a credential, a value the input would refuse, a value carrying a control character or a second line, or a call leaving out an input the capability cannot run without is refused when the plugin loads. Four at most, each with a one-line Title that says what the call is for.
Write Agent when Description is written for the person at a terminal and an agent would pay for every word of it. The two readers do not want the same text: a person reads which flag does what and what a pipe changes, and a model has no flags and no pipe — it reads the same sentences as context it pays for on every connection, before it has decided to call anything. Agent is what an MCP client is shown after the Summary, in place of Description (Capability.AgentText() is the choice); unset, Description serves both, as it always did. Word it for a reader with a tool list and an arguments schema: inputs by name, other capabilities by ID, nothing about flags, rta command lines or the TUI, and nothing a schema already says. It is held to the same rules as Description and refused on a HumanOnly capability, which is never a tool.
Set HostSpecific: true if what a capability returns describes the machine your plugin's process happens to run on — its own filesystem, its own network configuration — rather than a configured remote service or a pure computation. A remote, HTTP-transport rta mcp serve hides a capability marked this way from tools/list entirely, since a caller on another machine is never the one it would be describing. Most plugins reach somewhere the operator configured (a database, a cluster) or compute from their own arguments, and never need this — it exists for the same reason rta's own sys, fs and git built-ins declare it themselves.
If your plugin needs a credential location
Plugins run confined and rta denies them a standard list of credential directories — ~/.ssh, ~/.aws, ~/.kube and the rest. If yours cannot work without one, declare it:
plugin.Plugin{
Name: "clusters",
Needs: []plugin.Need{plugin.NeedKubeconfig},
...
}Declaring is asking. The operator runs rta plugin allow <name> to grant it, against your artifact's digest, and a rebuild asks again. Until then your plugin still loads and runs — it just fails at whatever call wanted the file, and rta doctor tells the operator which command fixes that.
rta plugin dev honours your declaration without any of that, for the same reason it skips trust: you compiled it from a directory you named in the command you just typed. Its report says what it allowed and what installing the plugin will need instead, so the difference is visible before somebody else hits it.
Ask for the least you need. Every location you declare is one the operator has to weigh, and a plugin that asks for four when it uses one is a plugin people stop granting anything to.
rta explain <capability> prints the exact flag an operator would need. So does rta plugin dev, in its Agents column, which is the fastest way to check you classified something the way you meant to.
What rta does to your process
Worth knowing before you debug something surprising:
- Your stdin is
/dev/null. The protocol owns the real one. Never prompt — declare aplugin.Secretinput and let the surface ask. - On macOS you are sandboxed. You cannot read or write rta's own config and data directories, and cannot read
~/.ssh,~/.aws,~/.kubeand the rest. Everything else is readable.rta doctorprints the set. Linux is not confined and says so. - Your environment is filtered to
PATH,HOME,TMPDIR,TZ,LANG,LC_*and the TLS cert variables. NoRTA_*, no cloud credentials, nothing your user exported. If you need a value, take it as an input. - Your socket is in a directory rta makes. The SDK's transport listens on a unix socket it creates where
PLUGIN_UNIX_SOCKET_DIRsays, and rta points that at a private directory insideTMPDIR, one per process, which it removes when the process ends however it ended. ATMPDIRtoo long for a socket's path — at most 103 bytes on macOS and 107 on Linux, of which the directory and the socket's name take 37 — is refused asplugin.tmpdir.toolongbefore your plugin starts, naming the length that would work, and one that does not exist or cannot be written to asplugin.tmpdir.unusable. - You are in your own process group and everything you spawn dies with you.
- One process serves every call, started on the first one and reused. Do not assume a fresh process per capability, and do not hold per-call state in a global.
- Your stderr reaches no terminal. A line you write there at error — one that starts
[ERROR], or a JSON log entry at that level — and a panic that takes the process down are kept, the last few lines of it and the first of a panic's, and shown only when your process fails: at the end of what rta says about a launch that failed — at startup, inrta plugin dev, in an install's check — and of the error a call ends in when your process stopped under it. Everything else you write there is dropped — except when your process exits before its handshake, where what it wrote last is kept whatever it says, as it wrote it, since that is all anyone learns of why: rta says the plugin exited before its handshake, how, the line it printed on stdout in the handshake's place if it printed one, and those last lines. Tell the operator what they must know in a refusal instead. - A panic in a handler is caught and returned as an error naming your capability. It does not take the process down, so the other capabilities keep working — but it is still a bug and it still says your name.
Declared text is checked
Your Summary, Description, Help and Options are published verbatim to AI agents as tool descriptions. rta refuses control characters, bidirectional overrides, invisible characters and its own framing markers at registration, and caps the lengths. If rta plugin dev refuses your plugin over a summary, that is why.
Write the Description for somebody deciding whether to call it. It is the text a model reads before choosing.
sdktest holds the same text to one more rule: it spells nothing only a terminal can act on. Your plugin's summary, and every capability's Summary, Description, input Help and action and toggle Label, is shown to an agent and in the TUI as well as at a terminal, so "raise --limit to see more" sends an agent looking for a flag its schema does not have. A flag anywhere in prose fails the rule — any long one, whichever program it belongs to, and of the short ones the host's own, -o, -y and -h, since a dash and another letter is as often a program's option named as a thing — and so does an rta … command line: any in a code span, and one naming one of your capabilities or your namespace wherever it stands. In a code span a flag counts only when the span opens on one or follows your capability's words and is an input it declares or one of the host's own switches (spelling.HostSwitches(): --detail, --dry-run, --help, --no-color, --output, --profile, --yes; -h, -o and -y count as --help, --output and --yes) — pg_restore --jobs in a span is another program's usage and passes. Name the input as limit and the capability by its ID, and word anything surface-specific at run time through req.Surface(). A HumanOnly capability is read at a terminal alone, so its text may name one of rta's own commands, rta doctor or rta mcp serve, and still no capability's command line, yours or a built-in's such as rta net dns, since the TUI reads the same text. A text that has to spell what the rule holds is waived with sdktest.Skip(sdktest.RuleSpelling, "<capability>", "<why>"), or with your plugin's name for its own summary, and the reason is printed on every run. The speller itself is pkg/sdk/spelling, for any other text your tests read: spelling.ForPlugin(Plugin()).Find(text, false) holds text every surface reads to the same rule, and Find(text, true) holds a sentence only a person at a terminal reads — one a handler words in the branch where req.Surface() said CLI — to the HumanOnly rule. Built over your plugin alone, the speller knows no other plugin's words, so in prose it cannot tell "run rta net dns" from a sentence that happens to start with rta; in a code span it can, and holds it.
sdktest also holds what an agent is shown, because a model reads all of it on every connection before it has decided to call anything: the summary and the agent text (Agent, or Description when there is none) together stay within 800 bytes, and the help of each input an agent can give within 160. That text speaks to nobody at a terminal (from a terminal, piping, on the dashboard, the TUI: spelling.TerminalWording), names no input you declare Local, which an agent's schema does not have, and does not open by saying its summary again — "Reveals a stored value" under "Reveal a stored value" is a model reading the same thing twice, once for every tool that does it. An input that is a length of time says its unit, or is a Duration. These fail the suite; a summary that starts with a lowercase letter or ends with a full stop, a max-results where the catalogue says limit, an Int named timeout that does say its unit, and an Agent text that is the Description are notes. A HumanOnly capability is never a tool, so only its input names are held. A text that has to differ is waived with sdktest.Skip(sdktest.RuleWording, "<capability>", "<why>").
What a handler words at run time is in no declaration, and sdktest.WithSource(".") has the suite read it from your package's Go source instead. Every sentence the source spells out — a string literal with a space in it, or a sum of literals and other values read as one sentence, so a code span opened in one literal and closed in the next is still a span — is held to the terminal-only rule above, since no scan can see which branch of req.Surface() a literal sits in: rta doctor passes, a flag or a capability's command line does not. What you hand plugin.AskOperator is left out, being the one command line an agent may read. And every call you name through CapabilityName, CapabilityWith or Call with a literal ID in your namespace has to be one its reader can make: a capability you declare, given inputs it declares or one of the host's switches, each by its place or as a flag the way the CLI takes it — Call spells whatever it is told, so a positional input given as a flag is a command line the CLI refuses. SettingsHint counts among those calls, its ID being the page it sends the reader to, and every name you give SettingName, SettingTo or CAHint has to be an input you declare Local: one you do not declare is a flag the CLI refuses, and one an agent's tool takes as an argument is not the operator's to set, which is what the helper tells the agent. The name you give InputTo is held the other way: an input you declare neither Local — an agent told to set it would pass an argument the bridge drops — nor positional, which has no flag for the CLI to take. Test files are not read, a directory with no source in it fails rather than passing, and there is no waiver: a sentence the scan holds is one to word through the naming helpers. For a rule of your own over the same text — a product your plugin must never call by another's name — sdktest.Sentences(".") hands your test the sentences the suite reads, each with its file and line.
Next
Testing, conventions and publishing — the conformance suite, the habits worth keeping, and how a plugin reaches an index.