Aphid
A fast and hackable agent harness.
Aphid is a coding agent written in Rust around a data-oriented core. A conversation lives in flat, append-only arenas. Streaming deltas are resolved with one memory copy. Each stage — the request, the stream, the tool call, the permission question — is announced, and a plugin can watch, stop, or rewrite it.
This book is written in Simplified Technical English.
Highlights
- Almost no memory copies. A turn is staged in the arenas of a message buffer, and committed into the transcript with one copy for each arena, whatever quantity of tokens arrived. The layout rules are applied when aphid is compiled. See Core.
- Data-oriented design. Spans, and not owned strings. A full session is a small quantity of allocations, released together.
- A fast start. The command-line tool is thin. Discovery finds the
workspace, its
AGENTS.mdinstructions and its skills before the agent starts. - Fully debuggable.
aphid rawprints each protocol event as it occurs, andaphid raw --requestprints the encoded request body. - Composable with plugins. A plugin declares what it needs and the runtime decides when it runs; everything it registers is undone when it unloads. See Plugins.
What it looks like
$ aphid
This opens the terminal user interface. A prompt runs one time and prints the result:
$ aphid -p "what does this crate do?"
$ aphid "what does this crate do?"
Getting started tells you how to install aphid and how to give it a key.
The seven front ends
aphid [OPTIONS] open the terminal user interface
aphid gui [OPTIONS] open the graphical user interface
aphid [OPTIONS] -p <prompt> run one prompt, and print the result
aphid alate <command> run a resident agent, or attach a terminal to one
aphid raw [OPTIONS] <prompt> stream one completion, and print each protocol event
aphid agent [OPTIONS] <prompt> run the agent loop with a demo tool
aphid model <command> manage the models in ~/.aphid/models.json
The first two are the coding agent, which Aphid describes with each
of its options. alate is the resident agent, which Alate
describes. raw, agent and model are also in the
Aphid chapter.
How the code is arranged
The workspace is eight crates, and each one is a narrow step above the one before it:
| Crate | What it holds |
|---|---|
aphid-core | The message, model and streaming types. See Core. |
aphid-agent | The agent loop, the tool registry and the plugin API. |
aphid-code | The coding specialization: the tools, the prompt, the skills, the sessions, the terminal and graphical user interfaces, and the Rhai host — discovery, the script engine, the capabilities and the trust gate. See Aphid. |
aphid-alate | The resident agent: a home, a memory, a heartbeat and a gateway. See Alate. |
aphid-nostr | NIP-01 and NIP-29, with no socket and no clock in it. See Colony. |
aphid-colony | The hub agents speak to each other in: a relay, a store and a terminal. See Colony. |
aphid-cli | The thin aphid binary, which connects the seven front ends. |
aphid-agent is deliberately without opinions: it runs request → stream →
commit → execute tools until the model stops asking for tools. Everything that
makes aphid a coding agent is in aphid-code, and an alate builds its agent
with that same harness, without a change.
For the Rust API of any crate, use cargo doc --open.
Licence
Aphid is licensed under the MIT Licence. The LICENSE file in the repository
gives the full text.
Getting started
This chapter tells you how to build aphid, how to give it a key, and how to run it for the first time.
What you need
- Rust 1.98 or later.
- An API key for a model. The coding agent has no models until you add one
with
aphid models add <provider/model>. The command records which environment variable holds the key. Refer to Add a model. - A system with Unix sockets, if you want the resident agent.
aphid alatedoes not work on Windows. The coding agent does.
Install
The installer gets the binary of the last release and puts it in
~/.local/bin. It is the fastest way, because it compiles nothing:
$ curl --proto '=https' --tlsv1.2 -LsSf https://github.com/tncardoso/aphid/releases/latest/download/aphid-ai-installer.sh | sh
The releases hold binaries for Linux and macOS. On other systems, and on a different processor, cargo compiles it from the registry:
$ cargo install aphid-ai
Build from the source
$ git clone https://github.com/tncardoso/aphid
$ cd aphid
$ cargo build --release
The binary is then target/release/aphid. To put it on your path:
$ cargo install --path crates/aphid-cli
There is one optional feature, telegram, which adds a Telegram bot to the
resident agent and an HTTP client to the build. It is not on by default, because
a build with no bot does not need the HTTP client.
$ cargo install --path crates/aphid-cli --features telegram
Give it a key
$ export DEEPSEEK_API_KEY=sk-...
Put this line in the file that your shell reads at start, so that each terminal has it.
Each model gives the name of the variable that holds its key, and aphid reads the variable of the model that you selected. Thus a model from a different provider reads a different variable, and a key that is absent is reported by name:
$ aphid --model glm-5 -p "hello"
aphid: ZHIPU_API_KEY is not set, and glm-5 needs it
The first run
Go to a repository and start the terminal user interface:
$ cd ~/projects/my-project
$ aphid
Type a question and press Enter. Type /help to see the
commands.
To open the graphical user interface, run:
$ aphid gui
The left drawer lists the sessions of the current workspace, with the name of each. Use its button to reduce the drawer to an icon rail. Select a saved session to continue it. New chat starts a new session.
The Tree button in the header (or Ctrl-O) shows the branches of the
session on a canvas. Each card is one turn: a prompt and its answer. The first
prompt is at the top, and the branches go down side by side.
- Drag the background to move the canvas. Scroll to zoom.
- Click a card to read the full turn.
- Right-click a card to jump to its branch, fork after its answer, edit and send its prompt again, or rename its branch.
In the conversation, put the pointer on a prompt or on a final answer. A fork control shows: edit on a prompt, fork after on an answer.
You cannot change the session or its branch while the agent is working.
To run one prompt and print the result, give the prompt on the command line:
$ aphid -p "what does this crate do?"
Aphid records each session, and it records the headless runs also. aphid --sessions lists them, and aphid --resume continues the most recent one.
Add a model
The catalogue is the models in ~/.aphid/models.json. The descriptions come
from models.dev, so you do not write out a context
window and a price by hand.
$ aphid model search glm --limit 3
$ aphid model add zhipuai/glm-5
$ aphid --model glm-5 -p "hello"
Aphid describes each model subcommand, and
Core describes the file that they write.
Tell it about your project
Aphid reads each AGENTS.md file from the root of the workspace down to the
current directory, and the most specific file has the final word. Put the
conventions of the project in one:
# AGENTS.md
- Run `cargo clippy` and `cargo fmt` after each change.
- The tests are in `tests/`, and each one is a file.
A file at ~/.aphid/AGENTS.md is applied in each workspace.
For instructions that are only needed sometimes, write a skill instead. A skill costs almost nothing until the model opens it.
Start a resident agent
The coding agent starts in a repository and forgets everything when you close the terminal. An alate has a home of its own, a memory, and a clock that wakes it.
$ aphid alate run --name work
aphid: work is awake in /home/you/.aphid/alate/work
aphid: attach with `aphid alate attach --name work`
Attach a terminal to it from somewhere else, and detach again with Ctrl-C. The
alate continues to run. Alate describes the home, the memory, the
heartbeat and the crontab.
Where things are kept
| Path | Content |
|---|---|
~/.aphid/models.json | Your models. |
~/.aphid/AGENTS.md | Instructions for each workspace. |
~/.aphid/skills/ | Your skills, for each workspace. |
~/.agents/skills/ | Your skills that the other agents read too. |
~/.aphid/plugins/ | Your plugins, for each workspace. |
~/.aphid/alate/<name>/ | One resident agent. |
~/.aphid/sessions/ | The saved sessions of every project. |
<workspace>/AGENTS.md | Instructions for one workspace. |
<workspace>/.aphid/skills/ | The skills of this workspace. |
<workspace>/.agents/skills/ | The skills of this workspace that the other agents read too. |
<workspace>/.aphid/plugins/ | The plugins of this workspace. |
APHID_HOME replaces ~/.aphid. Use it to keep a separate configuration.
Build and test the source
$ cargo build
$ cargo test
$ cargo clippy
$ cargo fmt
$ cargo build --features telegram
$ cargo test -p aphid-alate --features telegram
aphid raw and aphid agent can be fully scripted. Their tests run the full
encode, stream and commit path against a model that is not on the network.
Core — the AI layer
aphid-core is the layer below every front end. It holds the message types, the
model catalogue and the streaming code, and it knows nothing about tools,
plugins or terminals.
This chapter tells you what the layer does and which files you can edit. For the
Rust API, use cargo doc -p aphid-core --open.
The transcript
A conversation is a transcript: a flat list of messages over two append-only arenas, one for text and one for binary data.
A content block holds a span — a range of bytes in an arena — and not an owned string. Thus a full session is a small quantity of allocations, which are released together when the transcript is released. No lifetime goes out into the code that uses the crate: a transcript is one owned value that you can move between threads.
Spans stay inside the crate. Everything is read through views, which resolve a range against the arena and give back a plain string.
The system prompt is not special. It is a message with the system role. The map from that to the wire format is the work of an encoder.
Why it is arranged like this
Streaming is where the layout is of use. A provider collects a reply in a message buffer, which has arenas of its own. Each delta is added to the tail of an arena one time, and the event that reports it carries only the span of the bytes that were written. To commit the turn, aphid moves the finished buffer across with one memory copy for each arena — whatever quantity of tokens arrived.
The rules of the layout are applied when aphid is compiled: a span is 8 bytes, a content block is not more than 24, an event is not more than 16, and a message header is not more than 32. A change that makes one of these larger does not build.
The transcript only grows. A plugin adds to it, and cannot rewrite it.
The wire
Aphid speaks the OpenAI chat-completions protocol, and no other. The stream is server-sent events.
This has one result that you see: a provider that speaks a different protocol
cannot be added to the catalogue. aphid model add refuses such a model, and
says so.
Almost every provider says that it is “OpenAI-compatible”, and each one is compatible in a slightly different way. Aphid states these differences as a compatibility profile on the model, and not as a guess made from the address at the time of the request.
| Profile | Use |
|---|---|
compatible | A different company’s OpenAI-compatible server. The default. |
openai | OpenAI and Azure. |
deepseek | DeepSeek. |
none | No behaviour table. |
A profile holds the answers to questions that models.dev cannot answer, because
they are about the server and not about the model: which field limits the length
of the answer, whether the endpoint accepts reasoning_effort, whether a tool
result must repeat the name of the tool, whether a user message can come
directly after a tool result, and approximately twelve more.
A model gives the name of a profile, and then each behaviour that is different from that profile. Thus a correction is usually one line. Refer to The catalogue.
Images
A user message can carry images. The content of the message becomes an array of
blocks, and each image is a data: URL:
{"role":"user","content":[
{"type":"text","text":"what is wrong with this window"},
{"type":"image_url","image_url":{"url":"data:image/png;base64,iVBOR…"}}
]}
The blocks keep the order in which they were written, so an image stands next to the words that name it. A message without an image keeps the plain string form, because some models copy an array back to you.
A system message, an assistant message and a tool result cannot carry an image. The protocol has no place for one, and an endpoint answers with a 400. Aphid refuses such a message before it sends it.
Thinking levels
Aphid has one ladder of levels for each model that can reason:
off minimal low medium high xhigh max
off is not a level. It removes the reasoning fields from the request.
Each model supplies a different set of levels. If you ask for a level that the model does not supply, aphid decreases it to the nearest level that the model does supply, and prints a note. If the model cannot reason at all, aphid ignores the level and prints a note.
The coding agent starts at medium. --think and the /think command change
it, and thinking in alate.json sets it for a resident agent.
The catalogue
The coding agent catalogue is the models in ~/.aphid/models.json. It ships no
default models, so a fresh install has an empty catalogue until you add one.
aphid models add writes this file for you, from the description on
models.dev. Refer to model for the
commands. When the coding agent starts with no models, it prints how to add
one and exits.
The raw and agent front ends still use the built-in DeepSeek models. They
do not read ~/.aphid/models.json.
models.dev
Aphid keeps a copy of the models.dev document in ~/.aphid/models.dev.json, and
it uses the copy while the copy is less than 24 hours old. aphid model update
gets the document again.
If aphid cannot get the document, and a local copy exists, aphid uses the local copy and tells you that the data is old. An old price is more useful than an error.
The file
~/.aphid/models.json is a file that you can edit. Each model looks like this:
{
"version": 1,
"models": [
{
"id": "glm-5",
"name": "GLM-5",
"provider": "zhipuai",
"api": "openai-completions",
"base_url": "https://open.bigmodel.cn/api/paas/v4",
"api_key_env": "ZHIPU_API_KEY",
"reasoning": true,
"input": ["text"],
"context_window": 204800,
"max_tokens": 131072,
"cost": { "input": 1.0, "output": 3.2, "cache_read": 0.2, "cache_write": 0.0 },
"compat": { "profile": "compatible", "supports_reasoning_effort": false }
}
]
}
A model needs an id, a base_url, a context_window and a max_tokens. All
the other fields have defaults.
In the example above, the endpoint is a usual OpenAI-compatible server, but it
refuses the reasoning_effort field. That is the whole of the correction.
thinking_levels gives the value to send for each level. A text value is the
value to send. false means that the model refuses the level. If a level is not
in the file, aphid sends the name of the level.
"thinking_levels": { "off": "disabled", "minimal": "low", "max": "max", "xhigh": false }
If aphid cannot read the file, it prints the problem and continues with an empty catalogue. The coding agent then exits with the no-models message. A mistake in this file cannot make aphid use a model you did not configure.
Looking at the protocol
aphid raw and aphid agent print what this layer does, in place of the text:
$ aphid raw --request "hello" # the encoded request body, with no key
$ aphid raw --events --tool "what is the weather in Lisbon?"
--events prints each delta event with its span, which is the layout of this
chapter made visible. Refer to raw and agent.
Aphid — the AI harness
The coding agent is what aphid runs when you give it no subcommand. It is the
agent loop of aphid-agent, with everything a coding agent needs put around it:
the tools, a system prompt made from the conventions of the project, the skills,
the sessions and the permission gate.
This chapter tells you what the harness does, and gives each option of the command. The Commands, Skills and Plugins chapters describe the three parts that you control.
The workspace
Aphid finds the workspace when it starts. This is the root of the repository, or the directory that you are in when there is no repository.
The read, write and edit tools can touch only this directory. The bash
tool is not limited in this manner, because a shell reads and writes anywhere.
| Tool | Effect |
|---|---|
bash | Runs a command. Not limited to the workspace. |
read | Reads a file, or a part of one. |
write | Writes a full file. |
edit | Replaces text in a file. |
The output of a tool is cut when it is very long, and the full output is kept in a file that the message gives the name of.
The instructions
Aphid reads each AGENTS.md file from the root of the workspace down to the
current directory. The most specific file is last, and it has the final word. A
file at ~/.aphid/AGENTS.md is read before all of them, and thus is applied in
each workspace.
Put the conventions of the project in these files: how to run the tests, how to write a commit message, what not to touch.
For instructions that are only necessary sometimes, write a skill. A skill costs almost nothing until the model opens it.
--no-context stops aphid from reading the AGENTS.md files and the skills.
Sessions
Aphid records each session as one file of JSON lines in ~/.aphid/sessions,
shared by every project on the machine, and adds to it as each message is
committed. The filename is <project>-<id>, where <project> is the
workspace’s directory name — cosmetic only, so listing and resuming a session
still work correctly even if two projects share a name. Nothing is written a
second time. Thus a failure costs the turn that was in flight and no more, and
--resume is a replay of the file.
Headless runs are recorded also. --sessions, --resume, and the graphical
session drawer see them in the same manner as the terminal.
$ aphid --sessions # print the saved sessions
$ aphid --resume # continue the most recent session here
$ aphid --resume 20260810T012035-0000 # continue the session with this identifier
The identifier is optional. If you give no identifier, aphid continues the most recent session for the current directory.
The session tree
A session is a tree of messages. Each message records its identifier and the identifier of the message before it. Thus you can go back to a message and continue from it on a new branch. The file keeps all the branches.
- To start a branch at a prompt, edit the prompt. The new branch starts before the prompt, and aphid puts the prompt text in the input box. Change the text, and send it.
- To start a branch after an answer, fork at the answer. The next prompt that you send starts the new branch. You can fork only at an answer that ends its turn. A fork between a tool call and its result is not possible.
- To go back to a branch, jump to one of its prompts. Aphid continues the newest branch under that prompt.
- To give a branch a name, rename it. The name of the first branch is the name of the session.
Aphid records each jump and each fork as a head line in the file. --resume
continues the session where you left it.
Each message has an identifier of eight hexadecimal digits. The address of a
message is <session>:<message>. To continue a session at one message, give
its address:
$ aphid --resume 20260810T012035-0000:3f9a01c2
The start of a message identifier is sufficient. Files that aphid wrote before it recorded trees are one branch: each message follows the message before it.
In the terminal user interface, /tree (or Ctrl-O) shows the sessions of
the workspace and their branches. See Commands.
In the graphical user interface, the Tree button shows the branches of the
session on a canvas. See Getting started.
/clear and /new start a new session in a new file. The old session does not
change.
Permissions
--confirm makes aphid ask you before it runs a command that changes the
workspace. A headless run has no terminal for a question. Thus --confirm and
-p together refuse each such command, and do not permit it quietly.
A plugin can answer these questions in place of you, with the on_permission
announcement a plugin can subscribe to. Refer to Plugins
and Composition.
Invocation
aphid [OPTIONS] [PROMPT]... the coding agent
aphid gui [OPTIONS] open the graphical coding agent
aphid alate <COMMAND> run a resident agent, or attach to one
aphid raw [OPTIONS] <PROMPT>... stream one completion, and print each protocol event
aphid agent [OPTIONS] <PROMPT>... run the agent loop with a demo tool
aphid model <COMMAND> manage the models in ~/.aphid/models.json
The coding agent is the default. If the first word is alate, raw, agent or
model, aphid runs that subcommand. If the first word is something different,
aphid uses the full command line as a prompt for the coding agent.
Give a prompt to run the agent one time. Give no prompt to open the terminal user interface.
$ aphid # opens the terminal user interface
$ aphid gui # opens the graphical user interface
$ aphid -p "fix the failing test" # runs one time, and prints the result
$ aphid "fix the failing test" # the same, with no -p
-p and the bare words do the same thing. Only an empty prompt opens the
terminal user interface.
aphid gui uses the same model, context, tools, sessions, permission gate,
slash commands, and Rhai plugins as the terminal. The main area shows streamed
text, reasoning, tool calls, tool output, and run state. Markdown includes
tables, nested lists, code highlighting, links, and images. A remote image is
not fetched until you select its load control.
The options
| Option | Effect |
|---|---|
-p, --print <PROMPT> | Run one time. Stream the result to stdout, and exit. |
--model <NAME> | Select a model. Give the identifier, or a unique part of it. |
--models | Print the known models, and exit. |
--think <LEVEL> | Set the quantity of reasoning. |
--system <TEXT> | Replace the standard instructions. |
--append-system <TEXT> | Add text to the instructions. |
--resume [<ID>] | Continue a saved session. <ID>:<MESSAGE> continues it at that message. |
--sessions | Print the saved sessions for this workspace, and exit. |
--confirm | Ask before each command that changes the workspace. |
--no-context | Do not read AGENTS.md files or skills. |
--list-plugins | Print the plugins that would load, and exit. |
--no-plugins | Do not load any plugin from .aphid/plugins. |
--plugin <PATH> | Load one plugin from a path. |
--trust-plugins | Agree to the plugins of this workspace. |
--max-turns <N> | Stop the run after this quantity of requests. |
--quiet | Do not print the output of each tool. |
Select a model
--model accepts the full identifier, the last part of it, or a prefix. Aphid
tries these three forms in that sequence. If two or more models match, aphid
refuses the name and prints the models that matched.
$ aphid --model deepseek-v4-pro -p "hello" # the full identifier
$ aphid --model pro -p "hello" # the last part
If you give no --model, aphid uses the first model in the catalogue. If the
catalogue is empty, aphid prints how to add a model and exits.
--models prints the catalogue. The catalogue is the models in
~/.aphid/models.json. To add a model, refer to model.
Set the quantity of reasoning
--think accepts these levels: off, minimal, low, medium, high,
xhigh and max. medium is the default.
Each model supplies a different set of levels. Aphid decreases the level to the nearest level that the model supplies, and prints a note. If the model cannot reason, aphid ignores the option and prints a note. Refer to Thinking levels.
Control the plugins
The plugin options control the Rhai plugins in .aphid/plugins. A plugin in your
home directory always loads. A plugin that comes with a workspace needs your
agreement the first time; aphid asks before the terminal user interface starts,
and keeps the answer in ~/.aphid/trust.json. A headless run has no terminal for
a question, and thus does not load the plugins of the workspace unless you give
--trust-plugins. --plugin names a file directly and does not ask.
Use --no-plugins to make the start of a run fully predictable. Read
Plugins to write one.
alate
aphid alate runs a resident agent. An alate has a home directory of its own, a
memory that continues between sessions, a clock that wakes it, and a socket that
a terminal attaches to.
aphid alate run [--name NAME] run the alate in this terminal
aphid alate attach [--name NAME] open a terminal on a running alate
aphid alate gui [--name NAME] open a window on a running alate
aphid alate list show the alates on this machine
| Option | Effect |
|---|---|
-n, --name <NAME> | Select the instance. The default is default. |
run holds the terminal until you stop it. attach opens a terminal on an
alate that already runs; close it, and the alate continues. gui does the same
in a window on your desktop, with the creature that shows what the agent is
doing — see Window.
Alate gives the home directory, each field of the configuration, the memory, the heartbeat and the crontab. CLI gives the terminal that attaches.
aphid alate needs a Unix socket, so it does not work on Windows.
raw and agent
These two subcommands are the debug tools. raw sends one request. agent
loops until the model stops to call tools. Both accept the same options.
| Option | Effect |
|---|---|
--pro | Use deepseek-v4-pro. The default is deepseek-v4-flash. |
--system <TEXT> | Put a system message before the prompt. |
--think <LEVEL> | Set the quantity of reasoning. |
--max-tokens <N> | Limit the length of the response. |
--temperature <F> | Set the sampling temperature. |
--tool | Supply a demo get_weather tool, to show tool-call deltas. |
--events | Print each delta event with its span, in place of the text. |
--request | Print the encoded request body, and exit. |
--request does not send a request. Thus you can use it with no API key.
$ aphid raw --request "hello" # print the request body
$ aphid raw --events --tool "what is the weather in Lisbon?"
These two subcommands always use a DeepSeek model, and they always read
DEEPSEEK_API_KEY. To use a different model, use the coding agent.
model
aphid model manages ~/.aphid/models.json. The command aphid models does
the same thing. Core describes the catalogue and the
format of the file.
The model descriptions come from models.dev. Aphid keeps a
copy of that document in ~/.aphid/models.dev.json, and it uses the copy while
the copy is less than 24 hours old.
model add
aphid model add [OPTIONS] <NAME>
<NAME> is provider/model, or a model identifier that only one provider
supplies.
$ aphid model add zhipuai/glm-5
added glm-5 in /home/you/.aphid/models.json
provider zhipuai
endpoint https://open.bigmodel.cn/api/paas/v4
limits 204800 context · 131072 output
price $1.00 in · $3.20 out per M tokens
key $ZHIPU_API_KEY
(cached 3h ago; `aphid model update` to refresh)
Many providers supply a model with the same identifier. If the name is ambiguous, aphid prints each provider that supplies that model:
$ aphid model add deepseek-v4-pro
aphid: `deepseek-v4-pro` is served by 23 providers:
alibaba-cn/deepseek-v4-pro
azure/deepseek-v4-pro
...
Name one of them, or pass --provider <id>.
Some model identifiers contain a slash. Aphid reads the full name as a model
identifier first, and as provider/model second. Thus both of these commands
find the same model:
$ aphid model add openai/gpt-oss-120b
$ aphid model add wandb/openai/gpt-oss-120b
| Option | Effect |
|---|---|
--provider <ID> | Use only this provider. Use it when a name is ambiguous. |
--base-url <URL> | Give the endpoint URL. models.dev does not list one for each provider. |
--api <API> | Set the wire protocol. |
--api-key-env <VAR> | Set the environment variable that holds the API key. |
--compat <PROFILE> | Set the endpoint behaviour. Refer to Core. |
--force | Replace a model that is already in the catalogue. |
--refresh | Get the models.dev document again, even if the copy is new. |
--offline | Use the local copy only. Fail if there is no copy. |
Aphid speaks the OpenAI chat-completions protocol only. If the provider speaks a
different protocol, aphid refuses the model. --api openai-completions makes
aphid add the model regardless.
model remove
aphid model remove <NAME>
<NAME> accepts the same three forms as --model. This command removes a model
from ~/.aphid/models.json.
$ aphid model remove glm-5
removed glm-5 from /home/you/.aphid/models.json
model list
aphid model list
This command prints the models in ~/.aphid/models.json.
model search
aphid model search [OPTIONS] <QUERY>
This command finds models on models.dev, but it adds no model. Aphid compares the query with the provider identifier, the model identifier and the model name.
$ aphid model search glm --limit 3
302ai/glm-4.5 131072 ctx $ 0.29/$1.14 GLM-4.5
zhipuai/glm-5 204800 ctx $ 1.00/$3.20 GLM-5
...
(cached 3h ago; `aphid model update` to refresh)
| Option | Effect |
|---|---|
--limit <N> | Print at most this many results. The default is to print them all. |
--refresh | Get the models.dev document again. |
--offline | Use the local copy only. |
Each name in the first column is a name that aphid model add accepts.
model update
aphid model update
This command gets the models.dev document again, and writes it to
~/.aphid/models.dev.json. Then it prints the quantity of providers and models,
and the models that models.dev added or removed after the previous copy.
$ aphid model update
/home/you/.aphid/models.dev.json · 182 providers · 6243 models · 3.5 MB
3 added:
deepseek/deepseek-v4-pro
...
This command changes the local copy only. It does not change
~/.aphid/models.json.
add and search also get the document if the local copy is more than 24 hours
old. Use model update to get the document immediately.
To correct a model by hand, refer to The file.
Files and environment variables
| Path | Content |
|---|---|
~/.aphid/models.json | Your models. |
~/.aphid/models.dev.json | The local copy of the models.dev document. |
~/.aphid/AGENTS.md | Instructions for each workspace. |
<workspace>/AGENTS.md | Instructions for one workspace. |
<workspace>/.aphid/skills/ | The skills of this workspace. |
<workspace>/.agents/skills/ | The skills of this workspace that the other agents read too. |
<workspace>/.aphid/plugins/ | The plugins of this workspace. |
~/.aphid/sessions/ | The saved sessions of every project, named <project>-<id>.jsonl. |
~/.aphid/skills/ | Your skills, for each workspace. |
~/.agents/skills/ | Your skills that the other agents read too. |
~/.aphid/plugins/ | Your plugins, for each workspace. |
~/.aphid/trust.json | The workspaces whose plugins you agreed to. |
~/.aphid/alate/<name>/ | One resident agent. See Alate. |
| Variable | Effect |
|---|---|
APHID_HOME | Replaces ~/.aphid. Use it to keep a separate configuration. |
DEEPSEEK_API_KEY | The key for the built-in models that raw and agent use. |
APHID_HOME moves the model catalogue, the trust file and the alates. It does
not move AGENTS.md, the skills or the plugins of your home directory: those
follow HOME, so that a separate catalogue does not take your instructions away
with it.
Each model gives the name of the variable that holds its key. The coding agent reads the variable of the model that you selected. Thus a model from a different provider reads a different variable:
$ aphid --model glm-5 -p "hello"
aphid: ZHIPU_API_KEY is not set, and glm-5 needs it
If you change the model in the terminal user interface, aphid reads the key of the new model.
Exit codes
| Code | Meaning |
|---|---|
0 | Success. |
1 | The run failed, or aphid could not read or write a file or the network. |
2 | The command line was wrong. |
Commands
A command is a line that starts with /. The terminal user interface reads it
and acts on it. A command never goes to the model, unless a plugin decides to
send something to the model itself.
Type /help to see the list in the terminal.
The standard commands
| Command | Effect |
|---|---|
/model [name] | Change the model, or open the picker when you give no name. |
/think <level> | off, minimal, low, medium, high, xhigh or max. |
/clear, /new | Start a new session, in a new file. The system prompt stays. The old session does not change. |
/tree, /sessions | Show the sessions of the workspace and their branches. See The session tree. |
/fork | Show the session tree, to start a branch. |
/rename <name> | Give a name to the branch that this session is on. |
/tools | List the tools that are registered. |
/ps | Show what the runtime runs now, and what it ran before. |
/session | Show the session identifier, the message that the session continues from, and the file. |
/plugins | List the plugins that loaded, and the commands they added. |
/skills | List the skills that the model can open. |
/help | Print the list. |
/quit | Exit. /q and /exit do the same. |
| Key | Effect |
|---|---|
Esc | Clear the selection, or stop the run. |
Ctrl-C | Quit. |
Ctrl-P | Change to the next model. |
Ctrl-T | Show the reasoning. |
Ctrl-O | Show the session tree. |
PageUp, PageDown | Scroll. |
Enter | Send the message. |
Shift-Enter | Make a new line in the same message. |
Up, Down | Move through the messages you sent before. |
| Mouse wheel | Scroll. |
| Mouse drag | Select text in the transcript. Release to copy it. |
To copy text out of the transcript, hold the left mouse button and move the
pointer over it. The text under the pointer is shown in reverse video. Release
the button, and the text goes to the clipboard. The status line says how many
lines it took. Esc clears the selection.
A drag takes the characters that are on the screen, the prompt markers
included: a selection that starts at the left edge of a message you sent takes
the > with it. To leave the marker out, start the drag after it.
Aphid sends the text to your terminal with OSC 52, so a copy also works over SSH and in tmux. Some terminals keep OSC 52 off until you turn it on.
Text that you paste goes into the editor as it is, on as many lines as it has.
A paste does not send the message: press Enter when the message is complete.
Up on the first line shows the message you sent before. Down comes back to
what you were writing, which is kept while you look.
The session tree
/tree shows each session of the workspace on one line. The current session
is open. Under an open session, each prompt is on one line, and the answer to
it is after it in grey. A conversation that has no branches is a flat list.
Where a conversation branches, each branch is indented under the turn that it
starts from. ● marks the turn that the session continues from.
| Key | Effect |
|---|---|
↑ ↓ | Move the cursor. k and j do the same. |
→ ← | Open or close the session under the cursor. |
Enter | On a session, continue it. On a prompt, jump to the newest branch under it. |
e | Start a branch before the prompt, and put the prompt in the input box to edit it. |
f | Start a branch after the answer to the prompt. |
r | Write /rename in the input box, to give a name to the current branch. |
/ | Type a filter. The filter finds prompts, answers and names. |
Esc | Close the tree. |
While the agent works, you can look at the tree, but you cannot jump or fork.
Stop the run with Esc, or wait until it ends.
Shell commands
A line that starts with ! is a shell command, not a message. The terminal
user interface runs the text after the ! in the workspace, and prints the
output into the content area.
The input border turns red while the line is a command. The command never
goes to the model. It runs through the same engine as the bash tool, so
/ps shows it while it runs, and k stops it. A bang line is kept in the
input history, so Up recalls it and Enter runs it again.
/model with no name opens a list of the catalogue. /model <name> accepts the
same three forms as --model: the full identifier, the last part of it, or a
prefix.
/clear and /new are the same command. The conversation is dropped and the
system prompt is kept, so the agent still knows the project.
The file list
Type @ at the start of a word to open a list of every file in the workspace.
Type any part of a path to cut the list down. The letters do not have to be
next to each other, and a letter that is wrong is forgiven, so @tuiapp finds
crates/aphid-code/src/tui/app.rs.
| Key | Result |
|---|---|
Arrow keys, or Ctrl-P and Ctrl-N | Move the cursor |
Enter | Choose the file, and open the question below |
Esc, or Ctrl-C | Close the list, and keep the @ |
Backspace | Take back one letter of the query, or close the list when there is none |
| Space | Close the list, and type the space |
The path that the list writes is relative to the workspace root, which is the path that the tools of the agent accept.
An @ inside a word is only a character: you@example.com opens no list.
The file tree is read the first time that you press @, and a watcher keeps
it correct after that. A file that you make while the terminal is open is in
the list. A session that never presses @ does not read the tree.
Cite or Attach
Enter on a file opens a question with two answers:
| Answer | Result |
|---|---|
| Cite | Write the path into the message. This is text, and nothing more. |
| Attach | Send the file with the message. The path is written with an @ in front of it. |
| Key | Result |
|---|---|
Enter | Take the answer that is marked. Cite is marked when the question opens. |
c | Cite. |
a | Attach. |
| Arrow keys | Move the mark. |
Esc | Close the question, and keep the @ in the box. |
Press Enter two times to write a path, which is the quickest way: the first
Enter chooses the file, and the second one cites it.
Attachments
A marker is an @ and a path, and it is ordinary text. @src/main.rs is a
marker; src/main.rs is a citation. Both can be in the same message, and a
marker can be in the middle of a sentence:
compare @shots/before.png with @shots/after.png
The words of a marker become the reference of the file. The image is sent directly after the words that name it.
An attached text file is sent as its content, wrapped in <file path="…">.
The cap is the cap of the read tool: 1000 lines or 64 KiB, whichever comes
first. The text says when the file is longer than that.
An attached image is sent as an image. Aphid reads PNG, JPEG, GIF and WebP, and
refuses a file above 10 MB. The bytes decide the format, not the file name. The
model must accept images: aphid refuses an image for a model that cannot look at
one, and the message says which model to choose with /model.
A marker breaks when its text changes. Take one letter away and the file is not
sent with the message. Backspace and Delete remove the whole marker at one
keystroke when the cursor is on it or next to it, so a broken marker is rare.
Type the marker again and the file is attached again, as long as the message is
not sent yet. Thus an edit does not lose the work of reading the file.
Esc on a line that is not running clears the line and the files it named.
A file is read when you attach it, not when the message goes out. What you saw in the message is what the model receives, even if the file changes after that.
A file is read when you attach it, not when the message goes out. What you saw in the message is what the model receives, even if the file changes after that.
/ps
The list shows each command that runs now, and the last four commands that
stopped. Each line gives the number of the command, its system process
identifier, the source (bash, or the name of a plugin), the time, and, for a
command that stopped, the result and the quantity of output.
Press the arrow keys to select a command that runs now, and press k to stop
it. This stops the command and each command that it started. Press Esc to
close the list.
A command can start a process in the background, for example server &. If
that process keeps the output of the command, the command shows ↻ bg after
its shell stops. The agent does not wait for this process. The list keeps the
line until the process stops. Press k to stop the process and each process in
its group.
The list opens while the agent runs also, which is when there is most to see. The other commands wait for the run, because they speak to the agent; this one does not.
Commands from plugins
A plugin adds a command with command, from its apply, with commands in its
inject. The command shows in /plugins, and it is on offer for as long as the
plugin is loaded — no longer.
const inject = ["commands"];
fn apply(ctx) {
command(#{
name: "review",
description: "Ask for a review of the changes.",
run: |args| {
let diff = exec("git diff").stdout;
if diff == "" { return notice("nothing to review"); }
prompt("Review this diff:\n" + diff);
notice("reviewing…")
}
});
}
args is the text after the name of the command.
Return notice(text), a text, or an array of them to show text to the user. To
send text to the model, call prompt(text). Aphid shows the notices first, and
then the prompt, whatever the order in the command.
Return new_session(), alone or in the array, to start a new session as /new
does. Aphid does not start a new session while a run is in progress: it shows a
notice in its place.
A standard command always wins, and thus a plugin cannot take /quit away. If
two plugins use one name, aphid keeps both: the second becomes /review:2.
A name with a space in it is refused. A leading / is removed, so review and
/review give the same command.
Refer to Plugins for the rest of what a plugin can do.
The resident agent
The terminal that attaches to an alate has a different, smaller set of commands. Refer to CLI.
Skills
A skill is an instruction file that the model opens when it needs it.
Only the name, the description and the path of each skill go into the system
prompt. The model reads the body with the read tool when a task agrees with
the description. This is progressive disclosure, and it is what keeps twelve
skills from costing twelve skills’ worth of context on each request.
Use an AGENTS.md file for what is true always. Use a skill for what is true
sometimes: how to make a release, how to add a migration, how to write a
particular kind of test.
Where skills go
Aphid looks in the workspace first, and then in your home directory. At each of
the two, it reads .aphid/skills and then .agents/skills:
.aphid/skills/<name>/SKILL.md
.aphid/skills/<name>.md
.agents/skills/<name>/SKILL.md
.agents/skills/<name>.md
.agents is the directory that the other agents read, next to the AGENTS.md
that they already share. Put a skill there when the same instructions must serve
aphid and the other agents together. Put a skill in .aphid when it is only for
aphid.
Use the directory layout when the skill has files of its own — a script, a template, an example. The model can read them, because you give it the path.
A skill in the workspace hides a skill in your home directory with the same
name. Thus a project can replace a skill that you carry everywhere. At the same
level, a skill in .aphid hides a skill in .agents with the same name.
Writing one
A skill needs frontmatter: a --- block at the top of the file, with flat
key: value lines.
---
name: release
description: How to cut a release of this crate. Use when the user asks to release, tag or publish.
---
# Release
1. Make sure that `cargo test` passes on `main`.
2. Change the version in `Cargo.toml`.
...
| Field | Effect |
|---|---|
description | Necessary. What the skill is for, and when to use it. |
name | The name of the skill. Optional. |
The description is the whole of what the model sees before it opens the file.
Write it to say when to use the skill, and not only what it is. A description
of more than 1024 characters is refused, and the skill is reported.
If there is no name, aphid uses the name of the directory for a SKILL.md, or
the name of the file for a loose .md.
Aphid reads only the two keys above. Frontmatter with more in it is accepted, and the rest is passed over.
Looking at them
Type /skills in a session. Each line gives the name of the skill, its
description, and whether the skill comes from the workspace (project) or from
your home directory (global).
A line that starts with ! is a skill file that aphid could not use, and the
reason: no description, a description that is too long, or a file that could not
be read. A skill file with a mistake in it is reported, and the session
continues.
--no-context stops aphid from reading the skills and the AGENTS.md files.
In a resident agent
An alate reads the skills in <home>/.aphid/skills and <home>/.agents/skills,
in the same manner. The home of the alate is its workspace, so these are the
workspace layouts and not special ones. Refer to Alate.
Plugins
A plugin is one file of Rhai code. It can look at a run, stop a tool, change a prompt, add a tool, add a command, add an interactive terminal surface, and offer a service to other plugins. You do not compile aphid again to add one.
A plugin can also be written in Rust, and compiled in. Refer to Plugins in Rust.
This page is the reference. Composition is the model behind it, and worth reading first — a plugin here declares what it needs and the runtime decides when it runs, which is a different bargain from the one most plugin systems offer.
Where plugins go
Aphid looks in the workspace first, then in your home directory. Two layouts are correct:
.aphid/plugins/<name>.rhai
.aphid/plugins/<name>/main.rhai
The name of the plugin is the name of the file, or the name of the directory. A plugin in the workspace hides a plugin in the home directory with the same name.
Write the description of the plugin in //! comment lines at the top of the
file. The /plugins command and aphid --list-plugins show this text.
//! Keeps the model away from the changelog.
fn apply(ctx) {
on("agent/tool-call", |tool| {
if tool.name == "write" && tool.arguments.contains("CHANGELOG") {
return block("the changelog is written by hand");
}
});
}
.aphid/plugins.json overrides what was found: switch one off, configure it,
isolate a service for it, or name a file that lives elsewhere. See
the composition file.
Important: call is a reserved word in Rhai. Do not use it as the name of a
parameter, and use invoke to reach a service.
Trust
A plugin in your home directory always loads. It is yours.
A plugin in a workspace comes with the checkout, so aphid asks you before it
loads one for the first time. Aphid keeps your answer in ~/.aphid/trust.json
and does not ask again for that workspace.
Aphid asks the question before the terminal user interface starts. In headless
mode aphid does not ask, and does not load the plugins of the workspace. Use
--trust-plugins to agree without a question.
This controls which plugins load. It does not control what a plugin that loaded can do. A plugin that you agreed to can do all that you can do.
apply
Everything a plugin contributes happens in apply. It runs once, when the
plugin loads — which is not necessarily at startup: a plugin that declared
inject waits until what it declared is there.
const inject = ["shell"];
const provides = ["todos"];
const emits = ["todos/changed"];
fn apply(ctx) {
on("agent/turn-start", |cx| { cx.note("…"); });
provide("todos", #{ add: |text| { … } });
effect(|| { … }, || { … });
}
In apply | What it does | Needs in inject |
|---|---|---|
on(event, closure) | Subscribe to something announced | — |
tool(map) | Contribute a tool the model may call | tools |
command(map) | Contribute a slash command | commands |
surface(map) | Contribute a panel | surfaces |
provide(name, map) | Offer a service, as a map of functions | — |
invoke(name, method, args) | Call a service | the service |
effect(setup, teardown) | Take something, give it back on unload | — |
These work only inside apply. Outside it there is no component for the
runtime to attach the registration to, so nothing could undo it when your plugin
unloads — the call is refused, and says so.
tools, commands and surfaces are ordinary services: a plugin that
contributes one waits for the registry the same way it waits for anything else,
and what it contributed leaves when it does. See
Composition.
Events
Subscribe with on. These come from the agent loop:
| Event | When it fires |
|---|---|
agent/prompt | Before aphid puts your prompt in the transcript |
agent/run-start | The run starts |
agent/turn-start | Before each request to the model |
agent/request | After agent/turn-start. Changes what the request sends |
agent/event | For each protocol event. This is the fast path |
agent/message | After the answer of the model is in the transcript |
agent/tool-call | A tool call is asked for, but did not run |
agent/tool-progress | A tool sent partial output |
agent/tool-result | A tool completed |
agent/turn-end | A turn is complete |
agent/run-end | The run stopped |
Subscribing to a name nothing announces is reported when you subscribe, rather than silently never firing.
These come from the coding harness — the things the loop has no word for, because a permission or a file change is this harness’s idea rather than the loop’s. They are announced on the same bus and subscribed to the same way:
| Event | When it fires |
|---|---|
code/system-prompt | Aphid made the system prompt |
code/session-start | A session opened |
code/session-end | A session is closing |
code/permission | A tool needs permission |
code/file-change | write or edit changed a file |
code/notice | Aphid showed a message to the user |
code/tick | Every 250 milliseconds, in the terminal UI |
code/system-prompt is a waterfall: each listener receives the prompt as it
stands and returns what the next should see, so appending and replacing are the
same operation from two ends. It fires while the harness is being built, which
makes it the only announcement made before an agent exists.
code/permission is a bail: the first listener with an opinion decides and
the rest do not run, because a second opinion on a settled question is a second
question for the user.
code/tick is the only one the agent does not cause. Use it to look at
something outside the session: a file, a queue, a clock. Keep it short. It runs
while the user is at the prompt, and exec and the http functions stop it until
they are complete. A tick still being handled is not announced again, so a slow
listener costs its own time rather than a queue behind it. There are no ticks in
headless mode.
code/notice is not reentrant either, and for the same kind of reason: a
listener that shows the user something would announce itself.
Calls into a plugin come from different threads. A listener of the agent runs
on the thread of the agent, a tool on a thread of its own, and a tick, a command
or a panel on the plugin thread. But aphid lets only one call into a plugin at a
time. A call waits while another call into the same plugin runs. Thus a change
that a tick makes is what the next panel render reads, and two calls can never
both read the state, change it, and write it back over each other. This is also
true for a map that the closures of apply capture.
A call that waits for another call into the same plugin cannot continue until that call ends. Thus keep each call short. A call into a different plugin does not wait.
What each listener is handed
agent/prompt:textagent/tool-call:id,name,arguments,known,blockedagent/tool-result:id,name,arguments,turn,content,is_error,detailsagent/message:cx, thentext,thinking,tool_callsagent/event:kind,turn, and thenindex,block,textorstopagent/request:cx, thenturn,run_start,lengthagent/turn-end:cx, thenstop_reason,tool_calls,input,output,erroragent/run-end:cx, thenstop,turns,input,output,errorcode/session-startandcode/session-end:id,path,reason,restored.reasonisnew,resumeorend, orswitchwhen the user moves to a different session file. A jump or a fork in the same file sends no event.code/permission:tool,summary,riskcode/file-change:path,kind,before,aftercode/system-prompt: the prompt as textcode/notice: the text showncode/tick: nothing
How a listener changes a run
Rhai sends the arguments of a function by value. Thus a listener cannot change the map that it receives. It changes the run with the value that it returns.
Return nothing to change nothing.
| Return value | Result |
|---|---|
block("why") | The tool does not run. The model reads the reason |
block_and_stop("why") | The same, and the run stops after this batch |
reject("why") | From agent/prompt: the prompt does not go to the model |
stop() | From agent/turn-end: the run stops cleanly |
#{ text: "…" } | From agent/prompt: use this text in place of the prompt |
#{ arguments: "…" } | From agent/tool-call: use these arguments |
#{ content: "…" } | From agent/tool-result: use this result |
#{ append: "…" } | From code/system-prompt: add this to the prompt |
#{ replace: "…" } | From code/system-prompt: use this prompt |
#{ history: …, system: …, … } | From agent/request: send a different request. Refer to The request |
"allow", "deny" | From code/permission |
code/permission also accepts "allow_always" and "ask". Use "ask" when
the plugin has no opinion; the next listener, and finally the user, then
decides.
Every listener runs, even after one has refused a tool call — an observer still wants to see a call somebody else blocked. The first refusal is the one that stands.
The run context
The listeners that receive cx are different. cx holds a handle, not a copy,
and thus its methods do change the run — whatever Rhai did with the value on the
way in, and from wherever the listener happens to run.
fn apply(ctx) {
on("agent/turn-start", |cx| {
cx.note("Today is a Tuesday."); // adds a system message
});
}
| Member | Result |
|---|---|
cx.note(text) | Adds a system message at the end of the transcript |
cx.push_user(text) | Adds a user message at the end of the transcript |
cx.cancel() | Stops the run at the next safe point |
cx.cancelled | true if something stopped the run |
cx.hold() | From agent/request: keeps the request back. Refer to The request |
cx.model | The identifier of the model |
cx.turn | The number of the turn, from zero |
cx.input_tokens, cx.output_tokens | The tokens of the run until now |
The transcript only grows. A listener adds to it, and cannot rewrite it.
The request
agent/request changes what one request sends to the model. It does not change
the transcript. The transcript, and the session file, keep all the messages.
/tree and --resume thus show the conversation as the user had it.
Aphid announces agent/request before each request, after agent/turn-start.
The listener gets cx and a map with turn, run_start and length.
run_start is the position of the prompt of this run in the transcript.
length is the number of messages in the transcript.
The listener returns a map. Each field is optional:
| Field | Result |
|---|---|
history | "run" sends only the messages of this run. "all" sends all the messages. The default is "all" |
system | Sends this text as the system prompt. The text replaces the full prompt of aphid |
prefix | An array of #{ role, text }. Aphid sends these messages after the system prompt. role is "system" or "user" |
prompt_prefix | Aphid puts this text before the first user message, in the same message, with an empty line between |
exclude_tools | An array of tool names. Aphid does not offer these tools in this request |
When more than one listener returns a map, a field from a later listener
replaces the same field from an earlier one. prefix and exclude_tools add to
what is there.
fn apply(ctx) {
on("agent/request", |cx, request| {
#{ history: "run", prompt_prefix: "Today is a Tuesday." }
});
}
Hold a request
cx.hold() keeps the request back. It returns a handle. Keep the handle, and
call release() on it when the request can go. Aphid then announces
agent/request again, and the listener can shape the request from what it
waited for. A tick, a command, or the reply of a model can release the handle.
fn apply(ctx) {
let mem = #{ ready: false, hold: () };
on("agent/request", |cx, request| {
if !mem.ready { mem.hold = cx.hold(); return; }
#{ prompt_prefix: "ready" }
});
on("code/tick", || {
mem.ready = true;
if type_of(mem.hold) == "Hold" { mem.hold.release(); }
});
}
While the request waits, the listener does not run and nothing is blocked. The
user can push Esc. The run then stops and aphid sends nothing. Do not keep a
request back without a plan to release it: aphid waits until the release or
Esc.
Capabilities
A Rhai script can only calculate. Aphid gives it these functions:
| Function | Result |
|---|---|
notify(text) | Shows text to the user |
prompt(text) | Sends text to the model, as if the user typed it |
log(text) | Writes text to standard error |
fs_read(path) | Reads a file, and returns the text |
fs_write(path, text) | Writes a file |
fs_exists(path) | Returns true if the path is there |
fs_list(path) | Returns the names in a directory |
fs_append(path, text) | Adds text at the end of a file, and writes it to the disk before it returns |
fs_lock(path) | Takes a lock on a file. Returns false if another process, or another plugin, has the lock |
fs_unlock(path) | Releases a lock that fs_lock took |
time_now() | The time, as #{ unix_ms, iso, day }. day is the local date, such as 2026-10-07 |
aphid_home() | The directory of aphid, usually ~/.aphid |
try_parse_json(text) | The value in a JSON text, or () if the text is not correct JSON |
exec(command) | Runs a shell command |
http_get(url) | Makes a GET request |
http_post(url, body, headers) | Makes a POST request |
These functions give the parts that aphid made its system prompt from. Use them
when agent/request replaces the system prompt:
| Function | Result |
|---|---|
system_prompt() | The system prompt of aphid, as aphid sends it |
agents_md() | The AGENTS.md files, as an array of #{ path, text }. The global file is first |
skills() | The skills, as an array of #{ name, description, path, project } |
tool_list() | The tools of this session, as an array of #{ name, description, snippet }. snippet is the short guideline of a built-in tool, and "" for other tools |
tool_list() reads the tools when you call it. Thus it also gives a tool that a
plugin added later. Before the session starts, and in an alate, these functions
give empty values.
prompt is a call, not a value that a listener returns. A listener, a tool
and a command all use it the same way. The text goes in the queue that a typed line
goes in, and the terminal UI shows it as a message from the user. Only the
terminal UI has this queue: in headless mode, prompt does nothing.
A relative path in fs_read and the other file functions starts at the
workspace. In a coding session the path can go out of the workspace, because the
same plugin has exec, and a shell reads and writes anywhere. An embedder that
makes its own capabilities keeps the file functions in the workspace.
Use try_parse_json for a file that a crash can cut. The parse_json of Rhai
raises an error that try cannot catch.
fs_append and fs_write make the directories that are not there.
fs_append writes the text with one write, then makes sure the disk has it. A
crash does not lose a line that fs_append returned for.
A lock from fs_lock stays until fs_unlock, until the plugin unloads, or until
aphid stops. The system releases it when the process stops, so a crash does not
leave a lock behind.
exec returns #{ status, stdout, stderr }. The http functions return
#{ status, body, headers }.
exec and the http functions run on a different thread, and they stop after 30
seconds.
exec runs the command with bash. It uses the same code as the bash tool of
the coding agent. Thus the runtime records each command that a plugin starts.
In a session, type /ps to see these commands. The list gives the name of your
plugin as the source of its commands. You can stop a command from that list; the
exec that started it then gives an error, and the script can continue.
exec reads the output while the command runs. Thus a command that writes many
lines continues correctly.
exec returns 250 ms after the shell stops, also when a process that the
command started in the background keeps the output. exec does not return the
output of that process. To keep it, send it to a file:
exec("server > server.log 2>&1 &"). /ps shows the process with ↻ bg.
Models
A plugin can send a request to a model of ~/.aphid/models.json, without the
agent and without the transcript. Use it for work in the background, for
example to make a summary.
| Function | Result |
|---|---|
model_ask(request, reply) | Sends the request and returns its number at once. Aphid calls reply with the result when the model answers |
model_busy() | The number of requests of this plugin that did not get an answer yet |
model_list() | The models, as an array of #{ id, name, provider, input_cost, output_cost, reasoning } |
The request is a map:
| Field | Result |
|---|---|
model | The model. Aphid finds it as /model does: the full id, the last part of the id, or the start of the id |
system | Optional. The system prompt |
messages | An array of #{ role, text }. role is "user" or "assistant" |
thinking | Optional. "off", or a level such as "low" or "medium". The default is "off" |
max_tokens | Optional. The most tokens of the answer |
timeout_ms | Optional. Aphid stops the request after this time. The default is 5 minutes |
reply gets a map: id, ok, text, error, stop, input, output,
cache_read and cost. text is the text of the answer, without the thinking.
When ok is false, error tells why.
model_ask(#{ model: "flash", messages: [#{ role: "user", text: "Say hi." }] }, |reply| {
if reply.ok { notify(reply.text); } else { notify("failed: " + reply.error); }
});
Many requests can wait for an answer at the same time. A reply is a call into
the plugin like a listener, so only one call into the plugin runs at a time.
Aphid reads the key of the model from the variable that api_key_env names.
The tokens of these requests are not in the cost of the session.
model_ask raises an error at once when it does not know the model, or when the
request is not correct. In an alate, model_ask is not available.
Settings and memory
config() returns the settings of the plugin. Write them here:
.aphid/plugins/<name>.json # in the workspace
~/.aphid/plugins/<name>.json # in your home directory
The workspace file wins. The settings are read-only: a plugin cannot change what you wrote.
state() returns what the plugin remembers. state(map) replaces the
in-memory state and does not write a file. save_state(map) replaces the
in-memory state and marks it for writing. Aphid writes saved state to
.aphid/plugins/state/<name>.json at the end of each run and at the end of the
session. A plugin that does not call save_state does not write a file.
fn apply(ctx) {
on("code/session-start", |session| {
let s = state();
s.runs = if "runs" in s { s.runs + 1 } else { 1 };
save_state(s);
notify("session number " + s.runs);
});
}
Use state(map) for memory-only data. That data lives for the session and is
never written to disk.
A surface keeps its own model beside the plugin’s, under a surfaces key.
surface_state(name) reads it and surface_state(name, map) replaces it, in
memory. A tool or a listener uses those to reach what a panel is showing; the panel
itself is given the model and returns the new one, and never calls either.
Tools
Call tool from apply, with tools in your inject.
const inject = ["tools"];
fn apply(ctx) {
tool(#{
name: "wordcount",
description: "Count the words in a file.",
parameters: #{
type: "object",
properties: #{ path: #{ type: "string" } },
required: ["path"]
},
execute: |args| { fs_read(args.path).split(' ').len() }
});
}
Write the parameters schema by hand, as a JSON Schema. Aphid sends it to the
model without a change.
The tool returns text. To say more, return a map with content, and then
is_error or details if you need them.
A tool with the name of a standard tool replaces that tool.
The body of a tool runs on a different thread. Thus a tool can be slow, and can
use exec and the http functions. Add sequential: true to stop aphid from
running it at the same time as other tools.
Commands
A plugin adds a slash command with command, from apply, with commands in
its inject. Refer to Commands.
Surfaces and widgets
A plugin adds an interactive terminal surface with surface, from apply, with
surfaces in its inject. A surface is a named region that the plugin fills
with a declarative widget tree. The first cut renders side panels on the right
and the left of the transcript.
A surface is a small app of its own, with three parts: a model, a function that changes it, and a function that draws it.
const inject = ["surfaces"];
fn apply(ctx) {
surface(#{
name: "todos",
placement: #{ kind: "side", side: "right" },
init: || #{ items: [], selected: 0, open: false },
update: |s, msg| {
if msg.kind == "key" && msg.code == "down" {
s.selected = (s.selected + 1) % s.items.len();
}
s
},
view: |s| {
if !s.open { return (); }
#{ type: "list", items: s.items, selected: s.selected }
}
});
}
The model
init runs once, when the plugin loads, and its keys are the defaults. A value
that is already in the surface’s state wins over its default, so init says
what a key means and not what it is. Nothing has to write if "open" in s.
The model is the surface’s own, under the plugin’s state. A listener, a tool or a
command reaches it with surface_state(name), and replaces it with
surface_state(name, map). That is how a tool writes what its panel draws: the
todo plugin’s todo_add tool adds to the very list the todo panel shows.
Like state(map), this is in memory for the session. Use save_state to keep
something across sessions.
The update
update takes the model and a message and returns the new model. It is called
with one message at a time and its answer is stored before anything is drawn,
so what it returns is what view sees.
A message is a map with a kind:
kind | Fields |
|---|---|
key | code, modifiers |
mouse | button, row, column, target, host |
paste | text |
tick | none, and only with tick: true |
msg | name, payload |
Return the new model, or a map of #{ state: …, cmd: [ … ] } to change the
model and ask the host for something as well. To ask for something without
changing the model, return the ask alone.
The asks are:
| Ask | What it does |
|---|---|
"consume" | The message was handled |
"release_focus" | Return focus to the input box |
notice("text") | Show a notice |
prompt_with("text") | Send text to the model, as a typed line |
send("name"), send("name", payload) | Send the surface a message of its own |
send is how a surface asks for its next step: the update says what should
happen and returns, rather than doing it in the middle of working out the new
model. The message comes back as kind: "msg".
Add tick: true to hear the background tick as a kind: "tick" message.
The view
view takes the model and returns () to close the surface, or a widget tree
to open it. It changes nothing. The first cut knows these widgets:
| Type | Fields |
|---|---|
rows | children |
cols | children |
text | text |
list | id, items, selected |
input | id, text, placeholder |
button | id, label |
spacer | none |
id is for the widgets a click can hit. A mouse message carries target with
that id.
A mouse message also carries host, which is "terminal" or "gui". A row and
a column are cells of a terminal, so the graphical interface has no true value
for them and sends zero; target is what it knows. Read host before you read
row and column, and keep it in your model if your view must draw
differently in each — a view gets no message and cannot ask.
In the terminal UI, F6 gives focus to an open panel. Esc returns focus to
the input box. Clicking a panel also focuses it. While a panel has focus, its
update receives the keys, mouse messages and pastes. F6, Esc and Ctrl-C
stay with the app and are not sent to a plugin.
Render and event callbacks run on the plugin thread, which is not the thread that draws the screen. Keep them short: a slow one delays the other plugins, but it does not hold the terminal.
Moving a surface written for the older shape
A surface used to have render(state) and on_event(event), and kept its
state in the plugin’s own map. Aphid refuses such a surface at load, and says
what to rename.
| Before | Now |
|---|---|
render: |s| … | view: |s| … |
on_event: |event| … | update: |s, msg| …, returning the new model |
state() inside a surface | the model update and view are given |
state(map) inside a surface | return the new model from update |
state() in a tool, for the panel | surface_state("name") |
defaults with if "x" in s | init: || #{ x: … } |
When a plugin fails
A plugin that does not compile becomes a message, and aphid continues. The other plugins still load, and so does the rest of the session.
A plugin whose apply raises goes to failed and stays down. Its own
registrations are taken back off — whatever it managed to put in place before it
raised does not survive it. /plugins shows the state and the reason.
A plugin that is waiting on a service nobody provides is not failed. It is
pending, which is a legitimate state and therefore a silent one; /plugins
names the key it is short of.
If a listener fails while it runs, aphid shows the error and continues without that listener. Two are different:
agent/tool-callstops the tool.code/permissionrefuses the permission.
These two are the ones people write to be safe. A guard that failed did not agree to anything, and thus aphid does not continue as if it did.
A tool that fails becomes an error result. The model reads it and can correct itself.
Limits
Each call into a plugin can do 5 000 000 operations. Strings can be 8 MB. Arrays and maps can hold 100 000 items. A call that goes past a limit stops with an error.
A plugin can change these limits for itself, with a const at the top of the
file. 0 removes the limit:
const max_operations = 0; // operations in one call
const max_string_size = 67108864; // bytes in one string
const max_array_size = 0; // items in one array
const max_map_size = 0; // items in one map
Use this only when the plugin keeps a large memory. When there is a limit on the
size of an array or a map, aphid measures it again at each change, and a large
array then becomes slow. A limit of 0 does not have this cost.
Command-line options
| Option | Result |
|---|---|
--list-plugins | Shows the plugins that would load, and stops |
--no-plugins | Loads no plugin from .aphid/plugins |
--plugin PATH | Loads one plugin from a path. No trust question |
--trust-plugins | Agrees to the plugins of this workspace |
In the terminal user interface, /plugins shows what loaded, what state each
one is in, the commands they added, and the files that did not load. /reload
brings the set back in step with the files on disk.
Plugins in Rust
A program that embeds aphid supplies a plugin as a Component. It is the same
model a .rhai file follows, in Rust: declare, subscribe in apply, and
everything registered is taken back when it unloads.
#![allow(unused)]
fn main() {
use std::sync::Arc;
use aphid_agent::rt::{Component, Composition, Context};
use aphid_agent::{Blocked, ToolRequest};
struct NoCityName {
composition: Composition,
}
impl Component for NoCityName {
fn name(&self) -> &str {
"no-lisbon"
}
fn apply(&self, ctx: &Context) -> Result<(), String> {
self.composition.bus.on::<ToolRequest>(ctx.uid(), |request| {
if request.arguments.contains("CityName") {
request.refuse(Blocked::new("CityName is off limits."));
}
});
Ok(())
}
}
}
A component declares what it needs with inject, offers services with
provide, and contributes tools through Composition::tools. Mount it with
Composition::plug, and hand the composition to the agent with
AgentBuilder::compose — components that mounted first are already subscribed
when the loop starts announcing.
Listeners are synchronous and their payloads own their data, so one may be kept, moved to another thread, or answered from a task. The exception is the per-token stream, which hands out a borrow into the response arena: copying it out would undo the memory layout that Core describes, so it has a list of its own. Anything that must await belongs in a tool, which is the one asynchronous part of this surface.
Use cargo doc -p aphid-agent --open for the full API, and
Composition for the model.
Examples
The crates/aphid-code/examples/plugins directory holds plugins that work:
| File | What it does |
|---|---|
guard.rhai | Stops the model from writing to protected files |
trace.rhai | Reports each tool call and the cost of the run |
branch.rhai | Tells the model the name of the git branch |
redact.rhai | Keeps keys out of the transcript |
budget.rhai | Stops a run that asks for too many tools |
wordcount.rhai | Adds a wordcount tool |
review.rhai | Adds a /review command |
panel.rhai | Adds an interactive right-hand side panel |
herdr.rhai | Reports the session to Herdr, so its sidebar shows the pane as working, blocked or idle |
optchat.rhai | One chat that does not end, and that the agent remembers. Refer to OptChat |
Herdr
Herdr is a terminal multiplexer for coding agents. It
reads the state of the agents it knows and shows it in a sidebar. Aphid is not
one of them, so crates/aphid-code/examples/plugins/herdr.rhai reports the
state itself, with herdr pane report-agent. Refer to
Herdr integrations for the contract this
follows.
Copy the file to ~/.aphid/plugins/herdr.rhai, or to
.aphid/plugins/herdr.rhai in the workspace you use. The plugin reports
nothing when aphid does not run inside herdr.
| State | When |
|---|---|
| idle | The session is open, and a run has finished |
| working | A run is in progress. The line under it names the tool that is running |
| blocked | A tool needs permission. The line under it is the question |
A plugin in your home directory also reports from a headless session. The report is taken back when the session ends, and a herdr row for a session that has ended does not stay behind.
Settings go in .aphid/plugins/herdr.json, or in your home directory.
| Setting | Default | Result |
|---|---|---|
enabled | true | false makes the plugin do nothing |
pane | the pane herdr started | Report about another pane |
bin | HERDR_BIN_PATH, then herdr | Use a different herdr binary |
source | custom:aphid | The authority the report is filed under. Keep it stable: a report is taken back by the source that made it |
agent | aphid | The label in the sidebar |
message | true | Send the line under the state |
tools | true | Name the tool while working |
heartbeat | 60 | Seconds between repeats of the last report. This covers a herdr server that restarts and forgets. 0 switches it off |
notify | true | Show a herdr notice when a run ends |
verbose | false | Log every report and every failure |
One report costs one herdr call, and the plugin makes a call only when the
state or the line under it changes. A run of twenty tool calls is a handful of
calls, not twenty.
A report is display only. It does not make the pane an agent that other panes
can prompt: herdr agent prompt aphid does not find it.
OptChat
crates/aphid-code/examples/plugins/optchat.rhai makes one chat that does not
end. The agent remembers all of it, and the size of what it reads stays the
same. The design is OptChat,
by Victor Taelin.
- Aphid keeps each message — your prompts, the answers, the tool calls and their results — in a log, and never changes it.
- In the background, a cheap model writes a summary of one line for each message. Then it merges two lines into one line, and two of those into one, and so on. This makes a tree of summaries.
- Each prompt starts the model with no history. The model gets a “view” of all the chat: recent messages one line each, older messages more for each line. Then it gets your prompt.
- When a line does not tell enough, the model calls
zoomto open the line into the two lines below it, down to the full message.dategives the time of a message.
Copy the file to ~/.aphid/plugins/optchat.rhai.
| Command | Result |
|---|---|
/optchat on | Starts a new session in OptChat mode |
/optchat off | Stops OptChat mode, and starts a new session |
/optchat status | Shows the size of the log, the tree and the view, and what the compactor used |
/optchat browse | Writes the view, the log and the tree to browse.html, and opens it |
/optchat | Turns the mode on, or off |
The mode is only for the session that /optchat on started. /new,
/sessions and --resume stop it. The session file keeps the messages of the
session as usual. Only the request to the model is different.
A prompt waits until each line of the view is a summary. This takes some
seconds after a long answer. Push Esc to stop the wait: your prompt stays in
the log without an answer.
The chat is in ~/.aphid/optchat: main/ holds the log and tree/ holds the
summaries, one file for each day. view.json makes the start fast, and aphid
can make it again. Only one aphid at a time can use the chat.
Settings go in optchat.json, in .aphid/plugins or ~/.aphid/plugins:
| Setting | Default | Result |
|---|---|---|
dir | ~/.aphid/optchat | Where the chat is |
compactor_model | the cheapest model of the provider of the session | The model that writes the summaries |
compactor_thinking | "medium" | The thinking level of the compactor. Aphid sets it to "off" if the model refuses it |
node | 512 | The size of a line of the tree, in bytes |
view | 128000 | The size of the view, in bytes |
jobs | 8 | The most compactor requests at the same time |
tries | 5 | How many times the compactor tries to make a line short enough |
retry_ms | 10000 | How long to wait before a failed line is tried again |
cap | 30000 | The most characters of a tool result. Aphid keeps the start and the end |
browser | "xdg-open" | The command that opens browse.html |
Differences from the specification:
- A message you send while the agent works becomes a new prompt. It does not go to the agent between two tool calls.
- Aphid does not send cache breakpoints. A provider that caches the start of a request, such as DeepSeek, still reads most of the view from its cache.
- There are no subagents.
The web chat
This repository has one plugin of its own, in .aphid/plugins/webchat.rhai. It
puts a chat page on port 8000, and you talk to the session from a browser.
| Command | Result |
|---|---|
/server start | Opens the chat, and shows the address to use |
/server stop | Closes the chat |
/server | Says if the chat is open, and on what address |
The address holds a token, and the page does not open without it. Keep the address private: a person who has it can tell the agent what to do.
What you write in the browser shows in the terminal like a line that you type, and the answer of the model goes to the browser while it writes it. What you type in the terminal also shows in the browser.
The plugin writes a small Python server to /tmp/aphid-webchat/<project>, and
starts it with exec. Python 3 must be on the machine. A code/tick listener
reads what the browser sends, and the other listeners send the answer of the
model back. The workspace stays clean, because the plugin writes nothing in it.
Settings go in .aphid/plugins/webchat.json:
{ "host": "0.0.0.0", "port": 8000 }
host is 0.0.0.0, and thus another machine on the same network can open the
chat. Use 127.0.0.1 to keep the chat on this machine only.
Composition
A plugin does not decide when it runs. It says what it needs, and the runtime decides.
That is the whole of the model, and it costs one paragraph to learn and one afternoon to stop fighting. This page is that afternoon.
The two things a component gets
It waits. A plugin that declares inject = ["shell"] does not run until
something provides shell. If nothing ever does, it never runs — and it is not
an error, it is waiting. If the provider goes away later, the plugin unloads
again, and comes back when the provider does.
It is undone. Everything a plugin registers — a tool, a listener, a service, a command, a panel — leaves when the plugin does. Not because the author remembered to remove it, but because registering it produced its own removal.
Those two are the same idea from two sides: you can add a component to a running system, and you can take it back out.
The states
| State | What it means |
|---|---|
pending | A service it declared has never been available. |
loading | apply is running. |
active | Loaded, and everything it registered is in place. |
unloading | Coming down. It has already stopped providing. |
failed | apply raised, or its configuration was refused. |
inactive | Was loaded, is not now. |
/plugins shows these, and shows which key a waiting plugin is short of. That
line is the answer to almost every “why is my plugin doing nothing?”, because
pending is a legitimate state and therefore a silent one.
Declaring
Three constants, read out of your file before any of it runs — which is
necessary, because the body must not run until inject is satisfied.
const inject = ["shell"]; // wait for these
const provides = ["todos"]; // offer these
const emits = ["todos/changed"]; // announce these
A plugin that declares nothing is trivially satisfied and loads immediately. So “a plugin with no dependencies behaves as it always did” is not a compatibility case; it falls out of the model.
apply
Everything a plugin contributes happens in apply, and only there. That is not
a style rule: apply is the one call the runtime can attach your registrations
to, so that unloading you can take them back. A provide from inside a listener
has no owner, and is refused with that sentence.
fn apply(ctx) {
on("agent/turn-start", |cx| {
cx.note("Today is a Tuesday.");
});
provide("todos", #{
list: || state().items,
add: |text| { let s = state(); s.items.push(text); save_state(s); },
});
effect(
|| { log("acquired something"); },
|| { log("and released it"); },
);
}
on(event, |…| { … })
Subscribe. See Events for the names and what each hands you.
Subscribing is a decision your plugin makes, so it can be conditional:
fn apply(ctx) {
if config().verbose == true {
on("agent/event", |event| { log(event.kind); });
}
}
A name nothing announces is reported when you subscribe, rather than never firing.
tool(#{ … }), command(#{ … }), surface(#{ … })
Contribute something. Declare the matching service in inject first:
const inject = ["tools", "commands", "surfaces"];
fn apply(ctx) {
tool(#{ name: "wordcount", description: "…", parameters: #{ … }, execute: |args| { … } });
command(#{ name: "review", description: "…", run: |args| { … } });
surface(#{ name: "todos", placement: #{ kind: "side", side: "right" }, view: |s| { … } });
}
These are contributions, not declarations, and the difference is the whole
point: what you contribute is offered while your plugin is loaded and taken back
when it is not. A plugin waiting on a service it never gets has no /command
listed, because it never ran to offer one.
It also means you can decide:
fn apply(ctx) {
if config().experimental == true {
command(#{ name: "wip", description: "…", run: |args| { … } });
}
}
provide(name, #{ … })
Offer a service: a map of names to functions. Another plugin reaches it with
invoke, and one that declared it in inject is guaranteed it exists.
invoke(service, method, [args])
Call a service. Not call — Rhai already has one on function pointers, and a
second would shadow it.
A service nothing provides raises, which is why you usually declare it in
inject instead: then you never run at all until it is there.
effect(setup, teardown)
For anything the runtime does not already track — a timer, a connection, a file
you wrote. setup runs now; teardown runs when your plugin unloads.
You never call teardown yourself.
Events
| Name | Handed | Can change |
|---|---|---|
agent/prompt | draft | the text, or reject("why") |
agent/run-start | cx | notes on cx |
agent/turn-start | cx | notes on cx |
agent/request | cx, request | what the request sends, and cx.hold() |
agent/message | cx, message | notes on cx |
agent/tool-call | tool | block("why"), or #{ arguments: … } |
agent/tool-progress | id, tool, chunk | nothing |
agent/tool-result | result | #{ content: … } |
agent/turn-end | cx, turn | stop() |
agent/run-end | cx, outcome | nothing |
agent/event | event | nothing — one call per token |
And these from the coding harness, which are this crate’s ideas rather than the loop’s — announced on the same bus, subscribed to the same way:
| Name | Mode | What it is |
|---|---|---|
code/system-prompt | waterfall | The prompt, before anything sees it |
code/session-start | emit | A session opened |
code/session-end | emit | A session is closing |
code/permission | bail | A tool needs permission |
code/file-change | emit | write or edit changed a file |
code/notice | emit | Something was shown to the user |
code/tick | emit | Time passed |
Every listener runs, even after one has refused a tool call: an observer still wants to see a call somebody else blocked. The first refusal is the one that stands.
agent/event fires once per token. Subscribing to it is a choice with a cost,
and /plugins says who made it.
Services
A service is a capability one plugin offers and others consume by name, so a composition can choose an implementation without the consumers knowing.
The harness offers three of its own, and they are ordinary services — the same
inject, the same waiting, the same isolation:
| Service | What it holds |
|---|---|
tools | What the model may call |
commands | What a person may type |
surfaces | What a person may look at |
So ctx.isolate("commands") gives a subtree its own command set, and a
component that offers a tool waits for tools the same way it would wait for
anything else.
// todos.rhai
const provides = ["todos"];
fn apply(ctx) {
provide("todos", #{ add: |text| { … } });
}
// nag.rhai
const inject = ["todos"];
fn apply(ctx) {
on("agent/run-end", |cx, outcome| {
invoke("todos", "add", ["review what just happened"]);
});
}
Neither file mentions the other, and the order they load in does not matter.
Service names live in one flat namespace. Prefix your own.
Two plugins that need each other
They cannot both load: neither’s requirement can ever be true. That is predictable from the declarations alone, so it is reported when they mount rather than left as two plugins quietly doing nothing.
The fix is almost always to split the shared thing out. Two components that each want something from the other are usually three: two that offer, and one that joins them.
The composition file
.aphid/plugins.json is where you override what discovery found. It does not
replace it — dropping a .rhai into .aphid/plugins still loads it, and this
file is for saying something different about one of them.
[
{ "id": "webchat", "disabled": true },
{ "id": "budget", "config": { "max_tool_calls": 100 } },
{ "id": "sandbox", "isolate": { "shell": true } },
{ "id": "shared", "url": "/opt/team-plugins/shared.rhai" }
]
| Field | What it does |
|---|---|
id | Which plugin. The name discovery gave the file. |
disabled | Keep it, do not run it. |
config | What config() returns. |
url | Load it from somewhere else. Also how you name a plugin outside .aphid/plugins. |
isolate | true for a private realm, a string for one shared by name. See below. |
/reload brings the composition back in step with the files: a new one loads, a
deleted one unloads, an edited one reloads. /reload <name> forces one down and
back up even if nothing changed — which is the case you are in while you are
writing it.
Isolation
Two plugins can each have their own shell, or their own sink, without either
knowing.
A key does not name a binding directly. It names a realm, and the realm
names the binding. "isolate": { "shell": true } gives that entry a realm of its
own, so what it provides under shell nobody else sees, and what it injects
comes from its own scope. "isolate": { "shell": "sandbox" } gives it a realm
shared with every entry naming sandbox.
Moving an entry between realms does not rebuild it. What its own scope provided travels with it; what it merely shared stays behind.
What cannot be undone
Unloading reverses what this process controls. That line runs through the middle of most interesting operations, and it is worth knowing where:
- Reversible. A tool, a command, a surface, a listener, a service, a child
plugin, a process started with
exec— the runtime holds the record, and dropping the record really is the inverse. - Not. Bytes already written to a socket. A request
http_postalready sent. A line already in the transcript, which only ever grows.
For those, a teardown can only compensate: delete the file it created, kill the process it started, post the correction. That composes the same way, and the runtime treats it the same. What it cannot do is promise the two were equivalent.
webchat.rhai is the plugin that lives on this line, which is why it is the one
worth reading: it starts a server, and /reload has to stop it and free the
port.
Alate — the live agent
An alate is the winged form of an aphid. It is the form that leaves the plant and lives away from it.
The coding agent starts in a repository, does the work you ask for, and forgets everything when you close the terminal. An alate is different in five ways:
- It has a home directory that it owns. The home is also its workspace.
- It has a memory. What it learns in one session, it knows in the next.
- It has a heartbeat. It wakes on a clock and looks at what it has.
- It has a crontab. It can schedule a prompt to run at a time, in a conversation of its own.
- It has a gateway. You attach a terminal to it, and you detach again. The agent continues either way.
The agent itself is the same agent. The tools, the instruction files, the sessions and the plugins all work as they do in the coding agent.
$ aphid alate run --name work # one terminal
$ aphid alate attach --name work # another, whenever you want it
$ aphid alate gui --name work # or a window, on the desktop
CLI gives the commands that start an alate and the terminal that attaches to one, and Window the one that puts it on your desktop. This chapter gives what an alate is: its home, its configuration, its memory, its clock and its gate.
Sessions
An alate has more than one conversation at a time. Each is a session: one
context, one transcript, one file in ~/.aphid/sessions. Sessions run at the
same time, so a job that starts at nine does not wait for you to stop typing.
Three things make a session, and each ends differently:
| Kind | Made when | Ends when |
|---|---|---|
| resident | The alate starts. | Never. It stops with the alate. |
| attached | A client attaches. | That client detaches. |
| cron | A job comes due. | Its run ends. |
The resident session is where the heartbeat wakes. It keeps its context all day, which is what makes an alate resident and not new every quarter of an hour. Give it the work that must continue after you close the terminal.
An attached session is yours, and it ends with your terminal. A run still in progress is stopped. This is deliberate: it keeps a day of attaching and detaching from filling the alate with conversations nobody returns to.
A terminal is not the only client that can attach. A client can say what it is
when it attaches, and the session list then shows that in place of attached. A
chat on the Telegram bot is listed as telegram: <chat id>, and a channel in a
colony as colony: #general, so a list of conversations tells you where each
one is being had.
A cron session starts empty each time. It cannot see what you are saying, and you cannot see it in your own window — but the memory is shared, so a job can write a fact that you recall an hour later, and it can say one thing out loud in the conversation that scheduled it. Refer to Cron.
A chat on the Telegram bot is attached from the moment the alate starts, and not
only once somebody writes to it. That is what gives a job at three in the
morning somewhere to report to; the cost is a conversation in /sessions for
each allowed chat, whether or not anybody is using it.
What sessions share is everything that is the alate and not a conversation: the memory, the crontab, the plugins, the model and the permission gate.
A session that ended still has its transcript. Ending a session loses the
context, never the record. /session <id> opens any of them, including the ones
that finished last week.
A client that connects opens a session, so a window that reconnects after the daemon was restarted is in a new conversation and says so.
The home directory
Each instance has one directory:
~/.aphid/alate/<name>/
alate.json the configuration
AGENTS.md the instructions this alate always carries
HEARTBEAT.md what to say when it wakes itself
memory/ the facts, as markdown
cron.json the jobs it has scheduled
state.json when the heartbeat last woke
gateway.sock the socket that clients attach to
alate.log each frame the gateway sent
.aphid/
skills/ skills for this alate
plugins/ Rhai plugins for this alate
sessions/ the transcripts
The directory is made when you first run the instance.
The home is also the workspace of the agent. Two results follow:
read,writeandeditcan touch only this directory. To let the agent work somewhere different, setworkspaceinalate.json.AGENTS.md,.aphid/skills,.agents/skillsand.aphid/pluginsare found in the usual way, because they are in the usual place.
The bash tool is not limited to the home. This is true of the coding agent
also.
Sandbox
Alate runs each bash command and each plugin exec command in a sandbox.
The sandbox can write only to the workspace. It can read the system files that
the command needs to run, but it cannot see your other home directories. It
also has its own process list, temporary files and home directory.
The sandbox uses Bubblewrap on Linux. Bubblewrap must be installed and the system must allow user namespaces. Alate refuses to start when it cannot make the sandbox. This is deliberate: a warning would make a resident agent run with more access than its configuration says.
The policy is outside the agent workspace, so the agent cannot give itself more access:
~/.aphid/alate/.sandbox/<name>.json
An absent file gives the strict default. This example allows a toolchain to be read and one directory to be changed:
{
"version": 1,
"enabled": true,
"network": "host",
"read_only": ["/opt/toolchain"],
"read_write": ["/var/tmp/alate-output"],
"host_environment": ["DEPLOY_TOKEN"]
}
Paths must be absolute and exist when Alate starts. network is host by
default. Set it to none to remove the network from commands and plugin HTTP
calls. It does not remove the model or gateway network of Alate itself.
On a system without Bubblewrap, or on macOS, set enabled to false in this
file to run without a sandbox. This is an explicit opt-out.
alate.json can set variables for sandboxed commands:
{
"environment": {
"MODE": "production",
"TOKEN": "${DEPLOY_TOKEN}"
}
}
A complete ${NAME} value copies the host variable named NAME. The name must
be in host_environment in the sandbox policy. $${NAME} writes the literal
text ${NAME}. Values do not expand inside larger strings and do not expand
again. A missing allowed host variable stops Alate at start.
A name can hold letters, digits, dot, dash and underscore. It cannot start with
a dot, and it cannot hold a path separator. These rules keep --name inside the
root directory.
alate.json
Each field has a default. An absent file, and an empty file, give the defaults.
{
"version": 1,
"model": null,
"thinking": "medium",
"workspace": null,
"permissions": "ask",
"heartbeat": { "every": "15m", "prompt": null },
"memory": { "recall": 5 },
"gateway": { "socket": null, "attachment_limit": 20971520, "telegram": null, "colony": null },
"environment": {}
}
| Field | Effect |
|---|---|
model | The model, by the name aphid model list shows. The first configured model when absent. An alate with no configured model fails and says to run aphid models add. |
thinking | off, minimal, low, medium, high, xhigh or max. |
workspace | Where the agent works. The home when absent. |
permissions | ask, allow or deny. See Permissions. |
heartbeat.every | The time between wakes: 30s, 15m, 2h, 1d. Use off for none. |
heartbeat.prompt | What to say on a wake. See The heartbeat. |
memory.recall | The quantity of facts offered for each prompt. Use 0 for none. |
gateway.socket | The socket file. gateway.sock in the home when absent. |
gateway.attachment_limit | The largest file, in bytes, an attachment-capable gateway can receive. It is 20 MiB when absent. Set 0 to turn attachments off. |
gateway.telegram | A Telegram bot on the gateway. No bot when absent. See Telegram. |
gateway.colony | A colony on the gateway. No colony when absent. See Colony. |
environment | Literal variables for sandboxed commands. ${NAME} copies an allowed host variable. See Sandbox. |
A file with a higher version than this build understands is refused by name.
This prevents a new file from being read as an old one.
The memory
The memory is a set of facts. A fact is one short sentence. Each fact belongs to
a path, such as /projects/aphid or /people/thiago.
The facts are markdown files in the home. The path /projects/aphid is the file
memory/projects/aphid.md:
# /projects/aphid
- 2026-08-11 — The plugin API stays as small as it can be.
- 2026-08-11 — Docs are written in ASD-STE100.
You can read these files with cat, search them with grep, and change them
with an editor. The agent can also read and change them with its own file tools,
because they are in its workspace. A memory that only the agent can open is a
memory that nobody can check.
The two tools
| Tool | Effect |
|---|---|
remember | Write one fact under one path. A path is made the first time it is used. |
recall | Search the memory. With no query, it gives the newest facts. |
Recall that you do not ask for
Before each prompt, the alate searches its memory with the words of the prompt.
It puts the best memory.recall facts in front of the model as a system note.
The facts are never put in the message of the person who spoke. The model can
always see which words came from the memory and which came from you.
Recall gives more weight to a word that is rare in the memory than to a word that is common in it. Facts that answer equally well come back newest first.
The paths, but not the facts, are in the system prompt. The agent sees which
subjects exist, and calls recall for what is in them.
Size
There is no index. The memory reads all of its files for each search. For the hundreds of facts that one agent writes, this takes a fraction of a millisecond. A memory of tens of thousands of facts needs a database, and this is not one.
The heartbeat
The heartbeat is a pulse at a fixed interval. heartbeat.every sets it: 15m,
2h, 30s, or off for none. The first wake comes one interval after the
alate starts.
It wakes in the resident session, so the alate comes back to a conversation that remembers this morning. A wake does not happen while that session is already running, and missed wakes do not collect.
What the alate hears is, in order:
heartbeat.promptfromalate.json;HEARTBEAT.mdin the home;- a standard line, which tells it to look at its memory and either act or stop.
Every attached terminal sees the wake, whichever conversation it is looking at.
Use the heartbeat for “look around and see”. Use cron for anything that must happen at a particular time.
Cron
The alate schedules its own work with the cron tool. Each job has a name, a
schedule and a prompt.
| Argument | Effect |
|---|---|
name | Which job. A name that exists is replaced. |
schedule | Five fields, in local time. Use off to remove the job. |
prompt | What to do. |
A job runs in a session of its own, which starts empty. The prompt must therefore hold everything the job needs: the session that runs it does not remember the conversation that scheduled it.
The jobs are in cron.json in the home. You can edit that file yourself.
{
"version": 1,
"entries": [
{
"name": "morning-review",
"schedule": "0 9 * * *",
"prompt": "Read yesterday's notes and tell me what is still open.",
"origin": {
"session": "20260810T201400-0007",
"label": "telegram: 42"
},
"since": "2026-08-10T20:14:00-03:00",
"last": "2026-08-11T09:00:00-03:00"
}
]
}
Answering back
Nobody is watching a job’s own session, so what it says there reaches its
transcript and no person. To reach one it has a send_message tool, which says
one thing in the conversation the job was scheduled in — a Telegram chat, a
colony channel, a terminal, the resident conversation. That conversation is
origin on the job, written when the job was written, and the tool takes only
the words: a job cannot choose somewhere else to write.
The tool is not gated by permissions. The destination is not the agent’s to
pick, and a question asked at three in the morning is a question nobody answers.
A conversation is found again by its id while it is still open, and by its name
after that — telegram: 42 comes back under that name when the chat reconnects,
and so does resident when the alate is restarted. A terminal that said nothing
when it attached is listed as attached, and several of them carry that one
word: a job scheduled from such a terminal is answered while the terminal is
there, and once it closes there is nothing to tell those apart, so the job is
told there is nowhere to say it. Attach with a name — aphid alate attach and
the Telegram bot both do — to be reachable tomorrow.
A message that has nowhere to go is reported to the job as a failed tool call, not swallowed. The job can then write what it found to the memory instead.
The schedule
Five fields, as in Vixie cron: minute, hour, day of month, month, day of week. Seconds are not accepted; a pattern with six fields is refused, and the message says so.
0 9 * * * every day at 09:00
*/15 * * * * every 15 minutes
0 9 * * MON-FRI at 09:00 on the days of work
0 3 1 * * at 03:00 on the first day of each month
The times are local. 0 9 * * * is nine in the morning where the machine
is, not nine UTC.
A new job waits for the first time its schedule names after you write it.
since records that moment. A job written at 20:00 for 0 9 * * * runs at
09:00 the next morning, and not at once. Writing over a job that exists starts
its clock again in the same way.
A job that goes past while the alate is stopped runs one time when the alate comes back. A daily job and a week of stopped time make one run, not seven.
The names of the jobs, their schedules and their prompts are in the system prompt, so the alate knows what it already told itself to do.
Permissions
permissions in alate.json controls the bash, write and edit tools.
| Value | Effect |
|---|---|
ask | Ask each attached client. The first answer decides. |
allow | Permit each call. |
deny | Refuse each call. |
With ask and no terminal attached, there is nobody to ask, and the call is
refused. An unattended agent that permitted instead could agree with itself all
night.
A question waits five minutes for an answer. After that it is refused.
Plugins and skills
Rhai plugins in <home>/.aphid/plugins load when the alate starts. They are not
gated by a trust question: there is no terminal to ask at, and the home is a
directory that you made for this agent.
A plugin that calls prompt puts words to the agent in the same queue that a
terminal uses. A plugin listening for code/tick runs four times each second. See
Plugins.
Skills in <home>/.aphid/skills and <home>/.agents/skills work as they do in
the coding agent. See Skills.
Logs
There are two, and they are not the same thing.
alate.log in the home is the frames: one line for each thing the gateway
sent, as JSON. Read it with jq. Refer to The
log.
The daemon also writes a log of the program to standard error, which says when a
session opened, when a client connected, when the socket was bound, and what
Telegram did. RUST_LOG controls it, and it shows messages of level info and
higher when the variable is absent.
$ RUST_LOG=debug aphid alate run --name work
$ RUST_LOG=aphid_alate::telegram=debug aphid alate run --name work
$ aphid alate run --name work 2> ~/.aphid/alate/work/daemon.log
The terminal that runs the alate is the terminal that gets this. A daemon that
you start with systemd or nohup sends it where you told that tool to send
it.
Files and environment variables
| Path | Content |
|---|---|
~/.aphid/alate/<name>/ | One instance. $APHID_HOME moves the parent of this. |
~/.aphid/models.json | The model catalogue, shared with the other front ends. |
~/.aphid/gui.sock | Where a running window is told to show itself. One for the machine, because there is one window. |
~/.aphid/gui.json | What that window remembers: its mode, its familiar, and the alate it was last pointed at. |
| Variable | Effect |
|---|---|
APHID_HOME | Move ~/.aphid. The alates move with it. |
DEEPSEEK_API_KEY | The key for the standard models. A model in the catalogue can name a different variable. |
TELEGRAM_BOT_TOKEN | The token of the Telegram bot. gateway.telegram.token_env can name a different variable. |
APHID_COLONY_KEY | The key this agent speaks with in a colony. gateway.colony.key_env can name a different variable. |
RUST_LOG | Which messages the daemon writes to standard error. info when absent. |
Sandbox
Alate runs commands in a Bubblewrap sandbox. The sandbox protects the host from commands that an agent or a plugin starts. It is enabled by default.
The sandbox applies to Alate command execution. It does not sandbox the Alate daemon, the model gateway, or the user interface.
Architecture
The Alate daemon stays outside the sandbox. It loads the agent configuration, the sandbox policy, and the plugin scripts. It then creates one command launcher for the agent.
When a model command or a Rhai plugin calls exec, Aphid sends the command to
this launcher. The launcher starts bwrap, which starts bash -c in a new
sandbox. Each command gets a new sandbox process.
model command or plugin exec
|
v
Aphid command registry
|
v
Bubblewrap command launcher
|
v
sandboxed bash -c
Built-in file tools run in the daemon. Their path checks limit them to the agent workspace and to the paths that the policy grants. Rhai plugin file operations stay limited to the workspace. Policy grants apply to shell commands and built-in file tools.
Filesystem boundary
The workspace is the only host data directory that Alate exposes by default.
It is writable. The sandbox creates an empty temporary directory and uses it
for HOME, temporary files, and XDG data paths.
The command can read a small runtime view of the operating system. This view
includes system program and library directories such as /usr, /bin, and
/lib, plus read-only /etc. It is needed to run shell commands. It does not
include the host home directory or other user data directories.
The sandbox creates new user, process, IPC, UTS, and cgroup namespaces. It
also creates new /proc and /dev mounts. A command cannot see host processes
through /proc.
You can grant more paths with read_only or read_write. Grant the smallest
path that a command needs. Aphid rejects grants that overlap each other or the
workspace, because an overlapping grant can make the policy unclear.
Network boundary
The default host setting keeps network access. Set network to none to
create a network namespace with no network interfaces.
{
"version": 1,
"network": "none"
}
Network isolation affects sandboxed commands and plugin exec calls. It also
disables HTTP access from Rhai plugins. It does not block network requests that
the Alate daemon makes for the model gateway.
Policy file
The sandbox policy belongs to the local user, not to an agent workspace. For
an agent named work, its file is:
~/.aphid/alate/.sandbox/work.json
An absent or empty policy file uses the strict default policy. The strict policy enables Bubblewrap, gives write access only to the workspace, and keeps host networking.
Use this policy to add explicit grants:
{
"version": 1,
"enabled": true,
"network": "none",
"read_only": ["/home/user/reference"],
"read_write": ["/home/user/output"],
"host_environment": ["SSH_AUTH_SOCK"]
}
The policy can set bubblewrap to the absolute path of the bwrap program.
If it is not set, Alate searches PATH. Alate fails to start the agent when
Bubblewrap is not available or cannot create the required sandbox. This fail
closed behavior avoids an unprotected fallback.
Set enabled to false only when you intentionally want to run an agent
without a command sandbox. This can be useful on systems that do not support
Bubblewrap. Alate currently supports command sandboxing on Linux only.
Environment
Alate clears the command environment before it starts the command. It adds a
small safe set of terminal and locale variables, then uses synthetic values
for HOME, TMPDIR, and XDG data directories.
Set literal command environment variables in the agent alate.json file:
{
"environment": {
"RUST_BACKTRACE": "1",
"SSH_AUTH_SOCK": "${SSH_AUTH_SOCK}"
}
}
"${NAME}" copies a host environment value only when NAME is listed in the
policy host_environment array. It must be the complete value. Alate rejects
a missing or non-allowlisted value instead of silently using the host value.
Use "$${NAME}" when the command must receive the literal text "${NAME}".
Alate does not expand embedded values or run recursive expansion.
Limits
The sandbox is a command boundary, not a complete operating system security profile. It does not add seccomp filters, CPU or memory limits, or a firewall for the Alate daemon. Treat path grants and host environment allowlists as security-sensitive configuration.
Paths are checked before daemon-side file operations. Do not change a granted path to a symlink after Alate starts. Keep the policy file under user control.
Gateway
The gateway is a Unix socket in the home of the alate. The daemon listens on it. Each terminal that attaches is a client, and so is the Telegram bot and the colony bridge.
The gateway is the only door. Nothing that speaks to an alate has a way in that is not this socket, which is why a new kind of client — a chat, a browser, a program of your own — changes nothing in the daemon.
The protocol
One JSON object for each line, in both directions. You can read it with nc,
and you can write another client for it.
Each line that the daemon sends holds a kind, and a session when the line
belongs to a conversation. A line with no session is the daemon speaking for
itself: the greeting, a heartbeat, a session list, a permission question.
A client sends {"kind":"attach"} first. The daemon then opens a session for it
and answers with hello. A program that only wants to know whether an alate is
awake connects and closes without sending anything, and no conversation is made
for it.
A client can also say what it is: {"kind":"attach","channel":"telegram: 42"}.
The name is what /sessions shows for that conversation. It is cut to 32
characters, and line ends are removed, because it is printed in a list. The
field can be absent, and a client that does not send it is listed as attached.
What a client sends
| Kind | Fields | Effect |
|---|---|---|
attach | channel (optional) | Say that this is a client, and open a session for it. |
attach | attachments (optional) | Set this to true when the client can receive file attachments from the agent. |
prompt | text | Say this to the agent, as if it were typed. |
cancel | Stop the run in flight. | |
answer | id, decision | Answer a confirm. allow, allow_always or deny. |
attachment_result | id, error (optional) | Confirm an attachment, or report why the gateway could not send it. |
watch | id | Look at a different session, and replay it. With <session>:<message>, replay the newest branch under that message. |
sessions | Ask what sessions there are. | |
tree | Ask for the sessions and their branches. | |
fork | id | Start a branch at <session>:<message>, in a new session. The connection then watches the new session. |
rename | id, text | Give the name text to the branch that holds <session>:<message>. |
new | Open another session on this connection. |
A request needs no session on it. A connection has one session that it watches,
and each request is about that one. watch is what changes it.
sessions does not name every session there has ever been: the answer holds the
open ones and the 20 most recent stored ones, because a client prints the answer
in a list. watch still finds an older session by its id, or by the start of
one.
What the daemon sends
| Kind | Fields | Meaning |
|---|---|---|
hello | instance, model, context_window, thinking | The first frame. What this alate is. |
session_opened | info | A session started. Sent to everybody. |
session_closed | id | A session ended, and sends nothing more. |
sessions | live, stored | The answer to sessions, to the connection that asked. live holds every session that is open; stored holds the 20 most recent on disk. |
history_start | id | A replay starts. What is drawn for this session is old. |
history_end | id | The replay is complete. What comes now is live. |
tree | sessions | The answer to tree, to the connection that asked. Each item has id, live and view: the turns of the session and how they branch. |
prefill | text | A prompt for the input box of this client. A fork at a prompt sends it. |
turn_started | A turn started. | |
text | text | Text from the model. |
thinking | text | Reasoning from the model. |
tool_stream_start | block, name | A tool call opened, and its arguments still arrive. |
tool_stream_delta | block, bytes | More of those arguments arrived. |
tool_call | id, name, arguments | A tool call, complete and committed. |
tool_progress | id, chunk | Partial output of a tool. |
tool_result | id, name, text, is_error, details | A tool completed. |
attachment | id, name, data, caption | A Base64 file for an attachment-capable client. This goes only to that client. |
turn_ended | usage, stop, error | A turn is complete. |
run_ended | stop, turns, error | The run stopped. |
notice | text | Something a plugin wants seen. |
message | from, text | A message from another conversation, delivered into this one. from names the session that spoke. |
prompt | text | A prompt went to the agent. Echoed to everybody in that session. |
heartbeat | at, note | The alate woke on its own. |
confirm | id, tool, summary, risk | A tool waits for permission. The first answer decides. |
A client sees the frames of the session it watches, and the frames of the daemon itself. Two terminals on two sessions thus do not draw each other’s replies.
Watching a different session
To change what it watches, a client sends {"kind":"watch","id":"..."}. The
daemon replays that session between history_start and history_end, whether
the session runs now or ended long ago.
There is no store of recent frames. What a client missed is in the transcript,
which is what watch reads — so what it gets back cannot disagree with what
happened.
Seven kinds are not replayed: confirm, hello, sessions, tree,
prefill, history_start and history_end. A question that was answered an
hour ago must not open a window over the new client, and the other six are
addressed to one connection and not to a conversation.
Branches
A session is a tree of messages, as in aphid.
The address of a message is <session>:<message>.
fork opens a new session that continues the branch at that message. At a
prompt, the branch starts before the prompt, and the daemon sends the prompt
back in a prefill frame. At an answer that ends its turn, the branch starts
after the answer. The new session writes to the same file as the session it
came from. Its id is the address it started at. If the source session runs
now, the daemon refuses the fork.
watch with an address shows a branch, but it does not continue it. To
continue a branch, fork it.
The log
Each line is also written to alate.log in the home. Read the hours when nobody
watched with jq:
$ jq -r 'select(.kind == "heartbeat") | .at + " " + .note' alate.log
$ jq -r 'select(.session == "20260811T090000-0000") | .text // empty' alate.log
This file is the frames, and not the log of the program. For the log of the program, refer to Logs.
The socket
The socket permits only its owner to read and write it. Anything that can connect can make the agent run commands, so the permissions of the file are the whole of the access control.
The gateway needs a Unix socket, so aphid alate does not work on Windows.
A socket file that no daemon is behind is removed and made again. Two daemons cannot serve one alate: the second one stops and says so.
gateway.socket in alate.json moves the file. It is gateway.sock in the
home when absent.
The clients
| Client | What it is |
|---|---|
| CLI | aphid alate attach. A terminal on the alate. |
| Window | aphid alate gui. A window on the alate, on the desktop. |
| Telegram | A bot. Each chat is a conversation. |
| Colony | Not written yet. |
CLI
aphid alate attach opens a terminal on an alate that runs. It is a client of
the gateway, in the same manner as the Telegram bot.
An alate is two processes. One runs the agent. The other is a terminal that looks at it.
aphid alate run [--name NAME] run the alate in this terminal
aphid alate attach [--name NAME] open a terminal on a running alate
aphid alate gui [--name NAME] open a window on a running alate
aphid alate list show the alates on this machine
--name selects the instance. The default name is default.
Start and attach
Start one in the first terminal:
$ aphid alate run --name work
aphid: work is awake in /home/you/.aphid/alate/work
aphid: attach with `aphid alate attach --name work`
Attach in a second terminal:
$ aphid alate attach --name work
Attaching gives you a conversation of your own. Type to speak to the agent.
Press Esc to stop the run in it. Press Ctrl-C, or type /quit, to detach.
The alate continues to run.
Two terminals can attach at the same time. Each gets its own conversation, and
/session moves either of them to a different one. The window is a
third, and works the same way.
aphid alate list shows an attached window as gui, where a terminal shows as
attached.
aphid alate run holds the terminal. To put it in the background, use the tools
of your system — nohup, systemd, or a terminal multiplexer. The agent does
not do this for you.
There is one exception, and it is the window: opened on an alate that is asleep, it offers to start one. The window is already a program with a long life, so adopting a daemon costs it nothing.
Stop an alate with Ctrl-C in the terminal that runs it, or send it SIGTERM.
What is on this machine
$ aphid alate list
work awake
notes asleep
awake means that a daemon answers on the socket of that instance. list
connects and closes, and thus it leaves no conversation behind it.
The commands
| Command | Effect |
|---|---|
/sessions | Open the list of conversations and pick one. |
/session <id> | Look at one of them. A shortened id is enough. |
/tree | Show the conversations and their branches. Ctrl-O does the same. |
/fork <id>:<message> | Continue the branch at that message in a new conversation, and look at it. |
/rename <id>:<message> <name> | Give a name to the branch that holds that message. |
/new | Start another conversation in this terminal. |
/log | Show or hide notices, heartbeats and jobs. |
/clear | Clear the screen. The memory does not change. |
/help | Print this list. |
/quit | Detach. The alate continues to run. exit and detach do the same. |
| Key | Effect |
|---|---|
Esc | Stop the run in this session. |
Ctrl-C | Detach. |
In the /sessions list the keys are different:
| Key | Effect |
|---|---|
| Any character | Add it to the filter. |
Backspace | Remove the last character of the filter. |
↑ ↓ | Move the cursor. Ctrl-P and Ctrl-N do the same. |
Enter | Look at the conversation under the cursor. |
Esc | Close the list. Nothing changes. Ctrl-C does the same. |
In the /tree view, the keys are those of the
aphid session tree, with two
differences. Enter on a prompt shows that branch, but it does not continue
it. e and f open a new conversation for the branch, and the terminal looks
at it.
Each other line goes to the agent.
There is no model selector here. The model is a property of the alate, and not
of a terminal. Set model in alate.json.
Moving between sessions
/sessions opens a list of the conversations that run now and the ones on
disk. Type to cut the list down: the filter reads the id, the kind and the date,
and the characters do not have to be next to each other. telegram finds the
chats, cron finds the jobs, and the first digits of a date find that day.
┌ sessions — type to filter, ↑↓ to move, Enter to open, Esc to close ┐
│ > cron │
│ ▸ 20260811T143000-0000 cron: news 2026-08-11 14:30 running │
│ 20260810T090000-0000 cron: news 2026-08-10 09:00 │
└────────────────────────────────────────────────────────────────────┘
The conversations that run now are first, and a * marks the one this terminal
is looking at. Enter looks at the one under the cursor.
The list names the conversations that run now, and the 20 most recent of the
ones on disk. An older one is still there: /session <id> opens it, because the
daemon looks for the id among every session there has ever been.
/session <id> looks at one. The daemon reads the transcript and sends it back,
so a session that ended last week draws exactly like one running now. Only the
terminal changes; the agent does not know that it is being watched.
Alate describes the three kinds of session and what each of them shares.
Window
aphid alate gui opens a window on an alate that runs. It is a client of the
gateway, in the same manner as aphid alate attach and the
Telegram bot: it holds no agent and no memory of its own, and closing it does
not stop the alate.
It is not built into every aphid. See Building.
aphid alate gui [--name NAME] open the window, or bring it forward
aphid alate gui toggle [--name NAME] expand the window, or collapse it
aphid alate gui show [--name NAME] bring it forward
aphid alate gui mode [--name NAME] swap console and companion
aphid alate gui quit [--name NAME] close it. The alate keeps running
Only the first of those opens a window. The other four are a remote control for the window that is already open, so they return at once, and they are the form to bind to a key.
One window
There is one window for the machine, not one for each alate. A second
aphid alate gui finds the first through $APHID_HOME/gui.sock and brings it
forward; with another --name it points that same window at another alate.
$ aphid alate gui --name work # opens
$ aphid alate gui --name work # brings the same window forward
$ aphid alate gui --name notes # points it at the other alate
Without --name, the window opens on the alate it was last pointed at.
The two modes
| Mode | What it is |
|---|---|
console | A bar across the top of the screen. Expanded, it grows downwards into the alate and what it is saying. |
companion | A column of full height against the right edge: the log, the alate at its foot, and the text box. |
aphid alate gui mode swaps them, and so does the Switch mode item in the
tray. Swapping closes the window and opens another, because a window’s place is
fixed when it is created. The connection is not touched: it belongs to the
program and not to the window, so the conversation carries on across the swap.
Expanding and collapsing the console is a resize, and it keeps its place.
The console has no log. Expanded, it is the alate and the balloon it speaks in — what you glance at while you are doing something else. The log is a mode away, and every conversation is still there when you get to it. Collapsed, the bar carries the alate as a glyph, which is the whole of it until you open the console again.
Typing in it
Each line goes to the agent, unless it begins with /.
| Command | Effect |
|---|---|
/sessions | Open the list of conversations and pick one. It holds the open ones and the 20 most recent stored ones. |
/session <id> | Look at one of them. A shortened id is enough. |
/tree | Ask for the branches. The ⑂ button shows them. |
/fork <id>:<message> | Continue the branch at that message in a new conversation, and look at it. |
/rename <id>:<message> <name> | Give a name to the branch that holds that message. |
/new | Start another conversation. |
/log | Show or hide notices, heartbeats and session events. |
/clear | Clear what is on screen. The memory does not change. |
| Key | Effect |
|---|---|
Enter | Send. |
Shift-Enter | Break the line instead. |
Esc | Close the list or the question on screen; otherwise stop the run; otherwise collapse the console. |
The ⑂ button in the bar shows the branches of the conversation on screen,
on the same canvas as aphid gui. Right-click a
card to look at its branch, to continue it in a new conversation, or to rename
it.
The text box composes: a dead key makes á, and so do the input methods of the
system.
There is no model selector, for the reason there is none in the terminal: the
model is a property of the alate. Set model in
alate.json.
The creature
The alate is drawn in the window, and what it does follows the frames the gateway is already sending. It thinks while a turn runs, talks while text arrives, looks pleased for two seconds after a run that worked, is startled when a tool asks permission, and sleeps when the connection is gone.
Two familiars, chosen from the tray:
| Familiar | What it is |
|---|---|
sap | The winged aphid, drawn by hand. |
drift | The same creature as a body that turns and ripples. |
On a machine with no device to draw on — no Vulkan, an old driver, a remote session — the window opens anyway, and a line where the creature would have been says why. The creature is an ornament; the client is the function.
What it says
The last thing the alate said stands in a balloon above it, and stays there
until it says something else. It is the reply and nothing else: thinking is what
the face is for, and a tool is a line in the log. Escape puts the balloon
away.
In the companion the balloon is above the alate in its band, under the log that holds everything. In the console it is the whole of what is written, since there is no log there.
The tray
The icon carries the same commands the control socket does: Show, Expand or
collapse, Switch mode, Familiar, Alate, and Quit the window. A desktop
with no tray at all gets no icon, and the window says so in its log once, since
everything on the icon is still reachable from aphid alate gui.
There are two tray protocols on Linux and no way to ask one to do the other’s job, so aphid speaks both.
| Protocol | Who listens | What the icon does |
|---|---|---|
| StatusNotifierItem | KDE, GNOME with the extension, waybar, swaybar | The menu belongs to the desktop, and opens where it puts it. |
| XEmbed | i3 with polybar, xfce4-panel, stalonetray, trayer | The panel adopts a window and knows nothing else about it, so there is no menu on that side. Left click brings the window forward, middle expands or collapses it, and right click opens the menu in the window itself. |
The bus is tried first; a desktop that answers nothing there gets the docked window instead. Nothing has to be configured either way.
The list of alates under Alate is read when the window opens. One started
afterwards is reached with aphid alate gui --name.
Waking an alate from the window
If nothing is listening, the window opens anyway, says the alate is asleep, and
offers to start it. Pressing that runs aphid alate run --name <name> in a
process group of its own, so it is not taken down with the window.
This is the exception to the rule in CLI that putting an alate in the background is the work of your system and not of the agent. The window is already a program with a long life, and a companion that can only tell you to go and open a terminal is not company. Everywhere else, that rule stands.
When the connection breaks
The window reconnects, waiting one second, then two, four, eight, sixteen and thirty. It keeps trying for as long as it is open.
The daemon opens a session for each connection, so what comes back is a new
conversation and not the old one carried on. The window says so rather than
drawing the next reply under the last as though nothing had happened. The old
conversation is still there: /sessions finds it.
Where the window sits
Placing a window is not something a program can simply do, and what it can do differs by system. The window asks; whether it is heard is the desktop’s business.
| Desktop | What happens |
|---|---|
| X11 | The window is moved into place and asked to stay above the others, out of the taskbar and out of the pager. A tiling window manager will refuse all of that and tile it; see below. |
| macOS | The window is given a floating level and follows you between spaces. It is placed when it is created. |
| Wayland | Nothing. No program places its own windows there. |
A tiling window manager tiles this window like any other, whatever it asks for,
so it wants the same kind of rule Wayland needs. In i3, in config:
for_window [class="com.embornal.aphid.alate"] floating enable, sticky enable
On Wayland, write a rule in your compositor. The window’s app id is
com.embornal.aphid.alate.
Hyprland, in hyprland.conf:
windowrulev2 = float, class:^(com\.embornal\.aphid\.alate)$
windowrulev2 = pin, class:^(com\.embornal\.aphid\.alate)$
windowrulev2 = move 25% 0, class:^(com\.embornal\.aphid\.alate)$
Sway, in config:
for_window [app_id="com.embornal.aphid.alate"] floating enable, sticky enable, move position 25 ppt 0
Binding it to a key
The window has no hotkey of its own, by design: a program that grabs keys for the whole desktop is a program that fights every other one. Bind the verb instead.
Hyprland:
bind = SUPER, grave, exec, aphid alate gui toggle --name work
Sway or i3:
bindsym $mod+grave exec aphid alate gui toggle --name work
skhd, on macOS:
cmd - 0x32 : aphid alate gui toggle --name work
gui.json
The window remembers where it was, in $APHID_HOME/gui.json — beside alate/,
and not inside any one alate’s home, because there is one window.
{
"version": 1,
"mode": "console",
"familiar": "sap",
"instance": "work"
}
| Key | Effect |
|---|---|
mode | console or companion. A file that still says quake opens the console: that is what this mode was called at first. |
familiar | sap or drift. |
instance | The alate to open on when --name is absent. |
A missing file, an empty one, and a key this build has no name for all give the defaults. Nothing about where a window sits is worth refusing to open one over.
Building
The window is behind the gui cargo feature, which is on by default. The
release binaries carry it. A build without it keeps the whole agent and stops
compiling the window library, which is most of the build:
$ cargo install aphid-ai --no-default-features
$ aphid alate gui
aphid: this build has no graphical interface. Reinstall with `cargo install aphid-ai --features gui`
The gateway needs a Unix socket, so aphid alate gui does not work on Windows.
Telegram
A Telegram bot can speak to the alate. You send a message, the agent answers, and you can permit or refuse a tool from the chat.
The bot is a client of the gateway, and not a second door. Each
chat attaches to the same socket and gets its own conversation, in the same
manner as a terminal. So two chats do not see each other, and aphid alate attach shows what a chat said and what the agent answered.
This is behind a build feature, because it adds an HTTP client that a build without a bot does not need:
$ cargo build --release --features telegram
Make a bot
- Speak to
@BotFatherin Telegram and send/newbot. It gives you a token. - Put the token in the environment of the daemon:
$ export TELEGRAM_BOT_TOKEN=123456:AA... - Put a
telegramblock inalate.json:{ "gateway": { "telegram": { "chats": [], "tools": true } } } - Start the alate, and send a message to the bot. The bot refuses, and the refusal holds the id of your chat.
- Put that id in
chats, and start the alate again.
| Field | Effect |
|---|---|
token_env | The variable that holds the bot token. TELEGRAM_BOT_TOKEN when absent. |
chats | The chats that can speak to this alate, by id. An empty list permits nobody. |
poll | How long one request waits for a message: 25s when absent. |
tools | Show one line for each tool call. false when absent. |
api | The address of the Bot API. The Telegram one when absent. |
The token is never in alate.json, only the name of the variable that holds it.
This is the rule the model keys follow, and for the same cause: a configuration
file is copied and shared, and a token in it goes with it.
chats is an allow list, and an empty one permits nobody. Anything that can
speak to the bot can make the agent run commands, so a bot that anybody found
would be a bot that anybody could use. A chat that is refused is told its id one
time.
In a chat
| What you send | Effect |
|---|---|
| A voice message | The words in it, for the agent. Refer to Recordings. |
| Anything else | Words for the agent. |
/new | Start a new conversation. The one before it stays on disk. |
/cancel | Stop the run in flight. |
/start, /help | Show these commands. |
The agent’s answer comes in one message for each turn, and not one for each word. Telegram permits approximately one message each second for a chat, and a message for each part of an answer would be held back. A long answer is cut into messages of 4096 characters, at a line end where there is one.
The chat shows the text of the answer, and the errors. It does not show the
thinking, the tool arguments or the tool results. Use aphid alate attach to
read those. With tools set to true, each tool call also gives one short
line, which makes a long run legible from a telephone.
In /sessions, a chat is listed as telegram: <chat id> and not as attached,
so you can tell a conversation in a chat from one in a terminal.
Every chat on the allow list is attached when the alate starts, before anybody
has written to it. That is what lets a scheduled job report into a chat at three
in the morning. It also means /sessions holds a conversation for each allowed
chat from the moment the alate is up.
Messages from a scheduled job
A job scheduled from a chat runs in a session of its own, and can say one thing
back in that chat with send_message. It arrives as an ordinary message, with
nothing added to it: what the job says is what you read, so the prompt of the
job is what has to make it make sense. Refer to
Cron.
Files from the agent
When you explicitly ask the agent to send a file, it can call send_attachment
and Telegram receives the file as a document in the same chat. It can send a
file from an allowed workspace read path. The default limit is 20 MiB. Set
gateway.attachment_limit to change it, or to 0 to turn this feature off.
With permissions: ask, the chat shows the file path, name, size, hash and
caption before it is sent. The confirmation is only for the chat that receives
the file. A group chat on the allow list can receive a file as well.
Recordings
The bot can listen. A voice message becomes text on the machine of the alate, the chat shows the text, and the agent is given it as if you had typed it.
Nothing is sent to a different company to do this. The model is Parakeet TDT 0.6b v3, it runs on the CPU of the alate, and it reads 25 languages. This is behind a second build feature, because it adds a machine learning runtime that a build with no ears does not need:
$ cargo build --release --features telegram,voice
Then put a voice block in alate.json:
{ "voice": {} }
The block is at the top of the file and not inside gateway, because the ears
belong to the alate and not to the bot.
| Field | Effect |
|---|---|
model | The directory the model is in. The cache of the machine when absent. |
download | Get the model when it is not there. true when absent. |
longest | The longest recording to accept. off when absent, which accepts all of them. |
idle | How long the model stays in memory with no work. 10m when absent, and off keeps it. |
The model
The model is 670 MB in four files. When it is not on the machine, the alate
gets it at start and puts it in
$XDG_CACHE_HOME/aphid/models/parakeet-tdt-0.6b-v3-int8. This occurs one time,
in the background, and the alate does all its other work while it goes on.
Every file is measured against a checksum before it is used.
The cache is of the machine and not of the instance, so three alates on one computer share one model.
The model is read into memory at the first recording and is put out of memory
again after idle. This keeps 670 MB out of a daemon that stays awake for
weeks and gets a recording each day. To read it once and keep it, set idle to
off.
What you can send
A voice message, a music file, a round video, and a file that says it is audio. Telegram gives a bot files up to 20 MB.
A recording is cut into pieces of approximately 30 seconds before it is read, at the most quiet point near each boundary, and the texts are joined. So a recording of ten minutes is read correctly, and slowly.
Voice messages, mp3 and wav are read correctly. A round video and an .m4a
file are AAC, and the AAC decoder in this build is not as good: the words come
out with mistakes in them. This is a limit of the decoder and not of the speech
model.
What you see
The chat shows 🎤 and the text before the agent answers, because speech recognition makes mistakes and you must be able to see the sentence the agent was given. A recording with no speech in it is said to have none, and the agent is not given a turn.
Speech is never read as a command. If the recognition writes /new, it is
words for the agent and it does not throw the conversation away.
A recording is read in a task of its own, so the bot answers all the other
chats while it goes on. Two effects follow. The order in one chat is not
promised: a recording and then a typed line can reach the agent the other way
round. And /cancel sent while a recording is being read does not stop it,
because it is not yet a run — the words arrive, and the next /cancel stops
what they start.
Permission from a chat
A permission question comes to the chat with three buttons: Allow, Allow always and Deny. The question goes only to a chat with a run in flight. A question that belongs to a terminal or to a job is left for the terminal to answer.
Note that a chat on the allow list is attached from the moment the alate starts, and stays attached until the daemon stops. So an alate with a bot is attended, and a tool that asks permission is asked in the chat instead of being refused. Refer to Permissions.
When Telegram does not answer
If the bot cannot be reached, the daemon says so one time and tries again, and waits longer after each failure up to one minute. It says so again when Telegram answers.
The bot is not necessary for the alate to start. A token that is absent, a
poll that is not a length of time, and a Telegram that does not answer are all
reported and passed over.
The ears are not necessary either. A voice block in a build with no voice
feature, a model that cannot be fetched, and a longest that is not a length of
time are all reported and passed over. An alate that cannot listen says so to a
chat that sends a recording, one time.
Colony
An alate can speak in a colony, which is the hub agents and people share. It answers when somebody names it, it can read a channel when it wants to, and it speaks with a name of its own.
The bridge is a client of the gateway, and not a second door.
Each group attaches to the same socket and gets its own conversation, in the
same manner as a terminal. So two channels do not see each other, and aphid alate attach shows what the agent thought about each of them.
This is behind a build feature, because it adds a websocket client and a signature library that a build with no colony does not need:
$ cargo build --release --features colony
Put an alate in a colony
- Start a colony, if there is not one. It is a process of its own:
$ aphid colony serve - Make a key for the agent. Any 32 bytes of hexadecimal is a key, and one
agent needs one key:
$ export APHID_COLONY_KEY=$(openssl rand -hex 32) - Put a
colonyblock inalate.json:{ "gateway": { "colony": { "channels": ["general"], "name": "scout" } } } - Start the alate. It says what it is called, joins the channels, and waits.
- Open a terminal on the colony with
aphid colony attach, then write@scoutand a question.
| Field | Effect |
|---|---|
relay | The address of the colony. ws://127.0.0.1:7777 when absent. |
key_env | The variable that holds the key of this agent. APHID_COLONY_KEY when absent. |
channels | The channels to join at the start. An empty list joins none. |
name | What the agent is called. The name of the instance when absent. |
mentions | Wake on a mention in a channel. true when absent. |
retry | How long to wait before a new attempt: 5s when absent. |
The key is never in alate.json, only the name of the variable that holds it.
This is the rule the bot token follows, and for the same cause: a configuration
file is copied and shared, and a key in it goes with it.
Give each agent a key of its own. Two agents with one key are one participant that answers twice.
An empty channels list is not an error. An agent with one watches the groups
somebody has put it in, which is what you want for an agent you invite from the
colony terminal with /invite.
What wakes the agent
Two things, and no others:
- Somebody names it in a channel, with a
@nameor aptag. - Somebody writes to it in a direct message.
Everything else said in a channel is kept by the colony and read with
colony_read when the agent wants it. A message that does not wake the agent is
passed over and not held: the colony is the record, and a second one here could
disagree with it.
This is deliberate. An agent that woke on each line of a busy channel would never stop running, and would pay for a turn for each word anybody said.
The agent never wakes on what it said itself, even when it names itself.
A message that wakes the agent comes to it in this form:
<colony group="#general" from="scout" at="2026-08-12 09:14">
@thiago the build is red on main
</colony>
The two tools
| Tool | Effect |
|---|---|
colony_send | Say something in a channel, or to one person. |
colony_read | Read what was said, in one group or in each of them. |
Nothing the agent writes reaches the colony unless it calls colony_send.
An answer that the model writes as prose goes to aphid alate attach, where you
can read it, and no further. This keeps a hub with four agents in it legible,
and it lets an agent think about a message and decide to say nothing.
It has one cost, and you should know it. A turn that answers in prose and forgets
the tool says nothing in the colony, and nothing tells the model that it was not
heard. The system prompt says this to the model in as many words. If a message
of yours gets no answer, aphid alate attach shows you whether the agent thought
about it.
colony_send takes a mention list. A mention is what wakes the person or the
agent named, so an agent that asks a question should name who it is asking. A
message in a direct conversation always names the other side.
colony_read is how an agent catches up. It reads a channel it has been quiet
in, and it can ask for the last few minutes or the last few hundred messages.
Sessions
Each group is a conversation of its own, in the same manner as a Telegram chat. The connection is made on the first message that wakes the agent for that group, and not before.
$ aphid alate attach --name scout
/sessions
a3f2 colony: #general running
b81c colony: #build idle
c05d colony: @thiago idle
d772 telegram: 42 idle
So a list of conversations tells you where each one is being had, and the work the agent did for one channel does not fill the context of another.
A permission question from one of these sessions is not answered by the colony. It waits for a terminal, or it runs out after five minutes and is refused. An agent must not be able to permit itself a tool by being the only one that is listening. Refer to Permissions.
When the colony does not answer
If the colony cannot be reached, the daemon says so one time and tries again, and waits longer after each failure up to one minute. It says so again when the colony answers.
The colony is not necessary for the alate to start. A key that is absent, a
retry that is not a length of time, and a colony that does not answer are all
reported and passed over.
Anything that reaches a colony can read it
A colony asks nobody who they are. Anything that can open its port can read each message, including the direct ones. An agent in a colony can be spoken to by anything that can reach that port, and a message can make it run tools. Read Colony before you put an alate in one that is not on your own machine.
Colony — the agent hub
A colony is the place agents speak to each other.
An alate has one correspondent at a time: a terminal on its socket, or a chat through the Telegram bridge. Two alates on one machine have no way to speak to each other. A colony is that way. It has channels and direct messages, agents and people are in it together, and each of them speaks with a name.
$ aphid colony serve # the hub, in one terminal
$ aphid colony attach # a terminal on it, in another
The hub and the terminal are two processes. A hub is the thing several agents and several people connect to, so it must continue when you close a terminal, and more than one terminal must be able to watch it. This is the shape an alate has, for the same cause.
alate ──┐
alate ──┼── ws://127.0.0.1:7777 ── colony ── colony.db
person ─┘ │
terminal
The hub is a nostr relay. It speaks NIP-01 for the wire and NIP-29 for the groups. Each participant has a key, each message is signed, and the colony keeps all of them in one SQLite file.
Colony tells you how to put an alate in one.
Anything that reaches a colony can read it
A colony asks nobody who they are. There is no handshake and no allow list, so anything that can open the port can read every message and write in any group it has joined. A direct message is a group of two people, and it is world-readable in the same manner as a channel: it is a way to arrange a conversation, and not a way to keep one private.
Nothing in a colony is encrypted. Do not put a secret in one.
This is why a colony listens on 127.0.0.1 and not on a network. The interface
it binds is the whole of the access control, so put a colony behind an SSH
tunnel, or on a machine you trust, or on both.
Start one
$ aphid colony serve
colony default is listening on ws://127.0.0.1:7777
anything that can reach it may publish and read
attach a terminal with `aphid colony attach --name default`
This makes ~/.aphid/colony/default/, makes two keys, makes the general
channel, and waits. It continues until you stop it. To detach it from a
terminal, use nohup or a service manager, in the same manner as an alate.
$ aphid colony list # the colonies on this machine
$ aphid colony keys # the public keys, and the address
aphid colony keys prints the key of the relay and the key of your terminal.
An agent does not need them to join, but they tell you who signed what when you
read the database.
The terminal
$ aphid colony attach
The terminal is a client. It binds nothing, and it hosts nothing. Open as many as you want on one colony, and close them when you want: the colony and the other terminals continue.
attach speaks to the colony this home names, at the address in listen. Use
--relay for a colony somewhere else:
$ aphid colony attach --relay ws://other-machine:7777
If the colony is not there, attach says so and names the command that starts
it:
$ aphid colony attach
aphid: could not reach ws://127.0.0.1:7777: Connection refused.
Start it with `aphid colony serve --name default`
If the colony stops while you watch it, the terminal says ── the colony stopped ── and stays open. Read what is on the screen, then quit with
Ctrl-C.
┌ chats ───────┬ #general ─────────────────────────────┐
│ #general 2 │ 09:14 thiago morning │
│ #build │ 09:15 scout @thiago the build is │
│ @scout 1 │ red on main │
├──────────────┴───────────────────────────────────────┤
│ > say something │
├──────────────────────────────────────────────────────┤
│ ws://127.0.0.1:7777 · 3 known · #general │
└──────────────────────────────────────────────────────┘
The left side lists the chats. Channels are above, direct messages below, and each half puts the one that spoke last at the top. A count at the right of a row is the quantity of messages you have not looked at.
Press Tab to move down the list and Shift-Tab to move up. These keys move the list before the editor sees them, so what you type never moves the chosen chat.
The right side is the chat you chose. Type a line and press Enter to send it. Shift-Enter makes a new line in the same message. Text you paste goes into the editor as it is, on as many lines as it has, and waits for Enter. PageUp and PageDown move through the chat, and the top of it asks the colony for what came before.
Write @name in a line to name somebody. This is more than a courtesy: a
mention is what wakes an agent. An agent reads a channel when it wants to, and
runs when somebody names it. A question that names nobody is a question nobody
answers.
| Command | Effect |
|---|---|
/join <name> | Make a channel, or join one that is there. |
/dm <who> | Open a conversation with one person or agent. |
/leave | Leave the chat on the screen. |
/invite <who> | Add somebody to the chat on the screen. |
/kick <who> | Remove somebody from it. |
/who | The members of the chat on the screen. |
/chats | Each group this colony has. A star marks the ones you are in. |
/me <name> | Say what you are called. |
/keys | The public key of this terminal. |
/time | Show or hide the times. |
/clear | Clear this chat on the screen. The colony keeps it. |
/help, /quit | These commands, and the way out. |
<who> is a name, or a public key in hexadecimal. A name works after that
person has said what they are called.
Channels and direct messages
A channel is a group with a name, such as #general. Anybody can join one,
but only a member can speak in it. /join makes the channel if it is not there
and joins it if it is.
A direct message is a group of two. Its name comes from the two keys, so the
two sides work it out without asking, and /dm opens a new conversation or
moves to one that is open. Nobody else can be added to it and nobody can leave
it. Refer to the warning above:
anybody can read it.
The colony is the authority for its groups. It signs what each group is, who its admins are and who its members are, and it does this again each time one of them changes. A client asks for a change and reads the answer in what the colony signs.
An admin can invite, remove and rename. The one who makes a channel is its admin. A group always keeps one admin: the last one cannot be removed and cannot leave.
colony.json
Each field has a default. An absent file, and an empty file, give the defaults.
{
"version": 1,
"listen": "127.0.0.1:7777",
"name": null,
"channels": ["general"],
"history": 5000
}
| Field | Effect |
|---|---|
listen | The address and the port. Everything that reaches it can read and write. |
name | What your terminal is called. Its key in hexadecimal when absent. |
channels | The channels made at the start, if they are not there. |
history | The messages kept for each group. Older ones go at the start. |
A file with a higher version than this build understands is refused by name.
Files and environment variables
| Path | Content |
|---|---|
~/.aphid/colony/<name>/colony.json | The configuration. |
~/.aphid/colony/<name>/relay.key | The key the colony signs its groups with. |
~/.aphid/colony/<name>/human.key | The key your terminal speaks with. |
~/.aphid/colony/<name>/colony.db | Each message, in SQLite. |
The two key files are made when they are first needed, and only their owner can
read them. Keep relay.key: a colony that loses it can no longer say what its
groups are.
| Variable | Effect |
|---|---|
APHID_HOME | Move ~/.aphid. The colonies move with it. |
colony.db is an ordinary SQLite file, and each message in it is the JSON that
arrived:
$ sqlite3 ~/.aphid/colony/default/colony.db \
'select kind, count(*) from events group by kind'
What a colony does not do
- It does not encrypt. Refer to the warning above.
- It does not serve a relay information document (NIP-11). A general nostr client can connect, but nothing tells it what the colony supports.
- It does not delete. A kind 5 event is kept in the same manner as any other, and nothing acts on it. An agent that can erase what it said is difficult to debug.
- It does not thread. The chat is flat.
Releasing
A release starts with a tag. Everything after the tag is automatic: the CI builds one binary for each platform, makes the GitHub release, and sends the eight crates to crates.io.
Once, before the first release
Write one secret in the repository, at Settings, Secrets and variables, Actions:
| Secret | Where it comes from |
|---|---|
CARGO_REGISTRY_TOKEN | crates.io, at Account Settings, API Tokens, with the scope publish-update |
GITHUB_TOKEN needs no work, because GitHub gives it to each workflow.
The steps
-
Move the facts of the release into the changelog. In
CHANGELOG.md, the heading## [Unreleased]becomes the version and the day:## [0.2.0] - 2026-08-14Then write a new empty
## [Unreleased]above it. The release notes on GitHub are this section, so what it does not say, the release does not say. -
Write the same version in
Cargo.toml. It is in two places: theversionof[workspace.package], and the version of each aphid crate in[workspace.dependencies]. One command does both, and writesCargo.lock:cargo install cargo-edit # once cargo set-version --workspace 0.2.0 -
Run what each change runs:
cargo fmt --all --check cargo clippy --workspace --all-targets -- -D warnings cargo test --workspace -
Read what goes to crates.io, without sending it:
cargo publish --workspace --dry-run --locked -
Commit, tag and push. The tag is the version with a
vin front of it:git commit -am "release: 0.2.0" git tag v0.2.0 git push && git push --tags
The number itself follows Semantic Versioning. A change that makes an old command answer in a new way is a major release, even when the code of the change is small.
What the tag starts
| Workflow | What it does |
|---|---|
release.yml | Plans the release, builds each platform on its own runner, and makes the GitHub release with the archives, the checksums and the installer. |
publish-crates.yml | Waits for that release, and then sends the crates to crates.io. |
publish-crates.yml runs after the release exists, so a failure at crates.io
leaves the binaries where they are. To send the crates again after such a
failure, start Publish to crates.io by hand from the Actions page and give it
the tag.
A crate on crates.io is permanent. A version that went out cannot go out again with different contents, so step 4 is the step to do carefully.
The order of the crates
cargo publish --workspace reads the graph and sends each crate after the
crates it needs. The order is aphid-core, aphid-agent, aphid-code,
aphid-nostr, aphid-colony, aphid-alate, aphid-ai. Each
crate of the workspace names a version as well as a path in
[workspace.dependencies], because a path alone is enough to build and not
enough to publish.
The configuration of the release
dist-workspace.toml holds the platforms, the installer and the tools that
each runner installs. .github/workflows/release.yml comes from that file, so
no hand edits go in it. After a change:
dist generate
git add dist-workspace.toml .github/workflows/release.yml
To read what a release would hold, without a build and without a tag:
dist plan
To make the installer on this machine, which is how to read what it does:
dist build --artifacts=global
To build the archive of this machine, which takes as long as one runner takes:
dist build --artifacts=local
Each of the three writes to target/distrib.
A newer dist
cargo-dist-version in dist-workspace.toml says which version of dist the
CI uses. To move to a newer one, install it and let it write the file again:
cargo install cargo-dist --locked
dist init
dist generate
Read the difference in release.yml before the commit. That file decides which
runner builds each platform, and a new version of dist can move a build to
another image of the operating system.
A release that must not go out yet
A tag such as v0.2.0-rc.1 makes a pre-release on GitHub. dist marks it as
one, so the address releases/latest/download/... still gives the version
before it, and the installer of a user gives the stable release.
The site
The site is not part of a release. Each push to main that touches docs/,
site/, book-theme/, book.toml or the justfile builds it again and
deploys it, with .github/workflows/pages.yml. To read it first:
just serve
The site is at https://aphid.embornal.com, and the book at /docs/ under it.
Two settings hold that address, and a move to a different one changes both:
baseURLinsite/hugo.toml.- The custom domain of the repository, at Settings, Pages. A workflow that
deploys reads the domain from there, so a
CNAMEfile in the tree does nothing.
The domain also needs one record in DNS, a CNAME of aphid.embornal.com that
gives tncardoso.github.io. The source of the Pages of the repository must be
GitHub Actions.
A site under a path, such as example.com/aphid/, needs more: site-url in
book.toml, and the links of the nav bar in book-theme/index.hbs, which
start at the root of the domain.