Reference for your agent. The precise shape of things — signatures, fields,
rules — written so an agent can read it and act. You do not need to: ask your
agent for what you want in words, and if it needs this page, hand it the URL or
use Copy page. What to ask for, and how to check it, is in Ask;
how to see what changed is in Look inside.
The filesystem is the default
The sandbox belongs to the machine — not to the run, and not to the chat. A file written in one turn is there in the next turn, in the next chat, and tomorrow. Everything the agent produces goes there because there is nowhere else it needs to go./home/daytona. That same filesystem is shared by the parent run and every
sub-agent of the conversation, which is why a sub-agent’s files are readable at
their path the moment it answers — no fetch, no download, no artifact step.
A long-running program keeps its state here too. Write the loop so that starting
it twice is harmless and starting it again after a stop picks up where it left off,
and the state file is what makes that true.
A harness that keeps notes
Nothing about note-keeping is built into the platform. It is a pattern, and the starter harness implements it in about ten lines ofhooks/context.py: a file on
disk, read at the top of every turn and printed into the system prompt.
INDEX.md holds standing rules and
pointers: which files hold what, which past chats hold what, which searches are
worth running. The detail lives in the files it points at, and the agent reads one
when the task touches it. That is the difference between 8000 characters spent every
turn and a megabyte of notes that would never fit.
It is cut, not truncated silently. _read reads limit + 1 characters and, if
the file is longer, appends (Cut off here: INDEX.md is longer than 8000 characters.). An agent that can see it was cut can go and read the rest.
The prose around it says what kind of thing it is. The block above tells the
agent these are the user’s notes rather than instructions from the system, and that
anything said in this conversation overrides them. Without that sentence, a note
from three months ago outranks what the person just said.
The tree does not change inside a run, so the starter reads these files once per
sandbox rather than once per turn — with one exception: a file that was not there
is not cached, because parts of the tree are projected while the run is starting,
and a miss on the first turn is a file that exists on the second.
Writing the notes
The agent writes them, the way it writes any other file. If you want that narrowed to something the model cannot get wrong, it is a tool of half a page:t.store — the harness’s own key-value
A hook is the one place that cannot simply write a file and expect it to matter, because a hook’s decision often has to be the same on a machine it has never run on.t.store is addressed by harness rather than by run, so a note survives
the sandbox, the run and the next publish.
None, which is the first run of every hook that keeps a
note. A value may be at most 256 KB and a harness at most 1000 keys; over
either limit the write is refused and the hook is told which one, because a value
that came back trimmed is a note that lies.
It is a round trip to the API, so ask for it only when you need it.
Inside one run: the window
A conversation that outgrows the model’s context is compacted. The decision is made before a request is sent, from the provider’s own count of the last prompt it ingested — not an estimate, and not a refusal parsed out of an error message, because every vendor words those differently and some send none. A conversation with no measurement yet is left alone; so is one whose budget resolves to zero.context_memory on the agent declaration is where the numbers live:
context_tokens bounds what goes in; max_output_tokens, one line above it in
most declarations, caps what comes out. The two were one name until somebody
noticed nobody could tell which was which.
When the budget is passed, the oldest atomic groups are deleted down to
trim_target_percent and one summary message takes their place. Groups, not
messages: a tool call and its result are never separated. If the summarization call
comes back empty the slice is halved and retried, which needs no error
classification and cannot make the trim drop more than trim_target_percent
already decided.
max_last_images is a different mechanism and runs on every turn, budget or no
budget. Only the newest images survive; the older ones are replaced in place by a
stub naming the file. It has a default rather than being unlimited because
providers refetch every image URL on every request, so a screenshot loop with an
unbounded history turns one turn into hundreds of downloads. A configured 0 is
honored and drops them all.
Saying what a summary must not lose
summarize_prompt is used as rules, not as a conversation. The summarizer is
given You are given a chat log inside <chat> tags. Rules: followed by your text,
with the log itself appended in <chat> tags — so write instructions about what to
keep and what to drop, and not a greeting or a description of the task.
Leaving it out does not turn summarizing off: the platform’s own one-line prompt
runs instead.
The starter puts the same thing in a file so it can be edited without touching
main.py, through the memory.summarize hook:
hooks/memory.py’s keep, where it can see the turn it is deciding about.
Continuing a conversation
A sub-agent’schat_id is an ordinary chat. Keep it and the same sub-agent picks up
where it left off, with everything it already knows:
run_id instead and AgentRun(run_id) reattaches to that run later, even
from a different script.
Choosing
- Something this user always wants — a file under
~/memory, pointed at fromINDEX.md. - Something the work produced — a file where the work is:
~/workspace, or the worktree the sub-agent wrote it in. - Something a hook has to remember across machines —
t.store. - Something every agent of this harness should be told — not memory at all.
That belongs in a prompt file or in
hooks/context.py, where it is versioned and reviewed.

