Context between tasks

What each task in a plan is handed, and what it is not

A plan runs several sessions, one after another or several at once. None of them shares memory with the others. Each one starts with an empty context window, so the only thing that carries knowledge forward is what the orchestrator decides to hand over. That decision is worth reading, because it is the difference between a task that continues the work and a task that rediscovers it.

Disclosure: I build Ordewell, so what follows is how it does this rather than a survey. It is free and Apache 2.0. The shape of the problem is the same in any tool that runs more than one session, so the reasoning is useful even if you never install it.

The problem: an agent knows nothing about the session before it

It is tempting to imagine the plan as a conversation. It is not. Task two is a fresh process with a fresh prompt. Whatever task one learned about your codebase, whichever approach it abandoned, whatever it named a function, all of that is gone unless it was written into the working tree or handed back explicitly.

The lazy fix is to paste everything: the whole transcript of every previous task, or the planning conversation that produced the plan. That fails in a way you will not notice until it does damage. The transcript of an interactive runner contains the prompt it was given, and that prompt contains the token the watcher looks for to mark the task complete. Hand a predecessor's output to a dependent as raw text and you can settle the dependent on the predecessor's evidence, the moment the session starts.

What a task is actually handed

Three blocks, assembled in a fixed order and placed above the task's own prompt. First the plan map, then the outputs of the tasks it directly depends on, then the task itself. All three are composed from the plan state, not from anything a model wrote about itself.

The plan map is a numbered list of the tasks in the plan with their current status, the runner assigned to the current one if a runner is set, and a single marker on the line the session is working:

prompt context, renderPlanMap
## Plan map
 1. [done   ] Add the settings model
 2. [NOW    ] Wire the flag into the CLI   ← you are here
 3. [next   ] Update the reference docs

(4 tasks omitted from this view)

Statuses are read off the stored task state, so the map is a report of what already happened rather than a plan the model is free to reinterpret. The header states the one rule that matters: this list is for context, do only the task marked as current, and do not preempt the tasks that follow.

The outputs of direct dependencies, and nobody else's

Below the map comes a block per dependency the task declared, in plan order. Each block is short by construction: the review note recorded for that task, and the tail of its captured output.

Two details matter. The first is that the tail is a tail. It is capped at 500 characters, taken from the end, which is where a runner's final summary and its completion marker sit. The second is that only direct dependencies contribute. A task that depends on A, which depends on B, gets A's output and not B's. The dependency graph the planner emitted is also the context boundary, so the amount a task has to read grows with the shape of the graph rather than with the length of the run.

Two tasks running side by side in the same wave see nothing of each other. That is deliberate: they touched different files by design, and giving each one the other's in-progress output would invite it to edit the other's work.

A quoted completion marker would be evidence for the wrong task

Interactive runners echo the prompt into the terminal, and the watcher reads that terminal output to decide whether a task is done. So any text carried over from a predecessor has its markers broken before it is inserted. The token is rewritten into a hyphenated form that no scanner accepts, and the identifier inside it is dropped entirely.

The hyphen rather than a space is not cosmetic. The watcher also reads a whitespace flattened view of the terminal, where a spaced out token would be rejoined into a live one. And the dropped identifier matters for a second reason: transcripts are bound to a task by that identifier, so if a dependent quoted a predecessor's full token, the dependent's own transcript would later answer questions about the predecessor.

The same reasoning applies to the completion marker instruction itself. It is given to the model in two halves, with the assembled token never appearing in the prompt, so that echoing the prompt cannot complete the task on the first second of the run.

A resumed session gets the files and a warning

Continuing a task is a different case. The session still holds its original prompt, so the prompt is not resent. What changed is the working directory: it was recreated from the integration branch, so an earlier attempt's edits are present only if that work actually landed. The continuation reminder says exactly that, and tells the model to check the files before relying on them.

That is the whole protocol. No transcript pasted forward, no summary generated by a model, no memory file that quietly drifts out of date. Task state and captured output are the interface, and everything a task knows about its predecessors came through it.

Honest limits

  • The tail can miss the useful part. A test failure printed near the top of a long run is not in the last 500 characters. The review note is meant to carry that, and it is only as good as the attempt that wrote it.
  • Indirect dependencies are invisible. If B's result mattered to a task that only depends on A, it reaches that task only if A mentioned it. Declaring the dependency is the fix.
  • The map is a window. Past 30 entries the list is trimmed around the current task and the footer counts what was left out. A task cannot see the whole plan, by design.
  • It is text, not an enforcement. A model can ignore the instruction to stay in scope. The boundary keeps it informed, not obedient.
  • Defusing is a convention. It protects against a marker being read out of carried context. It is not a security boundary, since a runner could print a valid looking token on its own.
  • Captured output is bounded for a reason. The cap keeps a predecessor's noise from crowding out the task's own prompt. Raise it and you trade attention for detail.

Source and design notes

Everything above is in the repository rather than recalled: promptAugment.ts (the composition order, the 500 character tail, the 30 entry window and its look back, the marker defusing), the Task model, and github.com/ordewell/ordewell.