Context and long tasks

When a coding agent runs out of context mid task

Every coding agent works inside a fixed context window. On a short session you never notice it. On a long refactor it fills with file reads, test output and edit history, and then the session stops or quietly starts forgetting the rule you set twenty minutes ago. The usual advice is to start over and restate the task. That works. It also throws away the reasoning the session spent an hour building.

Disclosure: I build Ordewell, so the second half of this is how my tool approaches the problem rather than a survey. It is free and Apache 2.0. The first half is general: any agent session has a ceiling, and the question of what to do about it applies whether or not you ever install anything.

Why a long task runs out of room

The window holds everything at once: your messages, the model's replies, every file it read, every command it ran, and the before and after of every edit. There is no selective memory. The agent cannot quietly drop the file it no longer needs and keep the function you care about. The buffer just fills.

The fastest filler is not your typing. It is tool output. A single read of a large source file can put tens of thousands of tokens into the window, and three reads plus a search plus a test run can take half the budget before any code is written. That is why a long session feels sharp at the start and foggy near the end.

When the window gets tight the runner either stops or summarizes. A summary written under pressure is lossy. It tends to keep the goal and drop the specifics: the constraint you named, the path you said not to touch, the approach you already tried and abandoned. The agent is then working from a shorter memory of what you said.

What people do inside the session

The standard remedies are real, and worth knowing before you reach for anything structural. Compaction replaces the transcript with a summary. In Claude Code that is /compact, and you can steer it (/compact keep the no-GraphQL rule and the vitest convention) so what survives is what you named rather than what the automatic pass guessed. Restarting with --continue or --resume carries forward a summary, not the full history, so the same loss applies in a cleaner package.

Moving stable rules into a project file the agent reads at startup, such as CLAUDE.md or AGENTS.md, survives both. What it cannot hold is the in-flight state: which approach you abandoned, which files were already read, what the partial refactor looks like right now. Those live in the session, and the session is the thing that is running out of room.

Compaction is a patch. If you find yourself compacting the same task more than once, that is the signal, not the fix: the task is larger than one window can carry.

The structural fix: fewer things per window

The alternative is to stop needing one session to hold the whole job. Break the work into tasks small enough that each can finish inside a single window, run each task in a fresh session, and hand forward only what the next task actually needs. This is decomposition before the run starts. It is the same idea as delegating a noisy search to a subagent that has its own window, applied to the plan as a whole rather than to one step.

Once the work is split, the plan becomes the artifact that outlives any one session. When a session dies, its task is a line on a list rather than a conversation you have to reconstruct. You can read it, change it, and run it again. The reasoning that a compaction would have summarized away is written down somewhere it can be edited.

How Ordewell does this

The planner reads the repository read-only, asks about whatever is vague, and writes an ordered plan with dependencies. Then each task starts its own fresh coding-agent session, in its own git worktree, and is given the results of the tasks it depends on. Independent tasks run concurrently, three at a time by default.

So what sits in any one session is a single task plus the outcome of its prerequisites, not the entire goal and every file read along the way. A task that fails is a failed task on the plan rather than a session that has quietly degraded. Its dependents are blocked, you can see which ones, and you can edit the task and run it again without replaying the run.

The plan itself is stored on disk, so it survives closing the terminal and coming back later, and what each task is handed is bounded on purpose. Those two pages go into the detail.

A real run

From the command line, the flow is three commands:

export AI_PROVIDER=claude-code
ordewell plan --goal "Add rate limiting to the public API"
ordewell run
ordewell handoff review

Between the plan and the run, the tasks are yours to change. This is how you keep a single task from growing past what one window can hold, by splitting it or by moving work it does not need to do:

ordewell task-model 3 sonnet
ordewell task-deps 3 1,2

That second command makes task three wait for tasks one and two, which also decides what task three's session is handed: the outcome of those two tasks, and nothing from the rest of the plan.

Honest limits

Source and design notes

The mechanics above are in the repository rather than recalled: the README (each task starts a fresh session given the results of the tasks it depends on, three at a time by default), where the plan lives when the terminal closes, what each task is handed, and github.com/ordewell/ordewell.