Tech Reads
AI Engineering9 min read

Context Window Management in Long-Running Enterprise Workflows

A 12-step procurement workflow runs cleanly through step 6. At step 7 it starts approving purchase orders for vendors it rejected at step 3. The context window is full, and the agent can no longer see its own earlier decisions. Hardly anyone designs for this failure until it causes a production incident.

Context window limits are documented. Every model provider publishes them. Claude 3.5 Sonnet: 200K tokens. GPT-4o: 128K tokens. These numbers look enormous in isolation. In a long-running agentic workflow with tool calls, retrieved documents, accumulated conversation history, and a substantial system prompt, they evaporate faster than anyone expects.

It is easy to ship a multi-step agent without thinking about any of this, and then spend months retrofitting context management into a system that should have had it from the start. What follows is how we think about it.

Treat the context window as a budget

It is tempting to treat the context window as a large buffer that the agent fills up as it works. That model is wrong in one important way: a buffer is something you drain as you consume it. The context window works like a sliding window that silently truncates what falls off the left edge when it fills up. The agent does not know what got dropped. It cannot tell you. It keeps running with whatever portion of its history fits.

The better mental model: treat the context window as a budget with line items. System prompt: fixed cost, paid on every call. Conversation history: grows with each round-trip. Tool call inputs and outputs: highly variable, often large. Retrieved documents: can be enormous if you are not careful. Current task state: the thing you actually need the model to reason about.

When you model it as a budget, you start asking the right questions: what does this workflow cost per step? At what step does the budget run out? What gets evicted when it does, and what are the consequences?

200Ktokens in Claude 3.5 Sonnet's context window. A long agentic workflow with tool calls and retrieved documents can still fill it.

What actually happens when you exceed the limit

There are two failure modes, and they behave very differently.

The first is a hard error: the API call fails with a context length exceeded error, the workflow crashes, and you know exactly what happened. This is the good failure mode. It is recoverable if you have designed for it. Without a recovery path, it is still a production incident, but at least it is a visible one.

The second is silent truncation: the provider or the agent framework drops the oldest tokens to fit within the limit, and the model continues running with a degraded view of its history. This is the dangerous failure mode. The workflow does not crash. The agent keeps producing outputs. Those outputs are based on an incomplete picture of what happened in steps 1 through 6, and there is no error signal telling you any of this.

Watch out

Silent truncation is not a bug. Some provider configurations and many agent frameworks do it by design. Check whether your stack truncates silently or errors loudly, and don't assume you will get a hard error when you run over.

The procurement workflow at the top of this article is silent truncation at work. The vendor rejections from step 3 were in the conversation history. When the window filled, they fell off. The model didn't hallucinate anything. It re-approved vendors it had once rejected, because it could no longer see the rejection.

Summarization as a compression strategy

The most commonly recommended solution is progressive summarization: periodically compress older conversation history into a concise summary, replacing the full history with the summary in the context. This works and we use it. It also has real costs that people understate.

Summarization loses information. That is the point of it, and what gets lost is unpredictable. A detail that seemed minor at step 2 may be critical at step 10. Suppose a supplier's minimum order quantity, set at step 2, gets summarized as "initial supplier constraints were discussed." At step 9 the agent places an order below the minimum, because the number itself is gone.

The practical rule: summarize narrative context aggressively. Never summarize structured decisions, commitments or constraint values. Those go into the structured state store (see the checkpoint pattern section below), not the context window.

What to keep and what to evict

Not everything in the context window has equal value at every step. "Drop the oldest tokens" is what happens by default when you hit the limit. A good eviction policy is one you design around what the model needs in order to reason correctly at each step of the workflow.

Things that should almost never be evicted from the active context: the system prompt, the current step's input and task definition, structured decisions made in prior steps (the actual values, not the narrative), and any constraint or rule established earlier that governs later steps.

Things that compress well: intermediate reasoning traces, tool call chains where the final output is what matters and the intermediate steps are scaffolding, retrieved documents that informed a decision already made, and status updates that have since been superseded.

We build a context manifest — a structured description of what is currently in the window, what has been summarized, and what has been moved to the state store. The orchestrator consults this manifest before each step to decide whether the current context is sufficient for the next task. If it is not, it reconstructs the relevant pieces before calling the model.

Stateful memory patterns: pulling from the state store

The right architecture separates ephemeral context (what the model is reasoning about right now) from durable state (what the workflow has decided and written to external systems). The durable state lives in a database, not in the context window.

This connects directly to the checkpoint pattern we use for agentic workflow state management. Each step writes its structured outputs (decisions, extracted values, confirmations of external writes) to the checkpoint store before the next step starts. The context window carries the narrative. The state store carries the facts.

When context pressure builds, the orchestrator can reconstruct the critical facts from the state store and inject them back into the context as a compact structured summary. "Prior decisions: Vendor A rejected (lead time: 21 days, limit: 14 days). Vendor B selected. PO draft #PO-2024-1147 created. Approval threshold: $50,000. Current total: $43,200." That is 40 tokens conveying what might have been 8,000 tokens of conversation history.

Our take

The checkpoint store and the context window serve different purposes. The checkpoint store is for the system's recovery: it records what was decided and what was written. The context window is for the model's reasoning. Never use one as a substitute for the other.

Measuring and monitoring context pressure in production

Track context utilization in every agentic workflow. Before each model call, log the token count by category (system prompt, history, tool outputs, injected state, current task) against the model's limit. The thresholds we use: alert at 75% utilization, start active compression at 85%, and fail the step to human review at 95%.

Token counting is not free. The Anthropic and OpenAI tokenizers add overhead to every request if you call them client-side. We use approximation (4 chars ≈ 1 token for English text) for monitoring and reserve exact counting for the compression decision. Accuracy matters most at the boundary, not throughout.

Workflows fail silently from context overflow when nobody measured how much context a realistic 12-step run consumes. Measure it before you ship, and design the eviction policy before you hit the limit.

Share

Related reading