
The AI Agent Context Collapse: Why More Documentation Makes Your Coding Agent Dumber — and How to Fix It
More context doesn't mean better AI coding agents. Explore the context collapse problem, the attention dilution mechanics behind it, and proven architectural patterns to fix it.
You've added every doc, every comment, every runbook to your AI coding agent's context window. You've maxed out your RAG pipeline with comprehensive knowledge bases. And somehow, your agent is writing worse code than the one that only had the README.
This isn't a failure of model intelligence. It's a structural inevitability called context collapse — the degradation of output quality that occurs when an LLM's effective context is overloaded with redundant, conflicting, or irrelevant information. Understanding why this happens, and how to architect against it, is the difference between a coding agent that's genuinely useful and one that confidently produces garbage.
This deep-dive explores the mechanics of context collapse in AI coding agents, the attention-layer phenomena that drive it, and the architectural patterns — from selective retrieval to hierarchical context windows — that fix it.
Table of Contents
- 1. What Is Context Collapse?
- 2. The Attention Dilution Problem
- 3. Why Coding Agents Are Especially Vulnerable
- 4. The Retrieval Quality Spectrum
- 5. Fix #1: Tiered Context Architecture
- 6. Fix #2: Retrieval with Relevance Filtering
- 7. Fix #3: Summarization Before Injection
- 8. Fix #4: Context Budgeting and Token Accounting
- 9. Fix #5: Agent-Orchestrated Context Assembly
- 10. Measuring Context Collapse in Your Pipeline
- 11. Frequently Asked Questions
1. What Is Context Collapse?
Context collapse is the measurable degradation of an LLM's output quality as its input context grows beyond a useful threshold. It manifests in several observable ways:
- Lost-in-the-middle failure: The model attends to the beginning and end of the context but ignores critical information buried in the middle. Research has consistently shown that LLMs perform worst on information located at the center of long contexts.
- Conflicting signal confusion: When the context contains contradictory information — say, two different API signatures for the same function from different documentation versions — the model doesn't flag the ambiguity. It picks one, sometimes randomly.
- Attention dilution: As the context grows, the model's attention weights spread thinner across more tokens, reducing the effective signal-to-noise ratio for any single piece of information.
- Style contamination: Boilerplate, examples, and template code from documentation bleed into the model's output, causing it to generate documentation-style prose instead of production code.
The critical insight is that context collapse is non-linear. Going from 0 to 50% context utilization might improve performance. Going from 50% to 75% might help slightly. But pushing past 80% utilization often causes a sharp cliff — not a gradual slope — in output quality.
This matters enormously for coding agents because they operate in a fundamentally different regime than chat assistants. A chatbot answering "What's the weather?" needs very little context. A coding agent generating a module that must conform to your codebase's conventions, use your internal libraries correctly, and avoid your known anti-patterns needs precise context — not maximal context.
2. The Attention Dilution Problem
To understand context collapse, you need to understand what happens inside the transformer's attention mechanism when the context window fills up.
The Attention Weight Distribution
In a transformer, each token's representation is computed by attending to all other tokens in the context. The attention weight between any two tokens is proportional to the dot product of their query and key vectors, normalized by the square root of the embedding dimension, and softmaxed across all keys.
# Simplified attention computation
import torch
import torch.nn.functional as F
def attention(q, k, v):
"""
q: Query tensor [batch, heads, seq_len, d_k]
k: Key tensor [batch, heads, seq_len, d_k]
v: Value tensor [batch, heads, seq_len, d_v]
"""
d_k = q.size(-1)
# Raw attention scores
scores = torch.matmul(q, k.transpose(-2, -1)) / (d_k ** 0.5)
# Softmax normalizes across ALL keys in the context
weights = F.softmax(scores, dim=-1) # [batch, heads, seq_len, seq_len]
output = torch.matmul(weights, v)
return output, weights
The softmax normalization is the key mechanism. When your context has 1,000 tokens, each token's attention weight averages around 0.1%. When it has 10,000 tokens, each averages 0.01%. The model doesn't get "more attention budget" — it gets the same budget spread across more candidates.
The Practical Consequence
This means that if your RAG pipeline stuffs 40 relevant code snippets into a 128K context window, the model's attention on each snippet is roughly equivalent to what a 4K-context model would give a single snippet. You've spent your context budget buying coverage at the expense of depth.
For coding tasks, depth matters more. The model needs to deeply understand one API's signature, its constraints, and its usage patterns — not shallowly scan forty APIs.
The Middle-Position Penalty
Beyond dilution, there's a positional effect. Studies on long-context models have shown a consistent U-shaped performance curve: models attend well to the beginning (primacy) and end (recency) of the context, but poorly to the middle. This is partly architectural — positional encodings in most models encode position information that makes middle positions less distinguishable — and partly learned, as models trained on shorter contexts develop positional biases.
For a coding agent, this means the most important context (your system prompt, your conventions, your critical constraints) should be placed at the beginning or end of the context window, never buried in the middle of retrieved documentation.
3. Why Coding Agents Are Especially Vulnerable
Coding agents face a unique combination of pressures that make context collapse particularly damaging:
Code Is Denser Than Prose
Natural language has high redundancy — synonyms, filler words, restatements. Code is dense: every token carries meaning. A 200-token code snippet packs more semantic information than a 200-token prose paragraph. This means code contexts saturate the model's understanding capacity faster than you'd expect based on raw token count.
Conflicts Are Silent
When a chat assistant encounters conflicting information, it can hedge: "There are different opinions on this..." But a coding agent must produce executable code. If the context contains two contradictory patterns — say, one doc says use async/await and another shows callback-based calls — the agent must pick one. It usually picks the one that appears more recently or with more surrounding text, not the one that's actually correct for your codebase.
Hallucination Surface Area
Every piece of documentation in the context is a potential source of hallucination. If the docs describe a function signature that's been deprecated, the agent will confidently generate code using the deprecated signature. More documentation means more stale information means more hallucination surface area.
The Documentation Paradox
Documentation is written for humans. Human readers can skim, skip irrelevant sections, and focus on what matters. LLMs can't skim — they process every token equally (in terms of computational cost), and their attention mechanism doesn't have a "skip this paragraph" capability. Documentation is optimized for human comprehension, not machine attention.
4. The Retrieval Quality Spectrum
Not all context is created equal. When building a coding agent, you need to think about the quality of your retrieval, not just its quantity. Here's a spectrum:
| Quality Tier | Example | Effect on Agent | Token Efficiency |
|---|---|---|---|
| Signal | The exact function signature the agent needs | +30% correctness | 100% |
| Context | Related code patterns, conventions | +15% correctness | 60% |
| Noise | Full documentation file, unrelated modules | -10% correctness | 15% |
| Poison | Contradictory patterns, deprecated APIs | -40% correctness | -20% |
Most RAG pipelines for coding agents are designed to maximize retrieval volume, pulling in everything that's semantically similar. But the difference between a good agent and a great agent often lies in the gap between Signal and Context — and the ability to keep Noise and Poison out.
The problem compounds when you consider that retrieval is probabilistic. A vector search with top_k=10 doesn't return the 10 most useful results — it returns the 10 results with the highest cosine similarity, which may include stale documentation, irrelevant examples, and conflicting information.
5. Fix #1: Tiered Context Architecture
The most impactful structural fix for context collapse is to abandon the flat context window and adopt a tiered architecture where different information lives at different levels of the agent's context.
The Three Tiers
Tier 1 — System Context (Always Present, ~2K tokens)
This is the agent's "personality" and operational constraints. It includes:
- Role definition and output format requirements
- Code style conventions (naming, formatting, import order)
- Hard constraints ("Never use eval()", "All DB queries must use parameterized queries")
- The single most important architectural decision the agent must know about
SYSTEM_CONTEXT = """
You are a senior software engineer specializing in {language}.
## Hard Constraints
- All database queries MUST use parameterized queries
- Maximum function length: 50 lines
- No global mutable state
- Use {framework}'s built-in error handling, never bare except clauses
## Code Style
- snake_case for functions and variables
- PascalCase for classes
- All public functions must have docstrings
- Import order: stdlib, third-party, local
## Current Task Context
The user is working in the {module_name} module of the {project_name} project.
The project uses {framework} {version}.
"""
Tier 2 — Retrieved Context (Dynamic, ~4-8K tokens)
This is the RAG output, but curated. Instead of dumping all retrieved chunks, you select the most relevant ones and truncate the rest. The key principle: quality over quantity.
Tier 3 — Scratch Context (Per-Query, ~2-4K tokens)
This is the user's specific request plus any conversation history relevant to the current query. It's the smallest tier but the most task-specific.
Implementation Pattern
class TieredContextBuilder:
"""
Builds a context window with explicit tier separation.
Tier 1 (system) is always at the start.
Tier 2 (retrieved) is in the middle, sorted by relevance.
Tier 3 (task) is at the end, where recency bias helps it.
"""
def __init__(self, max_tokens=16384):
self.max_tokens = max_tokens
self.token_budget = {
"tier1_system": 2048,
"tier2_retrieved": 6144, # ~37% of budget
"tier3_task": 4096, # ~25% of budget
# Remaining 38% is reserved for model output + safety margin
}
def build(self, system_context: str, retrieved_chunks: list, task: str) -> str:
# Tier 1: System context (always first — primacy position)
tier1 = system_context[:self.token_budget["tier1_system"]]
# Tier 2: Retrieved context, filtered and ranked
tier2 = self._filter_retrieved(retrieved_chunks)
# Tier 3: Task context (always last — recency position)
tier3 = task[:self.token_budget["tier3_task"]]
# Assemble with explicit separators
context = (
f"<system_context>\n{tier1}\n</system_context>\n\n"
f"<reference_material>\n{tier2}\n</reference_material>\n\n"
f"<task>\n{tier3}\n</task>"
)
return context
def _filter_retrieved(self, chunks: list) -> str:
"""Sort by relevance score, keep top-N, truncate each."""
budget = self.token_budget["tier2_retrieved"]
sorted_chunks = sorted(chunks, key=lambda c: c.score, reverse=True)
selected = []
current_tokens = 0
for chunk in sorted_chunks:
chunk_tokens = len(chunk.text) // 4 # rough estimate
if current_tokens + chunk_tokens > budget:
break
selected.append(chunk.text)
current_tokens += chunk_tokens
return "\n\n---\n\n".join(selected)
The key insight: by using explicit XML-like tags (<system_context>, <reference_material>, <task>), you give the model structural cues about what each section is for. Research shows that labeled sections improve attention allocation even when the model has no special training for those tags.
6. Fix #2: Retrieval with Relevance Filtering
The default RAG pattern — vector search with top_k and no filtering — is actively harmful for coding agents. Here's how to fix it.
Pre-Filtering by Recency and Source Authority
Before semantic search, filter your knowledge base by:
- Recency: Prefer documentation from the last 6 months. A function signature from 2 years ago is more likely to be wrong than right.
- Source authority: Code from your actual codebase should rank above external documentation. External docs should rank above blog posts. Blog posts should rank above Stack Overflow answers (which are often outdated).
class AuthorityWeightedRetriever:
"""
Retrieves code context with authority and recency weighting.
"""
SOURCE_WEIGHTS = {
"codebase": 1.0, # Your actual code
"internal_docs": 0.85, # Internal documentation
"official_docs": 0.7, # Official framework docs
"blog_post": 0.4, # Community content
"stackoverflow": 0.3, # Often outdated
}
RECENCY_DECAY = 0.01 # 1% weight reduction per month old
def retrieve(self, query: str, top_k: int = 5) -> list:
# Step 1: Raw vector search (oversample)
candidates = self.vector_store.search(query, top_k=top_k * 4)
# Step 2: Score adjustment
scored = []
for chunk in candidates:
# Authority score
authority = self.SOURCE_WEIGHTS.get(chunk.source_type, 0.5)
# Recency score
months_old = (datetime.now() - chunk.last_updated).days / 30
recency = max(0.1, 1.0 - (months_old * self.RECENCY_DECAY))
# Combined score
adjusted_score = chunk.similarity_score * authority * recency
chunk.adjusted_score = adjusted_score
scored.append(chunk)
# Step 3: Return top-k by adjusted score
scored.sort(key=lambda c: c.adjusted_score, reverse=True)
return scored[:top_k]
The Contradiction Detection Filter
This is the most advanced filtering technique. Before injecting retrieved chunks into context, check for contradictions:
class ContradictionFilter:
"""
Detects and resolves contradictions between retrieved chunks.
If two chunks describe the same API differently, keep only the authoritative one.
"""
def __init__(self, llm_client):
self.llm = llm_client
async def filter(self, chunks: list) -> list:
if len(chunks) < 2:
return chunks
# Group chunks by the API/function they describe
groups = self._group_by_entity(chunks)
filtered = []
for entity, group_chunks in groups.items():
if len(group_chunks) == 1:
filtered.append(group_chunks[0])
continue
# Check for contradictions using a lightweight model
contradiction_result = await self._check_contradiction(group_chunks)
if contradiction_result.has_conflict:
# Keep only the most authoritative chunk
best = max(group_chunks, key=lambda c: c.adjusted_score)
filtered.append(best)
else:
# No contradiction — keep the most relevant
best = max(group_chunks, key=lambda c: c.similarity_score)
filtered.append(best)
return filtered
async def _check_contradiction(self, chunks: list) -> ContradictionResult:
"""Use a small, fast model to check for contradictions."""
prompt = (
"Are the following code documentation snippets contradictory? "
"Do they describe the same function/API with different signatures, "
"return types, or behavior?\n\n"
f"{chr(10).join(c.text for c in chunks)}\n\n"
"Respond with JSON: {\"has_conflict\": bool, \"reason\": str}"
)
response = await self.llm.complete(prompt, max_tokens=100)
return ContradictionResult.from_json(response)
The cost of this filter is one additional LLM call per retrieval batch. For a coding agent that makes 5-10 retrievals per session, this adds maybe 200-500ms of latency — negligible compared to the correctness improvement.
7. Fix #3: Summarization Before Injection
Raw documentation is optimized for human reading, not machine attention. Summarizing retrieved chunks before injection can dramatically improve the signal-to-noise ratio.
The Summarize-Then-Inject Pattern
class ContextCompressor:
"""
Compresses retrieved documentation into high-signal summaries
optimized for LLM consumption.
"""
COMPRESSION_PROMPT = """\
Extract the following from this code documentation snippet. Be terse.
No prose, no explanations, just facts:
1. Function/API name and signature
2. Return type
3. Key parameters and their types
4. Any constraints or gotchas
5. One minimal usage example (code only)
Do NOT include: introductions, background, alternatives, history, or prose.
SNIPPET:
{snippet}
"""
async def compress(self, chunks: list, target_ratio: float = 0.3) -> str:
"""
Compress chunks to ~target_ratio of their original size.
"""
tasks = [
self.llm.complete(
self.COMPRESSION_PROMPT.format(snippet=chunk.text),
max_tokens=chunk.text_length * target_ratio // 4
)
for chunk in chunks
]
summaries = await asyncio.gather(*tasks)
return "\n\n".join(summaries)
When Summarization Helps vs. Hurts
Summarization is not always beneficial. Here's the decision matrix:
| Chunk Type | Summarize? | Reason |
|---|---|---|
| API documentation | Yes | High redundancy, compresses well |
| Code examples | No | Code is already dense |
| Architecture descriptions | Yes | Prose-heavy, compresses well |
| Error messages | No | Must be exact |
| Configuration files | No | Must be exact |
| Design docs | Yes | Often verbose |
A practical heuristic: if a chunk is more than 60% prose, summarize it. If it's more than 60% code, keep it verbatim.
8. Fix #4: Context Budgeting and Token Accounting
The most disciplined approach to context management is explicit token budgeting — treating the context window as a finite resource with explicit allocation rules.
The Token Budget Framework
class TokenBudget:
"""
Explicit token budget management for coding agents.
Every piece of context must be justified by its value-per-token ratio.
"""
def __init__(self, model_max_context: int = 128000):
self.model_max = model_max_context
# Reserve 40% for output (code generation needs room)
# Reserve 10% for safety margin
self.available = int(model_max_context * 0.50)
self.allocations = {
"system_prompt": 0.10, # 10% — role, constraints, style
"conversation": 0.15, # 15% — recent conversation history
"retrieved_context": 0.45, # 45% — RAG results (the main pool)
"task_description": 0.10, # 10% — current user request
"output_reserve": 0.20, # 20% — model's generated output
}
def allocate(self, tier: str) -> int:
return int(self.available * self.allocations[tier])
def report(self, used: dict) -> dict:
"""Generate a budget report for monitoring."""
report = {}
for tier, budget in self.allocations.items():
allocated = self.allocate(tier)
spent = used.get(tier, 0)
report[tier] = {
"allocated_tokens": allocated,
"used_tokens": spent,
"utilization": round(spent / allocated * 100, 1) if allocated > 0 else 0,
"over_budget": spent > allocated
}
return report
# Usage example
budget = TokenBudget(model_max_context=128000)
# Before making the LLM call, verify you're within budget
usage = {
"system_prompt": len(system_prompt) // 4,
"conversation": len(conversation_history) // 4,
"retrieved_context": len(retrieved_chunks_text) // 4,
"task_description": len(user_request) // 4,
}
report = budget.report(usage)
if any(r["over_budget"] for r in report.values()):
# Trim retrieved context first — it's the most expendable
retrieved_text = trim_to_budget(retrieved_chunks_text, budget.allocate("retrieved_context"))
The Value-Per-Token Metric
The key insight from token budgeting is to evaluate every piece of context by its value-per-token ratio. A 200-token code snippet that directly answers the agent's question has a much higher value-per-token than a 2000-token documentation page that provides background context.
def value_per_token(chunk: RetrievedChunk, query_relevance: float) -> float:
"""
Estimate the value-per-token of a retrieved chunk.
Higher is better. Use this to rank chunks for inclusion.
"""
token_count = len(chunk.text) // 4
# Base value from retrieval relevance
base_value = query_relevance
# Penalty for redundancy (how much of this chunk overlaps with already-included context)
redundancy_penalty = chunk.overlap_with_existing / token_count
# Bonus for specificity (code chunks are denser than prose)
code_density = chunk.code_line_count / max(1, len(chunk.text.split()))
return (base_value * (1 - redundancy_penalty) * (1 + code_density * 0.5)) / token_count
9. Fix #5: Agent-Orchestrated Context Assembly
The most sophisticated approach is to let the agent itself decide what context it needs — a form of agentic retrieval where the agent issues targeted queries for specific information rather than receiving a pre-assembled context dump.
The Self-Query Pattern
class AgenticContextAssembler:
"""
The agent decides what context it needs, rather than
receiving a pre-assembled context dump.
This is fundamentally different from standard RAG:
- Standard RAG: retrieve everything similar → stuff into context
- Agentic: agent identifies gaps → retrieves specific info → decides if more needed
"""
def __init__(self, llm_client, vector_store, knowledge_base):
self.llm = llm_client
self.vector_store = vector_store
self.kb = knowledge_base
async def assemble_context(self, task: str, max_iterations: int = 3) -> dict:
"""
Iteratively build context by having the agent identify what it needs.
"""
context = {"base": task, "retrieved": [], "queries_made": []}
for iteration in range(max_iterations):
# Step 1: Agent identifies what information it needs
assessment = await self._assess_needs(task, context)
if assessment.is_sufficient:
break
# Step 2: Agent formulates specific queries
for query in assessment.needed_queries:
results = await self.vector_store.search(query, top_k=3)
# Step 3: Agent evaluates each result
for result in results:
is_useful = await self._evaluate_result(result, task, context)
if is_useful:
context["retrieved"].append(result)
context["queries_made"].append(query)
return context
async def _assess_needs(self, task: str, current_context: dict) -> NeedsAssessment:
"""
Have the agent evaluate whether it has enough context.
"""
prompt = f"""\
You are working on this task:
{task}
Current context you have:
{self._format_context(current_context)}
Questions:
1. Do you have enough information to complete this task accurately?
2. If not, what SPECIFIC information do you need? (Be precise — name the exact functions, APIs, or patterns you need)
3. How would you search for that information?
Respond as JSON:
{{
"is_sufficient": bool,
"missing_info": [str],
"needed_queries": [str],
"confidence": float
}}
"""
response = await self.llm.complete(prompt, max_tokens=500)
return NeedsAssessment.from_json(response)
async def _evaluate_result(self, result, task: str, context: dict) -> bool:
"""
Quick evaluation: is this result actually useful?
"""
prompt = f"""\
Task: {task}
Candidate result:
{result.text[:500]}
Is this result directly useful for the task? (Yes/No)
If no, what's missing?
"""
response = await self.llm.complete(prompt, max_tokens=50)
return "Yes" in response
Why This Works
This pattern works because it inverts the traditional RAG assumption. Standard RAG assumes the retriever knows what the agent needs. Agentic assembly assumes the agent knows what it needs — and the agent is better at this because it has the task context and can reason about information gaps.
The trade-off is latency. Agentic assembly makes 3-10 LLM calls during context construction, adding 5-30 seconds of latency. For interactive coding agents, this may be acceptable if the quality improvement is significant. For batch processing, it's clearly worth it.
10. Measuring Context Collapse in Your Pipeline
You can't fix what you can't measure. Here's how to instrument your pipeline to detect context collapse in real time.
The Context Health Dashboard
class ContextHealthMonitor:
"""
Monitors context quality metrics to detect collapse.
"""
def __init__(self):
self.metrics = {
"context_utilization": 0.0,
"retrieval_precision": 0.0,
"contradiction_rate": 0.0,
"output_hallucination_rate": 0.0,
}
def record_session(self, session: dict):
"""Record metrics for a completed agent session."""
# Context utilization: what % of the window was actually used
self.metrics["context_utilization"] = (
session["tokens_used"] / session["max_tokens"]
)
# Retrieval precision: what % of retrieved chunks were cited in output
if session["retrieved_chunks"]:
cited = session["chunks_cited_in_output"]
self.metrics["retrieval_precision"] = cited / len(session["retrieved_chunks"])
# Contradiction rate: how often we detected conflicting chunks
self.metrics["contradiction_rate"] = (
session["contradictions_found"] / max(1, session["retrieval_batches"])
)
# Output hallucination rate: code that references non-existent APIs
if session["output_lines"]:
hallucinated = session["hallucinated_references"]
self.metrics["output_hallucination_rate"] = (
hallucinated / max(1, session["output_lines"])
)
def health_score(self) -> float:
"""
Composite health score (0-100). Lower is better for some metrics.
"""
m = self.metrics
# Optimal utilization: 40-70% (not too sparse, not too full)
utilization_penalty = abs(m["context_utilization"] - 0.55) * 100
# Higher precision is better
precision_score = m["retrieval_precision"] * 40
# Lower contradiction rate is better
contradiction_penalty = m["contradiction_rate"] * 30
# Lower hallucination rate is better
hallucination_penalty = m["output_hallucination_rate"] * 30
score = 100 - utilization_penalty - contradiction_penalty - hallucination_penalty
return max(0, min(100, score))
def alert_if_degrading(self, threshold: float = 60.0):
"""Alert when health score drops below threshold."""
score = self.health_score()
if score < threshold:
return {
"status": "degraded",
"score": score,
"metrics": self.metrics,
"recommendation": self._recommend_fix()
}
return {"status": "healthy", "score": score}
def _recommend_fix(self) -> str:
m = self.metrics
if m["context_utilization"] > 0.85:
return "Context is over-utilized. Reduce retrieval top_k or add summarization."
if m["contradiction_rate"] > 0.2:
return "High contradiction rate. Add contradiction detection filter."
if m["retrieval_precision"] < 0.3:
return "Low retrieval precision. Improve embedding model or add re-ranking."
if m["output_hallucination_rate"] > 0.1:
return "High hallucination rate. Reduce context volume and increase specificity."
return "Investigate context quality. Consider tiered architecture."
Key Metrics to Track
| Metric | Healthy Range | Collapse Signal |
|---|---|---|
| Context utilization | 40-70% | >85% or <20% |
| Retrieval precision | >60% | <30% |
| Contradiction rate | <5% | >20% |
| Hallucination rate | <3% | >10% |
| Output correctness | >80% | <50% |
If you're building a coding agent today, the single most important thing you can do is reduce your context volume by 40% and measure whether output quality improves. If it does, you've found your collapse threshold. Then work backward from there, optimizing your retrieval quality to fill the gap with higher-signal context.
11. Frequently Asked Questions
How do I know if my coding agent is suffering from context collapse?
The most reliable signal is a correlation between context size and output quality. Run a controlled experiment: take 20 representative tasks, run them with your full RAG pipeline, then run them with top_k reduced by 50%. If the smaller-context runs produce equal or better output, you're experiencing context collapse. The second signal is a high retrieval precision score — if only 20-30% of your retrieved chunks end up being cited or referenced in the output, the other 70-80% is noise.
Should I use a larger context window model to solve this?
No — this is the most common mistake. Larger context windows make context collapse easier to trigger, not harder. With a 128K window, you're tempted to stuff in everything. The right approach is to use a smaller effective context regardless of the model's maximum. A 32K window with carefully curated context will outperform a 128K window stuffed with everything. Consider using a model with a smaller context window (32K or 64K) as a forcing function to keep your retrieval lean.
How does this compare to the "lost in the middle" problem?
The lost-in-the-middle problem is a specific mechanism of context collapse — the positional bias where middle-of-context information gets less attention. Context collapse is the broader phenomenon that includes attention dilution, contradiction confusion, style contamination, and the lost-in-the-middle effect. Fixing lost-in-the-middle (by placing critical info at the start/end) is necessary but not sufficient to prevent context collapse. You also need retrieval quality filtering, contradiction detection, and token budgeting.
For more on AI agent architecture patterns and context management strategies, see Tamiz's Insights.
Bottom line: Your coding agent doesn't need more context. It needs better context. The path from a mediocre agent to an excellent one isn't through larger models or bigger context windows — it's through the disciplined architecture of retrieval quality, contradiction filtering, summarization, and token budgeting. Less is more, and in AI agent design, "less" means more signal per token.