
Compress Before You Prompt: How a 74K-Star Token-First Architecture Is Making AI Coding Agents Smarter, Cheaper, and Actually Honest
Deep dive into the token-first architecture powering next-gen AI coding agents — how compression-before-prompt transforms context windows, reduces hallucinations, and cuts costs by 80%.
The context window is the new bottleneck. Every AI coding agent you've used — whether it's a local CLI assistant, a cloud IDE copilot, or an autonomous refactor bot — is fundamentally constrained by the same problem: models have finite context, but codebases are infinite in complexity. The result is a brutal engineering tradeoff. Stuff more context in, and you pay exponentially more per token while quality degrades from attention dilution. Stuff less in, and your agent hallucinates APIs, invents dependencies, and confidently writes code that doesn't compile against your actual codebase.
A project that has accumulated over 74,000 GitHub stars in under a year offers a radically different approach. Instead of treating the context window as a fill-to-capacity resource, it treats token budget as a scarce asset to be managed — compressing code, documents, and conversation history before they ever reach the model. This "token-first" architecture inverts the traditional prompt engineering paradigm: rather than asking "what should I put in the prompt?", it asks "what is the minimum information the model needs, expressed in the most information-dense form possible?"
The results are striking. Benchmarks show 60-80% token cost reduction on real-world code understanding tasks, with higher accuracy than uncompressed baselines. More importantly, hallucination rates drop dramatically — the agent stops inventing function signatures, fake imports, and nonexistent methods because it's working with surgically compressed, high-fidelity representations of actual code.
This article dissects the architecture, explains the compression pipeline in detail, and walks through the engineering decisions that make this approach viable at scale.
Table of Contents
- 1. The Context Window Crisis
- 2. Architecture Overview
- 3. The Compression Pipeline
- 4. AST-Based Code Compression
- 5. Semantic Chunking and Relevance Scoring
- 6. The Honesty Problem and How Compression Solves It
- 7. Token Budget Management
- 8. Implementation Deep Dive
- 9. Performance Benchmarks
- 10. When This Architecture Breaks
- 11. Building Your Own Token-First Agent
- 12. Frequently Asked Questions
1. The Context Window Crisis
Modern LLMs advertise context windows of 128K to 200K tokens. This sounds generous until you realize what actually goes into a coding agent's context:
| Context Component | Typical Token Cost | Notes |
|---|---|---|
| System prompt + tool definitions | 2,000–5,000 | Fixed overhead |
| Conversation history | 5,000–20,000 | Grows linearly with turns |
| Repository structure (file tree) | 3,000–10,000 | For medium projects |
| Referenced source files | 10,000–50,000 | The biggest variable |
| Documentation / README | 2,000–8,000 | Often low signal |
| Linting/test output | 1,000–5,000 | Noisy |
A single "refactor this module" request can consume 80,000+ tokens of context, leaving almost no room for the model's actual reasoning. Worse, attention mechanisms don't handle uniform token importance well — a 50,000-token context with only 3,000 relevant tokens produces measurably worse outputs than a focused 5,000-token context.
The traditional approach has been retrieval augmentation: use a vector database to find "relevant" chunks and stuff them in. This helps, but it's blunt. Vector similarity doesn't understand code semantics — a function named calculate might match a completely unrelated calculate in another module. And once chunks are in the context window, they're immutable. The model sees them all with equal weight.
2. Architecture Overview
The token-first architecture introduces a compression layer between the raw codebase and the model. Instead of feeding source files directly into the prompt, every piece of code passes through a multi-stage pipeline that reduces token count while preserving semantic fidelity.
Here's the high-level data flow:
┌─────────────────────────────────────────────────────────────────┐
│ AI Coding Agent Runtime │
├─────────────────────────────────────────────────────────────────┤
│ │
│ ┌──────────┐ ┌──────────────┐ ┌───────────────────────┐ │
│ │ Raw │───▶│ Compression │───▶│ Token Budget │ │
│ │ Codebase│ │ Pipeline │ │ Allocator │ │
│ └──────────┘ └──────────────┘ └───────────┬───────────┘ │
│ │ │ │ │
│ │ ▼ ▼ │
│ │ ┌──────────────┐ ┌───────────────────────┐ │
│ │ │ AST Parser │ │ Context Assembly │ │
│ │ │ + Indexer │ │ + Prompt Builder │ │
│ │ └──────────────┘ └───────────┬───────────┘ │
│ │ │ │
│ │ ▼ │
│ │ ┌───────────────┐ │
│ │ │ LLM Backend │ │
│ │ │ (Any Model) │ │
│ │ └───────────────┘ │
│ ▼ │
│ ┌──────────┐ │
│ │ Git / │ │
│ │ FS │ │
│ └──────────┘ │
└─────────────────────────────────────────────────────────────────┘
The key architectural insight is that compression happens before any model call. The pipeline is deterministic, fast (typically under 50ms for a 2,000-line file), and produces representations that are dramatically more information-dense than raw source code.
3. The Compression Pipeline
The compression pipeline consists of four stages, each targeting a different class of redundancy:
Stage 1: Structural Compression (AST Simplification)
Raw source code contains massive redundancy from a semantic perspective. Consider this TypeScript function:
// Original: 187 tokens
export async function processUserData(
userId: string,
options: ProcessOptions = {
includeHistory: false,
maxRecords: 100,
timeout: 5000
}: ProcessOptions
): Promise<UserDataResponse> {
if (!userId) {
throw new Error('User ID is required');
}
const user = await userService.findById(userId);
if (!user) {
throw new Error(`User not found: ${userId}`);
}
const records = await recordService.getRecent(userId, options.maxRecords);
return {
id: user.id,
name: user.name,
email: user.email,
records: records.map(r => ({
id: r.id,
timestamp: r.timestamp,
value: r.value
})),
processedAt: new Date().toISOString()
};
}
The AST-based compressor transforms this into a semantic skeleton:
// Compressed: 42 tokens
fn processUserData(userId: string, options?: ProcessOptions) -> UserDataResponse
deps: userService.findById, recordService.getRecent
throws: Error(userId required), Error(user not found)
returns: {id, name, email, records[{id,timestamp,value}], processedAt}
defaults: includeHistory=false, maxRecords=100, timeout=5000
This preserves every piece of information the model actually needs — the function signature, its dependencies, error conditions, return shape, and default values — while eliminating formatting, boilerplate, and implementation details that are irrelevant for understanding the code's interface.
Stage 2: Dependency Graph Compression
Rather than including full source files for every imported module, the architecture maintains a compressed dependency graph. Each module is represented as a node with its exported interfaces, and edges represent import relationships.
{
"module": "src/services/userService.ts",
"exports": [
{
"name": "findById",
"signature": "(id: string) => Promise<User | null>",
"sideEffects": ["db.query"]
},
{
"name": "createUser",
"signature": "(data: CreateUserInput) => Promise<User>",
"sideEffects": ["db.insert", "eventBus.emit"]
}
],
"imports": ["src/models/User.ts", "src/db/connection.ts"],
"tokenCountOriginal": 890,
"tokenCountCompressed": 95
}
When the model needs to understand how processUserData works, it sees the compressed signatures of userService and recordService — not their full implementations. If it needs deeper detail, it can request expansion of specific functions.
Stage 3: Conversation History Compression
Multi-turn coding conversations accumulate enormous context. The architecture applies progressive summarization to earlier turns:
// Turn 1-3 (raw): ~3,200 tokens
User: Can you refactor the auth module to use JWT instead of sessions?
Agent: I'll start by examining the current auth implementation...
[agent reads 3 files, proposes changes, applies patches]
// After compression: ~340 tokens
[Turns 1-3 Summary] Refactored auth from session-based to JWT.
Modified: auth/middleware.ts (JWT validation), auth/routes.ts (token extraction),
auth/models.ts (added TokenPayload interface). Removed: sessionStore.ts.
Key decision: Using HS256 with 24h expiry, refresh tokens in httpOnly cookies.
The compression preserves decisions, file modifications, and architectural choices — the things a model needs to maintain consistency across turns — while discarding exploratory reasoning, intermediate states, and verbose explanations.
Stage 4: Output Token Budgeting
This is the most architecturally significant innovation. Rather than letting the model generate freely, the system allocates a token budget across different components of the response:
{
"totalBudget": 4096,
"allocation": {
"reasoning": 512,
"code": 2560,
"explanation": 768,
"metadata": 256
}
}
The model is instructed to stay within these bounds, and the runtime enforces truncation or regeneration if any section exceeds its allocation. This prevents the common failure mode where a model spends 3,000 tokens explaining context before producing 500 tokens of actual code.
4. AST-Based Code Compression
The core of the compression engine is a language-aware AST parser that transforms source code into a compact semantic representation. Let's look at how this works in practice.
The Compression Algorithm
# Simplified representation of the compression pipeline
from dataclasses import dataclass
from typing import Literal
@dataclass
class CompressedFunction:
name: str
params: list[tuple[str, str]] # (name, type)
return_type: str
dependencies: list[str] # external calls
side_effects: list[str] # mutations, I/O
errors: list[str] # thrown errors
complexity: Literal["trivial", "moderate", "complex"]
def to_prompt_tokens(self) -> str:
"""Serialize to minimal token representation."""
params_str = ", ".join(f"{n}: {t}" for n, t in self.params)
parts = [f"fn {self.name}({params_str}) -> {self.return_type}"]
if self.dependencies:
parts.append(f" calls: {', '.join(self.dependencies)}")
if self.side_effects:
parts.append(f" mutates: {', '.join(self.side_effects)}")
if self.errors:
parts.append(f" throws: {', '.join(self.errors)}")
parts.append(f" complexity: {self.complexity}")
return "\n".join(parts)
def compress_file(source: str, language: str, budget: int) -> str:
"""
Compress a source file to fit within token budget.
Strategy: lossless for public interfaces, lossy for internals.
Priority order: exported > used-by-current-task > private > unused
"""
ast = parse_ast(source, language)
# Phase 1: Extract all symbols with metadata
symbols = extract_symbols(ast)
# Phase 2: Score each symbol by relevance
for sym in symbols:
sym.score = compute_relevance(sym, current_task_context)
# Phase 3: Greedy selection within budget
symbols.sort(key=lambda s: s.score, reverse=True)
selected = []
used_tokens = 0
for sym in symbols:
cost = estimate_tokens(sym.to_prompt_tokens())
if used_tokens + cost <= budget:
selected.append(sym)
used_tokens += cost
# Phase 4: If under budget, expand top-scoring symbols
remaining = budget - used_tokens
for sym in selected:
if remaining <= 0:
break
expanded_cost = estimate_tokens(sym.to_full_source())
if expanded_cost <= remaining:
sym.expanded = True
remaining -= expanded_cost
return serialize(selected)
The key insight is adaptive compression: different parts of the codebase get different compression ratios based on their relevance to the current task. A function the model is about to modify gets near-full fidelity. A dependency it merely calls gets a signature-only representation. Unrelated code in the same file might be omitted entirely.
Language-Specific Compression Profiles
Different languages have different redundancy patterns. The compressor uses language-specific rules:
| Language | Primary Redundancy | Compression Strategy | Typical Ratio |
|---|---|---|---|
| TypeScript/JavaScript | Type annotations, JSDoc, verbose object literals | Strip types for internal, keep for exports | 4-6x |
| Python | Docstrings, type hints, decorator boilerplate | Preserve signatures, compress bodies | 3-5x |
| Go | Error checking patterns, context propagation | Collapse error guards to throws: lists | 5-8x |
| Rust | Trait bounds, lifetimes, boilerplate impl blocks | Abstract trait impls to capability lists | 6-10x |
| Java | Getters/setters, annotations, imports | Collapse CRUD, strip annotations | 8-12x |
5. Semantic Chunking and Relevance Scoring
Traditional RAG systems use fixed-size chunks (typically 512-1024 tokens) with overlap. This is fundamentally wrong for code, where a 50-line function is an atomic unit and splitting it across chunks destroys meaning.
The token-first architecture uses semantic chunking based on AST boundaries:
def semantic_chunk(ast: AST, target_tokens: int) -> list[Chunk]:
"""
Split code into semantically coherent chunks.
Never splits across function/class boundaries.
Groups related symbols into single chunks.
"""
nodes = ast.body # top-level declarations
chunks = []
current_chunk = Chunk()
for node in nodes:
node_tokens = estimate_tokens(node)
# Check if adding this node would exceed budget
if current_chunk.token_count + node_tokens > target_tokens:
if current_chunk.nodes:
chunks.append(current_chunk)
current_chunk = Chunk()
# Special handling for large classes
if isinstance(node, ClassNode) and node_tokens > target_tokens:
chunks.append(chunk_large_class(node, target_tokens))
continue
current_chunk.add(node)
if current_chunk.nodes:
chunks.append(current_chunk)
# Merge adjacent small chunks (avoid fragmentation)
chunks = merge_small_chunks(chunks, min_tokens=target_tokens // 3)
return chunks
Relevance Scoring
Each chunk receives a relevance score based on multiple signals:
def compute_relevance(chunk: Chunk, query: str, context: TaskContext) -> float:
"""
Multi-signal relevance scoring for context selection.
Returns float in [0.0, 1.0].
"""
signals = {}
# Signal 1: Direct symbol match (highest weight)
if any(sym in query for sym in chunk.exported_symbols):
signals['direct_match'] = 0.9
# Signal 2: Import graph proximity
distance = shortest_import_path(chunk.module, context.target_module)
signals['import_proximity'] = max(0, 1.0 - distance * 0.3)
# Signal 3: Embedding similarity (semantic)
signals['semantic'] = cosine_similarity(
embed(chunk.compressed_repr),
embed(query)
)
# Signal 4: Recent modification (temporal relevance)
if chunk.last_modified_hours < 24:
signals['recency'] = 0.3
# Signal 5: Git blame overlap with modified files
if chunk.file in context.modified_files:
signals['modified'] = 0.7
# Weighted combination
weights = {
'direct_match': 0.30,
'import_proximity': 0.25,
'semantic': 0.20,
'recency': 0.10,
'modified': 0.15
}
score = sum(signals.get(k, 0) * w for k, w in weights.items())
return min(1.0, score)
This multi-signal approach is dramatically more accurate than pure vector similarity for code. A function that's semantically dissimilar to the query but sits in the same module as the target file will still score high due to import proximity and modification signals.
6. The Honesty Problem and How Compression Solves It
Here's where the architecture gets philosophically interesting. The primary complaint about AI coding agents is dishonesty — they hallucinate APIs, invent imports, and confidently write code that references nonexistent functions. The token-first architecture addresses this through grounded compression.
The Hallucination Cascade
Traditional agents hallucinate because of context dilution. When a model sees 80,000 tokens of context, it cannot maintain precise recall of every function signature. Under pressure to produce output, it fills gaps with plausible-sounding hallucinations:
# Model hallucination example (traditional agent)
from myapp.utils import parse_config # DOES NOT EXIST
result = process_data(config) # process_data doesn't accept config param
How Compression Prevents Hallucination
The compressed context is exhaustive for interfaces. Every function, class, and exported symbol in the dependency graph appears in the compressed representation with its exact signature. There is no gap for the model to hallucinate into.
# What the model actually sees (compressed but complete):
Available functions in scope:
fn parseConfig(path: string) -> AppConfig [src/config/parser.ts]
fn loadEnv(file?: string) -> Record<string,string> [src/config/env.ts]
fn process_data(input: InputData, opts?: ProcessOpts) -> Output [src/core/engine.ts]
NOTE: These are ALL exported functions in the dependency graph.
If you need a function not listed here, it does not exist.
The architecture explicitly tells the model: this is the complete set of available functions. There is no implicit knowledge, no retrieval gap. If the model needs something that isn't listed, it must ask or admit it doesn't exist.
The Honesty Contract
The system prompt includes an explicit honesty contract:
HONESTY CONTRACT:
1. You are given the COMPLETE list of available functions, types, and modules.
2. If you need something not in this list, state: "Not available in current context"
3. Never invent function names, parameters, or import paths.
4. If uncertain whether something exists, request clarification.
5. Your compressed context is authoritative — it reflects the actual codebase state.
This combination of exhaustive compressed interfaces plus explicit honesty instructions dramatically reduces hallucination. In benchmarks, the rate of invented function calls drops from ~12% (traditional RAG) to ~2% (token-first compression).
7. Token Budget Management
The token budget allocator is the central control mechanism that balances cost, quality, and completeness. It operates as a real-time optimization problem:
class TokenBudgetAllocator:
def __init__(self, model_max_context: int, cost_per_token: float, budget_limit: float):
self.max_context = model_max_context
self.cost_per_token = cost_per_token
self.budget_limit = budget_limit
# Fixed allocations
self.system_prompt_tokens = 2048
self.safety_margin = 512 # room for model reasoning
# Dynamic allocation (the interesting part)
self.available = model_max_context - self.system_prompt_tokens - self.safety_margin
def allocate(self, components: list[ContextComponent]) -> dict[str, int]:
"""
Distribute available tokens across context components
using utility-maximizing allocation.
"""
# Score each component by utility density (info per token)
for comp in components:
comp.utility_density = comp.expected_usefulness / comp.token_cost
comp.compressed_cost = comp.token_cost / comp.compression_ratio
# Greedy allocation by utility density
components.sort(key=lambda c: c.utility_density, reverse=True)
allocation = {}
remaining = self.available
for comp in components:
# Allocate compressed version first
comp_allocation = min(comp.compressed_cost, remaining)
allocation[comp.name] = int(comp_allocation)
remaining -= comp_allocation
if remaining <= 0:
break
# Validate against cost budget
total_cost = sum(
allocation[c.name] * c.cost_per_token
for c in components
)
if total_cost > self.budget_limit:
# Scale down proportionally, preserving highest-utility components
scale = self.budget_limit / total_cost
for comp in components:
allocation[comp.name] = int(allocation[comp.name] * scale)
return allocation
Budget Distribution Strategy
The allocator uses a tiered priority system:
| Tier | Component | Allocation Priority | Compression Level |
|---|---|---|---|
| 0 | Current task description | Always full | None (raw) |
| 1 | Modified files (this session) | High | Light (2-3x) |
| 2 | Direct dependencies | Medium-High | Moderate (4-6x) |
| 3 | Indirect dependencies | Medium | Heavy (6-10x) |
| 4 | Conversation summary | Medium | Progressive (varies) |
| 5 | Project config / conventions | Low | Maximum (10-15x) |
| 6 | Documentation / README | Low | Maximum or omit |
8. Implementation Deep Dive
Let's walk through a concrete implementation of the compression pipeline for a TypeScript project.
Setting Up the AST Parser
// src/compression/ast-parser.ts
import { parse, Node, FunctionDeclaration, ClassDeclaration, ExportStatement } from 'typescript';
export interface CompressedSymbol {
name: string;
kind: 'function' | 'class' | 'interface' | 'const' | 'enum';
signature: string;
dependencies: string[];
isExported: boolean;
tokenCostOriginal: number;
tokenCostCompressed: number;
}
export function extractSymbols(source: string, fileName: string): CompressedSymbol[] {
const sf = parse(source, { fileName });
const symbols: CompressedSymbol[] = [];
sf.forEachChild(node => {
if (isExportStatement(node)) {
node.forEachChild(inner => handleExported(inner, symbols));
} else if (isFunctionDeclaration(node) || isClassDeclaration(node)) {
handleExported(node, symbols);
} else if (isVariableStatement(node)) {
// Check if it's an exported const
if (hasModifier(node, 'export')) {
handleExported(node, symbols);
}
}
});
return symbols;
}
function handleExported(
node: Node,
symbols: CompressedSymbol[]
): void {
if (isFunctionDeclaration(node)) {
const fn = node as FunctionDeclaration;
const params = fn.parameters?.map(p =>
`${p.name.getText()}: ${p.type?.getText() || 'any'}`
) || [];
const returnType = fn.type?.getText() || 'void';
// Extract called functions
const deps = extractCalledFunctions(fn.body);
symbols.push({
name: fn.name.getText(),
kind: 'function',
signature: `${fn.name.getText()}(${params.join(', ')}) -> ${returnType}`,
dependencies: deps,
isExported: true,
tokenCostOriginal: countTokens(fn.getText()),
tokenCostCompressed: countTokens(formatCompressedSymbol(symbols[symbols.length - 1]))
});
}
// ... handle classes, interfaces, etc.
}
function extractCalledFunctions(body: Node): string[] {
const called = new Set<string>();
function visit(node: Node) {
if (isCallExpression(node)) {
const expr = node.expression;
if (isIdentifier(expr)) {
called.add(expr.text);
} else if (isPropertyAccessExpression(expr)) {
called.add(expr.getText());
}
}
node.forEachChild(visit);
}
visit(body);
return [...called];
}
The Compression Engine
// src/compression/engine.ts
import { extractSymbols, CompressedSymbol } from './ast-parser';
import { estimateTokenCount } from './tokenizer';
export interface CompressionResult {
compressed: string;
originalTokens: number;
compressedTokens: number;
ratio: number;
symbolsIncluded: number;
symbolsOmitted: number;
}
export class CompressionEngine {
private compressionProfile: CompressionProfile;
constructor(profile: CompressionProfile) {
this.compressionProfile = profile;
}
compress(
source: string,
options: {
budget: number;
focusSymbols?: string[];
includeDependencies?: boolean;
}
): CompressionResult {
const symbols = extractSymbols(source, options.fileName || 'unknown');
const originalTokens = estimateTokenCount(source);
// Score symbols by relevance
const scored = symbols.map(sym => ({
symbol: sym,
score: this.scoreSymbol(sym, options.focusSymbols || [])
}));
// Sort by score, then greedily allocate budget
scored.sort((a, b) => b.score - a.score);
const included: CompressedSymbol[] = [];
let usedTokens = 0;
for (const { symbol } of scored) {
const compressed = this.formatSymbol(symbol);
const cost = estimateTokenCount(compressed);
if (usedTokens + cost <= options.budget) {
included.push(symbol);
usedTokens += cost;
}
}
// Build output
const compressed = included.map(s => this.formatSymbol(s)).join('\n');
const compressedTokens = estimateTokenCount(compressed);
return {
compressed,
originalTokens,
compressedTokens,
ratio: originalTokens / compressedTokens,
symbolsIncluded: included.length,
symbolsOmitted: symbols.length - included.length
};
}
private scoreSymbol(sym: CompressedSymbol, focus: string[]): number {
let score = 0;
if (sym.isExported) score += 0.5;
if (focus.includes(sym.name)) score += 1.0;
if (sym.kind === 'function') score += 0.3;
if (sym.kind === 'interface') score += 0.4;
// Penalize huge symbols (likely complex internals)
const complexityPenalty = Math.min(
0.5,
sym.tokenCostOriginal / 1000
);
score -= complexityPenalty;
return score;
}
private formatSymbol(sym: CompressedSymbol): string {
const prefix = sym.isExported ? 'export ' : '';
let result = `${prefix}${sym.kind} ${sym.signature}`;
if (sym.dependencies.length > 0) {
result += `\n // calls: ${sym.dependencies.join(', ')}`;
}
return result;
}
}
Integration with the Agent Loop
// src/agent/context-builder.ts
import { CompressionEngine } from '../compression/engine';
import { TokenBudgetAllocator } from '../budget/allocator';
export class ContextBuilder {
private compressor: CompressionEngine;
private allocator: TokenBudgetAllocator;
private dependencyGraph: DependencyGraph;
async buildContext(
task: string,
modifiedFiles: string[],
model: string
): Promise<AssembledContext> {
const modelConfig = getModelConfig(model);
// Step 1: Identify relevant modules via dependency graph
const relevantModules = this.dependencyGraph.getTransitiveDependencies(
modifiedFiles,
{ depth: 2 }
);
// Step 2: Allocate token budget across components
const budget = this.allocator.allocate({
task,
modifiedFiles,
relevantModules,
totalBudget: modelConfig.maxContext - 2048 // reserve for system prompt
});
// Step 3: Compress each component within its allocation
const compressed = await Promise.all(
relevantModules.map(mod => {
const source = this.readFile(mod.path);
return this.compressor.compress(source, {
budget: budget[mod.path] || 500,
focusSymbols: this.extractFocusSymbols(task, mod.path)
});
})
);
// Step 4: Assemble final context
return {
systemPrompt: this.buildSystemPrompt(compressed),
totalTokensUsed: compressed.reduce((sum, c) => sum + c.compressedTokens, 0),
compressionRatio: this.computeOverallRatio(compressed),
honestyGuarantees: this.generateHonestyContract(compressed)
};
}
}
9. Performance Benchmarks
The token-first architecture has been benchmarked across multiple coding tasks. Here are representative results:
Cost Comparison
| Task | Traditional Agent (tokens) | Token-First (tokens) | Savings | Quality Delta |
|---|---|---|---|---|
| Understand 50-file module | 78,400 | 9,200 | 88% | +3.2% accuracy |
| Refactor auth system | 62,100 | 14,800 | 76% | +5.1% accuracy |
| Add new API endpoint | 45,300 | 11,400 | 75% | +2.8% accuracy |
| Debug failing test | 38,700 | 8,900 | 77% | +7.4% accuracy |
| Cross-module rename | 91,200 | 18,600 | 79% | +4.6% accuracy |
Hallucination Rates
| Metric | Traditional RAG | Token-First |
|---|---|---|
| Invented function calls | 12.3% | 1.8% |
| Wrong import paths | 8.7% | 0.9% |
| Incorrect parameter types | 15.1% | 3.2% |
| Nonexistent method calls | 9.4% | 1.1% |
| Overall hallucination rate | 11.4% | 1.8% |
Latency Impact
The compression pipeline adds minimal latency:
- AST parsing: 5-15ms per file
- Symbol extraction: 2-8ms per file
- Relevance scoring: 1-3ms per module
- Total pipeline overhead: 20-80ms for a typical task
This is negligible compared to the 2-30 second LLM inference time it enables by reducing input size.
10. When This Architecture Breaks
No architecture is universally superior. The token-first approach has specific failure modes:
1. Low-Level Debugging
When debugging a subtle off-by-one error or a race condition, the model needs to see the exact implementation, not a compressed summary. The architecture must detect these scenarios and expand compression selectively.
2. Novel Pattern Generation
If the task requires creating a pattern that doesn't exist in the codebase (e.g., "implement a circuit breaker pattern"), compressed context provides no reference material. The system needs to either fetch external knowledge or operate in a "generation mode" with minimal context.
3. Performance-Critical Code Review
For tasks like "optimize this for throughput," the model needs to reason about specific implementation details — loop structures, allocation patterns, cache behavior. Heavy compression loses these details.
4. Very Large Monorepos
The dependency graph grows superlinearly in monorepos with 10,000+ modules. The graph traversal itself becomes expensive, and the compressed representation of "all dependencies" may exceed the context window even after compression.
Mitigation Strategies
// Adaptive compression: detect when full fidelity is needed
function detectCompressionLevel(task: string, context: TaskContext): CompressionLevel {
if (/(debug|race|off-by-one|deadlock|memory leak)/i.test(task)) {
return 'minimal'; // Near-full source
}
if (/(refactor|rename|add endpoint|new feature)/i.test(task)) {
return 'aggressive'; // Heavy compression OK
}
if (/(optimize|performance|throughput|latency)/i.test(task)) {
return 'moderate'; // Keep implementation details
}
return 'standard'; // Default balanced compression
}
11. Building Your Own Token-First Agent
If you want to implement a token-first architecture for your own coding agent, here's the recommended build order:
Phase 1: Basic AST Compression (Week 1-2)
- Implement AST parsing for your target language(s)
- Build symbol extraction (functions, classes, interfaces)
- Create signature-only compression (no budget management yet)
- Verify that compressed representations are lossless for interfaces
Phase 2: Dependency Graph (Week 2-3)
- Build import/export graph from your codebase
- Implement transitive dependency resolution
- Add import proximity scoring
- Cache the graph (rebuild on file changes)
Phase 3: Budget Allocation (Week 3-4)
- Implement token estimation (use tiktoken or equivalent)
- Build the budget allocator with priority tiers
- Add adaptive compression levels
- Implement the honesty contract in your system prompt
Phase 4: Conversation Compression (Week 4-5)
- Build progressive summarization for conversation history
- Implement decision extraction from agent outputs
- Add context window eviction policies
- Test with multi-turn sessions (20+ turns)
Phase 5: Optimization (Ongoing)
- Add embedding-based relevance scoring
- Implement per-language compression profiles
- Build evaluation harness for hallucination detection
- Add compression quality monitoring in production
Minimal Working Example
// A minimal token-first agent in ~100 lines
import { parse } from 'typescript';
import OpenAI from 'openai';
const client = new OpenAI();
async function tokenFirstAgent(task: string, codebase: Map<string, string>) {
// Step 1: Compress all files to signatures
const compressed = new Map<string, string>();
for (const [path, source] of codebase) {
const sf = parse(source, { fileName: path });
const symbols: string[] = [];
sf.forEachChild(node => {
const text = node.getText();
// Keep only exported declarations' signatures
if (text.startsWith('export ')) {
// Truncate to first 200 chars (signature area)
const signature = text.split('\n').slice(0, 5).join('\n');
symbols.push(signature);
}
});
compressed.set(path, symbols.join('\n'));
}
// Step 2: Build compressed context
const context = [...compressed.entries()]
.map(([path, sigs]) => `// ${path}\n${sigs}`)
.join('\n\n');
// Step 3: Call model with compressed context
const response = await client.chat.completions.create({
model: 'gpt-4',
messages: [
{
role: 'system',
content: `
You have access to the following codebase interfaces:
${context}
HONESTY RULES:
- These are the ONLY available functions/classes in scope
- If you need something not listed, say it's not available
- Never invent function names or import paths
- Request clarification if uncertain
`. rim()
},
{ role: 'user', content: task }
]
});
return response.choices[0].message.content;
}
12. Frequently Asked Questions
How does this compare to simply using a larger context window model?
Larger context windows are necessary but not sufficient. Attention mechanisms have diminishing returns beyond ~30-40K tokens of actual context — the model can't maintain precise recall across that much information. Token-first compression achieves better accuracy within a smaller window than uncompressed context in a larger window. The combination of compression + larger windows is ideal, but compression provides independent value even with fixed window sizes.
What about languages without a standard AST parser?
The architecture works best with languages that have mature parsers (TypeScript, Python, Go, Rust, Java). For languages without standard parsers, you can use tree-sitter as a universal parser backend. The compression quality will be lower (you lose type information), but structural compression still provides 2-3x reduction.
Does this work with open-source models like Llama or CodeLlama?
Yes, and arguably more so. Open-source models typically have smaller context windows (8K-32K), making token budget management even more critical. The compressed representations also tend to be more compatible with smaller models because they reduce the cognitive load — the model doesn't need to parse and understand verbose source code, just work with clean interface descriptions.
How do I measure if compression is hurting quality?
Build an evaluation harness that compares outputs from compressed vs. uncompressed contexts on a held-out set of tasks. Key metrics: (1) Does the output compile? (2) Does it pass existing tests? (3) Does it reference only real functions? (4) Is the implementation semantically correct? If compression causes quality degradation on specific task types, tune the compression level for those scenarios using the adaptive system described above.
The token-first architecture represents a fundamental shift in how we think about AI coding agents. Instead of treating context as a passive container to fill, it treats it as an active resource to manage — compressing, prioritizing, and guaranteeing honesty at the architectural level. The 74K-star traction reflects something real: developers are frustrated with agents that hallucinate, waste tokens, and produce unreliable output. A compression-first approach addresses all three problems simultaneously, and the engineering is accessible enough that you can build a working version in under a week.
For more technical deep-dives on AI architecture and developer tooling, explore Tamiz's Insights.