
Building Agent-Native Apps Without Burning Your Budget: Practical Patterns for Sandboxed AI Agent Architectures
Deep dive into cost-controlling architectural patterns for AI agent systems: sandboxed execution, token budgeting, circuit breakers, progressive escalation, and caching layers with production-ready code.
AI agents are the most powerful and most expensive primitive in modern software. A single agentic loop that explores, plans, calls tools, retries, and synthesizes can burn through a month's API budget in hours — and your monitoring won't catch it until the invoice arrives. The core engineering challenge isn't capability; it's containment. How do you give agents autonomy to accomplish complex tasks while ensuring a runaway loop doesn't liquidate your runway?
This deep dive covers five production-proven patterns for building sandboxed, budget-aware AI agent architectures, with complete implementation code, architectural diagrams, and the operational metrics you need to keep costs predictable while preserving agent effectiveness.
Table of Contents
- 1. The Agent Cost Problem: Where Budgets Actually Burn
- 2. Pattern One: Tiered Compute Isolation
- 3. Pattern Two: Token Budgeting with Circuit Breakers
- 4. Pattern Three: Sandboxed Tool Execution
- 5. Pattern Four: Semantic Caching and Deduplication
- 6. Pattern Five: Progressive Escalation
- 7. Reference Architecture: Putting It All Together
- 8. Implementation: The Sandbox Orchestrator
- 9. Operational Metrics and Alerting
- 10. Frequently Asked Questions
1. The Agent Cost Problem: Where Budgets Actually Burn
Before diving into solutions, you need a precise mental model of where agent costs accumulate. Unlike a simple request-response API call, an agent performs a loop: plan → act → observe → replan. Each iteration multiplies cost through several channels simultaneously.
The Cost Multiplication Taxonomy
| Cost Channel | Typical Amplifier | Example Scenario |
|---|---|---|
| LLM inference | Iteration count × context length | 15-step reasoning loop with growing conversation history |
| Tool execution | Retry storms, unbounded loops | Agent retries a failed API call 47 times |
| Context accumulation | O(n²) token growth | Each loop appends prior outputs to the prompt |
| Parallel exploration | Branching factor × depth | Agent spawns 5 sub-agents, each spawning 3 |
| Silent failures | Error recovery loops | Agent keeps retrying a misconfigured tool |
The worst-case scenario is a feedback loop: a tool returns a confusing error, the agent re-plans using a larger context, the new plan calls more tools, which produce more errors, and the context grows quadratically. A single user request can cascade into thousands of tokens and multiple tool calls per second.
The solution isn't to reduce agent capability — it's to build containment layers that let agents work autonomously within hard financial boundaries.
2. Pattern One: Tiered Compute Isolation
Concept
Not every agent task deserves the same compute resources. A simple summarization task and a multi-step research task have fundamentally different cost profiles. Tiered compute isolation assigns each agent invocation to a resource tier based on task complexity, with hard ceilings per tier.
┌─────────────────────────┐
│ Task Complexity │
│ Classifier │
└────────┬────────────────┘
│
┌──────────────┼──────────────┐
▼ ▼ ▼
┌────────────┐ ┌────────────┐ ┌────────────┐
│ Micro │ │ Standard │ │ Extended │
│ (1 model │ │ (1 model │ │ (multi- │
│ call, │ │ + tools, │ │ model, │
│ ≤2K tok) │ │ ≤15K tok)│ │ ≤50K tok)│
└────────────┘ └────────────┘ └────────────┘
Implementation
The classifier is itself a lightweight model call — a single prompt asking a small/fast model to categorize the task. This costs a fraction of what the actual task will cost, so it's economically sound.
from dataclasses import dataclass
from enum import Enum
import time
class ComputeTier(Enum):
MICRO = "micro" # Single LLM call, no tools
STANDARD = "standard" # LLM + tools, bounded iterations
EXTENDED = "extended" # Multi-model, multi-step reasoning
@dataclass
class TierConfig:
tier: ComputeTier
max_iterations: int
max_tokens_total: int
max_tool_calls: int
allowed_models: list[str]
timeout_seconds: int
max_concurrent_subagents: int
TIER_REGISTRY: dict[ComputeTier, TierConfig] = {
ComputeTier.MICRO: TierConfig(
tier=ComputeTier.MICRO,
max_iterations=1,
max_tokens_total=2_000,
max_tool_calls=0,
allowed_models=["gpt-4o-mini", "claude-haiku"],
timeout_seconds=30,
max_concurrent_subagents=0,
),
ComputeTier.STANDARD: TierConfig(
tier=ComputeTier.STANDARD,
max_iterations=8,
max_tokens_total=15_000,
max_tool_calls=12,
allowed_models=["gpt-4o", "claude-sonnet"],
timeout_seconds=120,
max_concurrent_subagents=2,
),
ComputeTier.EXTENDED: TierConfig(
tier=ComputeTier.EXTENDED,
max_iterations=25,
max_tokens_total=50_000,
max_tool_calls=40,
allowed_models=["gpt-4o", "claude-sonnet", "claude-opus"],
timeout_seconds=300,
max_concurrent_subagents=5,
),
}
class TierClassifier:
"""Classifies incoming agent tasks into compute tiers."""
def __init__(self, model_client):
self.model_client = model_client
def classify(self, user_request: str, context: dict) -> ComputeTier:
classification_prompt = f"""Classify this task into one of three tiers:
- MICRO: Simple single-step tasks (summarize, translate, extract)
- STANDARD: Multi-step tasks requiring tool use (search, research, code review)
- EXTENDED: Complex multi-phase tasks (system design, deep research, multi-file refactoring)
Request: "{user_request[:500]}"
Context summary: {context.get('summary', 'none')}
Return ONLY the tier name."""
response = self.model_client.complete(
model="gpt-4o-mini",
prompt=classification_prompt,
max_tokens=20,
)
tier_str = response.strip().upper()
if "EXTENDED" in tier_str:
return ComputeTier.EXTENDED
elif "STANDARD" in tier_str:
return ComputeTier.STANDARD
return ComputeTier.MICRO
The key insight: the classification call itself costs roughly 50-100 tokens, which is negligible compared to the 2,000-50,000 token budgets it governs. You spend pennies to avoid burning dollars.
Tier-Based Model Routing
Each tier restricts which models the agent can use. Micro tasks go to the cheapest capable model. Extended tasks can access premium models but only when the agent's reasoning actually benefits from it — enforced by a quality gate that compares model outputs.
3. Pattern Two: Token Budgeting with Circuit Breakers
Concept
A circuit breaker is a reliability pattern from distributed systems that prevents cascading failures. Applied to agent token budgets, it monitors consumption velocity and trips when spending exceeds safe thresholds, halting the agent loop before the budget is exhausted.
The circuit breaker has three states:
- CLOSED — Normal operation. Tokens are tracked but no intervention occurs.
- HALF-OPEN — Budget threshold crossed (e.g., 70%). The agent is allowed to continue but with reduced capabilities (fewer tools, smaller context windows).
- OPEN — Hard budget exceeded. The agent loop terminates immediately. The partial result is returned with a status flag.
Implementation
import time
from dataclasses import dataclass, field
from enum import Enum
from typing import Callable
class CircuitState(Enum):
CLOSED = "closed"
HALF_OPEN = "half_open"
OPEN = "open"
@dataclass
class TokenBudget:
"""Token budget with circuit breaker enforcement."""
total_budget: int
consumed: int = 0
state: CircuitState = CircuitState.CLOSED
warning_threshold: float = 0.70 # 70% triggers HALF_OPEN
hard_threshold: float = 1.00 # 100% triggers OPEN
velocity_window_seconds: int = 60
max_velocity_tokens_per_second: int = 500
_token_log: list[tuple[float, int]] = field(default_factory=list)
def consume(self, amount: int) -> bool:
"""Attempt to consume tokens. Returns False if budget exceeded."""
if self.state == CircuitState.OPEN:
return False
self.consumed += amount
now = time.time()
self._token_log.append((now, amount))
# Clean old log entries
cutoff = now - self.velocity_window_seconds
self._token_log = [(t, a) for t, a in self._token_log if t >= cutoff]
ratio = self.consumed / self.total_budget
if ratio >= self.hard_threshold:
self.state = CircuitState.OPEN
return False
elif ratio >= self.warning_threshold:
self.state = CircuitState.HALF_OPEN
return True
def get_velocity(self) -> float:
"""Calculate current token consumption rate."""
if len(self._token_log) < 2:
return 0.0
elapsed = self._token_log[-1][0] - self._token_log[0][0]
if elapsed == 0:
return 0.0
total_in_window = sum(a for _, a in self._token_log)
return total_in_window / elapsed
def is_velocity_exceeded(self) -> bool:
return self.get_velocity() > self.max_velocity_tokens_per_second
class BudgetEnforcedAgent:
"""Agent loop with token budget circuit breaker."""
def __init__(self, model_client, budget: TokenBudget, tools: list):
self.model_client = model_client
self.budget = budget
self.tools = tools
self.iteration = 0
def run(self, task: str, max_iterations: int = 10) -> dict:
messages = [{"role": "user", "content": task}]
result = {"status": "completed", "output": None, "iterations": 0}
for self.iteration in range(1, max_iterations + 1):
# Check circuit breaker state
if self.budget.state == CircuitState.OPEN:
result["status"] = "budget_exhausted"
result["partial_output"] = messages[-1].get("content", "")
break
# In HALF_OPEN state, reduce context by trimming older messages
effective_messages = messages
if self.budget.state == CircuitState.HALF_OPEN:
effective_messages = self._trim_context(messages, max_messages=4)
# Check velocity
if self.budget.is_velocity_exceeded():
result["status"] = "velocity_throttled"
result["partial_output"] = messages[-1].get("content", "")
break
# Make LLM call
response = self.model_client.chat(
messages=effective_messages,
tools=self.tools if self.budget.state != CircuitState.HALF_OPEN else self.tools[:2],
)
# Track token consumption
tokens_used = response.get("usage", {}).get("total_tokens", 0)
if not self.budget.consume(tokens_used):
result["status"] = "budget_exhausted"
result["partial_output"] = response.get("content", "")
break
messages.append({"role": "assistant", "content": response.get("content")})
# Check if agent is done
if not response.get("tool_calls"):
result["output"] = response.get("content")
result["iterations"] = self.iteration
break
return result
def _trim_context(self, messages: list, max_messages: int) -> list:
"""Reduce context in HALF_OPEN state to slow consumption."""
if len(messages) <= max_messages:
return messages
# Keep system message + last N messages
if messages[0].get("role") == "system":
return [messages[0]] + messages[-(max_messages - 1):]
return messages[-max_messages:]
Velocity-Based Throttling
The velocity check adds a second axis of protection. Even if the total budget hasn't been reached, a sudden spike in token consumption (indicating a runaway loop) triggers throttling. This catches the scenario where an agent enters a pathological retry loop that would exhaust a large budget in seconds.
4. Pattern Three: Sandboxed Tool Execution
Concept
Agent tools are the highest-risk component for cost overruns. A tool that returns large payloads (e.g., a web search returning 50KB of HTML), a tool that triggers external side effects (e.g., sending emails), or a tool that loops internally (e.g., a recursive file processor) can each independently blow past budget limits.
Sandboxed tool execution wraps every tool call with resource constraints: output size limits, execution time limits, retry limits, and side-effect gates.
Implementation
import asyncio
import hashlib
import json
from dataclasses import dataclass
from typing import Any, Callable
@dataclass
class ToolBudget:
max_calls: int
calls_made: int = 0
max_output_tokens: int = 2_000
max_execution_seconds: float = 10.0
max_retries: int = 2
class SandboxedTool:
"""Wraps a tool function with resource constraints."""
def __init__(self, name: str, func: Callable, budget: ToolBudget):
self.name = name
self.func = func
self.budget = budget
self._call_log: list[dict] = []
async def execute(self, arguments: dict) -> dict:
# Enforce call limit
if self.budget.calls_made >= self.budget.max_calls:
return {
"error": "tool_call_limit_exceeded",
"tool": self.name,
"max_calls": self.budget.max_calls,
}
# Enforce execution time limit with asyncio timeout
for attempt in range(self.budget.max_retries + 1):
try:
self.budget.calls_made += 1
start = asyncio.get_event_loop().time()
result = await asyncio.wait_for(
self.func(**arguments),
timeout=self.budget.max_execution_seconds,
)
elapsed = asyncio.get_event_loop().time() - start
# Enforce output size limit
result_str = json.dumps(result) if not isinstance(result, str) else result
if len(result_str) > self.budget.max_output_tokens * 4: # ~4 chars per token
result = self._truncate_output(result, self.budget.max_output_tokens)
self._call_log.append({
"attempt": attempt,
"elapsed": elapsed,
"output_size": len(result_str),
"success": True,
})
return {"result": result, "metadata": {"elapsed": elapsed}}
except asyncio.TimeoutError:
self._call_log.append({"attempt": attempt, "error": "timeout"})
if attempt == self.budget.max_retries:
return {"error": "tool_timeout", "tool": self.name}
except Exception as e:
self._call_log.append({"attempt": attempt, "error": str(e)})
if attempt == self.budget.max_retries:
return {"error": "tool_failed", "tool": self.name, "detail": str(e)}
return {"error": "unexpected_exit", "tool": self.name}
def _truncate_output(self, result: Any, max_tokens: int) -> Any:
"""Truncate tool output to prevent context bloat."""
max_chars = max_tokens * 4
if isinstance(result, str):
return result[:max_chars] + "\n[TRUNCATED]"
result_str = json.dumps(result)
if len(result_str) <= max_chars:
return result
# For structured data, keep the first N items
if isinstance(result, list):
truncated = result[:10]
return {"items": truncated, "truncated": True, "total_count": len(result)}
return json.loads(result_str[:max_chars])
@property
def stats(self) -> dict:
return {
"tool": self.name,
"calls_made": self.budget.calls_made,
"max_calls": self.budget.max_calls,
"avg_elapsed": sum(c["elapsed"] for c in self._call_log if "elapsed" in c) / max(len([c for c in self._call_log if "elapsed" in c]), 1),
}
class ToolRegistry:
"""Central registry of sandboxed tools with per-tool budgets."""
def __init__(self):
self._tools: dict[str, SandboxedTool] = {}
def register(self, name: str, func: Callable, budget: ToolBudget) -> None:
self._tools[name] = SandboxedTool(name, func, budget)
async def execute(self, name: str, arguments: dict) -> dict:
if name not in self._tools:
return {"error": "tool_not_found", "tool": name}
return await self._tools[name].execute(arguments)
def get_all_stats(self) -> dict:
return {name: tool.stats for name, tool in self._tools.items()}
Side-Effect Gating
For tools that produce irreversible side effects (sending emails, making payments, modifying databases), add a confirmation gate. The agent must produce a structured confirmation that matches the expected action before the tool executes:
class SideEffectGate:
"""Requires explicit confirmation before executing side-effect tools."""
def __init__(self, allowed_actions: set[str]):
self.allowed_actions = allowed_actions
async def execute(self, action: str, payload: dict, confirm: dict) -> dict:
if action not in self.allowed_actions:
return {"error": "action_not_allowed", "action": action}
# Verify the agent's confirmation matches the actual payload
if not self._confirmations_match(payload, confirm):
return {
"error": "confirmation_mismatch",
"expected": confirm,
"actual": payload,
}
# Now actually execute
return await self._do_execute(action, payload)
def _confirmations_match(self, payload: dict, confirm: dict) -> bool:
"""Check that agent's stated intent matches actual payload."""
for key, expected in confirm.items():
actual = payload.get(key)
if actual != expected:
return False
return True
5. Pattern Four: Semantic Caching and Deduplication
Concept
Agents frequently make redundant LLM calls with semantically similar prompts — especially in multi-step workflows where the agent re-asks similar questions after gathering new information. Semantic caching intercepts these calls and returns cached results when the incoming prompt is sufficiently similar to a previously answered one.
Unlike exact-match caching (which fails when prompts differ by a single word), semantic caching uses embedding similarity to match semantically equivalent queries.
Implementation
import hashlib
import time
from collections import OrderedDict
from dataclasses import dataclass
from typing import Optional
import numpy as np
@dataclass
class CacheEntry:
embedding: np.ndarray
response: str
model: str
timestamp: float
hit_count: int = 0
tokens_saved: int = 0
class SemanticCache:
"""Embedding-based cache for LLM responses."""
def __init__(
self,
embedding_model: str = "text-embedding-3-small",
similarity_threshold: float = 0.92,
max_entries: int = 1000,
ttl_seconds: int = 3600,
):
self.embedding_model = embedding_model
self.similarity_threshold = similarity_threshold
self.max_entries = max_entries
self.ttl_seconds = ttl_seconds
self._cache: OrderedDict[str, CacheEntry] = OrderedDict()
self._hits = 0
self._misses = 0
async def get_or_compute(
self,
prompt: str,
model: str,
compute_fn: Callable[[], str],
token_count: int = 0,
) -> str:
"""Check cache first, compute if miss."""
embedding = await self._embed(prompt)
# Find best match
best_match: Optional[CacheEntry] = None
best_score = 0.0
for entry in self._cache.values():
if time.time() - entry.timestamp > self.ttl_seconds:
continue
score = np.dot(embedding, entry.embedding) / (
np.linalg.norm(embedding) * np.linalg.norm(entry.embedding)
)
if score > best_score:
best_score = score
best_match = entry
if best_match and best_score >= self.similarity_threshold:
best_match.hit_count += 1
best_match.tokens_saved += token_count
self._hits += 1
self._cache.move_to_end(best_match.key)
return best_match.response
# Cache miss — compute and store
self._misses += 1
response = await compute_fn()
cache_key = hashlib.md5(prompt.encode()).hexdigest()
self._cache[cache_key] = CacheEntry(
embedding=embedding,
response=response,
model=model,
timestamp=time.time(),
)
# Evict oldest if over capacity
while len(self._cache) > self.max_entries:
self._cache.popitem(last=False)
return response
async def _embed(self, text: str) -> np.ndarray:
"""Generate embedding for text (delegated to embedding API)."""
# In production, this calls your embedding provider
# Placeholder for illustration
return np.random.randn(1536) # Replace with actual embedding call
@property
def hit_rate(self) -> float:
total = self._hits + self._misses
return self._hits / total if total > 0 else 0.0
@property
def stats(self) -> dict:
return {
"hits": self._hits,
"misses": self._misses,
"hit_rate": self.hit_rate,
"entries": len(self._cache),
"total_tokens_saved": sum(e.tokens_saved for e in self._cache.values()),
}
Deduplication Across Sub-Agents
When an orchestrator spawns multiple sub-agents working on parallel tasks, they often query the same external APIs or ask similar questions. A shared semantic cache at the orchestrator level prevents redundant work:
class SharedContextCache:
"""Cross-agent cache for shared knowledge."""
def __init__(self, semantic_cache: SemanticCache):
self.semantic_cache = semantic_cache
self._agent_registry: dict[str, dict] = {}
async def agent_query(
self, agent_id: str, question: str, compute_fn: Callable[[], str]
) -> str:
"""Query with cross-agent deduplication."""
result = await self.semantic_cache.get_or_compute(
prompt=f"[agent:{agent_id}] {question}",
model="shared",
compute_fn=compute_fn,
)
# Track what each agent has queried
if agent_id not in self._agent_registry:
self._agent_registry[agent_id] = {"queries": [], "cache_hits": 0}
self._agent_registry[agent_id]["queries"].append(question)
return result
def get_agent_stats(self, agent_id: str) -> dict:
return self._agent_registry.get(agent_id, {})
6. Pattern Five: Progressive Escalation
Concept
Progressive escalation is the agent equivalent of starting with a small investment and scaling up only when justified. Rather than committing to a full multi-step, multi-tool, premium-model execution from the start, the agent begins with the cheapest approach and escalates only when the initial result is insufficient.
The escalation ladder:
- Direct answer — Single LLM call with no tools
- Augmented — Single LLM call with tool retrieval (RAG)
- Reasoned — Multi-step reasoning with tool use
- Ensemble — Multiple model calls with synthesis
Implementation
class EscalationLevel(Enum):
DIRECT = 0 # One LLM call
AUGMENTED = 1 # RAG + one LLM call
REASONED = 2 # Multi-step with tools
ENSEMBLE = 3 # Multi-model synthesis
class ProgressiveEscalationAgent:
"""Escalates through complexity levels until result quality is sufficient."""
def __init__(self, model_client, retriever, quality_evaluator, budget: TokenBudget):
self.model_client = model_client
self.retriever = retriever
self.quality_evaluator = quality_evaluator
self.budget = budget
async def execute(self, task: str) -> dict:
last_result = None
for level in EscalationLevel:
if not self.budget.consume(0): # Check budget still available
return {"status": "budget_exhausted", "output": last_result}
# Execute at current escalation level
result = await self._execute_at_level(task, level)
last_result = result.get("output")
# Evaluate quality
quality = await self._evaluate_quality(task, result)
if quality >= self._threshold_for_level(level):
return {
"status": "completed",
"output": result.get("output"),
"escalation_level": level.name,
"quality_score": quality,
}
# Exhausted all levels — return best result
return {
"status": "max_escalation_reached",
"output": last_result,
"escalation_level": EscalationLevel.ENSEMBLE.name,
}
async def _execute_at_level(self, task: str, level: EscalationLevel) -> dict:
if level == EscalationLevel.DIRECT:
return await self._direct_answer(task)
elif level == EscalationLevel.AUGMENTED:
return await self._augmented_answer(task)
elif level == EscalationLevel.REASONED:
return await self._reasoned_answer(task)
elif level == EscalationLevel.ENSEMBLE:
return await self._ensemble_answer(task)
async def _direct_answer(self, task: str) -> dict:
response = await self.model_client.chat(
messages=[{"role": "user", "content": task}],
model="gpt-4o-mini",
max_tokens=500,
)
return {"output": response["content"], "model": "gpt-4o-mini"}
async def _augmented_answer(self, task: str) -> dict:
context = await self.retriever.search(task, top_k=5)
prompt = f"Use this context to answer:\n\nContext:\n{context}\n\nQuestion: {task}"
response = await self.model_client.chat(
messages=[{"role": "user", "content": prompt}],
model="gpt-4o",
max_tokens=1000,
)
return {"output": response["content"], "model": "gpt-4o", "context_used": True}
async def _reasoned_answer(self, task: str) -> dict:
# Multi-step with tool use
agent = BudgetEnforcedAgent(
self.model_client,
TokenBudget(total_budget=self.budget.total_budget // 3),
tools=self.model_client.get_tools(),
)
result = await agent.run(task, max_iterations=6)
return {"output": result.get("output"), "iterations": result.get("iterations")}
async def _ensemble_answer(self, task: str) -> dict:
# Multiple models, synthesized
responses = await asyncio.gather(
self._direct_answer(task),
self._augmented_answer(task),
)
synthesis_prompt = f"""Two responses to the same question:
Response 1: {responses[0].get('output')}
Response 2: {responses[1].get('output')}
Question: {task}
Synthesize the best answer, noting where responses agree and disagree."""
response = await self.model_client.chat(
messages=[{"role": "user", "content": synthesis_prompt}],
model="claude-sonnet",
max_tokens=1500,
)
return {"output": response["content"], "model": "claude-sonnet", "ensemble": True}
def _threshold_for_level(self, level: EscalationLevel) -> float:
thresholds = {
EscalationLevel.DIRECT: 0.85,
EscalationLevel.AUGMENTED: 0.80,
EscalationLevel.REASONED: 0.75,
EscalationLevel.ENSEMBLE: 0.70,
}
return thresholds[level]
async def _evaluate_quality(self, task: str, result: dict) -> float:
"""Evaluate result quality using a lightweight scoring model."""
eval_prompt = f"""Score this response from 0.0 to 1.0 for quality, accuracy, and completeness.
Question: {task}
Response: {result.get('output', '')}
Return ONLY a number between 0.0 and 1.0."""
response = await self.model_client.chat(
messages=[{"role": "user", "content": eval_prompt}],
model="gpt-4o-mini",
max_tokens=10,
)
try:
return float(response["content"].strip())
except ValueError:
return 0.0
The quality evaluator itself uses a cheap model, so the cost of evaluation is minimal. The key economic principle: spend $0.001 evaluating to decide whether to spend $0.05 or $0.50 on a more expensive approach.
7. Reference Architecture: Putting It All Together
All five patterns compose into a layered architecture:
┌──────────────────────────────────────────────────────────────┐
│ Client / User │
└─────────────────────────┬────────────────────────────────────┘
│
┌─────────────────────────▼────────────────────────────────────┐
│ Request Gateway + Rate Limiter │
│ (per-user/per-org budgets, request dedup, auth) │
└─────────────────────────┬────────────────────────────────────┘
│
┌─────────────────────────▼────────────────────────────────────┐
│ Task Classifier (Micro/Standard/Extended) │
│ (lightweight model call, tier assignment) │
└─────────────────────────┬────────────────────────────────────┘
│
┌─────────────────────────▼────────────────────────────────────┐
│ Progressive Escalation Engine │
│ (starts cheap, escalates on quality failure) │
└─────────────────────────┬────────────────────────────────────┘
│
┌─────────────────────────▼────────────────────────────────────┐
│ Token Budget + Circuit Breaker │
│ (global budget, velocity limits, state machine) │
└─────────────────────────┬────────────────────────────────────┘
│
┌───────────────┼───────────────┐
▼ ▼ ▼
┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ Semantic │ │ Sandboxed │ │ Sub-Agent │
│ Cache │ │ Tool Registry│ │ Spawner │
│ (dedup) │ │ (limits) │ │ (concurrency)│
└──────┬───────┘ └──────┬───────┘ └──────┬───────┘
│ │ │
▼ ▼ ▼
┌────────────────────────────────────────────────────────────┐
│ LLM Provider(s) + Tool Backends │
└────────────────────────────────────────────────────────────┘
The request flows through each layer, and each layer has the authority to short-circuit, reduce scope, or reject the request. This defense-in-depth approach ensures that no single failure mode can cause unbounded spending.
8. Implementation: The Sandbox Orchestrator
Here's the complete orchestrator that wires all patterns together:
import asyncio
import time
import uuid
from dataclasses import dataclass, field
from typing import Any, Callable, Optional
@dataclass
class AgentRequest:
id: str = field(default_factory=lambda: str(uuid.uuid4()))
user_id: str = ""
task: str = ""
context: dict = field(default_factory=dict)
priority: str = "normal" # low, normal, high, critical
created_at: float = field(default_factory=time.time)
@dataclass
class AgentResult:
request_id: str
status: str # completed, budget_exhausted, timeout, error
output: Optional[str] = None
tier: str = ""
escalation_level: str = ""
tokens_used: int = 0
tools_called: int = 0
elapsed_seconds: float = 0.0
cost_estimate_usd: float = 0.0
metadata: dict = field(default_factory=dict)
class SandboxOrchestrator:
"""Main orchestrator combining all budget-control patterns."""
def __init__(
self,
model_client,
retriever,
tool_registry: ToolRegistry,
semantic_cache: SemanticCache,
org_budget_per_month: float = 1000.0,
):
self.model_client = model_client
self.retriever = retriever
self.tool_registry = tool_registry
self.semantic_cache = semantic_cache
self.org_budget_per_month = org_budget_per_month
self.tier_classifier = TierClassifier(model_client)
self._monthly_spent = 0.0
self._request_count = 0
self._cost_log: list[dict] = []
async def process(self, request: AgentRequest) -> AgentResult:
start = time.time()
# Step 1: Check organizational budget
if self._monthly_spent >= self.org_budget_per_month:
return AgentResult(
request_id=request.id,
status="org_budget_exhausted",
metadata={"monthly_spent": self._monthly_spent, "limit": self.org_budget_per_month},
)
# Step 2: Check semantic cache
cached = await self.semantic_cache.get_or_compute(
prompt=request.task,
model="cache-check",
compute_fn=lambda: "CACHE_MISS",
)
if cached != "CACHE_MISS":
return AgentResult(
request_id=request.id,
status="completed",
output=cached,
metadata={"cache_hit": True, "elapsed": time.time() - start},
)
# Step 3: Classify tier
tier = self.tier_classifier.classify(request.task, request.context)
tier_config = TIER_REGISTRY[tier]
# Step 4: Create per-request token budget
# Scale budget by priority
priority_multiplier = {"low": 0.5, "normal": 1.0, "high": 1.5, "critical": 2.0}
request_budget = TokenBudget(
total_budget=int(tier_config.max_tokens_total * priority_multiplier.get(request.priority, 1.0)),
max_velocity_tokens_per_second=200 if tier == ComputeTier.MICRO else 500,
)
# Step 5: Execute with progressive escalation
escalation_agent = ProgressiveEscalationAgent(
model_client=self.model_client,
retriever=self.retriever,
quality_evaluator=self.model_client,
budget=request_budget,
)
try:
result = await asyncio.wait_for(
escalation_agent.execute(request.task),
timeout=tier_config.timeout_seconds,
)
except asyncio.TimeoutError:
result = {"status": "timeout", "output": None}
# Step 6: Calculate and track cost
elapsed = time.time() - start
cost = self._estimate_cost(request_budget.consumed, tier)
self._monthly_spent += cost
self._request_count += 1
self._cost_log.append({
"request_id": request.id,
"user_id": request.user_id,
"cost": cost,
"tokens": request_budget.consumed,
"tier": tier.value,
"elapsed": elapsed,
})
return AgentResult(
request_id=request.id,
status=result.get("status", "error"),
output=result.get("output"),
tier=tier.value,
escalation_level=result.get("escalation_level", ""),
tokens_used=request_budget.consumed,
tools_called=self.tool_registry.get_all_stats(),
elapsed_seconds=elapsed,
cost_estimate_usd=cost,
metadata=result.get("metadata", {}),
)
def _estimate_cost(self, tokens: int, tier: ComputeTier) -> float:
"""Rough cost estimation based on tier's typical model."""
# Approximate pricing (update for your provider)
cost_per_1k_tokens = {
ComputeTier.MICRO: 0.0006, # gpt-4o-mini
ComputeTier.STANDARD: 0.005, # gpt-4o
ComputeTier.EXTENDED: 0.015, # claude-sonnet
}
return (tokens / 1000) * cost_per_1k_tokens.get(tier, 0.005)
def get_budget_report(self) -> dict:
"""Generate budget utilization report."""
avg_cost = sum(c["cost"] for c in self._cost_log) / max(len(self._cost_log), 1)
return {
"monthly_spent": self._monthly_spent,
"monthly_limit": self.org_budget_per_month,
"utilization_pct": (self._monthly_spent / self.org_budget_per_month) * 100,
"total_requests": self._request_count,
"avg_cost_per_request": avg_cost,
"cache_hit_rate": self.semantic_cache.hit_rate,
"projected_monthly_cost": avg_cost * 30 * 24 * 60, # naive projection
}
9. Operational Metrics and Alerting
Budget control is useless without observability. These are the metrics you should instrument and alert on:
Key Metrics
| Metric | Alert Threshold | Action |
|---|---|---|
agent.tokens_per_request.p95 | > 2x median | Investigate runaway loops |
agent.budget_exhausted.rate | > 5% of requests | Review tier classification |
agent.tool_call_failures.rate | > 10% | Check tool health |
agent.circuit_breaker.open.count | Any | Immediate investigation |
org.monthly_budget.utilization | > 80% | Warn stakeholders |
semantic_cache.hit_rate | < 10% | Cache may need tuning |
agent.velocity.throttled.count | Any | Check for retry storms |
Alerting Configuration Example
# Prometheus alerting rules for agent budget monitoring
groups:
- name: agent-budget-alerts
rules:
- alert: HighBudgetUtilization
expr: |
agent_monthly_spend / agent_monthly_budget > 0.80
for: 5m
labels:
severity: warning
annotations:
summary: "Agent monthly budget at {{ $value | humanizePercentage }}"
- alert: RunawayAgentLoop
expr: |
rate(agent_tokens_consumed_total[5m]) > 2000
for: 2m
labels:
severity: critical
annotations:
summary: "Token velocity exceeded: {{ $value }} tokens/sec"
- alert: BudgetExhaustionRate
expr: |
rate(agent_budget_exhausted_total[1h]) / rate(agent_requests_total[1h]) > 0.05
for: 10m
labels:
severity: warning
annotations:
summary: "5%+ of requests hitting budget limits"
Cost Attribution Dashboard
For multi-tenant systems, cost attribution per user/team/project is essential. The _cost_log in the orchestrator provides the raw data. Aggregate it into a dashboard that shows:
- Cost per user, per day
- Cost per feature/workflow
- Token consumption trends over time
- Cache effectiveness by request type
- Tier distribution (what % of requests hit each tier)
10. Frequently Asked Questions
How do I set the right token budget per tier?
Start by profiling your actual workloads. Run 100-200 representative requests through each tier and record the p50, p90, and p99 token consumption. Set your tier budgets at approximately the p95 of actual usage. This allows 95% of requests to complete within budget while containing the 5% of outliers that would otherwise spiral. Adjust over time as your workload distribution shifts.
What's the impact of semantic caching on answer quality?
With a similarity threshold of 0.92+, semantic caching has minimal quality impact because it only returns cached results for nearly identical queries. The main risk is returning a stale answer when the underlying data has changed. Mitigate this with TTL-based expiration (1 hour for volatile data, 24 hours for stable data) and by excluding queries that contain timestamps or version-specific references from caching.
Should I use these patterns for internal tools or only customer-facing apps?
Use them everywhere. Internal tools often have less budget scrutiny, which makes them the most likely place for silent cost overruns to accumulate. An internal agent that runs unattended overnight with no budget limits can generate more waste than a customer-facing agent with per-request limits. The same architectural patterns apply; the only difference is that internal tools can afford higher per-request budgets and more generous escalation thresholds since they're not competing for a shared organizational budget with paying customers.
The fundamental principle across all five patterns is defense in depth for compute resources. No single mechanism is sufficient — the circuit breaker catches velocity spikes, tier classification prevents over-provisioning, semantic caching eliminates redundancy, progressive escalation minimizes unnecessary complexity, and sandboxed tool execution contains individual failure modes. Together, they form a system where agent autonomy and budget safety coexist.
For more patterns and architectural deep dives on AI systems, see Tamiz's Insights and explore the broader engineering landscape at tamiz.pro.