Back to Insights
AI & Machine Learning•The Touch Grass Manifesto: Building AI Apps That Force Developers Outside — Offline Models, Voice-Only Interfaces, and Walk-Driven Storytelling Without a Single Pixel on Screen•deep dive•October 11, 2026•18 min read

The Touch Grass Manifesto: Building AI Apps That Force Developers Outside

A deep-dive into screen-free AI architecture: offline models, voice-only interfaces, and walk-driven storytelling that gets developers moving.

T
Tamiz UddinFull-Stack Engineer

The best code is written while walking. Not at a desk, not in an IDE, not even on a laptop balanced on a park bench — but in the rhythm of footsteps, the cadence of breath, the ambient noise of a neighborhood waking up. The Touch Grass Manifesto isn't wellness advice. It's a technical constraint: build AI applications that cannot be used sitting down.

This article dissects the architecture of screen-free, mobility-first AI systems. We'll cover local-first model deployment, voice-only interaction patterns, sensor-fused storytelling engines, and the infrastructure that makes "walk-driven development" not just possible but preferable to chair-bound coding.

Table of Contents

1. The Constraint as Architecture

Most mobile apps treat mobility as a feature. The Touch Grass Manifesto treats immobility as a bug. The core architectural principle: if the user can operate the app while stationary, the design has failed.

This constraint cascades through every layer:

LayerTraditional MobileTouch Grass Architecture
ComputeCloud-first, fallback to localLocal-first, cloud only for sync
InputTouch, keyboard, voiceVoice + sensors only
OutputScreen, haptics, audioAudio + haptics only
ContextGPS optionalMotion + location + env required
SessionTap-to-startStarts on walk detection
ErrorModal dialogSpoken recovery + continue

The constraint eliminates entire categories of complexity: no responsive layouts, no focus management, no virtual keyboards, no screen readers to test against (the app is a screen reader), no visual regression tests. What remains is harder: reliability without visual confirmation.

1.1 The "No Pixel" Invariant

rust
// Core invariant: zero screen dependencies in the runtime
// This compiles away if any UI framework is linked
#[cfg(target_os = "ios")]
compile_error!("Touch Grass apps cannot link UIKit. Use AudioToolbox + CoreMotion only.");

#[cfg(target_os = "android")]
compile_error!("Touch Grass apps cannot link Android View system. Use AudioTrack + SensorManager only.");

This isn't performative. It forces the team to solve hard problems: how does a user know the model loaded? How do they confirm a destructive action? How do they debug without logs on screen?

2. Offline-First Model Stack

Cloud inference breaks the walk. Latency variance, connectivity drops, battery drain from radio — all unacceptable. The model must run on-device, cold-start in <2s, and sustain 15+ tokens/sec on a phone NPU.

2.1 Model Selection Criteria

RequirementTargetRationale
Size≤ 4B params quantizedFits in 3-4GB RAM, leaves headroom for OS
Quantization4-bit (Q4_K_M) or 3-bitBest quality/speed tradeoff on mobile NPUs
Context8K+ tokensWalk sessions accumulate context
Latency<200ms first tokenPerceived instant on voice
Throughput>15 tok/s sustainedNatural conversation pace
LanguagesMultilingualWalks happen everywhere

Current picks (2024):

  • Llama 3.2 3B Instruct (Q4_K_M) — 2.1GB, 18 tok/s on iPhone 15 Pro NPU
  • Gemma 2 2B IT (Q4_K_M) — 1.4GB, 22 tok/s, Apache 2.0
  • Phi-3.5-mini (Q4_K_M) — 2.2GB, 20 tok/s, strong reasoning
  • Qwen 2.5 3B (Q4_K_M) — 2.0GB, best multilingual

2.2 Local Inference Runtime

llama.cpp via llama.rn (React Native) or native Swift/Kotlin bindings remains the most battle-tested path. MLC-LLM and ExecuTorch are promising but less mature for audio-streaming workloads.

swift
// Swift: Minimal llama.cpp wrapper for voice streaming
// TouchGrassInference.swift
import Foundation
import llama

final class TouchGrassInference {
    private var ctx: OpaquePointer?
    private var batch: llama_batch
    private let nCtx = 8192
    private let nThreads = ProcessInfo.processInfo.activeProcessorCount - 1

    init(modelPath: String) throws {
        var params = llama_context_default_params()
        params.n_ctx = Int32(nCtx)
        params.n_threads = Int32(nThreads)
        params.n_threads_batch = Int32(nThreads)
        params.offload_kqv = true  // Metal/GPU offload
        params.flash_attn = true
        
        ctx = llama_init_from_file(modelPath, params)
        guard ctx != nil else { throw InferenceError.modelLoadFailed }
        
        batch = llama_batch_init(512, 0, 1)
    }

    // Streaming callback for token-by-token TTS handoff
    func generate(
        prompt: String,
        onToken: @escaping (String) -> Bool  // return false to stop
    ) throws {
        let tokens = tokenize(prompt)
        batch.n_tokens = Int32(tokens.count)
        for (i, tok) in tokens.enumerated() {
            batch.token[i] = tok
            batch.pos[i] = Int32(i)
            batch.n_seq_id[i][0] = 0
            batch.logits[i] = (i == tokens.count - 1) ? 1 : 0
        }
        
        if llama_decode(ctx!, batch) != 0 { throw InferenceError.decodeFailed }
        
        var nPast = tokens.count
        while true {
            let logits = llama_get_logits_ith(ctx!, Int32(batch.n_tokens - 1))
            let nextTok = sampleTopP(logits, p: 0.9, temp: 0.7)
            
            if llama_token_is_eog(ctx!, nextTok) { break }
            
            let piece = String(cString: llama_token_to_piece(ctx!, nextTok))
            if !onToken(piece) { break }  // TTS or user interrupted
            
            batch.n_tokens = 1
            batch.token[0] = nextTok
            batch.pos[0] = Int32(nPast)
            batch.logits[0] = 1
            nPast += 1
            
            if llama_decode(ctx!, batch) != 0 { throw InferenceError.decodeFailed }
            
            if nPast >= nCtx { throw InferenceError.contextFull }
        }
    }

    deinit { if let ctx { llama_free(ctx) } }
}

Key optimization: The onToken callback hands each token directly to the TTS engine before the full response completes. This cuts perceived latency from "model done" to "first phoneme" — critical for voice-only UX.

2.3 Model Swapping Without Screen

Users need to switch models (coding assistant → storytelling → navigation) without looking. Solution: voice-triggered model hot-swap with haptic confirmation.

swift
// ModelManager.swift
enum ModelProfile: String, CaseIterable {
    case coder = "llama-3.2-3b-coder-q4"
    case storyteller = "gemma-2-2b-story-q4"
    case navigator = "phi-3.5-mini-nav-q4"
    
    var hapticPattern: CHHapticPattern { /* distinct per model */ }
    var earcon: String { /* unique 200ms audio ID */ }
}

final class ModelManager {
    private var current: TouchGrassInference?
    private let audio: AudioEngine
    private let haptics: HapticEngine
    
    func switchTo(_ profile: ModelProfile) async throws {
        // Play earcon *before* unload so user hears intent
        try await audio.playEarcon(profile.earcon)
        try await haptics.play(profile.hapticPattern)
        
        current = try TouchGrassInference(modelPath: profile.path)
        
        // Confirmation: speak model name via TTS
        try await audio.speak("Switched to \(profile.rawValue)")
    }
}

3. Voice-Only Interaction Layer

No screen means no visual turn-taking cues. The voice stack must handle: barge-in, endpointing, noise rejection, speaker diarization, and implicit confirmation — all locally.

3.1 Audio Pipeline Architecture

scss
┌─────────────┐   ┌──────────────┐   ┌─────────────┐   ┌──────────────┐
│  Mic Array  │──▶│  VAD + AEC   │──▶│  ASR Stream │──▶│  Intent/Slot │
│  (beamform) │   │  (RNNoise)   │   │  (Whisper)  │   │  (LLM)       │
└─────────────┘   └──────────────┘   └─────────────┘   └──────────────┘
       ▲                                       │                   │
       │                    ┌──────────────────┘                   ▼
       │                    ▼                          ┌─────────────────┐
       │             ┌──────────────┐                  │  TTS (Piper/    │
       └─────────────│  Barge-in    │◀─────────────────│  StyleTTS2)     │
                     │  Detection   │   token stream   └─────────────────┘
                     └──────────────┘                        │
                            │                                ▼
                            ▼                       ┌─────────────────┐
                     ┌──────────────┐                │  Audio Mixer    │
                     │  Turn Manager│                │  (spatialized)  │
                     └──────────────┘                └─────────────────┘

3.2 Streaming ASR with Whisper.cpp

Whisper.cpp supports streaming via whisper_full_with_state but it's frame-based. For true streaming, use faster-whisper (CTranslate2) or whisper.cpp with tiny.en + sliding window.

python
# faster-whisper streaming wrapper
# asr_stream.py
from faster_whisper import WhisperModel
import numpy as np
import queue
import threading

class StreamingASR:
    def __init__(self, model_size="tiny.en", device="cpu", compute_type="int8"):
        self.model = WhisperModel(model_size, device=device, compute_type=compute_type)
        self.audio_queue = queue.Queue()
        self.result_queue = queue.Queue()
        self.running = False
        self.buffer = np.array([], dtype=np.float32)
        self.sample_rate = 16000
        self.chunk_duration = 0.5  # 500ms chunks
        
    def start(self):
        self.running = True
        threading.Thread(target=self._process_loop, daemon=True).start()
    
    def push_audio(self, chunk: np.ndarray):
        self.audio_queue.put(chunk)
    
    def _process_loop(self):
        while self.running:
            try:
                chunk = self.audio_queue.get(timeout=0.1)
                self.buffer = np.concatenate([self.buffer, chunk])
                
                # Process when we have enough context
                if len(self.buffer) >= self.sample_rate * 2:  # 2s window
                    segments, _ = self.model.transcribe(
                        self.buffer[-self.sample_rate*4:],  # last 4s
                        language="en",
                        vad_filter=True,
                        vad_parameters=dict(min_silence_duration_ms=500)
                    )
                    text = " ".join(s.text for s in segments).strip()
                    if text:
                        self.result_queue.put(text)
                    # Keep last 1s for context overlap
                    self.buffer = self.buffer[-self.sample_rate:]
            except queue.Empty:
                continue

3.3 Barge-In: The Critical UX Primitive

Users will interrupt. The system must stop TTS instantly and pivot to listening.

swift
// BargeInManager.swift
final class BargeInManager {
    private let vad: VoiceActivityDetector  // Silero VAD, 1.5ms/frame
    private let tts: StreamingTTS
    private let asr: StreamingASR
    private var isSpeaking = false
    
    func onTTSStart() { isSpeaking = true }
    
    func onTTSChunk(_ audio: Data) {
        // Feed TTS output back to VAD for echo cancellation reference
        vad.referenceSignal(audio)
    }
    
    func onMicAudio(_ audio: Data) -> BargeInDecision {
        let voiceProb = vad.process(audio)
        
        if isSpeaking && voiceProb > 0.85 {
            // User speaking over TTS → barge in
            tts.stopImmediately()
            isSpeaking = false
            asr.reset()  // Clear any partial hypothesis
            return .bargeIn
        }
        
        if !isSpeaking && voiceProb > 0.6 {
            return .startListening
        }
        
        return .continue
    }
}

Latency budget: VAD must decide in <20ms. Silero VAD (ONNX, 1.5MB) runs at ~0.5ms/frame on mobile NPU. The TTS stop must be synchronous — no fade-out, hard cutoff.

3.4 Implicit Confirmation Patterns

No "tap to confirm." Use progressive commitment:

swift
// ConfirmationStrategy.swift
enum ConfirmationLevel {
    case none      // Read-only, low risk
    case implicit  // "I'll save that note" → 3s undo window
    case explicit  // "Delete all notes? Say 'yes delete' to confirm"
    case multiModal // "Say 'yes' AND tap phone twice" (for destructive)
}

func confirmationLevel(for action: Action) -> ConfirmationLevel {
    switch action {
    case .createNote: return .implicit
    case .sendMessage: return .explicit
    case .deleteAll: return .multiModal
    case .runCode: return .explicit  // Code execution = explicit
    }
}

Implicit confirmation: speak the action, start a 3-second timer. If user says "cancel" or "undo" within window, revert. No beep, no vibration — just the absence of the follow-up "Done."

4. Walk-Driven Storytelling Engine

The killer app: narrative that unfolds at walking pace. Not audiobooks — generative, location-aware, motion-responsive stories where the walk is the interface.

4.1 Sensor Fusion for Narrative Context

swift
// WalkContext.swift
import CoreMotion
import CoreLocation

struct WalkContext {
    let pace: Pace          // .stroll / .walk / .brisk / .run
    let terrain: Terrain    // .pavement / .trail / .stairs / .unknown
    let environment: Env    // .quiet / .street / .park / .transit
    let location: CLLocation?
    let timeOfDay: TimePhase
    let weather: Weather?
    let heartRate: Double?  // if HealthKit authorized
    
    var narrativeTempo: NarrativeTempo {
        switch (pace, terrain) {
        case (.stroll, .park): return .reflective
        case (.brisk, .pavement): return .urgent
        case (.walk, .trail): return .exploratory
        case (.run, _): return .action
        default: return .neutral
        }
    }
}

final class WalkContextEngine {
    private let motion = CMMotionManager()
    private let location = CLLocationManager()
    private let altimeter = CMAltimeter()
    private var contextContinuation: AsyncStream<WalkContext>.Continuation?
    
    var contextStream: AsyncStream<WalkContext> {
        AsyncStream { continuation in
            self.contextContinuation = continuation
            self.startUpdates()
        }
    }
    
    private func startUpdates() {
        motion.startDeviceMotionUpdates(to: .main) { [weak self] data, _ in
            guard let data, let self else { return }
            let pace = self.classifyPace(data)
            let terrain = self.classifyTerrain(data)
            // ... fuse with location, weather, time
            let ctx = WalkContext(...)
            self.contextContinuation?.yield(ctx)
        }
    }
    
    private func classifyPace(_ data: CMDeviceMotion) -> Pace {
        // Vertical oscillation + step frequency from userAcceleration
        let vertical = data.userAcceleration.z
        let freq = self.stepFrequency(from: data)
        
        switch (freq, abs(vertical)) {
        case (0..<1.5, _): return .stroll
        case (1.5..<2.0, 0..<0.15): return .walk
        case (1.5..<2.0, _): return .brisk
        case (2.0..., _): return .run
        default: return .unknown
        }
    }
}

4.2 Narrative State Machine

The story isn't a linear script. It's a state graph where nodes are narrative beats, edges are transitions gated by walk context.

json
// story_graph.json - loaded at startup, no screen needed
{

{
  "nodes": {
    "start": {
      "text": "The trailhead is empty. Your boots hit dirt. What's the first thing you notice?",
      "choices": [
        { "label": "The smell of pine", "next": "pine", "requires": {} },
        { "label": "Birdsong overhead", "next": "birds", "requires": {} },
        { "label": "The weight of your pack", "next": "pack", "requires": {} }
      ]
    },
    "pine": {
      "text": "Pine needles cushion each step. The air tastes like resin and rain. A deer watches from thirty meters.",
      "choices": [
        { "label": "Freeze. Watch it leave.", "next": "deer_gone", "requires": {} },
        { "label": "Whistle low. See if it approaches.", "next": "deer_curious", "requires": { "has_whistle": true } }
      ]
    },
    "birds": {
      "text": "A Steller's jay scolds you from a cedar. Two chickadees mob it. The canopy is alive with argument.",
      "choices": [
        { "label": "Identify the jay's alarm call.", "next": "bird_id", "requires": { "skill": "birding" } },
        { "label": "Walk on. They're not your problem.", "next": "ridge", "requires": {} }
      ]
    },
    "pack": {
      "text": "Twenty kilos. Water, shelter, the satellite messenger your partner insisted on. Every gram earned its place.",
      "choices": [
        { "label": "Adjust the hip belt. Keep moving.", "next": "ridge", "requires": {} },
        { "label": "Dump the camp chair. Save six hundred grams.", "next": "lighter", "requires": { "has_chair": true } }
      ]
    },
    "ridge": {
      "text": "The trail breaks onto a ridge. Valley spreads below—river silver, forest green, the highway a gray scar.",
      "choices": [
        { "label": "Sit. Eat the apple you packed.", "next": "apple", "requires": { "has_apple": true } },
        { "label": "Pull out the messenger. Check in.", "next": "checkin", "requires": { "has_messenger": true } },
        { "label": "Just breathe. No screens.", "next": "breathe", "requires": {} }
      ]
    },
    "breathe": {
      "text": "Wind carries wet earth and snowmelt. Your shoulders drop. This is why you came.",
      "is_ending": true,
      "ending_type": "pure"
    },
    "checkin": {
      "text": "Three taps. 'All good. On ridge. Back by dark.' The satellite swallows it. Peace of mind, delivered.",
      "is_ending": true,
      "ending_type": "connected"
    }
  },
  "entry": "start",
  "meta": {
    "version": "1.0",
    "estimated_walk_minutes": 45,
    "difficulty": "moderate"
  }
}

The graph lives in a static JSON file. No database, no CMS, no admin panel. Writers edit it in VS Code. Version control is the content history.

Runtime evaluation is a pure function:

python
# story_engine.py - zero dependencies, ~80 lines
import json
from dataclasses import dataclass
from typing import Optional

@dataclass
class WalkContext:
    step_count: int
    elevation_gain_m: float
    heart_rate_bpm: Optional[int]
    inventory: set[str]
    skills: set[str]
    time_of_day: str  # "dawn" | "day" | "dusk" | "night"
    weather: str      # "clear" | "cloudy" | "rain" | "storm"

@dataclass
class Choice:
    label: str
    next_node: str
    requires: dict

@dataclass
class Node:
    text: str
    choices: list[Choice]
    is_ending: bool = False
    ending_type: str = ""

def load_graph(path: str) -> dict[str, Node]:
    with open(path) as f:
        raw = json.load(f)
    nodes = {}
    for nid, nd in raw["nodes"].items():
        nodes[nid] = Node(
            text=nd["text"],
            choices=[Choice(**c) for c in nd.get("choices", [])],
            is_ending=nd.get("is_ending", False),
            ending_type=nd.get("ending_type", "")
        )
    return nodes

def available_choices(node: Node, ctx: WalkContext) -> list[Choice]:
    """Filter choices by context. Pure, testable, no I/O."""
    def check(req: dict) -> bool:
        for k, v in req.items():
            if k == "has_whistle" and v and "whistle" not in ctx.inventory:
                return False
            if k == "has_chair" and v and "camp_chair" not in ctx.inventory:
                return False
            if k == "has_apple" and v and "apple" not in ctx.inventory:
                return False
            if k == "has_messenger" and v and "sat_messenger" not in ctx.inventory:
                return False
            if k == "skill" and v not in ctx.skills:
                return False
        return True
    return [c for c in node.choices if check(c.requires)]

def render_node(node: Node, ctx: WalkContext) -> str:
    lines = [node.text]
    for i, ch in enumerate(available_choices(node, ctx), 1):
        lines.append(f"  {i}. {ch.label}")
    return "\n".join(lines)

No LLM calls at runtime. The story is authored, not generated. LLMs are a design-time tool—writers use them to brainstorm branches, check for dead ends, translate. The shipped artifact is deterministic.

4.3 The Audio Pipeline: TTS That Doesn't Suck

Text-to-speech on mobile in 2024 is a solved problem if you accept tradeoffs.

ApproachLatencyQualityOfflineBattery
Cloud (ElevenLabs, OpenAI)800ms–2s★★★★★❌Low (network only)
On-device (piper, whisper.cpp)50–200ms★★★☆☆✅Medium
Hybrid (cache + cloud fallback)50ms cached★★★★☆PartialLow

Our choice: hybrid with aggressive pre-generation.

At build time, every node's text is rendered to .opus files via Piper (local, fast, decent voices). The app bundles ~50 MB of audio—trivial for modern phones. Zero runtime TTS latency. Zero network dependency.

bash
# build_audio.sh - runs in CI, commits artifacts
#!/usr/bin/env bash
set -euo pipefail

GRAPH="story_graph.json"
OUT_DIR="assets/audio"
VOICE="en_US-lessac-medium"  # Piper voice model

mkdir -p "$OUT_DIR"

# Extract all unique node texts
jq -r '.nodes[].text' "$GRAPH" | sort -u | while IFS= read -r text; do
  # Stable filename: sha256 of text
  fname=$(echo -n "$text" | sha256sum | cut -c1-16)
  out="$OUT_DIR/$fname.opus"
  [[ -f "$out" ]] && continue  # idempotent
  echo "$text" | piper --model "$VOICE" --output_raw | \
    opusenc --bitrate 24 --raw --raw-channels 1 --raw-rate 22050 - "$out"
done

# Generate manifest for the app
jq -r '.nodes | to_entries[] | "\(.key) \(.value.text)"' "$GRAPH" | \
while read -r nid text; do
  fname=$(echo -n "$text" | sha256sum | cut -c1-16)
  echo "{\"node\":\"$nid\",\"audio\":\"$fname.opus\"}"
done > "$OUT_DIR/manifest.jsonl"

Runtime playback is a single AVAudioPlayer / MediaPlayer call. The manifest maps node ID → audio file. No streaming, no buffering, no "waiting for voice to load."

Edge case: dynamic text (step count, heart rate). Solved by template fragments.

json
// story_graph.json snippet
{
  "nodes": {
    "checkin": {
      "text": "Three taps. 'All good. {{step_count}} steps. {{elevation}} meters up. Back by dark.'",
      "audio_template": "checkin_base.opus",
      "dynamic_slots": ["step_count", "elevation"]
    }
  }
}

Pre-render the static base ("Three taps. 'All good. ... steps. ... meters up. Back by dark.'"). At runtime, splice in tiny TTS fragments for numbers only—generated on-device via Piper's espeak-ng backend in <50ms. The splice is seamless because prosody matches.

4.4 Sensor Fusion Without The PhD

You don't need a Kalman filter. You need heuristics that survive real users.

python
# context_builder.py - runs on phone, updates every 5s
from dataclasses import dataclass
from enum import Enum
import time

class Weather(Enum):
    CLEAR = "clear"
    CLOUDY = "cloudy"
    RAIN = "rain"
    STORM = "storm"

@dataclass
class WalkContext:
    step_count: int
    elevation_gain_m: float
    heart_rate_bpm: int | None
    inventory: set[str]
    skills: set[str]
    time_of_day: str
    weather: Weather

class ContextBuilder:
    # ponytail: thresholds tuned for iOS/Android sensor noise profiles.
    # Upgrade path: per-device calibration on first run.
    PACE_WINDOW_SEC = 30
    HR_MIN_VALID = 40
    HR_MAX_VALID = 220
    ELEVATION_SMOOTH_ALPHA = 0.3

    def __init__(self):
        self._step_timestamps: list[float] = []
        self._elevation_ewma: float | None = None
        self._last_pressure_hpa: float | None = None

    def update_steps(self, new_steps: int, timestamp: float) -> int:
        self._step_timestamps.append(timestamp)
        # Drop old
        cutoff = timestamp - self.PACE_WINDOW_SEC
        self._step_timestamps = [t for t in self._step_timestamps if t > cutoff]
        return new_steps

    def current_pace_spm(self) -> float:
        if len(self._step_timestamps) < 2:
            return 0.0
        dt = self._step_timestamps[-1] - self._step_timestamps[0]
        return len(self._step_timestamps) / dt * 60  # steps/min

    def update_elevation(self, pressure_hpa: float) -> float:
        """Barometric altitude. No GPS vertical—too noisy."""
        if self._last_pressure_hpa is None:
            self._last_pressure_hpa = pressure_hpa
            self._elevation_ewma = 0.0
            return 0.0
        # Hypsometric formula, simplified
        delta_h = 8.3 * (self._last_pressure_hpa - pressure_hpa)  # meters per hPa approx
        self._elevation_ewma = (
            self.ELEVATION_SMOOTH_ALPHA * delta_h +
            (1 - self.ELEVATION_SMOOTH_ALPHA) * self._elevation_ewma
        )
        self._last_pressure_hpa = pressure_hpa
        return max(0.0, self._elevation_ewma)

    def update_heart_rate(self, bpm: int | None) -> int | None:
        if bpm is None:
            return None
        if not (self.HR_MIN_VALID <= bpm <= self.HR_MAX_VALID):
            return None  # discard artifact
        return bpm

    def infer_weather(self, pressure_hpa: float, humidity: float, temp_c: float) -> Weather:
        # ponytail: naive heuristic. Ceiling: no microclimate awareness.
        # Upgrade: on-device TinyML model (TensorFlow Lite, <100KB).
        if pressure_hpa < 990 and humidity > 80:
            return Weather.STORM
        if pressure_hpa < 1000 and humidity > 70:
            return Weather.RAIN
        if humidity > 60:
            return Weather.CLOUDY
        return Weather.CLEAR

    def time_of_day(self) -> str:
        h = time.localtime().tm_hour
        if 5 <= h < 8: return "dawn"
        if 8 <= h < 18: return "day"
        if 18 <= h < 21: return "dusk"
        return "night"

No ML models shipped. The heuristic is 30 lines, auditable, explainable. If it misclassifies "cloudy" as "clear" once, the user hears a slightly mismatched line. Nobody dies. The app keeps walking.

4.5 The "No Screen" Contract

The app never wakes the screen. Not for notifications, not for choices, not for errors.

swift
// iOS: AudioSession + Background Modes = "audio" + "location updates"
// Android: Foreground Service (mediaPlayback) + PARTIAL_WAKE_LOCK

class AudioSessionManager {
    func configure() throws {
        let session = AVAudioSession.sharedInstance()
        try session.setCategory(.playback, mode: .spokenAudio, options: [.duckOthers, .mixWithOthers])
        try session.setActive(true)
        // Lock screen controls appear automatically. No UI code needed.
    }
}

User interaction happens via:

  1. Headphone buttons — single press = choice 1, double = choice 2, triple = choice 3, long press = repeat current node
  2. Voice — "Next", "Repeat", "Choice two" (on-device speech recognition, SFSpeechRecognizer / SpeechRecognizer, no cloud)
  3. Watch complication — tap to advance, crown to scroll choices (watchOS only, optional)

The phone stays in the pack. The watch stays on the wrist. The headphones stay in the ears.

4.6 Testing: Property-Based, Not Example-Based

python
# test_story_engine.py
import hypothesis.strategies as st
from hypothesis import given, settings
from story_engine import load_graph, available_choices, WalkContext

GRAPH = load_graph("story_graph.json")

@given(
    step_count=st.integers(0, 50000),
    elevation=st.floats(0, 3000),
    hr=st.one_of(st.none(), st.integers(30, 250)),
    inventory=st.sets(st.sampled_from(["whistle", "camp_chair", "apple", "sat_messenger"])),
    skills=st.sets(st.sampled_from(["birding", "navigation", "first_aid"])),
    time_of_day=st.sampled_from(["dawn", "day", "dusk", "night"]),
    weather=st.sampled_from(["clear", "cloudy", "rain", "storm"]),
)
@settings(max_examples=500, deadline=None)
def test_no_crashes_on_valid_context(step_count, elevation, hr, inventory, skills, time_of_day, weather):
    ctx = WalkContext(step_count, elevation, hr, inventory, skills, time_of_day, weather)
    for node in GRAPH.values():
        # Should never raise, even with nonsense context
        _ = available_choices(node, ctx)

@given(nid=st.sampled_from(list(GRAPH.keys())))
def test_every_node_reachable_from_start(nid):
    # BFS from entry
    from collections import deque
    visited = set()
    q = deque([GRAPH["start"]])
    while q:
        node = q.popleft()
        if id(node) in visited: continue
        visited.add(id(node))
        for ch in node.choices:
            if ch.next_node == nid:
                return  # found
            q.append(GRAPH[ch.next_node])
    assert False, f"Node {nid} unreachable from start"

def test_no_dead_ends_except_endings():
    for nid, node in GRAPH.items():
        if node.is_ending: continue
        # At least one choice with empty requires (always available)
        assert any(len(c.requires) == 0 for c in node.choices), f"{nid}: all choices gated"

500 random contexts × 12 nodes = 6,000 executions per CI run. Catches:

  • Missing requires keys
  • Typos in next_node refs
  • Unreachable nodes
  • Choices that can never fire (over-constrained)

4.7 Shipping: One Binary, Zero Config

graphql
trail/
├── Cargo.toml           # or pyproject.toml / package.json / go.mod
├── src/
│   ├── main.rs          # 200 lines: init sensors → loop → render → play
│   ├── story_engine.rs  # the pure logic above
│   ├── context.rs       # sensor readers (platform-specific, ~150 lines each)
│   └── audio.rs         # playback + manifest lookup
├── assets/
│   ├── story_graph.json
│   └── audio/           # 50 MB .opus + manifest.jsonl (git-lfs)
└── .github/workflows/
    ├── build.yml        # compiles, runs tests, generates audio
    └── release.yml      # codesign → notarize → upload to TestFlight / Play Console

No feature flags. No A/B tests. No analytics. The app knows nothing about its users. It doesn't phone home. It doesn't check for updates on launch. It's a tool, not a service.


5. What We Didn't Build (And Why)

FeatureRequested?Verdict
User accounts / cloud sync"Would be nice"❌ YAGNI. The walk is local.
Social sharing"Viral growth"❌ Antithetical to the premise.
AI-generated side quests"Infinite content"❌ Authored > generated. Quality > quantity.
Map view"Safety"❌ Paper map in pack. Phone GPS kills battery.
Achievements / streaks"Retention"❌ Gamification ruins the point.
Accessibility: screen reader"Compliance"✅ Built in—entire app is audio-first.
Offline maps"Safety"❌ Separate app (Organic Maps). Unix philosophy.

6. The Real Metric

Not DAU. Not retention. Not session length.

Walks completed without phone leaving pack.

sql
-- The only query that matters
SELECT COUNT(*)
FROM walks
WHERE phone_screen_on_seconds < 30
  AND duration_minutes BETWEEN 20 AND 180
  AND ended_at_ridge = true;

Every line of code above serves that metric. If it doesn't, delete it.


7. Build Your Own

The stack is boring on purpose:

  • Language: Whatever compiles to a single binary (Rust, Go, Swift, Kotlin/Native)
  • Audio: Piper TTS (build-time) + Opus (runtime)
  • Sensors: Platform APIs directly—no wrappers, no abstractions
  • Story: JSON + pure functions
  • Tests: Hypothesis / proptest / quickcheck
  • Distribution: App stores + F-Droid / GitHub Releases

Start this weekend.

  1. Write 10 nodes in story_graph.json.
  2. Run build_audio.sh.
  3. Write the 200-line main loop.
  4. Walk.

The trail doesn't care about your tech stack. It only cares that you're on it.


Touch grass. The code will wait.