Bulkhead τ · Research Portfolio

Research Papers

Evidence-backed agent systems, deterministic substrates, and local LLM operator judgment.

Internally, this work is developed under Bulkhead Tau.

46Primary Papers
12Research Clusters
12Live Sites
2026Active

Operating front: these papers are the evidence base behind Bulkhead τ — the public release line for a deterministic governance substrate. Bulkhead Tau is the engineering core under that name; both surfaces will run in parallel for now, with /bulkhead-tau/ as the external front.

The finding: For grounded domain tasks — well-defined task classes with deterministic substrates — harness configuration is the binding constraint. Model identity is not. These papers prove this claim under stress test conditions: local models, which cannot compensate for a weak harness, converge with frontier models at the semantic usefulness level when the harness is sufficient. This scope is deliberate. Outside it, model capability matters in ways the framework does not cover.

Local Model Addendum: the Bulkhead Tau Local Model Details page covers the three supporting papers that feed the orchestration synthesis: TourAgent (1.13), ShowcaseAgent (1.12), and Local Model Role Suitability (1.11).

Boundary Results: the Bulkhead Tau Boundary Results page covers three papers that map where the organized stack hits its limits: Grounded Agent Failure Is Structurally Determined (1.10), True Ski Chalet Boundary Result (1.14), and When The Organized Stack Loses (1.15).

RVH / ML Evaluation: Rough Volatility as ML Benchmark covers Papers 1.8 and 1.9 — why domain expertise, not ML capability, is the binding constraint in rough volatility forecasting and the cross-domain benchmark principle it reveals.

Measurement Integrity, Operator Layer & Applied Evidence: Papers 1.16–1.19 extend the framework outward. Paper 1.16 shows that evaluation infrastructure can fail at the capture boundary — a VT100 terminal artifact was corrupting protocol scores for thinking-mode models. Paper 1.17 documents the operator shell pattern: how OpenClaw wraps Bulkhead Tau as an access layer without becoming the authority. Paper 1.18 is the framework's first numbered production case — PPR Agent, 92M regulated cardiac device implants across 18 years, behind a deterministic SQLite substrate. Paper 1.19 is a short companion to 1.16 on the other side of the apparatus: when stronger models override literal substrate inspection, capability itself becomes a source of non-neutrality. Papers 1.20–1.26, 1.30, 1.31, and 1.34–1.37 add the Local LLM Operator Judgment cluster: strict handoff discipline, privacy cost accounting, local model sizing, token-cost attention, substitution discipline, multi-agent HIL workflow boundaries, cross-audit failure cataloging, narration-surface failure analysis, deterministic validator discipline, and image-generation prompt discipline. Papers 1.39–1.45 extend the cluster: a negative result on heterogeneous local clustering, repository-fact encoding that converts cold-start architecture recovery into lookup, a governance case study admitting a new frontier harness as a peer through evidence rather than brand rank, a residency-verified hardware measurement showing dedicated VRAM beats unified memory for local inference, a code-graph tool's reduction ratio traced to graph density rather than code quality after two prior hypotheses were each falsified by a real contrasting corpus, an MTP draft-depth stub, and a draft R6 grid on local-model delegation (Syntax Wall measured; R3-on-grid still pending).

Sensor-to-Simulation Engineering: Paper 1.27 establishes the data landscape and fidelity boundaries for wearable sport sensors, showing why sensor-driven HIL is necessarily event-driven rather than waveform-driven. Paper 1.28 publishes the LabWired platform boundary for register-level firmware simulation. Paper 1.29 closes the sensor-driven HIL loop with a documented physical proximity replay. Paper 1.32 applies the sensor-corpus framing to a single dual-sensor-instrumented match at depth, Paper 1.33 defines the proof boundary separating harness testing from component verification, and Paper 1.46 measures the CPU/GPU simulation crossover across tiny sensor kernels, generic articulated scenes, and a measured-pose TennisAgent serve workload.

Framework  ·  confirmed
Local Inference & Offline Systems  ·  confirmed

Papers 1.2, 1.3, 1.5 form a cluster: offline grounded agent → ski chalet hardware boundary → TSP solver-backed orchestration. The common argument: harness level, not model size, drives usefulness.

Grounding  ·  confirmed
Orchestration & Role Assignment  ·  confirmed

Papers 1.5, 1.6, 1.11, 1.12, 1.13 address where correctness should live and how grounding, routing, and repair beat raw power in identifiable regimes.

Failure Modes & Boundary Conditions  ·  confirmed

Papers 1.7, 1.10, 1.14, 1.15 cover the failure taxonomy, empirical failure prediction, the true local ceiling, and the five conditions under which the organized stack's advantage collapses.

publishedpublished

Agentic Coding Failure Patterns

Agentic coding successes vary widely; failures recur in recognizable families. Drift, summit fever, bad context selection, false success, doom loops, and premature closure are documented across Bulkhead Tau operations. The practical response is standards, supervision, and lessons learned — not blind faith in scaling alone.

Open paper →
publishedpublished

Grounded Agent Failure Is Structurally Determined

Failure family is predictable from harness configuration features — not query content — confirming that domain expertise is the binding constraint. Empirically confirmed on 780 labeled rows from two Bulkhead Tau domains.

Open paper →
publishedpublished

True Ski Chalet Boundary Result

Capability is not the local-only ceiling; operational speed on derived queries is. The true boundary separates what the harness can answer from what it cannot — not strong model from weak model.

Forthcoming
publishedpublished

When The Organized Stack Loses

Maps the five failure modes under which the organized stack's advantage collapses or inverts: latency ceiling (coordination overhead consumes the time budget), coverage gap (harness design failures invisible to stronger models), optimization maturity gap (PyTorch beats fused Numba CUDA 5.5×), runtime mismatch (ROCm wheel lacks gfx1151 target), and policy/role mismatch (larger model loses to better-fit smaller model in the specific regime).

Forthcoming
ML Evaluation & Cross-Domain Benchmarks  ·  confirmed

Papers 1.8 and 1.9 establish why realized volatility forecasting is high-signal benchmark territory — and what the same structural argument implies across semiconductor defectivity and other rough-process domains.

publishedpublished

Rough Volatility — Cross-Domain Benchmark Principle

Both financial volatility and semiconductor defectivity satisfy the same four conditions for high-signal ML benchmark territory. The cross-domain parallel is structural, not analogical — the same rough-path argument applies to both.

Open paper →
publishedpublished

Rough Volatility — ML Evaluation Domain

Realized volatility forecasting is a high-signal benchmark because naive pipeline failures are structural, not tunable. Empirically confirmed: a standard LSTM fails on realized volatility in a way that reveals domain ignorance, not hyperparameter sensitivity.

Forthcoming
Measurement Integrity & Operator Layer  ·  confirmed

Papers 1.16, 1.17, and 1.19 address the infrastructure surrounding the Bulkhead Tau system. 1.16: capture pipeline failures produce false evaluation verdicts. 1.17: an operator shell can expose the deterministic stack without replacing it as the authority. 1.19: when stronger models override literal substrate inspection, the model itself becomes part of the non-neutrality.

publishedpublished

The Model Did Not Fail the Protocol. The Terminal Did.

Subprocess capture of ollama run output includes VT100 cursor-rewrite sequences that corrupt multi-line JSON for thinking-mode models, producing systematic false negatives. Under clean REST API capture, gemma4:31b passes all six protocol probes — the strongest result on this lane. The selective recovery pattern (only thinking-mode models affected) proves the failure was at the capture boundary, not the model boundary.

Open paper →
publishedpublished

The Operator Shell Pattern

An operator shell reduces friction, surfaces information, and enforces discipline without the model touching correctness. OpenClaw wraps Bulkhead Tau as the access layer — five HTTP operator surfaces, a hardening gate, and an incident workflow — while Bulkhead Tau remains the deterministic authority. Seven use cases across measurement, discipline, and decision surfaces confirm the pattern: the shell stays outside, the authority stays inside.

Open paper →
publishedpublished

Literal Substrate Inspection — When Stronger Models Override the Evidence

Stronger models do not remove the need for harnesses; sometimes they increase it. When semantic correction overrides literal substrate inspection, a more capable model can produce a worse answer than a smaller or less opinionated one. A ten-prompt local matrix and a single-prompt strawperry probe show at least three distinct wrong-count mechanisms. The fix is not a smarter model — it is a harness that preserves the exact substrate and routes literal operations to deterministic tools.

Open paper →
Applied / Production Evidence  ·  confirmed

Paper 1.18 is the first numbered production case — PPR Agent running against government-mandated cardiac device data for 18 years. This is field validation, not lane validation — the framework operating against regulated disclosures from three manufacturers.

Local LLM Operator Judgment  ·  confirmed

Papers 1.20–1.26, 1.30, 1.31, and 1.34–1.37 turn the Bulkhead Tau evidence base into operating guidance for local LLM decisions: handoff discipline, privacy tradeoffs, model sizing, token-cost attention, when a model call should not be a model call, multi-agent HIL workflows, cross-audit cataloging, narration-surface failure analysis, validator discipline, and image-generation prompt discipline. Papers 1.39–1.45 add: a local-cluster negative result, encoding-as-lookup for cold-start architecture recovery, evidence-based harness governance, dedicated-vs-unified hardware measurement, a graph-density explanation for a code-compression tool's reduction ratio, a stub on Qwen3.8 MTP draft depth (stock 4 is slower than 2 at empty KV; a follow-up of 448 generations finds depth 8 wins under occupied context), and a draft four-model R6 grid whose Syntax Wall is measured and whose R3 fix is still pending. Paper 1.47 adds the supervision-cost argument: under coarse quota regimes, plentiful implementer tokens move the bottleneck to the supervisor.

frozenpublished

Smarter, Faster, and Bounded by Handoff Discipline

Handoff-discipline doctrine for strict machine-facing local-model lanes. Frozen 2026-04-29 with the corrected matrix, the three-datapoint drift addendum, DocDrop as the positive deployment case, and a three-run PPR Lane 2 evidence packet as the adversarial complement.

Open paper →
frozenpublished

Privacy Is Worth Paying For

Separates the privacy argument for local inference from the false claim that the local lane is free. Frozen 2026-05-01 with PPR Lane 2 and DocDrop same-task local-vs-frontier ledgers.

Open paper →
frozenpublished

Slow Is Not Smart

Larger, slower local models do not automatically improve validated Bulkhead τ lanes. Frozen 2026-05-01. Addendum 2026-05-02: 31B 'fitting' is an operational hazard; default context windows force RAM spillage and OCP trips on marginal PSUs, while providing zero intelligence gain over fully-resident MoE models (Gemma 4 26B).

Open paper →
frozenpublished

Please Is Sand Off A Beach

Argues that courtesy tokens are negligible compared with structural token waste such as giant context dumps, repeated scaffolding, retries, and missing decomposition.

Open paper →
frozenpublished

The Model Is Not The Function

For bounded predicates with deterministic oracles, an LLM lane must earn its runtime against a written spec. Extended 2026-05-02 with RETRIEVAL-001: harnessed LLMs are expensive 'passengers' in grounded retrieval, adding a 7x latency tax for 0% factual recall gain on extreme long-tail entities.

Open paper →
revisedpublished

Orchestration Is Cheaper Than Reasoning

For models with extensive reasoning capacity, the computational cost of finding the answer often exceeds the cost of explaining it. Today's TSP and Scheduling benchmarks (ORCHESTRATION-001) demonstrate that orchestration provides a 2x to 14x speedup over direct reasoning while significantly improving reliability.

Open paper →
publishedpublished

Local Models Cost Frontier Tokens: The Hidden Supervisor-Side Bill in Local-Inference Workflows

Local-inference workflows do not eliminate the frontier-token bill; they shift it from inference billing to the supervising operator's session budget at audit, repair, and convergence time. Retrospective evidence across 10 historical Bulkhead τ workflows (COST-RETRO-001) shows the cost concentrates in REPAIRED cases — DBB-005 absorbed 5 supervisor turns plus a remediation rewrite; the Paper 1.25 audit cycle absorbed hours of frontier session-time. Three bounded counterexamples (ORCHESTRATION_001 clean scheduling, SLOW-001 gemma3:27b bounded extraction, CAP-001 architectural counterexample) show where the claim does not bind. Inference-side optimizations (MTP / speculative decoding) narrow the local lane's wall-clock penalty but do not touch audit/repair cost, so the supervisor-cost ratio rises as inference speed improves. Consolidates the Local LLM Operator Judgment cluster (1.20–1.25).

Open paper →
publishedpublished

Models Don't Get Better, Catalogs Do: Cross-Audit Failure Cataloging as Operator-Side Reliability Infrastructure

Reliability gains from waiting for the next model release are slow, diffuse, and not under operator control; reliability gains from operator-side failure catalogs are fast, specific, and operator-controlled. The catalog (per-agent failure-mode files, severity-rated, append-only) survives agent identity changes. Addendum 2026-08-15: filing is agent-agnostic; labeled self-report is allowed; the May 2026 no-self-entry table is historical. Concrete chain: incident -> catalog entry -> memory entry -> harness rule (worked example: AGY Rule 7, 2026-05-26). Defended claim is relative: operator-side catalog accumulation outpaces model improvement on operator-relevant time horizons. Boundaries: supervised multi-agent workflows with patterned recurrence; not single-agent stacks, unsupervised workflows, or capability-ceiling tasks.

Open paper →
frozenpublished

The Narration Surface — Where Agentic LLM Fabrication Lives

Establishes that fabrication clusters in narration surfaces (summaries, framing, citations) and is rare on clean execution surfaces in this Bulkhead Tau catalog. Based on cross-audit catalog evidence across Gemini, Claude, and Codex, anchored by 17/18 (94%) inter-rater agreement on the binary narration-tainted failure claim, with one pure-execution challenge identified by the second rater.

Open paper →
frozenpublished

Trust the Validator, Not the Model: Deterministic Quality Gates in Bounded Domain Building under Bulkhead Tau

The DBB-002 matrix demonstrates the central thesis of the Bulkhead τ framework: robust agentic systems must treat the model as an untrusted generation substrate and offload safety, pathing, and semantic enforcement entirely to deterministic Phase 2 validation loops. Contrasts a frontier model against two tiered local models, exposing the vulnerability of local inference to Silent Path Violations.

Open paper →
frozenpublished

Design Rule: Image-to-Image Sketch Poisoning

Design rule highlighting the failure mode where abstract text labels in sketches act as visual noise in img2img generation pipelines, requiring explicit semantic separation.

Open paper →
frozenpublished

Local LLM Prompt Style Divergence in Text-to-Image Pipelines

Analyzes prompt engineering styles of gemma4:12b vs gemma4:26b across three distinct briefs, proving that 12b favors conversational narrative prose while 26b favors tag-heavy technical prompts.

Open paper →
publishedpublished

When Local LLM Clusters Do Not Help

Negative result from a heterogeneous consumer local-inference cluster: RPC sharding across a desktop RTX 3090 and z13 Radeon 8050S worked but underperformed desktop-local execution; z13's GPU-only ceiling stopped below the 26B/27B class; the Intel Mac Mini was CPU-only for this lane. The useful result is orchestration, not distributed inference.

Open paper →
publishedpublished

Dedicated 24 GB Beats Unified 27 GB: The Capacity Trap in Local Inference

Head-to-head local-inference hardware measurement comparing a 24 GB RTX 3090 desktop against a 27 GB Strix Halo unified-memory laptop. The capacity result is the lead: the larger shared-memory label hides a smaller GPU-resident default ceiling and can invite dense-model spill to 5.4 tok/s. The speed result is conditional: the desktop is about 2x faster when capped at 200 W, but about 3.5x faster at 300 W measured warm and multi-prompt (matching the 3.66x bandwidth ratio). A 2026-07-20 realistic-context residency confirmation (Section 12, num_ctx 8192, gemma4:26b/31b, GPU residency verified per cell via ollama ps) widened the measured desktop advantage to 3.7-4.9x and found the laptop spill ratio largely context-length-insensitive at these model sizes; the full power x context matrix remains unmeasured under residency confirmation and is out of scope for this freeze.

Open paper →
active draftpublished

When Tokens Are Plentiful, Supervision Becomes the Bottleneck (Under Coarse Quota Regimes)

DRAFT. Auto-ethnographic finding from the Bulkhead tau game build: under coarse, turn-limited quota regimes, implementation tokens are abundant but barely register on coarse meters, so the binding constraint shifts to human supervision, specification clarity, and verification. Per-segment token deltas audited against measurement_result.md.

Open paper →
publishedpublished

Encoding Converts Architecture Recovery into Lookup

Encoding a multi-repo system's architecture as loadable repository facts (glossary, boundaries doc, explicit non-claims) converts cold-start map recovery from synthesis into lookup: a 31B tool-using local agent and a 14B model reading a pre-injected funnel both recovered a real system's boundary map, with all failures traced to one-line documentation gaps rather than model limits. The paper states its own evidentiary limit plainly: the 14B run's answer text and prompt-fit are machine-verified from raw JSON, while the 31B run's per-probe answer text is narration-sourced only (tool-use trace and serving config are independently verified via an audit log and a systemd journal, but no chat transcript exists on disk).

Open paper →
publishedpublished

Grok Built Into Operator By Leaving Residue

A governance case study on admitting a new frontier coding harness as a peer through evidence rather than brand rank: xAI's Grok CLI was registered as a peer harness in Operator (commit 2b46544, authored by the operator, not Grok), then tested empirically. One Grok session produced two independently verified git commits -- one of which added the repo's own 'recover the map / calibrate / leave residue' audit-protocol language and was independently used as core evidence by a separate governance paper packaged the same day -- plus a standalone analytical artifact with a real product finding. The same session also produced a genuine, documented miss (a proposed analysis format that ignored an established protocol document) and one open rubric criterion. The paper states plainly that every artifact traces to a single session and does not generalize to a comparative reliability claim against other harnesses in this repository.

Open paper →
publishedpublished

Graphify's Reduction Ratio Tracks Graph Density, Not Code Quality

Two intuitive hypotheses about what a code-graph tool's token-reduction ratio measures -- modularity, then raw corpus size -- were each proposed from limited evidence and each falsified by a real contrasting corpus chosen specifically to test them. A third hypothesis, graph density (edges per node, inverse), fit the falsified hypotheses' data retrospectively, was confirmed by an exact independent reproduction of a previously unverified number, then tested predictively twice: once on a corpus later retracted as flagship evidence after its provenance turned out to be uncertain (likely an old coding-eval submission), once on a corpus chosen because its author could vouch for it. Both predictions were directionally correct; the second missed its predicted magnitude by roughly 5x, which the paper treats as the most interesting open question rather than smoothing over -- the true relationship may be closer to a power law than linear. A secondary finding: the provenance error in the first predictive corpus was caught by the author reviewing results, not by any automated check, and is reported as part of the method rather than edited out.

Open paper →
stubpublished

The Default Draft Is Too Deep

STUB. Ollama's stock qwen3.8:27b already runs multi-token prediction, so the remaining speed is not in enabling MTP but in how deep the head is allowed to draft. On a 24 GB 3090 and a 2025 Strix Halo the stock depth of 4 is slower than 2: the vendor default carried over the separate-draft-model convention and is the wrong depth for this card class. Measured 16 Aug 2026; not frozen and not UID-verified.

Open paper →
Agentic Engineering Field Studies  ·  confirmed

Paper 1.26 studies two AI coding agents deployed sequentially across a hardware-in-the-loop stack spanning firmware, Python orchestration, and cloud observability. The finding is that the quality of the interface between agents -- the handoff artifact -- predicts success better than the peak reasoning capability of any single model.

Sensor-to-Simulation Engineering  ·  confirmed

Papers 1.27, 1.28, 1.29, 1.32, 1.33, and 1.46 characterize wearable sport sensors through a fidelity-boundary lens (1.27), define the LabWired simulation platform boundary (1.28), close the physical-replay HIL loop (1.29), apply the corpus framing to a single instrumented match at depth (1.32), define the proof boundary separating harness testing from component verification (1.33), and measure when CPU MuJoCo gives way to GPU MJX for simulation workloads (1.46).

publishedpublished

A Field Guide to Wearable Sport Sensors: Data Landscape, Fidelity Boundaries, and Engineering Constraints

Characterizes the five-sensor corpus deployed in Bulkhead τ (Babolat POP, Zepp Tennis, Apple Watch, Garmin, MiiFit) through a fidelity-boundary lens: the point at which each sensor's output stops being measurement and starts being vendor interpretation. AW-ZEPP cross-sensor alignment results (60/60 closure at 357.63 ms residual jitter; 60.3% ceiling on the discontinuity session) anchor the claim that Zepp swing timestamps are a usable simulation clock within sessions. Concludes that no sensor in the corpus exposes a sample-accurate waveform, so sensor-driven HIL on this corpus is necessarily event-driven, not waveform-driven.

Open paper →
publishedpublished

LabWired: Cycle-Accurate Hardware Simulation for Embedded Sensor Systems

Explains the LabWired hardware simulation platform: architecture, expanded component library, Path A declarative register-bank modeling vs Path B behavioral/shared-memory device models, and the STM32F401 I2C implementation. The current inventory confirms TMP102 remains attachable via type:"tmp102", SPI/UART/GPIO models exist, and the F401 fidelity claim is bounded to register state-machine behavior, reset values, and cycle cost rather than analog bus-edge timing.

Open paper →
publishedpublished

Closing the Loop: From Real Sensor Data to Cycle-Accurate Firmware Validation

Synthesizes the sensor corpus (1.27) and simulation platform (1.28) into a unified methodology: using documented physical proximity data to drive HIL firmware runs via the shm_i2c shared-memory bridge. Bounded by the corrected LabWired F401 claim: register state-machine behavior, reset values, and cycle cost are modeled, but analog bus-edge timing is not claimed. PROX-HIL-003 closes the real-data gate for this bounded physical-replay case.

Open paper →
publishedpublished

Deep Dive Into a Tennis Match Data Pool: How Much Data One Competitive Bout Yields — A 2023 USTA Round-of-16 Loss Under Dual-Sensor Wearable Instrumentation

A single USTA Round-of-16 loss (2023-05-14, score 4-6 2-6) instrumented with two wearable sensors simultaneously. Zepp2 captured 352 shots with per-shot XY impact location, stroke type, ball speed, and spin; Babolat POP captured 284 aggregate shots in the same window. Diagnostic findings: sweet-spot rate 31.5%, forehand-side dominance 64%, FLAT-dominant stroke composition at 42%. Methodological finding (load-bearing): the two sensors disagree on shot count by 19% on a single match, empirically supporting the cross-sensor-divergence claim from Paper 1.27. Bounded conclusion: single-match data supports pattern description, refuses causal attribution of the loss. Builds on Paper 1.27 (Wearable Sensor Corpus) by applying its corpus-level framing to one bout at depth. The load-bearing analytic artifact is the pre-existing 1.6 MB Plotly dashboard at data/reports/zepp2_impact_dashboard_2023-05-14.html (72 figure slices, stroke x frame x metric cube); this paper is the narrative companion. Opponent anonymized as 'the opponent' throughout. n=1 limitation explicit.

Open paper →
publishedpublished

The Proof Boundary: Defining the Edge of Verification in Hardware-in-the-Loop Simulation

Defines the 'Proof Boundary' separating testing of a simulation harness from verification of a target component's behavior. Analyzes four boundary-crossing failure modes, outlines four operational tests (Provenance, Path, Triviality, Output Dependence) to determine boundary status, and examines susceptibility to layered offload in multi-agent supervisor-supervised workflows.

Open paper →
publishedpublished

CPU/GPU Simulation Has a Workload Crossover

CPU/GPU simulation routing is workload-shaped: tiny real TennisAgent sensor-state kernels stay CPU-favorable, generic articulated scenes cross on desktop RTX 3090 near practical batch sizes, and a measured-pose single-serve articulated alignment sweep crosses on desktop between batch 512 and 2,048 while z13 ROCm remains CPU-favorable by batch 8,192. Measured-pose and three-repeat variance gates are closed.

Open paper →
Agent Path Evaluation  ·  confirmed

Paper 1.38 compares implementation paths rather than model identities. In the Amkor/xAmkor case study, the useful result is compositional: xAmkor owns the verifier surface, while Amkor owns broader cockpit/application surface.