Bad Theory Labs · Research index

Research should end in
something you can run.

We publish the question, the method, the artifacts, the failures, and the code needed to reproduce the result.

Updated July 2026
Research outputStatusArtifactsOpen
01
Research paper · released

Behavior Before Perplexity

The complete BTL-3 Compact compression recipe: failed routes, behavioral cliff localization, measured precision allocation, repair, packing, and native proof under an exact byte ceiling.

  • Academic paper
  • Engineering article
  • Recipe + evidence
  • Open source
02
Benchmark paper · released

Context Integrity

An auditable benchmark for whether long-running agents preserve, retrieve, update, and use evidence correctly across sessions.

  • 250 deterministic tasks
  • Dataset
  • Evaluation harness
  • PDF
03
Thesis paper · runtime implemented

ESP: Echo-Skeleton Perception

A stateful perception architecture that lets text-only models operate graphical interfaces through structure, OCR, affordance probes, and change events.

  • Pre-registered hypotheses
  • Working runtime
  • Browser tasks
  • PDF
04
Model systems report · released

BTL-3 Compact

The engineering record behind a complete 27B agentic coding model in one 8.39 GB native GGUF: allocation, packing, behavior repair, kernels, and artifact-faithful validation.

  • 8.39 GB artifact
  • Native CUDA + Metal
  • 92.2% tool retention
  • Open model
05
Evaluation program · released

The Reasoning Gap

Controlled tasks for separating observational pattern completion from interventional causal reasoning, with exact inference baselines.

  • Interventional tasks
  • Exact baselines
  • Public test
  • Reproducible

HOW WE WORK

Pre-register when the claim is uncertain.
Execute when the claim is mechanical.
Report both wins and failures.

Benchmarks stay sealed until representation and training choices freeze. Model claims carry denominators and protocol. Runtime claims come from the exact deployed artifact.

Experiments that fail remain part of the record. Research is useful when another builder can see exactly where the method worked—and exactly where it stopped.