BTL · Research index

Research should end in
something you can run.

We publish the question, the method, the artifacts, the failures, and the code needed to reproduce the result.

Updated July 2026
Research outputStatusArtifactsOpen
01
Technical report · research lane

Behaviour-Relearned Quantization

The BRQ paper: why static one-bit MoE quantization failed, how binary routed experts recovered teacher-forced structure, and why the result did not yet promote as a release artifact.

  • 11-page paper
  • OLMoE receipts
  • Qwen controls
  • V1b contract
02
Research paper · released

Range Before Representation

Behavior-gated two-bit quantization of a 35.1B-parameter mixture-of-experts model into a 9.96 GB stock-format GGUF, with controlled range-selection and expert-level ablations.

  • 35.1B MoE
  • 9.96 GB GGUF
  • 94.1% conditional retention
  • PDF
03
Research paper · released

Behavior Before Perplexity

The complete BTL-3 Compact compression recipe: failed routes, behavioral cliff localization, measured precision allocation, repair, packing, and native proof under an exact byte ceiling.

  • Academic paper
  • Engineering article
  • Recipe + evidence
  • Open source
04
Benchmark paper · released

Context Integrity

An auditable benchmark for whether long-running agents preserve, retrieve, update, and use evidence correctly across sessions.

  • 250 deterministic tasks
  • Dataset
  • Evaluation harness
  • PDF
05
Thesis paper · runtime implemented

ESP: Echo-Skeleton Perception

A stateful perception architecture that lets text-only models operate graphical interfaces through structure, OCR, affordance probes, and change events.

  • Pre-registered hypotheses
  • Working runtime
  • Browser tasks
  • PDF
06
Model systems report · released

BTL-3 Compact

The engineering record behind a complete 27B agentic coding model in one 8.39 GB native GGUF: allocation, packing, behavior repair, kernels, and artifact-faithful validation.

  • 8.39 GB artifact
  • Native CUDA + Metal
  • 92.2% tool retention
  • Open model
07
Evaluation program · released

The Reasoning Gap

Controlled tasks for separating observational pattern completion from interventional causal reasoning, with exact inference baselines.

  • Interventional tasks
  • Exact baselines
  • Public test
  • Reproducible

HOW WE WORK

Pre-register when the claim is uncertain.
Execute when the claim is mechanical.
Report both wins and failures.

Benchmarks stay sealed until representation and training choices freeze. Model claims carry denominators and protocol. Runtime claims come from the exact deployed artifact.

Experiments that fail remain part of the record. Research is useful when another builder can see exactly where the method worked—and exactly where it stopped.