Independent AI lab · Lagos

We ship the model,
the system, and the proof.

Bad Theory Labs builds open models, agent infrastructure, runtimes, and benchmarks. The work is public, runnable, and measured against the thing it claims to do.

Latest release · July 2026RL-0013
BTL3

27B agentic coding + tool-use model

Full editionOpen weights
Compact edition8.39 GB
Native runtimeCUDA + Metal
EvidencePublic artifacts
Release dossier
OPEN WEIGHTSNATIVE INFERENCEAGENT RUNTIMESEXECUTION-VERIFIED EVALSREPRODUCIBLE RESEARCH
88.5%BFCL v4 AST1,240 cases · BTL-3
95.12%HumanEval pass@1156 / 164 · thinking
8.39 GBComplete 27B editionone native GGUF
43.16 t/sCompact generationRTX PRO 6000

Selected output

The lab is already
in production.

Models are one layer. We also build the runtime, memory, agent scaffolding, evaluation, and deployment path around them.

02Native model system

BTL-3 Compact

The complete 27B text model, packed into 8.39 GB.

A byte-verified AVQ2/UniSVQ GGUF with its own CUDA and Metal runtime, OpenAI-compatible server, and Ollama and LM Studio bridges. No BF16 checkpoint is loaded behind the scenes.

  • 2,416 tensors
  • 92.2% tool retention
  • CUDA + Metal
Get Compact
03Inference infrastructure

Runtime

One production API for models across providers.

OpenAI-compatible chat and responses APIs with provider routing, usage accounting, billing, caching, rate limits, and self-serve workspace keys.

  • Multi-provider
  • Streaming + tools
  • Usage ledger
Open Runtime
04Agent memory

RetainDB

Persistent context with evidence, scope, and retrieval built in.

A memory layer for agents that need to remember across sessions without turning every old fact into current truth.

  • 79% LongMemEval
  • 0% stored-fact hallucination
  • Managed API
Open RetainDB

Research with artifacts

No hand-waving.
Build the test.

Our research produces papers, datasets, environments, runtimes, model artifacts, and explicit failure reports—not just a thesis page.

Models, agents, infrastructure

A lab should leave
working systems behind.

Each project attacks a different failure mode: weak tool mechanics, forgotten context, unsafe autonomy, expensive inference, or perception that starts over every frame.

THE OPERATING PRINCIPLE

Build ambitious systems.
Measure them without mercy.
Ship what survives.
Read the research Join the lab