Bad Theory Labs builds open models, agent infrastructure, runtimes, and benchmarks. The work is public, runnable, and measured against the thing it claims to do.
OPEN WEIGHTSNATIVE INFERENCEAGENT RUNTIMESEXECUTION-VERIFIED EVALSREPRODUCIBLE RESEARCH
88.5%BFCL v4 AST1,240 cases · BTL-3
95.12%HumanEval pass@1156 / 164 · thinking
8.39 GBComplete 27B editionone native GGUF
43.16 t/sCompact generationRTX PRO 6000
Selected output
The lab is already in production.
Models are one layer. We also build the runtime, memory, agent scaffolding, evaluation, and deployment path around them.
01Open-weight model
BTL-3
A 27B model trained to act, verify, recover, and know when to stop.
Our frozen RL-0013 release combines agentic coding with structured tool use. It ships with complete evaluation evidence, a full-quality adapter, and a native compact edition.
A byte-verified AVQ2/UniSVQ GGUF with its own CUDA and Metal runtime, OpenAI-compatible server, and Ollama and LM Studio bridges. No BF16 checkpoint is loaded behind the scenes.
Each project attacks a different failure mode: weak tool mechanics, forgotten context, unsafe autonomy, expensive inference, or perception that starts over every frame.