Small Model Research
An ongoing program asking how far small, cheap models can be pushed under one fixed testing protocol. Findings will be published here as they land.
This is the home base for our core research program. It is a living document: as findings firm up, they get published here, and the posts get linked.
The question
The industry puts all its money and attention into models that cost hundreds of millions to train. The question we are asking is quieter: how much of what people actually use AI for really needs that scale… and how much of it can a small model running on ordinary hardware handle just as well?
Nobody answers this honestly at large scale, because the answer is expensive to publish. So we answer it at small scale, where the experiments are cheap enough to run properly and repeat.
The method
A few rules keep the work honest:
- One protocol. Every model and architecture is compared on the same data, the same number of training steps, and the same schedule. The harness controls everything except the thing under test.
- Hypothesis first. Every run states what it expects to find before it trains. No post-hoc storytelling.
- Everything is journaled. Each run logs its full curve, config, and provenance. Results are reproducible or they do not count.
- Null results count. A failure list is part of the paper. Most experiments do not work, and we report that.
Findings
Nothing to report yet. The experiments are running. Check back, or read the blog for the thinking as it develops.