Project Sprout

Measured
intelligence.

A family of language models trained on a single scaling curve — where every result predicts the next. Three models, reproduced within 1%, for $180 of compute.

Results · the line

Four points. One line.

Each model is a controlled measurement on the same data recipe. Same method, rising scale — so every point predicts the next one.

$180
to train the family — twice, from scratch
<1%
reproduction gap across different hardware
15.1
perplexity of the 500M — for ~$51 of GPU time
3→4
measured points; the 1.5B is next
ModelRun 1Run 2 (from scratch)Δ
Seed · 45M3.66163.6072−1.5%
Sprout · 150M3.06793.0603−0.2%
Sapling · 500M2.71492.7116−0.1%

Final validation loss. The family retrained on independent hardware landed within 1% every time — the signature of a repeatable recipe, not a lucky run.

Live · Seed 45M on CPU

Watch the smallest one think.

Real, unedited output from the 45-million-parameter model — running on a laptop CPU, no GPU. It's tiny, and it shows: grammar works before knowledge does.

MODEL omnigenius-45m · DEVICE CPU
what color is the sky?
answer: the sky is blue. it is rand ng.
3960ms · 45M params
One plus one equals
two. Two plus two equals four.
2825ms · 45M params
A B C D E F
G H I J K L M N O P Q R S T U V W X Y Z
3395ms · 45M params

The alphabet and simple addition emerge cleanly; ask it a fact and it improvises. That gap is exactly what more scale closes — measured on the curve above.

Method · what scale buys

The same question, growing up.

Prompt: "To bake bread, you first need" — asked to each model as it grew. Watch coherence arrive before, then with, correctness.

SEED · 45M
"…you can learn how to play music again. When you start playing music, you can start to learn how to play it…"
Fluent grammar, wrong topic entirely. It answered a different question.
SPROUT · 150M
"…to consume 2.5–3.5 times more fiber, vitamins and minerals…"
Found the food aisle, not the recipe. Right neighborhood, wrong shelf.
SAPLING · 500M
"…mix flour and water together. Then make a dough. Let it rise… Bread is made of flour, water, and yeast."
An actual recipe — then it quizzes itself and answers correctly. That's what scale buys.
Roadmap

Built to grow up.

An architecture designed for staged development — and for decoupled memory, planning, and reasoning as it scales. Here's what's measured, and what's next.

NOW

The base family, reproduced

45M → 500M trained twice on independent hardware, within 1%. The 1.5B is next — the curve predicts a final loss of 2.37.

measured · on the record

The school phase

A teacher-graded curriculum that rewards an honest "I don't know" over a confident wrong answer — models that know when to look things up.

in development
FUNDED

Reasoning depth & world models

Recurrent test-time compute and a world-model interface — the architecture the codebase is built around, activated at the scale where it pays off.

designed · not yet shipped
Stay in the loop

Watch the next model train

One update when a new model finishes — with samples and numbers. Nothing else.

No spam, ever.
© OMNIGENIUS measured, not promised