
Models that keep learning, and a world to learn in.
Continual MI is solving continual learning: weights that keep changing with use instead of freezing the day training ends. Two things make that reachable — MGPT, an architecture small enough that a model can keep being trained on hardware you own, and Monte Lua, a persistent universe whose characters cannot work unless they genuinely remember. The game does not only fund the research. It is where the research runs.
Everyone is building bigger.
Nobody is building something that learns.
Frontier models are impressive, enormous, closed — and frozen. Everything that looks like memory in them is arranged outside the model: a longer prompt, a retrieved document, a summary of what happened before. The weights themselves never move again after training.
We think that is the thing to fix, and that two problems sit under it. The first is size: weights can only keep changing if changing them is cheap, so the model has to be small enough to finetune on hardware someone can own and govern. That is what MGPT is for — masking attention that pushes 4B–9B parameter models toward the capability of models many times their size.
The second is a place to work. Continual learning is hard to study on benchmarks, because a benchmark ends. So we built a world that does not: Monte Lua is a persistent universe where characters carry one history across hundreds of hours of real play, and where a frozen model visibly fails — it forgets people, contradicts what happened, loses the thread. Those failures are the measurements. The game is the environment continual learning needs, and it happens to be a product we can sell.
How the parts need each other.
Small enough to keep training
Masking attention gets larger-model capacity out of 4B–9B weights. Weights can only keep changing if a machine you can afford can change them.
A world that demands memory
A persistent universe whose characters have to remember what happened, for as long as it runs. Nothing in it works on a frozen model.
Weights that keep changing
Where it forgets is the research. The methods that fix it are the goal the company is named for, and they go back into the architecture.
↻ and back to the architecture
How the pieces line up.
- The problem · Continual learning
Every model you use is frozen.
A model finishes training and stops. It can be given a longer prompt, a bigger context, a database to look things up in — but the weights never change again. Continual learning is weights that keep changing with use, and it is the problem this company exists to solve.
- MGPT · The prerequisite
Small enough to keep training.
Weights that keep changing have to be cheap to change. MGPT reworks attention through masking so 4B–9B parameter models reach the capability of models many times their size — which is what puts finetuning on hardware a person or a company can actually govern, rather than a datacenter they rent.
- Monte Lua · The environment
A world that cannot run on a frozen model.
Monte Lua is a persistent universe generated as you play, and it is our research environment as much as our product. Its characters have to remember — across hundreds of hours, one continuous history, no reset. That is a continual-learning problem by construction, running against real players rather than a benchmark, and it is where our methods are tried.
- EMWaver · Continual School
Where the people come from.
EMWaver is a way into technology at all: a board on a USB-C plug, infrared and radio, something physical that answers. Continual School picks that person up and teaches them to do AI research, free, from the beginning — ending at a supervised project. One is the door, the other is the path.
- The end goal · Machine intelligence
A model that keeps learning.
Efficient enough to run and tune anywhere, open enough to own, and no longer fixed at the moment training stopped. Reaching that is what the company is named for.



Real play from Monte Lua. Everything a frozen model has to fake here — who these people are, what was said, what it cost — a continually learning one would simply know.
The bet is simple: a model that keeps learning beats a bigger model that cannot — and you only find out inside a world that refuses to forget.
Monte Lua is both halves of that: the environment where continual learning is required to operate, and the product whose revenue pays for the work. Small is the prerequisite — continual learning only works off the datacenter — and reaching it is what the company is named for.
Where the work is, right now.
Mask-Generative Pretrained Transformer — masking attention behind small 4B–9B models that rival much larger ones. The efficiency that makes continual finetuning affordable. No public API yet.
An endless visual novel, and the environment the continual-learning work runs in: one history, hundreds of hours, characters that have to remember.
Society’s teaching layer, and how the research gets more hands: one continuous track from how a computer works up to a supervised project, every lesson filmed by us and free to watch.
A board the size of a USB-C plug that does infrared and 433 MHz radio from your phone. For most people it is the first thing that makes technology feel touchable — the door the school opens onto.
Society is where the progress, the builds, and the open-model discussion live. The conversation happens on Discord — and Continual School is where you can learn to do the work itself, free, from the beginning.