You pay for SURVIVAL
It hides behind a pillar and never fires. Scores brilliantly. Loses every single fight — a flawless coward.
FREE · WINDOWS · MACOS · LINUX · MADE IN GODOT 4
Learn real reinforcement-learning techniques by playing a game. You don't code the fighter — you move nine sliders that say what it gets paid for, and a real search spends thousands of simulated matches finding a mind that maximises exactly that. Including the parts you didn't mean.
01 — THE LOOP
Nine weights — win, hit, taken, time, survive, accuracy, aggression, ricochet, flank. That is your entire input. You never touch behaviour.
Sixteen senses to pick from: aim error, range, incoming disc, time to impact, line of sight. Switch one off and it is blind to that fact forever — no amount of training brings it back.
Press TRAIN and watch the learning curve climb live. When you like it, SAVE & FIGHT drops the trained policy straight into the Grid against a champion.
02 — THE CATCH
Not what you meant. What you wrote. Every time.
You pay for SURVIVAL
It hides behind a pillar and never fires. Scores brilliantly. Loses every single fight — a flawless coward.
You punish TIME too hard
It walks out of the ring to end the round early. It solved the problem you actually gave it.
Closing the gap between the reward you wrote and the outcome you wanted is reinforcement learning. The coach panel watches your two curves and calls it out the moment they diverge.
03 — HONEST MACHINERY
The cross-entropy method: sample a population of whole minds, score each over simulated matches, keep the elite few, re-aim at those, narrow as the budget runs down. No gradients. No backpropagation.
Which is why you will not find a learning rate or a discount factor anywhere in this game — this method does not have them, and inventing knobs that do nothing is how tutorials teach people wrong. In advanced mode you drive the search yourself: population, elite fraction, exploration, generations.
04 — WHY WE BUILT IT
Ask an AI, and it hands you the answer in four seconds. What it cannot hand you is the ten minutes you spent training a coward—watching a machine do exactly what you asked, and lose anyway.
Nobody remembers the textbook definition of reward hacking. Everybody remembers their own fighter walking out of the ring to stop losing points.
That is the bet MINDFORGE is built on. Explanations are free and everywhere now; what is scarce is the experience of being wrong yourself, on purpose, with something you built. So MINDFORGE is a game first—but underneath, it runs on a real cross-entropy search, real reward functions, and real failure modes. Because a lesson you cannot reproduce isn't a lesson at all.
05 — THE GRID
Your trained policy plays the exact behaviour it learned — no scripting, no cheating. Three cameras, a contracting ring, best of five. Or take the disc yourself and find out how hard your own reward really was.




No maths homework, no ML background. The game explains itself as you go, and every number it quotes is the real default from the code.
WINDOWS · MACOS (UNIVERSAL) · LINUX · NO PAYMENT, NO ACCOUNT