Evolved Frogs
Nobody taught them to hop.
Each frog is a physics body with a small neural-network brain. The brains were not trained on examples. They were evolved by a genetic algorithm, the same trick nature uses.
We started with 200 random brains and let each drive a frog for up to ten seconds. The best were kept, the rest were replaced by their mutated offspring, and the whole thing ran again, 300 times. Nobody wrote a line of hopping code. Every score we tried before this one was gamed. This one also ends the run for any frog that stops making progress, and the frogs found a hop they keep up for the full ten seconds.
The ghost race
Champions from different generations, all recorded on the same flat course under the training rules and run together. Each is the best frog of its generation. A frog that falls, or moves less than 0.5 m in 2 seconds, stops where it is. The solid grey frog wearing Froggy's face is the one we ship; the coloured ghosts are champions of other generations. The camera follows the leader.
Why isn't the shipped frog always in front? On this one flat course, generation 299 edges it by about a quarter of a metre. We didn't choose on one course. We re-tested the last 50 champions on ten separate validation courses and shipped the best there. One race is exactly the kind of test the luckiest frog wins.
Paused because your device asks for reduced motion. Press Play to watch.
Watch the shipped champion drive
This frog is running live in your browser, under the same rules as training: the run ends if it falls or stops making progress. Its brain reads the frog's senses 60 times a second and sets a target angle for each of its six leg joints. It starts on the shipped champion; the slider picks any generation's champion.
Paused because your device asks for reduced motion. Press Play to watch.
Left: the 21 senses. Middle: two layers of 16 neurons. Right: the six joint commands. Green is positive, orange is negative, brighter is stronger. Lines are the strongest connections.
Getting better, one generation at a time
The score of the best frog in each generation, the typical (median) frog, and the best frog re-tested on five bumpy courses it never trained on. Click the chart to watch that generation.
- Best
- Best, on unseen courses
- Median
- Shipped champion
Introducing generation 298
The first frog we have shipped that keeps hopping. On our flat showcase course it covers 26.1 m in ten seconds, moving steadily at about 2.5 m/s and in the air about 35% of the time, and it is still going when the clock runs out. It stayed upright for all ten seconds on the flat course and on all five holdout courses.
Distance on the flat showcase course: the previous release (generation 277, which fell at 2.8 s) against this one (still hopping at 10 s).
Courses completed upright for the full ten seconds: the flat course and all five holdout courses.
Validation score, 10 courses: the latest champion (generation 299) against the shipped one. We ship the best of the last 50 generations on validation.
Holdout score, 5 courses never used for any decision. This is our honest estimate. The latest champion scored 23.74 here; we never pick on holdout, so that changes nothing.
Select on validation, report on holdout: the same discipline as our image classifier. It still earns its keep. Over the last 50 generations the champion's holdout score ranged from 11.66 to 23.97 (mean 20.04), because the top scorer on three training courses is often just the luckiest.
Evolution does exactly what you score. Not what you mean.
The score is the only thing evolution sees. Each time, the frogs found the cheapest way to earn it. Once, they fooled us as well: we described a trick as a gait and published it. Here is what happened, in order.
-
The divers
Our first score was simply distance travelled. Diving forward as hard as possible, then crashing, beat any careful hop.
14.8 m, then a crash at 2.9 sGeneration 29 champion. 99–100% of the population fell in every generation.
Fix: multiply the distance by the share of the ten seconds spent upright. The dive stopped within about 10 generations.
-
Lunge and freeze
With falling fixed, champions settled at 7–9 m per episode. We first thought they were scooting on their stiff front leg, and said so. They weren't. Measured second by second, they covered nearly all of it in the first 2–3 seconds, then stood still.
8.7 m by 3 s, 8.6 m at 9 sA champion from our earlier main run (generation 200) on the flat course. The front-leg contact we had measured was standing, not scooting.
Lesson: bank the distance early, then stand still and keep the full survival bonus. Reward hack number three, and we had misreported it.
-
Airborne-only didn't help
To stop paying for "scooting", we counted only distance covered in the air. There was no scooting to stop, and standing still after the lunge still kept the survival bonus.
8.9 m by 3 s, then nothingThe selected champion of that run (generation 120). Its generation 50 did the same: 7.1 m by 3 s.
Lesson: we had concluded that the hard part was controlling the hop. Wrong again. The limit was still the score.
-
The stall rule
End the episode if the frog moves less than 0.5 m in 2 seconds, and scale its score exactly as for a fall. Standing still after a lunge now costs the survival bonus.
26.1 m, still hopping at 10 sGeneration 298: about 2.5 m/s, airborne about 35% of the time, upright for all ten seconds on the flat course and all five holdout courses.
Status: shipped. A steady, repeated hop emerged within about 150–200 generations.
This is called reward hacking. Any optimiser, from a genetic algorithm to a large AI system trained on feedback, exploits the gap between the score and the intention. Sometimes the gap is in how you read the results, too.
“We asked for distance and got divers. We asked them to stay upright, and they lunged and then stood there, and we called it scooting. We told them standing still no longer counts. Now they hop, all ten seconds of it. I have never been prouder of this company, or more careful about what it measures.”
How it works
A body
Seven rigid parts and six motorised joints in a 2D physics world. The frog is not animated: motors push, and gravity, friction and momentum decide what happens.
A brain
A neural network of 726 numbers. In: joint angles, tilt, speed, whether each foot is down, and a ticking clock. Out: a target angle for each joint.
A score
Distance travelled, scaled by how much of the ten seconds the run lasted, minus a little for wasted effort. A run ends early if the frog falls, or if it moves less than 0.5 m in 2 seconds: the stall rule. The score is the only thing evolution sees.
Evolution
Keep the best five. Fill the rest with children of tournament winners, sometimes mixing two parents, then nudge about one number in ten at random. Repeat.
Teaching frogs to hop without telling them how
The full method, every experiment, the lunge-and-freeze correction, the stall rule, the winner's curse, and everything that could still be wrong with our conclusions.
Read the white paper