Teaching frogs to hop without telling them how
Abstract
We evolved the 726 weights of a small neural network that drives a two-dimensional physics frog, using a genetic algorithm with a population of 200 and no gradients, demonstrations or hand-written gait. The only teacher is the score, and the frogs gamed each score we gave them. Scoring distance alone bred divers: by generation 29 the champion covered 14.8 m and crashed 2.88 s in, and 99–100% of the population fell in every generation. Scaling distance by the time spent upright stopped the dive within about 10 generations. We originally reported that the survivors then scooted on their stiff front leg. That was wrong. Measured second by second, they lunged 7–9 m in the first 2–3 s and then stood still, which the survival-scaled score rewards. A follow-up that scored only airborne distance (E4) showed the same freeze.
Adding a stall rule, which ends the episode if the frog moves less than 0.5 m forward in 2 s, took away the reward for standing still. In a 300-generation run (E5) a continuous hop emerged. The shipped champion, generation 298, was selected on 10 validation courses (validation 24.12, holdout 23.04). On the flat showcase course it covers 26.1 m in 10 s at about 2.5 m/s, airborne about 35% of the time, and it stayed upright for all 10 s on that course and on all five holdout courses. The winner's curse remains: over the last 50 generations the champion's holdout score ranged from 11.66 to 23.97. Every experiment is a single run with one seed.
1. Results at a glance
Distances are for the frog's torso over one 10-second episode. Scores depend on the scoring rule in force, so scores from different experiments are not comparable; distances and airborne fractions are.
| Question | Answer | Evidence | § |
|---|---|---|---|
| Does scoring distance alone produce hopping? | No. It produces divers. | Generation 29 champion: 14.8 m, fell after 2.88 s. 99–100% of the population fell in every generation. | 5 |
| Which fix for falling works better? | Scaling by survival (A) | Both cures stopped the dive within about 10 generations; A gave a smoother signal than a flat 10-point penalty (B). | 6 |
| Do the survivors hop? | No. They lunge, then stand still. | E3 generation 200 on the flat course: 4.4, 7.7 and 8.7 m at 1, 2 and 3 s, and still 8.6 m at 9 s. First misreported as scooting. | 7, 9 |
| How far after 300 generations of E3? | About 9 m, nearly all in the first 2–3 s | Champion distance 9.1 m at generation 299; median score −0.29 to 6.77. | 8 |
| Is the latest champion the best one to ship? | No, in E3 or E5 | E3: generation 299 validation 4.36 against generation 277's 8.26. E5: generation 299 validation 21.97 against generation 298's 24.12. | 8, 11 |
| Does scoring only airborne distance fix it? | No | Same lunge and freeze: the selected champion reached 8.9 m by 3 s and did not move after that. | 10 |
| Does a stall rule fix it? | Yes | E5 generation 298: 26.1 m on the flat course, about 2.5 m/s for the full 10 s, about 35% airborne, upright for all 10 s on the flat course and all 5 holdout courses. | 11 |
| Is the winner's curse gone? | No | E5 holdout over the last 50 generations: min 11.66, mean 20.04, max 23.97. | 11 |
2. The setup
2.1 Body
The frog is seven rigid pieces: a torso (with a head and a stiff, fixed front leg), plus two back legs each made of a thigh, a shin and a foot. Six revolute joints (hip, knee and ankle on each back leg) connect them. Each joint has a position-target motor with a strength limit. Nothing is animated: the motors push, and gravity, friction, momentum and collisions decide what happens.
| Joint (each back leg) | Range (radians) | Maximum torque |
|---|---|---|
| Hip | −2.2 to 0.5 | 30 Nm |
| Knee | −0.3 to 1.9 | 30 Nm |
| Ankle | −2.4 to 0.3 | 15 Nm |
2.2 Senses
The brain receives 21 inputs, each scaled to roughly −1 to 1:
- the angle and angular speed of each of the six joints (12 inputs);
- body tilt and spin;
- forward and vertical speed;
- height above the ground;
- whether each back foot is touching the ground;
- a clock: the sine and cosine of a 1.5 Hz tick, which gives the brain a built-in rhythm to work with.
2.3 Brain
A fully connected network of 21 inputs, two hidden layers of 16 neurons with tanh activation,
and 6 outputs: 726 weights and biases in total. Each output picks a target angle within its joint's range,
and that joint's motor drives towards it. Choosing target poses is far easier to evolve than choosing raw
forces. Evolution only ever sees the 726 numbers as a flat list; it never knows they form a network.
2.4 Physics
planck.js 1.4.2, a port of the Box2D engine, with a fixed step of 1/60 s, 8 velocity and 3 position iterations, and gravity of −10 m/s².
2.5 Episode
One episode is 600 steps (10 s). It ends early if the torso or head touches the ground; we call that a fall. In E5 it also ends early if the frog stalls (§3.2). One episode takes about 16 ms of computer time.
3. Method
3.1 Genetic algorithm
Each generation, every brain drives a frog on the same three courses and scores the average. Then:
| Step | Setting | Purpose |
|---|---|---|
| Population | 200 brains | Initial weights drawn at random, standard deviation 0.5. |
| Elitism | best 5 copied unchanged | The best result so far is never lost. |
| Tournament selection | best of 4 picked at random | Good brains breed often; weaker ones sometimes, which keeps variety. |
| Crossover | 30% of children | Uniform: each weight comes at random from one of two parents. |
| Mutation | 10% of weights nudged | Gaussian nudge, standard deviation 0.3 × 0.99generation, never below 0.05: exploring first, fine-tuning later. |
| Courses | 3 per generation | The same three for every brain in a generation, so the comparison is fair. |
3.2 Scoring rules
The score is the only thing evolution sees, so it is where the real design work happens. We used four scoring rules over the course of the project, and for E5 added a stall rule to rule A:
- Distance only (E1):
score = distance − 0.002 × motor work (J)A fall ends the episode but costs nothing extra. - A, scale by survival (E2, E3, E5; the default):
score = distance × (seconds upright / 10) − 0.002 × motor work - B, flat fall penalty (E2):
score = distance − 10 (if fell) − 0.002 × motor work - Airborne only (E4): as A, but only forward distance covered while no part of the frog touches the ground counts.
- Stall rule (E5, added to A): the episode also ends if the torso has not moved at least
0.5 m forward in the last 2 s. A stalled episode is scaled by survival exactly like a fall, so standing still
after a lunge costs the survival multiplier:
score = distance × (seconds before a fall, a stall or the end / 10) − 0.002 × motor work
3.3 Courses
Courses are generated from a seed, so only the seed needs storing and the browser can rebuild the identical ground. We keep four kinds apart, following the same discipline as our image classifier (select on validation, report on test):
| Set | Seeds | Used for |
|---|---|---|
| Training | new random seeds each generation, 3 per generation | Selection during evolution. Gently bumpy. |
| Validation | 2001–2010 (10 courses) | Choosing which champion to ship, from the last 50 generations. Nothing else. |
| Holdout | 1001–1005 (5 courses) | Scoring every generation's champion. Never used for any selection, so the shipped champion's holdout score is an unbiased estimate. |
| Showcase | 0 (perfectly flat) | The site's recordings and the gait analysis. Fixed before any results were seen. |
4. E0: probes before evolution
Before evolving anything, we checked the body and physics with simple controllers.
| Controller | Result |
|---|---|
| Hold the crouch pose | Stands for 10 s; moves −0.16 m |
| Hand-coded "crouch, then kick every 0.8 s", joint torque 40/40/20 Nm | 5.6 m, then flipped (torso at 168°) and fell after 1.3 s |
| Same, torque 25/25/12 | 4.1 m, fell after 1.4 s |
| Same, torque 15/15/7 | 1.9 m, fell after 0.9 s |
| 20 random brains | All fell within 0.4–1.8 s, travelling 0–5.5 m (mostly diving forward) |
- Open-loop kicking always rotates the body in flight. Staying upright needs feedback from the tilt and spin senses.
- With distance as the only score, one uncontrolled leap beats any cautious gait. We expected reward hacking before training started.
- Joint torque was reduced to 30/30/15 Nm (hip/knee/ankle) for all later experiments.
5. E1: the reward hack
A 30-generation smoke run with the distance-only score (§3.2).
| Generation | Best score | Median | Champion |
|---|---|---|---|
| 0 | 4.70 | 0.77 | fell after 1.55 s, 5.6 m |
| 10 | 9.00 | 3.97 | fell after 2.95 s, 11.4 m |
| 29 | 12.02 | 6.82 | fell after 2.88 s, 14.8 m |
The population fall rate was 99–100% in every generation.
Finding. Evolution optimised the score and ignored the intent. It produced ever-longer dives followed by a crash: the reward hack predicted in E0. Diving is not a failure of the algorithm. It is the correct answer to the question we asked.
6. E2: fixing the score
Two cures, 100 generations each, seed 1, both keeping the energy term: A scales distance by the fraction of the episode survived; B subtracts a flat 10 for falling (§3.2).
| Gen | A best | A median | A holdout | A champion | A fell | B best | B median | B holdout | B champion | B fell |
|---|---|---|---|---|---|---|---|---|---|---|
| 0 | 1.56 | −0.29 | 2.89 | 4.8 m, 4.8 s | 99% | 1.04 | −9.23 | 1.04 | 1.5 m, 10 s | 99% |
| 10 | 5.95 | 0.63 | 4.93 | 6.9 m, 10 s | 62% | 4.47 | −5.75 | 4.42 | 4.9 m, 10 s | 80% |
| 50 | 6.30 | 1.05 | 6.06 | 7.3 m, 10 s | 64% | 6.77 | 3.13 | 6.74 | 7.2 m, 10 s | 46% |
| 99 | 6.79 | 5.47 | 5.33 | 7.1 m, 10 s | 34% | 7.44 | 2.87 | 6.12 | 8.0 m, 10 s | 42% |
Fell: share of all the population's episodes that ended in a fall. Holdout: the champion's score on the 5 holdout courses. Champion: average over that generation's 3 training courses. Scores under A and B are not directly comparable: A scales distance, B subtracts a penalty.
- Both cures stopped the suicidal dive within about 10 generations. Champions then stayed upright for the full 10 s (standing still for most of it, as §7 explains).
- Both plateaued at about 7–8 m per 10 s.
- B was noisier. Its median stayed negative until about generation 40, and its holdout score was unstable (1.37 at generation 85 while the training best was 6.84).
Decision. A (scale by survival) is the default for everything that follows. It is simpler to explain and gave a smoother signal.
7. Gait analysis (corrected)
With falling fixed, we measured how the champions move on the flat showcase course: the share of time with no part of the frog touching the ground (airborne), the number of separate flights, and the share of time the stiff front leg is on the ground.
| Champion | Distance | Airborne | Flights | Front leg on ground |
|---|---|---|---|---|
| A, gen 10 | 6.8 m | 14% | 2 | 81% |
| A, gen 99 | 7.0 m | 15% | 5 | 65% |
| B, gen 99 | 8.9 m | 19% | 5 | 73% |
Original finding (superseded): second reward hack. The champions mostly scoot on the fixed front leg with occasional short hops. Nothing in the score says "hop", and scooting is a stable way to cover distance without falling.
The champions do not scoot. Measuring distance second by second showed that they cover nearly all their distance in the first 2–3 s, then stand still for the rest of the episode. The high front-leg contact is standing, not scooting. Airborne share, flights and front-leg contact are totals over the whole episode, so they could not tell a steady gait from a lunge followed by standing still. The distance-by-second table is in §9.2.
The exhibits in this paper are stills from the recordings used by the site. Each still is centred on the frog with the ground fixed, so height above the ground is to scale; gridlines are 1 m apart. Unless a caption says otherwise, stills are taken every 0.4 s for the first 2.4 s, followed by the final recorded frame. A red border on the final still marks a recording that ended early (a fall or a stall); a green border, one that lasted the full 10 s.
8. E3: the main run
300 generations with rule A, seed 1. The run took about 7 minutes. Its selected champion, generation 277, was the one the previous preview of this site shipped.
| Generation | Best | Median | Holdout | Champion distance | Population fell |
|---|---|---|---|---|---|
| 0 | 1.56 | −0.29 | 2.89 | 4.8 m | 99% |
| 10 | 5.95 | 0.63 | 4.93 | 6.9 m | 62% |
| 50 | 6.30 | 1.05 | 6.06 | 7.3 m | 64% |
| 100 | 6.98 | 5.33 | 6.94 | 7.3 m | 35% |
| 150 | 8.19 | 5.89 | 6.12 | 8.7 m | 34% |
| 200 | 8.44 | 6.00 | 6.98 | 8.9 m | 32% |
| 250 | 8.93 | 6.45 | 5.15 | 9.4 m | 31% |
| 280 | 9.24 | 6.28 | 2.48 | 9.7 m | 28% |
| 299 | 8.62 | 6.77 | 6.08 | 9.1 m | 24% |
- Best (training courses)
- Champion on holdout courses
- Median
- Selected champion
8.1 The winner's curse
- From about generation 120, the champion's holdout score often dropped far below its training score: by more than 3 points in 57 of 300 generations.
- Each generation is scored on only 3 courses, so the top scorer is often the individual that was luckiest on those 3, not the most capable.
- The population median kept improving steadily (−0.29 to 6.77), which shows the population as a whole got better even though individual champions were noisy.
8.2 Model selection
The last generation's champion is not necessarily the best one to ship. The exporter scores the champions of the last 50 generations on the 10 validation courses (seeds 2001–2010) and ships the best. The holdout courses (1001–1005) are never used for selection, so the shipped champion's holdout score is still an unbiased estimate. The same procedure is used for E4 and E5.
| Champion | Validation (10 courses) | Holdout (5 courses) |
|---|---|---|
| Latest (gen 299) | 4.36 | 6.08 |
| Selected (gen 277) | 8.26 | 7.62 |
9. Two behaviours, and lunge and freeze
9.1 Two behaviours, as first measured
Late E3 champions switch between two behaviours depending on small differences in the terrain. Each cell is one 10-second episode: distance, time survived, and share of time airborne.
| Champion | Flat (0) | 1001 | 1002 | 1003 | 1004 | 1005 |
|---|---|---|---|---|---|---|
| gen 100 | 7.3 m, 10 s, 14% | 6.9 m, 10 s, 17% | 7.3 m, 10 s, 14% | 7.4 m, 10 s, 16% | 7.4 m, 10 s, 16% | 7.5 m, 10 s, 17% |
| gen 200 | 8.6 m, 10 s, 18% | fell: 9.9 m, 3.1 s, 61% | 8.4 m, 10 s, 19% | 8.4 m, 10 s, 18% | 8.8 m, 10 s, 18% | 8.7 m, 10 s, 18% |
| gen 277 (selected) | fell: 10.0 m, 2.8 s, 62% | 9.4 m, 10 s, 17% | fell: 10.1 m, 2.8 s, 63% | 9.4 m, 10 s, 17% | 9.4 m, 10 s, 17% | 9.5 m, 10 s, 18% |
| gen 299 | fell: 10.9 m, 3.3 s, 55% | 9.3 m, 10 s, 16% | fell: 11.0 m, 3.0 s, 64% | 8.8 m, 10 s, 17% | fell: 9.8 m, 2.6 s, 61% | 8.9 m, 10 s, 16% |
- The safe behaviour is about 17% airborne, reaches about 9.4 m and does not fall. We first called it a "scoot"; see §9.2.
- The hop is about 60% airborne. It covers about 10 m in under 3 s, then crashes.
Original interpretation (superseded). Evolution has found real hopping but cannot yet keep it stable, and the survival-scaled score rewards the safe scoot about as much as a short, fast hop.
9.2 Correction: lunge and freeze
We found this while investigating why the site's ghost race looked static. The "safe behaviour" is not a gait. It is a lunge of 7–9 m in the first 2–3 s, followed by standing still. Our earlier "scoot" description was wrong.
Distance at each whole second on the flat course (seed 0) and holdout course 1001:
| Champion | Course | Distance at 1 s, 2 s, …, 9 s (m) | Outcome |
|---|---|---|---|
| E3 gen 10 | flat | 4.1, 6.9, 6.8, 6.8, 6.8, 6.8, 6.8, 6.8, 6.8 | upright at 10 s |
| E3 gen 50 | flat | 2.7, 4.8, 4.8, 5.4, 6.0, 6.2, 6.1, 6.4, 6.9 | upright at 10 s |
| E3 gen 100 | flat | 4.4, 6.9, 7.3, 7.3, 7.3, 7.3, 7.3, 7.3, 7.3 | upright at 10 s |
| E3 gen 200 | flat | 4.4, 7.7, 8.7, 8.6, 8.6, 8.6, 8.6, 8.6, 8.6 | upright at 10 s |
| E3 gen 277 | 1001 | 4.3, 8.6, 9.4, 9.4, 9.4, 9.4, 9.4, 9.4, 9.4 | upright at 10 s |
| E3 gen 277 | flat | 4.3, 8.5 | fell at 2.8 s |
- Under the survival-scaled score this is optimal: distance is banked early, and standing still keeps the full survival multiplier.
- This is the third reward hack, and the same behaviour we had misreported as the second.
- The "hop" is the same lunge carried on too far, which ends in a fall.
- E5 (§11) addresses this directly.
9.3 A note on the showcase
The course was never changed. The flat showcase course (seed 0) was fixed before any of these results were seen. We did not change it to one where a champion looks better. On it, the E3 selected champion hops for 2.8 s and falls, and the previous preview said so on its front page. The E5 champion (§11) is shown on the same course.
10. E4: airborne-only scoring Complete · corrected
E4 was designed while we still believed the champions scooted. It counts only forward distance covered while no part of the frog touches the ground, still scaled by survival: 150 generations, seed 1. The aim was to stop rewarding scooting.
| Generation | Best | Median | Holdout | Champion distance | Champion airborne | Population fell |
|---|---|---|---|---|---|---|
| 0 | 1.60 | −0.33 | 1.60 | 1.6 m | 9% | 99% |
| 10 | 3.91 | 0.26 | 3.90 | 4.9 m | 10% | 67% |
| 25 | 5.19 | 2.44 | 5.02 | 6.5 m | 18% | 46% |
| 50 | 5.70 | 2.96 | 5.63 | 7.1 m | 18% | 51% |
| 75 | 6.14 | 3.72 | 4.29 | 7.9 m | 20% | 41% |
| 100 | 6.56 | 4.39 | 4.68 | 9.1 m | 20% | 36% |
| 125 | 6.41 | 4.37 | 2.40 | 8.6 m | 17% | 30% |
| 149 | 6.24 | 4.91 | 4.36 | 8.3 m | 18% | 24% |
Scores count airborne distance only, so they are not comparable with E2 or E3 scores. Distances and airborne fractions are comparable.
10.1 Selection
Using the same procedure as E3, generation 120 was chosen, with validation 6.00 and holdout 5.88. The latest champion (generation 149) had validation 4.26 and holdout 4.36.
| Champion | Flat (0) | 1001 | 1002 | 1003 | 1004 | 1005 |
|---|---|---|---|---|---|---|
| gen 50 | 7.1 m, 10 s, 16% | 7.1 m, 10 s, 17% | 7.2 m, 10 s, 17% | 6.9 m, 10 s, 15% | 7.1 m, 10 s, 17% | 7.2 m, 10 s, 17% |
| gen 100 | fell: 9.6 m, 3.3 s, 56% | 9.4 m, 10 s, 23% | fell: 9.6 m, 3.3 s, 55% | 8.6 m, 10 s, 18% | 8.3 m, 10 s, 22% | fell: 6.0 m, 5.1 s, 55% |
| gen 120 (selected) | 8.9 m, 10 s, 20% | 8.9 m, 10 s, 19% | 8.9 m, 10 s, 20% | 6.9 m, 10 s, 16% | 7.2 m, 10 s, 30% | 8.5 m, 10 s, 24% |
| gen 149 | 9.2 m, 10 s, 21% | fell: 8.8 m, 2.5 s, 80% | 8.1 m, 10 s, 28% | 8.6 m, 10 s, 20% | fell: 8.7 m, 2.5 s, 79% | 9.4 m, 10 s, 20% |
10.2 Findings
- Changing the score did not change the behaviour. Champions stayed around 16–30% airborne, much as in E3. We first reported this as a "skip gait"; see the correction below.
- The fast hop appeared again and still was not stable. It reached 55–80% airborne, then crashed within 2.5–5.1 s. Scoring airtime more directly did not make it stable.
- The selected champion (gen 120) is the most robust one tested. It completed all 10 s on the flat course and on all 5 holdout courses.
Original interpretation (superseded by E5). The limiting factor is not what the score rewards but whether the fast hop can be controlled at all.
Distance by second shows the same lunge and freeze as E3. Generation 50 reached 7.1 m by 3 s and generation 120 reached 8.9 m by 3 s, and neither moved after that. Airborne-only scoring still banks the lunge, and survival scaling still rewards standing still. E5 shows that the score was the limit: once standing still was penalised, a continuous hop emerged.
E4 was never shipped, and no E4 recordings were exported, so there are no E4 exhibits.
11. E5: the stall rule Shipped
Rule: the E2-A score (distance × fraction of episode survived, minus the energy term) plus a stall rule. The episode ends if the torso has not moved at least 0.5 m forward in the last 2 s. A stalled episode is scaled by survival exactly like a fall, so standing still after a lunge now costs the survival multiplier. 300 generations, seed 1.
Why: E3 and E4 champions lunged and then froze (§9.2). The stall rule makes continuous forward progress the only way to keep the multiplier.
| Gen | Best | Median | Holdout | Champion distance | Champion upright | Champion airborne | Population fell | Population stalled |
|---|---|---|---|---|---|---|---|---|
| 0 | 0.88 | −0.29 | 0.86 | 5.0 m | 2.5 s | 39% | 98% | 2% |
| 10 | 2.08 | 0.34 | 1.92 | 9.9 m | 3.0 s | 73% | 84% | 16% |
| 25 | 4.24 | 0.57 | 4.02 | 13.7 m | 4.0 s | 66% | 90% | 10% |
| 50 | 11.89 | 0.90 | 5.07 | 18.2 m | 6.5 s | 52% | 88% | 12% |
| 100 | 15.58 | 4.62 | 11.47 | 18.0 m | 10 s | 37% | 66% | 10% |
| 150 | 21.51 | 13.73 | 16.32 | 23.4 m | 10 s | 35% | 41% | 3% |
| 200 | 22.60 | 18.01 | 11.39 | 24.6 m | 10 s | 40% | 31% | 1% |
| 250 | 23.29 | 20.04 | 22.48 | 25.5 m | 10 s | 35% | 22% | 1% |
| 299 | 24.03 | 20.02 | 23.74 | 26.2 m | 10 s | 36% | 27% | 5% |
Champion columns are averages over that generation's 3 training courses. Population fell / stalled: share of all the population's episodes that ended that way.
- Best (training courses)
- Champion on holdout courses
- Median
- Shipped champion
11.1 Selection
Holdout over the last 50 generations: min 11.66, mean 20.04, max 23.97. Selection on the validation courses (seeds 2001–2010, last 50 generations) chose generation 298.
| Champion | Validation (10 courses) | Holdout (5 courses) |
|---|---|---|
| Latest (gen 299) | 21.97 | 23.74 |
| Selected (gen 298) | 24.12 | 23.04 |
The latest champion scored slightly higher on holdout. Holdout is never used for selection, so that does not change the choice; the selected champion's holdout (23.04) is close to its validation (24.12).
11.2 The selected champion, second by second
Measured with the stall rule switched off, to show what the frog does when nothing stops it:
| Course | Outcome | Distance | Airborne | Distance at 1 s … 9 s (m) |
|---|---|---|---|---|
| flat | upright at 10 s | 26.1 m | 35% | 3.5, 6.5, 8.8, 11.4, 13.7, 16.2, 18.9, 21.5, 23.8 |
| 1001 | upright at 10 s | 23.8 m | 37% | 3.5, 6.6, 9.1, 11.8, 13.2, 14.9, 17.2, 18.6, 21.0 |
| 1002 | upright at 10 s | 25.7 m | 35% | 3.5, 6.5, 8.9, 11.5, 14.5, 17.1, 18.7, 21.0, 23.2 |
| 1003 | upright at 10 s | 25.5 m | 37% | 3.5, 6.2, 8.5, 10.6, 13.2, 15.0, 17.5, 20.1, 22.7 |
| 1004 | upright at 10 s | 25.6 m | 31% | 3.5, 6.5, 8.7, 11.1, 13.3, 15.9, 18.3, 20.8, 23.2 |
| 1005 | upright at 10 s | 25.4 m | 39% | 3.5, 6.5, 8.6, 11.2, 13.9, 16.1, 18.0, 20.7, 23.4 |
11.3 Recorded champions on the flat showcase course
With the stall rule on, as shipped on the site's ghost race:
| Generation | Outcome |
|---|---|
| 0 | stalled, 4.18 m at 3.0 s |
| 5 | stalled, 4.48 m at 3.5 s |
| 10 | fell, 9.85 m at 3.0 s |
| 25 | stalled, 15.44 m at 6.8 s |
| 50 | stalled, 16.21 m at 7.8 s |
| 100 | fell, 11.91 m at 5.7 s |
| 200 | upright at 10 s, 25.49 m |
| 298 (selected) | upright at 10 s, 26.12 m |
| 299 | upright at 10 s, 26.38 m |
11.4 Findings
- The stall rule produced a continuous hopping gait. The selected champion moves about 2.5 m/s steadily for the full 10 s, is about 35% airborne, and completed all 10 s on the flat course and on all 5 holdout courses. It is still moving at 10 s.
- The winner's curse is still present (holdout min 11.66 over the last 50 generations), so validation-based selection remains necessary.
- The E3 and E4 limit was the score, not control. Once standing still stopped paying, evolution found a stable repeated hop within about 150–200 generations.
Decision. E5 generation 298 replaces E3 generation 277 as the published champion.
12. Limitations and future work
12.1 Threats to validity
- One seed per experiment, one run each. E2 to E5 all used seed 1, and no repeat runs with other seeds are reported. E5 ran once. How much outcomes vary between runs is unmeasured, so differences such as A against B, or E5 against E3, could partly be luck.
- The stall thresholds were not tuned. 2 s and 0.5 m were chosen once and never varied. We do not know how sensitive the result is to them.
- Three courses per generation. Selection during evolution is noisy by design, which is what produces the winner's curse (§8.1). It is still present in E5.
- Small evaluation sets. The holdout estimate rests on 5 courses and the validation choice on 10, and selection only considered the last 50 generations. Validation and holdout disagree noticeably (for E3 gen 299, 4.36 against 6.08), which is a reminder of how noisy a handful of courses is.
- Summary gait metrics misled us. Airborne share, flights and front-leg contact are totals over an episode, and they did not reveal the lunge and freeze. Distance over time did. Other summaries in this paper may hide behaviour in the same way.
- Gait figures are single episodes. Airborne share and distances by second come from one episode per champion per course. The flat showcase course alone decides what the site shows.
- Two dimensions, one body. The frog is a 2D body with a stiff front leg and a fixed 1.5 Hz clock. Results may not carry over to other bodies, to 3D, or to real animals.
- Energy term untested. The 0.002-per-joule energy cost was kept throughout and never varied or removed, so we do not know how much it shapes the gaits.
- Scores are rule-specific. Scores under A, B, airborne-only scoring and the stall rule measure different things. Only distances and airborne fractions compare across experiments.
- E1 was short. The distance-only run was a 30-generation smoke run; we did not test whether divers persist over longer runs.
12.2 Candidate next steps (not yet tested)
- Repeat E5 with other seeds.
- Vary the stall thresholds.
- More courses per generation, to reduce the winner's curse.
- Longer runs.
- An articulated front leg that can absorb landings.
- An evolvable clock frequency.
- A curriculum from short hops to long ones.
13. Reproducibility
- Deterministic. Every random choice in training comes from one seeded random number generator, and the physics is deterministic. The same configuration and seed produce the same frogs.
- Seeded courses. Every course is generated from its seed; the seeds used for validation, holdout and showcase are listed in §3.3.
- Resumable. Training can be stopped at any time. It finishes the current generation and saves a checkpoint holding the population, the random number generator's exact state and the generation number. The checkpoint is written to a temporary file and then renamed, so a crash mid-save cannot corrupt it.
- Tested. The project's test suite checks that a run stopped and resumed produces exactly the same frogs as one that ran straight through.
- Checked in the browser. The live panel on the front page re-runs a champion's brain in the same physics code, under the same rules including the stall rule, and compares every frame with the training recording.
- Figures from data. Every exhibit in this paper is drawn in your browser from the exported run data (configuration, score history, champion genomes and recordings) of E5 and, for the historical exhibits, E3. None are hand-made images.