Series: Reinforcement Learning from scratch, with Agentic Racing as the lab

  1. Part 1 · How Does an Agent Learn? Teaching a Car to Drive
  2. Part 2 · How Do We Improve a Policy Without Breaking It? PPO Explained from Scratch
  3. Part 3 · Can We Teach by Imitation? Behavioral Cloning Explained from Scratch
  4. Part 4 · Can We Learn to Behave Like an Expert? GAIL Explained from Scratch
  5. Part 5 · Can We Combine Imitation and RL? BC, GAIL, and PPO as One System
  6. Part 6 · How Do We Tell an Agent What We Want? Reward Engineering Explained from Scratch (this article)

Where we left off

Parts 3, 4, and 5 were about imitation: BC, GAIL, and how to combine them with PPO. All three came with the same warning: in Agentic Racing those techniques were configured but never run.

This part is different. The reward does exist, it’s in the code, and it has a history: eight PPO training runs and, according to the devlog, seven variants of the reward between the first and the last. Part 1 showed the final version and told the story of the project’s most famous shortcut. Here we use it as a case study for the underlying question:

How do we turn what we want into a signal a machine can optimize?

And there’s a debt to pay. Part 5 ended with an uncomfortable finding: the project’s reward doesn’t read the strategist’s directives, which are the pilot’s whole reason to exist. We’ll come back to that at the end, with a proposal.

The usual honesty: no version of this reward produced a pilot that completed laps consistently. Part of what this article explains is exactly why.


1. The agent optimizes the signal, not the intention

In part 1 we saw that the agent doesn’t maximize a single step’s reward but the return, the discounted sum of future rewards:

Gt=∑k=0∞γk rt+k+1G_t = \sum_{k=0}^{\infty} \gamma^k\, r_{t+k+1}

In Agentic Racing, γ=0.995\gamma = 0.995 per decision, and the agent decides 10 times per second.

The central idea of this article fits in one line:

The agent optimizes the signal we define, not the intention we had in mind.

If the reward represents the goal badly, the agent can learn a strategy that is excellent according to the formula and disappointing to us. Not because it “cheats”: it simply does what we asked, more rigorously than we asked it.


2. From a human goal to a list of terms

Agentic Racing’s goal, in words, was something like this:

Get around the circuit fast, without leaving the track, without crashing, and without stalling.

The final reward translated that into nine components (the full table is in section 6 of part 1), plus the episode end conditions. Grouped, without claiming an exact mapping, by the part of the goal they try to represent:

Part of the goalReward terms
get aroundprogress per metre, completed-lap bonus
fasttarget speed per second, extra bonus for a fast lap
without leaving the trackcloseness to the edge, off track
without crashingwall hit
without stallingslowness, stopped, no-progress cut
in the right directionwrong-way cut
drive wellfollow the racing line

Each row is a design decision, and each weight is a statement about priorities: “a metre of progress is worth 0.02”, “leaving the track is worth −1”. None of those numbers comes from physics or from the goal; we chose them. The question is what they say when you look at them together.


3. Scales: what each term says over a full lap

The first useful exercise is to add things up. Take the project’s fixed circuit (about 2 km) and a lap of about 95 seconds (the devlog’s estimate of that pace; the heuristic pilot measured 100 s with random directives). With the constants in the code, this is the most each term could pay over that lap:

What one lap of the fixed circuit could pay, term by term: progress about +40, target speed up to about +24, lap bonus a flat +12 plus a fast-lap bonus of only 1.7 at 95 seconds, racing line under +2, leaving the track −1, and each wall hit −0.1. These are bounds computed from the code's constants, not measurements.

A warning: these are bounds computed from the constants, not measurements. No RL policy ever drove this lap; the fixed circuit came after the last training run.

Even so, the sum says things the weight table doesn’t:

  • Progress dominates. About 40 points per lap, almost as much as everything else combined.
  • “Fast” is barely in the reward. The fast-lap bonus scales with how much of the episode’s time limit (120 s) is left. A 95 s lap earns about +1.7; a 79 s lap, about +2.7. Sixteen seconds of difference are worth one point, on a lap that can pay up to about 80.
  • The racing line is almost decorative. Under 2 points per lap, whatever the lap time: since it’s multiplied by the car’s speed, in practice it pays per metre.
  • Leaving the track costs the same as 50 metres of progress. Seen that way, −1 is cheap. But its real cost isn’t the −1: it’s everything the agent stops collecting over the rest of the lap. We’ll come back to that in section 6.

The project learned this exercise the hard way too. After the second run, the devlog notes that the newly added speed and racing-line terms were, per step, 10 to 30 times smaller than the progress term and about 1000 times smaller than the off-track penalty. They were in the reward, but they didn’t weigh anything.

The ML-Agents documentation gives a concrete scale rule: the reward between two decisions should be in the range [−1,1][-1, 1], because larger values can destabilize training. The off-track penalty respects it. The lap bonus, 12 to 20 in a single decision, doesn’t. In practice it barely mattered, for a worse reason: the agent almost never got to collect it.


4. Sparse and dense rewards, with numbers

A sparse reward pays only when the goal happens: for example, +1 for completing the lap and nothing else. A dense one pays small amounts along the way, like progress per metre.

The sparse one is closer to the goal. The problem is whether the agent can see it. With γ=0.995\gamma = 0.995 and a lap of about 950 decisions, a prize at the end of the lap, seen from the start, weighs:

0.995950≈0.0090.995^{950} \approx 0.009

less than 1% of its value. And that’s assuming the agent ever reaches the finish, which is exactly what wasn’t happening: from the fourth to the eighth run, only 2% to 4% of episodes ended with a completed lap. The devlog for the second run sums it up in two words: the lap bonus was “dead weight”.

That’s why the project’s reward is dense. Most likely, progress per metre is what taught the agent the first 10% of the lap: start moving, go forward, don’t leave the track on the first straight.

But every dense term is a new incentive, and the useful question isn’t “sparse or dense?”, but:

Does the intermediate signal help the agent discover the final goal, or does it teach it a different goal?


5. Units: per metre, per second, per event

The reward’s terms aren’t measured in the same units, and that changes what they mean.

  • Per metre (progress): pays the same for a metre whether it’s covered in one second or in ten. It doesn’t depend on how long the episode lasts.
  • Per second (target speed, edge, slowness, stopped): pays more the longer the situation lasts. A positive per-second term is, at bottom, a reward for staying alive.
  • Per event (lap, off track, hit): pays once, when something happens.

(The racing line is an in-between case: it’s computed per second, but multiplied by speed, so in practice it pays per metre.)

A good consequence of the design: since the per-second terms are multiplied by the physics step’s duration, they don’t change if the simulation frequency changes. Not everything is like that: the episode limit, the decision cadence (and with it γ\gamma), and the fast-lap bonus are counted in steps. A less visible one: the fast-lap bonus depends on the episode’s step limit, MaxStep. When the project moved to the fixed circuit, that limit went from 4000 to 6000 steps (from 80 to 120 seconds), because 80 seconds wasn’t enough to close a lap. With that change, the same lap time became worth a different bonus, without touching a single reward constant.

The most interesting interaction is between progress (per metre) and target speed (per second). The speed term peaks exactly at the target speed, which the code computes from the upcoming corner: 42% of top speed on a straight (23.1 m/s) and less in corners. You’d think the total reward also peaks there. Not necessarily.

Per second, progress pays 0.02 v0.02\,v and the speed term loses 0.25×1.3/vtarget0.25 \times 1.3 / v_{\text{target}} for every m/s above the target. The two balance out when vtarget=0.25×1.3/0.02=16.25v_{\text{target}} = 0.25 \times 1.3 / 0.02 = 16.25 m/s (a bit less if you add the racing line):

  • With a target below 16.25 m/s, there’s a local peak at the target: braking down to it pays more than overshooting a little. But it’s only local: at much higher speeds the speed term reaches zero and progress alone beats it. With a 10 m/s target, for example, going above about 22.5 m/s would pay more per second than braking.
  • With a target above it, there’s no peak: the reward keeps rising past the target.

And here’s the finding. A 10 m/s target corresponds to tight corners like those on the procedural tracks. On the fixed circuit, with 120 m radius corners, the target speed never drops below about 18.1 m/s, above the threshold. In other words, on the circuit where the project ended up, the speed term never tells the agent to brake: per second, more speed always pays more, on straights and in corners (in the corners, only barely).

Reward per second from progress plus target speed, as a function of the car's speed, for three targets. With 10 m/s, a tight corner on the procedural tracks, there's a local peak at the target, but above about 22.5 m/s progress alone pays more. With 18.1 m/s, the fixed circuit's corners, and 23.1 m/s, the straights, the curve keeps rising past the target.

So what would make braking right in a corner isn’t the speed term. It’s the chance of leaving the track, which ends the episode and cuts off everything that would have come after. If the car can take the corner faster without leaving the track, the reward prefers it that way. That doesn’t make the speed term useless (in tight corners it does create a zone where braking pays), but it shows that the final behavior comes from the combination, not from reading each term on its own.


6. Ending the episode is also a reward

This is the most underestimated point, and in Agentic Racing it was the origin of the most famous shortcut.

When an episode ends, the agent doesn’t just receive that moment’s penalty. It loses everything it would have collected afterwards. So the agent’s real decision isn’t “is the −1 worth it?”, but “is the −1 worth more than the future?”. And that depends on the sign of what the future pays:

  • If the expected per-step reward is negative, ending early is a reward.
  • If it’s positive, surviving is a reward.

Both sides have shown up in this series:

The negative side: the fourth run. The reward had a small flat per-step penalty so the agent would finish quickly, and stopping ended the episode with −1. The agent learned to use the rolling start, cover about 245 metres collecting progress, stop, and accept the −1:

The fourth run's shortcut: the car covers about 245 metres, stops on purpose, and the episode ends with a return of +4.9 from progress, −0.08 from the time penalty, and −1 for stopping, about +3.8, exactly where the training curve plateaued.

The curious part is that the per-step penalty is exactly what the ML-Agents documentation recommends to get an agent to finish a task quickly, with one condition: completing the task should coincide with the end of the episode. Here failing also ended the episode, and failing cost less than continuing to try. The generic recipe, applied to an environment where failure also ends the episode, rewarded failing fast.

The positive side: GAIL. In part 4 we saw that GAIL’s reward is never negative, so it stretches episodes out. The same goes for any positive per-second term, like the target speed from the previous section.

Ending the episode is also a reward: with a negative per-step reward, ending early pays (the fourth run); with a positive per-step reward, surviving pays (GAIL's bias from part 4). In the middle, the project's episode end conditions and what each one pays.

That’s why the episode end conditions are part of the reward design, as much as the weights. The project’s, after several versions:

Episode endConditionReward
off track2 m past the edge−1
stopped8 s stopped, after having started−1 (plus −0.3 per second while stopped)
wrong way4.5 s reversing faster than 7.5 m/s, only after 15 m of progress−1
no progressunder 8 m of track in 5 s−1
completed lap99% of the course+12 to +20
time limitMaxStep0

Almost every row is a scar. The fourth run’s fix was that stopping no longer ended the episode and instead cost something every second, with a hard cut only at 8 seconds. Wrong way became more tolerant because it ended the episode on the only maneuver that gets a car off a wall: a short reverse. And the no-progress cut arrived when the heuristic pilot’s evaluation showed cars in a “sawtooth” near the start: they hit, backed up a little, hit again, and since every reverse reset the “stopped” timer, they survived 80 seconds without moving forward.


7. Reward hacking: what you can check in the code

Reward hacking is when the agent gets a lot of reward through a path that isn’t the one we wanted. The practical rule for hunting it before training is one question:

What other behaviors could produce this same reward?

With the code in hand, several of the classic suspicions can be answered without training anything:

SuspicionIs it possible in the code?
Go back and forth to collect progress twiceNo: progress is signed, going back subtracts what going forward added
Collect a fake jump when crossing the finish lineNo: the progress calculation corrects for the circuit’s wraparound
Cut corners off the trackNo: 2 m beyond the edge the episode ends
Repeat a collision eventYes, but it’s a penalty: repeating it only subtracts
Sit still and aligned to collect the racing-line termNo: that term only pays if the car is moving, in proportion to its speed
Drive in circles collecting the speed termPartly: that term uses the car’s speed toward its own front, not along the track, so a car driving in circles collects it without moving forward. At the target speed a circle doesn’t fit between the walls, but at low speed it does, for a small gain. The no-progress cut ends it within 5 to 10 seconds with −1

The last row is the interesting one. There’s no evidence the policy ever did it (the circles in the middle of the track showed up in the heuristic pilot, not in RL), but the hole exists in the definition, and what plugs it is an end condition, not the reward. It’s the kind of thing worth knowing before a training run discovers it.

And then there are the shortcuts that aren’t shortcuts but conflicts between terms, which the project did live through:

  • The racing line against the walls. Rewarding following the racing line seemed reasonable, but on this circuit the racing line hugs the edge in every corner. The term pushed the car into the wall. The fix was to lower its weight and add a soft penalty for getting close to the edge.
  • Slowness against braking. The fifth version penalized going slow to keep the agent from crawling. But braking before a corner is going slower, so the penalty landed right on the maneuver the agent needed to learn. The seventh version replaced it with a target speed based on the corner, with a slowness penalty now measured against that target.

8. Reward shaping: helping without changing the goal

Reward shaping means adding intermediate signals to make learning easier. The risk is that they change what the agent considers optimal. There is a form of shaping that, under certain conditions, is proven not to change it (Ng, Harada, and Russell, 1999). It’s based on a potential function Φ(s)\Phi(s) of the state:

F(s,s′)=γ Φ(s′)−Φ(s)F(s, s') = \gamma\,\Phi(s') - \Phi(s)

which is added to the original reward:

r′(s,a,s′)=r(s,a,s′)+F(s,s′)r'(s, a, s') = r(s, a, s') + F(s, s')

The intuition: whatever the agent gains by getting closer to the goal it gives back if it moves away, so there’s no way to pile up shaping by going around in circles.

How close is Agentic Racing’s reward to this?

  • Progress is very similar. It’s the difference of a potential, Φ=0.02×\Phi = 0.02 \times (metres of the lap covered), between one step and the next. But it lacks the γ\gamma factor, and the potential doesn’t go back to zero when the episode ends; both are conditions of the theorem. Besides, if you take the completed lap and the terminations as the “true” reward, progress isn’t a small adjustment on top of it: it’s most of what pays. The guarantee doesn’t apply as is.
  • Target speed, racing line, and edge are not. They aren’t differences of a potential: they pay according to how the car moves, not just where it is. Those terms do change what’s optimal, and on purpose: they tell the agent how fast to go and where.

That last point has an interesting reading in light of parts 3 to 5. The reward’s target speed is a variant of the heuristic pilot’s braking rule: the code comment says it “mirrors the heuristic”, and both compute a speed from the upcoming corner, with similar constants. In other words, the reward already contained a form of imitation, written by hand instead of learned from demonstrations. The ML-Agents documentation advises the opposite: reward results, not the actions you think will lead to them. The project chose to reward an action (a speed) because the result (completing the lap) wasn’t showing up.


9. The reward can’t be the only metric

If we evaluate the agent with the same signal we trained it on, we inherit its blind spots. The project learned this through three concrete episodes:

A rising curve wasn’t an improvement. In the third run, the cumulative reward rose to about 10, two to three times the previous runs’ plateau, and episodes got five times longer. It looked like the hoped-for breakthrough. The fourth run, with more steps and the same configuration, fell back into the usual pit. The devlog closed it like this: the third run had been “favorable variance from one run”. A single run isn’t evidence.

A stable number wasn’t a good number. The fourth run’s plateau, around +3.8, matched exactly the return of section 6’s shortcut. The curve wasn’t saying “it learned little”: it was saying “it learned perfectly to do something else”.

Reward curves from different versions can’t be compared. The fifth run started at −34; the first, at −1.1. Not because the fifth was worse, but because the reward had changed. Each version has its own scale, and comparing their curves means comparing different units.

That’s why, after the fourth run, the project built an independent evaluator (eval.exe) that doesn’t look at the reward. It runs the policies and reports other things:

Evaluator metricWhat it revealed
episode end reasonsthe fourth run ended “stopped” 86% of the time
lap fraction coveredevery run stayed around 10%
mean brake0.02: the policy never braked, and without braking you can’t take a corner
mean speed, throttle, and steeringtwo opposite styles with the same lap fraction: crawling at 7 m/s, or speeding up and ending stopped or off the track

The mean brake of 0.02 was the most useful number of the whole phase, and it wasn’t on any reward curve.


10. What no reward could fix

There’s one last lesson, and it’s the most expensive one. The usual intuition, and the one the project’s plan also held, is that when an agent learns something weird, the reward is almost always to blame. In Agentic Racing the devlog counts seven variants of the reward. The number that mattered, the lap fraction, didn’t budge from between 8 and 13%.

Meanwhile, the devlog kept naming “root causes”, one after another, that weren’t the reward: the walls (zero-thickness meshes the car got stuck on), lateral grip (which erased lateral velocity instead of redirecting it, so every hard turn was a sudden stop), and the 12 m radius corners on the procedural tracks. Some were ruled out later: three different versions of the walls gave the same result, and the devlog concludes “the walls weren’t the cause”. Others were fixed before the eighth run, which still plateaued in the same place.

What unblocked the problem was something else: replacing the procedural tracks with the fixed circuit. There the heuristic pilot, for the first time, drove clean laps, and the devlog attributes the eight runs’ “paralysis” to the procedural tracks (degenerate centreline segments, malformed walls in corners, corners nobody could take), not to a deep bug in the car’s physics. Whether PPO learns to complete laps on the fixed circuit was never tested.

The experiment that could separate “reward problem” from “environment problem” was run halfway through, after the fifth run: putting the hand-written heuristic pilot in the same environment. But it was read in the environment’s favor: its best episode reaching 82% of a lap was taken as proof that the environment was learnable, and after the seventh run the devlog went as far as writing “the reward is already correct”. The full answer only came with the fixed circuit, when the heuristic pilot went from one good isolated episode to complete laps. The lesson isn’t just to run that experiment: it’s to demand a clear answer from it. If a hand-written controller can’t complete the lap consistently, no reward is going to teach a neural network to do it.

The ML-Agents documentation suggests something similar with a different intent: use the agent’s heuristic to drive it and watch how it accumulates reward. Applied here, that would have served two purposes at once: knowing whether the environment was drivable, and knowing whether the reward ranked the expert above the policies that failed. The project did the first; the devlog doesn’t record the second ever being measured.


11. Part 5’s debt: a reward that knows nothing about directives

In part 5 we saw that the pilot exists to follow the strategist, and that the heuristic pilot does: on the same track, it takes 112 s per lap with “conserve” and low aggression, and 79 s with “attack” and high aggression. The reward, on the other hand, computes its target speed from the corner alone. No term reads the directive channels.

The heuristic pilot does scale its target speed by the directive, through a single map (StrategyDirectiveMap) that the strategist also uses. The most direct proposal would be for the reward to use that same map:

vtargetdirective=vtarget(corner)×scale(directive)v_{\text{target}}^{\text{directive}} = v_{\text{target}}(\text{corner}) \times \text{scale}(\text{directive})

(In that map, the speed scale comes from aggression, between 0.86 and 1.16, with 10% less when the directive is “conserve”; the “attack” type mostly changes the racing line.) With that, the reward would tell the agent “with high aggression, the right speed is higher”, and a policy optimizing it would have a reason to respond to the directive.

It’s a proposal, not something tested, and it comes with its own questions: whether the scale should also apply to braking margin and racing line, as the heuristic does; whether the agent would learn to tell apart directives that differ only in speed in the reward; and whether it would hurt training stability. But it illustrates this article’s point: part 5’s finding isn’t fixed with an algorithm, it’s fixed by saying more precisely what we want.


12. A reproducible experiment

If the RL pilot is picked up again on the fixed circuit, the natural experiment compares versions of the reward, changing one thing at a time:

VariantRewardQuestion
A. Sparsecompleted lap and terminations onlycan it learn without a dense signal on this circuit?
B. Currentthe eighth run’sthe baseline: never trained on the fixed circuit
C. No target speedB without that termhow much does section 8’s hand-written imitation contribute?
D. With directivesB with the target speed scaled by the directivedoes directive response appear?

Comparisons: A vs. B, B vs. C, and B vs. D. For them to mean anything:

  • Same build, same circuit, same MaxStep. The fast-lap bonus depends on it (section 5).
  • Rebuild the player for every variant. The reward’s weights come from the code’s default values, not from a scene; changing them requires rebuilding the training executable.
  • Log every reward component separately. Today the project doesn’t log them separately: TensorBoard only sees the total reward. ML-Agents can publish custom metrics to TensorBoard from the agent’s code; without that, there’s no way to know which term is moving the curve.
  • Several seeds per variant, with mean and spread. The third run already showed what a single one is worth.
  • Evaluate with eval.exe, not with the reward. End reasons, lap fraction, mean brake and, for D, lap time under each forced directive.
  • Also measure the expert under each reward. If the heuristic pilot doesn’t earn more reward than a failing policy, the reward is measuring something else.
MetricABCD
Completed laps????
Lap fraction????
Mean brake????
Off-track exits????
Attack / conserve gap????
The expert’s reward under that function????

None of this was run. The cells stay at ”?“.


13. A workflow for designing rewards

Putting it all together, the loop the project would have liked to follow from the start:

0. Check that the environment can be solved with a simple controller
            ↓
1. Write the goal down in words
            ↓
2. Translate it into measurable events and terminations
            ↓
3. Start with a minimal reward
            ↓
4. Add up each term over a lap: scales, units, signs
            ↓
5. Train, logging each component separately
            ↓
6. Evaluate with metrics that aren't the reward, across several seeds
            ↓
7. Look at concrete episodes, not just curves
            ↓
8. Change one thing, and go back to step 4

Step 0 doesn’t usually show up in reward engineering guides. In Agentic Racing it’s the one that would have saved the most time.


Conclusion

Reward engineering isn’t handing out positive and negative points. It’s translating a goal we understand into a signal the agent can optimize, and then checking that optimizing that signal produces what we wanted.

Agentic Racing’s reward leaves several concrete lessons:

  • weights say nothing until you add them up over a full episode;
  • units (per metre, per second, per event) change what each term rewards;
  • ending the episode is a reward, and its sign depends on what the future pays;
  • terms can fight each other, like the racing line against the walls or slowness against braking;
  • a reward can contain hand-written imitation without anyone calling it that;
  • the reward can’t be the only metric;
  • and no reward fixes an environment where the problem can’t be solved.

The question worth asking before every training run is still the same:

What exactly are we rewarding, and what else could the agent learn to do to get that reward?

If you landed directly on this article, the five earlier parts are linked at the top, and Agentic Racing: A Pilot That Never Learned to Drive tells the project’s full story.


References