Series: Reinforcement Learning from scratch, with Agentic Racing as the lab
- Part 1 · How Does an Agent Learn? Teaching a Car to Drive
- Part 2 · How Do We Improve a Policy Without Breaking It? PPO Explained from Scratch
- Part 3 · Can We Teach by Imitation? Behavioral Cloning Explained from Scratch (this article)
- Part 4 · Can We Learn to Behave Like an Expert? GAIL Explained from Scratch
- Part 5 · Can We Combine Imitation and RL? BC, GAIL, and PPO as One System
Where we left off
Part 2 ended with an uncomfortable reading of what happened in Agentic Racing. PPO stably optimized the objective we gave it, but it can only reinforce what the policy already tries now and then. The full corner maneuver (brake, turn, accelerate again) is a coordinated sequence of one or two seconds, and if the policy almost never tries it by chance, no policy gradient is going to invent it.
In section 10 of that part we mentioned, in a couple of lines, a way out of that problem: starting from an expert’s demonstrations. This article develops it.
The idea is simple. What if we already have a driver who knows how to drive? Instead of forcing the agent to discover from scratch what to do, we can watch the driver and record their decisions:
observation → action
observation → action
observation → action
...
And then train a neural network to learn that relationship. That’s Behavioral Cloning (BC):
Instead of learning only by trial and error, we can learn by imitating examples of behavior.
An honest warning, as in the previous parts
In Agentic Racing, BC was prepared but never run. The project has an expert (the heuristic pilot that drives the demo’s cars today), a demonstration recorder, and it once had a training configuration with BC. But the project’s log never records a demonstration being recorded, and the last training run was done with plain PPO.
It fell by the wayside for two reasons. The first is a good lesson about imitation in itself: every time the plan was “record demonstrations”, a flaw showed up in the expert or the environment that was worth fixing first. The teacher had to learn to drive before it could teach. The second was a deliberate decision: with the lateral grip already fixed and the tracks smoothed, the eighth run was done with plain PPO, without imitation, to find out whether the algorithm alone was enough (we covered this in section 10 of part 2). It wasn’t, and shortly afterward the project set the RL pilot aside; BC was never tried again.
So in this article the BC concepts come with what the repository actually has, what it doesn’t, and what would have happened if it had been used with that particular expert.
1. From reinforcement learning to imitation learning
In PPO, the fundamental loop is the one from part 1:
observation
↓
policy
↓
action
↓
environment
↓
reward
↓
learning
↺
The agent has to gradually discover which actions lead to good outcomes.
In Behavioral Cloning we change the problem. We have an expert and we record what it does:
EXPERT
│
▼
DEMONSTRATIONS
│
▼
DATASET
│
▼
NEURAL POLICY
We can write it as:
The policy tries to reproduce the expert’s behavior. We no longer tell it:
“Discover the best action.”
We tell it:
“When you see something like what the expert saw, do what it did.”
Notice what disappears: the reward. BC doesn’t use it. It doesn’t know whether the expert drove well or badly; it only knows what it did.
2. What exactly is a demonstration?
A demonstration is a sequence of experiences produced by the expert:
t = 0 observation → [...] action → [...]
t = 1 observation → [...] action → [...]
t = 2 observation → [...] action → [...]
Each step contains a pair:
where is the observation and is the action the expert took.
A demonstration isn’t a video. A video shows what the driver did; a demonstration for learning contains exactly the data the policy will receive and produce, so that observation can be linked to action.
In Unity ML-Agents, demonstrations are recorded with the DemonstrationRecorder component, which
stores in a .demo file the agent’s observations and the actions it took at each decision (in our
case, 10 times per second).
In Agentic Racing, the project’s evaluator has a mode meant for this:
eval.exe -record
That mode does three things: it drives every car with the heuristic pilot, adds a
DemonstrationRecorder to each one, and runs for 300 seconds, writing one .demo file per car.
Because the recorder sits on the same agent that gets trained, each demonstration contains the same
42 observation numbers and the same 3 continuous actions we saw in part 1. Expert and policy speak
exactly the same language.
What we don’t know is whether a .demo file was ever generated. They’re data, so they’re
deliberately excluded from git, and the project’s log doesn’t record a recording ever being run.
3. The expert doesn’t have to be human
When we think of imitation, we tend to picture a person driving. It doesn’t have to be that way. The expert can be a person, a traditional controller, a rule-based system, an already-trained policy, or another agent.
In Agentic Racing, the natural expert is the heuristic pilot, RaceAgent.Heuristic(): the same
hand-written controller that drives the demo’s cars today, and whose simplified version we saw in
section 17 of part 1. At each step it:
- looks for the sharpest corner over the next 60 meters and computes a target speed;
- accelerates if it’s below it, brakes if it’s clearly above;
- aims the steering at a point between the racing line and the track’s centerline (closer to one or the other depending on the directive), 12 to 36 meters ahead depending on speed.
And it produces its actions in exactly the same format as the policy: steer and throttle between
−1 and 1, brake between 0 and 1.
Heuristic pilot
│
▼
demonstrations
│
▼
Behavioral Cloning
│
▼
Neural Policy
One detail makes this expert especially interesting: its behavior changes with the strategist’s directive. More aggression means a higher target speed and braking later; the more conservative directives pull it toward the center. When recording without forcing a directive, the evaluator assigns a random one each episode, just as in training. So demonstrations recorded with it today wouldn’t just teach how to drive: they’d teach how to drive differently depending on the directive, which is exactly what the six directive channels in the observation needed to learn.
With one caveat about dates: when the BC configuration was prepared, the directive wasn’t yet wired into the heuristic pilot; that happened a day later. Demonstrations recorded at that moment would have taught the policy exactly the opposite: to ignore the directive channels.
Having an expert in the repository doesn’t mean a BC pipeline was ever run. This is the actual state:
| Piece | Does it exist? | Where |
|---|---|---|
| Expert | Yes | RaceAgent.Heuristic(), in RaceAgent.cs |
| Demonstration recorder | Yes | eval.exe -record, in EvalRunner.cs |
| Recording and training instructions | Yes, but outdated | training/README.md, section 6 (says the configuration includes the BC and GAIL blocks; it no longer does) |
| BC and GAIL configuration | It existed | added for the eighth run and removed before it was launched |
.demo files | None recorded | the log doesn’t record a recording; they’re data and aren’t versioned |
| A run with BC | No | the eighth run was done with plain PPO |
| BC results | No | there are no metrics to compare |
4. BC turns imitation into supervised learning
Suppose we have a set of demonstrations:
Each example contains an observation and the action the expert took. We want to train a neural network so that:
(Here we write as the action the policy produces, to keep it simple; in reality, as we saw in part 1, the policy produces a distribution and the action is sampled from it.)
The process is ordinary supervised learning:
observation
│
▼
neural network
│
▼
predicted action
│
▼
compare with the expert
│
▼
loss
│
▼
backpropagation
│
▼
update θ
The difference lies in what the dataset represents. In a classic problem we’d have image → cat.
Here we have environment observation → expert action.
5. Which loss does Behavioral Cloning use?
It depends on the action space.
Continuous actions. If the expert produces values like:
steer = -0.35
throttle = 0.72
brake = 0.00
the most common formulation penalizes the distance between the expert’s action and the policy’s (up to a constant factor, depending on how it’s averaged over the action dimensions):
Discrete actions. If the actions were categories (LEFT, STRAIGHT, RIGHT), the problem is
classification and a loss like cross-entropy is used.
ML-Agents 1.1.0 does exactly that: mean squared error for continuous actions and cross-entropy for
discrete ones. For continuous actions it compares the action the policy samples from its
distribution with the expert’s action. That has an interesting side effect: since the sampling noise
also adds to the error, BC doesn’t just pull the policy’s mean toward the expert’s, it also shrinks
its spread, that is, how much it explores. TensorBoard shows it as Losses/Pretraining Loss.
In our car, with three continuous actions, it would be the mean squared error over steer,
throttle and brake.
The key idea:
The policy receives an observation and is penalized when its action moves away from the one the expert took.
6. Training the imitator
The basic loop is:
- take a batch of demonstrations;
- pass the observations through the policy;
- get the predicted actions;
- compare them with the expert’s;
- compute the loss;
- backpropagate;
- update the parameters;
- repeat.
The advantage over RL from scratch is huge. The agent doesn’t have to stumble onto how to take a corner: we show it how. It’s like teaching a student: “when you’re in this situation, this is what I’d do”.
In ML-Agents there’s an implementation detail that changes how it’s best thought about. BC isn’t a
separate phase before PPO training. It runs inside the same training: after each PPO update, the
policy also gets a BC update, which by default makes several passes over the demonstrations in
minibatches, with its own optimizer. That update uses a learning rate equal to PPO’s initial learning
rate multiplied by strength, decaying linearly to almost zero over the first steps steps.
The configuration that was written for Agentic Racing’s eighth run used strength: 0.5 and
steps: 2000000. In other words, BC would have pushed hard at the start and faded out over the first
2 million steps of a 10-million-step run, leaving the rest to PPO.
7. So… have we solved the problem?
No. Here comes Behavioral Cloning’s most important limitation.
Imagine the expert always drives near the center of the track:
────────────────────────────────
🚗
↓
expert trajectory
────────────────────────────────
The policy mostly learns those states. But when it runs, it may make a small mistake:
────────────────────────────────
🚗
↘
↘
────────────────────────────────
Now the car is in a state that may never have appeared in the demonstrations, and the policy has to decide what to do. If it doesn’t know how to recover either:
small mistake
↓
unseen state
↓
wrong action
↓
bigger mistake
↓
another unseen state
↓
...
This phenomenon is known as covariate shift and as compounding errors.
8. The expert and the policy visit different distributions
During training, the policy sees states visited by the expert:
But when we run it, it visits the states its own decisions lead it to:
And there’s no guarantee that:
At first they may look alike. After a small mistake, they start to drift apart. The expert never had to demonstrate what to do in a really bad situation, because the expert doesn’t usually get there. The policy, on the other hand, can.
Ross and Bagnell’s classic analysis puts numbers on it: if the imitator makes a mistake with probability in the expert’s states, then over a task of steps its performance can fall behind the expert’s by an amount that grows with , not . Each mistake doesn’t just cost its own share: it also takes the imitator to states where it will make more mistakes. Across Agentic Racing’s eight runs, a training episode could last up to 800 decisions.
That’s why we can get:
an excellent result on the dataset and a bad policy at run time.
9. An example with our car
In Agentic Racing this problem isn’t hypothetical: it’s written into the configuration itself.
The demonstrations would always start in the best possible state. Recording mode turns on the
evaluator’s CleanSpawn option: each car appears on the centerline, perfectly aligned with the
track, already at 14 m/s. During training, on the other hand, every episode deliberately starts with
noise: up to ±10° of heading, up to ±2 m of lateral offset, and at 8 m/s. From the very first step,
the trained policy would visit states the demonstrations never show.
The expert has its favorite zones. One of the observations is the car’s lateral offset from the center, divided by half the track width: 0 is the center, and −1 and +1 are the edges. The heuristic pilot aims at a point between the racing line and the center, and also corrects toward the center, so on the straights and with conservative directives its demonstrations would cluster near 0; in corners and with aggressive directives, more toward the racing line. Nobody measured that distribution, but the code does make one thing clear: the expert has no reason to spend time pressed against the wrong wall or sideways on the track. If the policy gets there, it has to extrapolate, and extrapolating outside the training distribution is much harder than interpolating inside it.
The expert itself was fragile outside its nominal state. In an earlier version, spawning on the racing line instead of the center was enough to send the heuristic pilot into an oscillation that ended in a spin: its correction toward the center was too strong for that starting point. And the evaluator’s code says so explicitly: training’s noisy spawn stalls the heuristic before it settles on the line, which is why the expert was only evaluated, and would only be recorded, from clean spawns. An expert that can’t recover from a situation can’t teach how to recover from it.
And today there would be one more problem: all the arenas use the same fixed circuit, so the demonstrations would cover a single layout.
Having many demonstrations doesn’t mean having a robust policy. Coverage matters as much as quantity.
10. The expert’s quality matters too
BC learns from the expert, and that includes its flaws. If the expert brakes too much, takes inefficient lines, makes mistakes, or produces noisy actions, BC can learn all of that.
Not necessarily:
Agentic Racing had a very concrete example. Before recording, someone took a closer look at how the heuristic pilot drove on the procedural tracks of the time. Only 2 or 3 out of every 9 cars drove well; the rest got stuck in a back-and-forth near the starting point: they hit the edge, did a short reverse, slowed down again, and repeated, for the full 80 seconds of the episode, without making progress. The log says it bluntly: those trajectories “would fill the demos with garbage”. The fix was to add a no-progress cutoff (less than 8 meters in 5 seconds ends the episode), so that stuck cars would be reset instead of recording minutes of an expert that wasn’t actually driving.
It was the first of several times that recording the demonstrations was postponed because the teacher wasn’t ready yet.
BC is imitation. It isn’t, by itself, an algorithm whose goal is to beat the expert.
11. So what is BC good for?
Despite those limitations, BC can be very useful, above all as a starting point:
Expert
│
▼
demonstrations
│
▼
Behavioral Cloning
│
▼
initial policy
│
▼
PPO
│
▼
improved policy
PPO no longer starts from scratch: it starts with a policy that can already do something reasonable. And that attacks part 2’s problem head-on. The agent doesn’t need to stumble onto the brake–turn–accelerate sequence by chance if someone shows it. It’s enough for BC to leave it in a zone where the policy tries that sequence often, so PPO can reinforce it.
That’s why imitation learning and reinforcement learning complement each other.
12. BC + PPO: teach first, optimize later
We can think of two teachers:
- The expert says: “this is how I drive”.
- The environment says: “this is what actually works to reach the goal”.
BC listens to the expert. PPO listens to the reward.
In ML-Agents, as we saw in section 6, both listen at the same time, with a volume that changes over time: BC’s weight starts high and fades out, while PPO keeps going for the whole training.
It’s a smooth form of “teach first, optimize later”: there’s no hard cut between the two stages, just a transition. In Agentic Racing that pipeline existed in the configuration, but it was never run.
13. And where does GAIL come in?
BC tries to learn directly.
GAIL (Generative Adversarial Imitation Learning) asks a different question:
Does my agent’s behavior look like the expert’s?
It trains a discriminator that tries to tell the expert’s experiences apart from the agent’s. The agent gets an extra reward when its experiences are hard to tell apart from the expert’s. In ML-Agents, GAIL is literally that: one more reward signal, added to the environment’s reward and optimized by PPO like any other.
The eighth run’s configuration also had a GAIL block, with a weight of 0.15 and comparing actions as well as observations.
So each technique answers a different question:
| Technique | Main question |
|---|---|
| Behavioral Cloning | Can you do what the expert did in this state? |
| GAIL | Does your behavior look like the expert’s? |
| PPO | Can you maximize the task’s reward? |
GAIL deserves its own article, and it’s part 4 of the series.
14. The real challenge: what are we teaching?
BC forces us to ask a deeper question:
Does the observation contain enough information to make the right decision?
Suppose two different situations:
situation A → the expert turns left
situation B → the expert turns right
If both produce exactly the same observation for the policy, the problem isn’t the neural network: the information simply isn’t there. And with a squared-error loss, the policy doesn’t pick one of the two answers: it learns the average, which may be not turning at all.
In Agentic Racing this is concrete, because the heuristic pilot doesn’t decide from the 42 observations alone that the policy would receive:
- It has memory. Its steering goes through a smoothing filter: each command blends the new calculation with the previous command. The policy from part 2 is a network with no memory, which only sees the current observation.
- It uses timers. If the car has been slow and against the wall for a while (the code counts it as more than a second), the expert does a 0.6-second reverse to break free. Same visible state, different actions depending on how long it has been stuck. For the policy, that would be exactly situations A and B above: sometimes accelerate, sometimes reverse, for the same observation.
- It looks at the track differently. The expert looks for the sharpest corner over the next 60 meters and aims at a specific point between the racing line and the center. The policy doesn’t do that scan: it gets three curvature numbers (0–22, 18–45 and 40–75 meters ahead), the raycasts against the walls, and its relation to the racing line, but not the exact point the expert is aiming at.
None of this stops BC from working. But it does mean the policy couldn’t reproduce the expert exactly, and that the imitation would be worst precisely in recovery maneuvers, where the expert relies on its memory.
This connects to section 3 of part 1: designing the observations is also part of the problem. A huge dataset can’t make up for an observation that doesn’t contain the information the expert used.
15. Demonstrations are an engineering problem too
Creating demonstrations isn’t about recording many hours. You have to ask:
| Question | In Agentic Racing |
|---|---|
| Is the expert actually good? | On the fixed circuit, yes: it completes full laps. On the procedural tracks of the time, no, although much of the blame lay with the environment (corners too tight and a lateral grip that braked the car on every turn). |
| Does it cover different situations? | Yes for directives (a random one per episode). No for layouts: today there’s only one circuit. |
| Are there easy and hard corners? | The fixed circuit has four corners, all with the same radius. |
| Are there recovery situations? | Hardly: recordings start centered and aligned, and the expert rarely gets into situations it has to get out of. |
| Are there inconsistent actions? | Yes: the ones that depend on the expert’s memory and timers. |
| Do the states resemble the ones the policy will see? | Not entirely: training starts with heading and position noise that the recording doesn’t have. |
If all the demonstrations are perfect, we may teach perfect driving very well, but not how to recover when something goes wrong. For driving, recovery demonstrations can be as important as ideal ones.
There’s a famous precedent. ALVINN, one of the first vehicles to learn to drive itself with a neural network (Pomerleau, 1989), ended up being trained by imitating a human driver. To teach it to correct, Pomerleau (1991) also generated shifted and rotated images, as if the vehicle were off to one side, labeled with the correction that would have been needed. It manufactured the recovery examples the driver had never needed to show.
16. What should we measure?
BC shouldn’t be judged by its loss alone. A policy can have a very small imitation error and still drive badly.
Imitation metrics:
- the BC loss (
Losses/Pretraining Lossin ML-Agents); - the error per action component:
steer,throttleandbrakeseparately.
Behavior metrics:
- fraction of the lap completed;
- laps completed and lap time;
- why each episode ends (off track, stopped, no progress);
- speed and brake usage.
Agentic Racing already has almost all of the behavior metrics: they’re what its evaluator reports, the same evaluator that in part 2 showed the mean reward going up while the lap fraction didn’t move.
A low loss doesn’t guarantee a good policy. What we want to measure is behavior outside the dataset.
17. The experiment worth running
For Agentic Racing, the interesting comparison would be:
PPO from scratch demonstrations → BC + PPO
│ │
▼ ▼
result A result B
comparing lap fraction, laps completed, track exits, stability, and how many steps it takes to reach each level, and then testing against situations that weren’t in the demonstrations.
Then the question stops being “does BC work?” and becomes:
How much easier does learning get when you start from a policy that can already imitate reasonable behavior?
This experiment wasn’t run. It would be easier to run today than back then: on the fixed circuit, where the track is no longer an obstacle, the expert completes laps, and the recorder and the evaluator exist. The BC blocks would need to be added back to the configuration, since the instructions assume they’re there. It’s still pending.
18. DAgger: teaching how to recover too
The classic solution to covariate shift is DAgger (Dataset Aggregation), by Ross, Gordon and Bagnell:
- train on the initial demonstrations;
- run the policy;
- record the states the policy visits;
- ask the expert what it would have done in each one;
- add those examples to the dataset;
- train again;
- repeat.
initial demonstrations
│
▼
BC
│
▼
policy
│
▼
run policy
│
▼
out-of-distribution states
│
▼
ask the expert
│
▼
new examples ──────► BC
DAgger is interesting because the expert teaches precisely in the situations where the policy is struggling.
With a human expert, DAgger is expensive: someone has to label thousands of states. With an expert that is a program, like Agentic Racing’s heuristic pilot, it’s much cheaper: you just run the heuristic “in the shadow” alongside the policy, in the same states the policy visits, and store what it would have done. It has to run alongside rather than as a one-off query because, as we saw in section 14, the heuristic has memory. ML-Agents doesn’t include DAgger, so it would have to be built. That wasn’t done either, but it’s probably the best answer to the problems in sections 9 and 15.
19. BC doesn’t replace reinforcement learning
BC can work very well when there are good demonstrations, the problem is well covered, and the policy doesn’t need to stray far from the examples.
But if we want the agent to discover better strategies, beat the expert, adapt to new situations, or optimize a specific reward, we need something more. In Agentic Racing’s case, beating the expert isn’t a detail: the heuristic pilot brakes according to a fixed rule, and an agent learning from the reward could find something better.
Expert
│
▼
Demonstrations
│
▼
Behavioral Cloning ──► initial policy
│
├──► GAIL: does it look like the expert?
│
▼
PPO: does it maximize the reward?
│
▼
optimized policy
In ML-Agents, all three pieces can be combined in a single run, which is what the eighth run’s configuration was preparing. That sequence is a proposal for Agentic Racing, not something that was run.
Conclusion: yes, we can teach by imitation
Behavioral Cloning transforms part of the problem. Instead of saying “explore until you discover what works”, we say “look at these examples and learn to behave like the expert”. That can be very powerful.
But it has a fundamental weakness: the agent can run into situations that never appeared in the demonstrations. A small mistake takes it out of the expert’s distribution, that mistake leads to another, and that one leads to another.
Agentic Racing adds a second, less obvious lesson: imitation requires a reliable expert and an observation that contains what the expert uses to decide. The project had an expert, a recorder and a configuration, but every time it was about to record it found something to fix: an expert that got stuck, tracks nobody could drive, physics that turned every turn into a sudden brake. And even with everything fixed, that expert decides with memory and timers the policy couldn’t see.
For Agentic Racing, the interesting question isn’t just whether we can teach by imitation, but:
Can we use imitation as a starting point and then get the agent to overcome the limitations of its demonstrations?
That’s where Behavioral Cloning connects with GAIL and PPO.
Next in the series: part 4, “Can We Learn to Behave Like an Expert? GAIL Explained from Scratch”.
If you landed directly on this article, part 1 explains the core concepts, part 2 opens up PPO, and Agentic Racing: A Pilot That Never Learned to Drive tells the project’s full story.
References
- Pomerleau, D. (1989). ALVINN: An Autonomous Land Vehicle in a Neural Network. Advances in Neural Information Processing Systems 1.
- Pomerleau, D. (1991). Efficient Training of Artificial Neural Networks for Autonomous Navigation. Neural Computation, 3(1).
- Ross, S. and Bagnell, J. A. (2010). Efficient Reductions for Imitation Learning. AISTATS.
- Ross, S., Gordon, G. and Bagnell, J. A. (2011). A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning. AISTATS.
- Ho, J. and Ermon, S. (2016). Generative Adversarial Imitation Learning. NeurIPS.
- Unity Technologies. Unity ML-Agents Toolkit, in
particular the BC module (
ml-agents/mlagents/trainers/torch_entities/components/bc/module.py) and the GAIL reward (.../reward_providers/gail_reward_provider.py) from themlagents1.1.0 release. - The expert (
RaceAgent.cs), the recorder (EvalRunner.cs), the imitation instructions (training/README.md) and the project’s log are at github.com/alulema/agentic-racing.