Series: Reinforcement Learning from scratch, with Agentic Racing as the lab
- Part 1 · How Does an Agent Learn? Teaching a Car to Drive
- Part 2 · How Do We Improve a Policy Without Breaking It? PPO Explained from Scratch
- Part 3 · Can We Teach by Imitation? Behavioral Cloning Explained from Scratch
- Part 4 · Can We Learn to Behave Like an Expert? GAIL Explained from Scratch
- Part 5 · Can We Combine Imitation and RL? BC, GAIL, and PPO as One System (this article)
Where we left off
We have asked four questions so far:
| Part | Question | Short answer |
|---|---|---|
| 1 | Can an agent learn by trial and error? | yes, from a reward; and it will find any shortcut that reward leaves open |
| 2 | How do we improve the policy without breaking it? | PPO: clipped updates that stay close to the previous policy |
| 3 | Can we teach it by showing what an expert does? | BC: supervised learning on pairs; fragile outside the demonstrations |
| 4 | Can we make it behave like the expert? | GAIL: a discriminator turns imitation into a learned reward |
Part 4 ended with the next one:
Can we combine all of these ideas?
The most concrete version of that question is this: can we use an expert’s demonstrations to start from a useful policy, and then let reinforcement learning improve it?
The usual warning, bigger this time
No combination was ever run in Agentic Racing. BC and GAIL were configured for the eighth run and removed before it launched; no demonstration was ever recorded; there is no policy trained with imitation and no later fine-tuning. So this article doesn’t report a pipeline: it designs the experiment the project never got to run, using what the repository does have (the expert, the recorder, the evaluator, the configuration that was written) and what the ML-Agents code says about how the pieces would combine.
As in the earlier parts, every results table is a template full of ”?”, not a result.
1. Why combine: the Agentic Racing case
The motivation isn’t abstract. In parts 1 and 2 we saw what happened with PPO alone: eight runs on procedural tracks, and the policy stayed around 10% of a lap (between 8 and 13%). In the eighth, the most careful one, 91% of episodes ended with the car off the track and only 4% with a completed lap.
Meanwhile the expert, the heuristic pilot from part 3, does drive. When the project moved to a fixed circuit (a rounded rectangle about 2 km long), the expert drove it without leaving the track or getting stuck, and once the directives were wired into its driving, every evaluation car completed full laps (always from clean starts, which is how the evaluator runs the expert).
That asymmetry is exactly the case that justifies combining:
PPO alone: learns from the reward, but can't find how to complete a lap
Expert: completes the lap, but learns nothing
The hypothesis is that the demonstrations spare PPO the hardest part of exploration (getting to the point of driving a lap) and that the reward then does the rest. It’s a reasonable hypothesis. But there’s also a competing explanation nobody ruled out: the fixed circuit made the problem much easier, and PPO alone was never tried on it. It may not need help anymore. We’ll come back to that.
2. The intuitive chain: BC → GAIL → PPO
The natural way to picture the combination is a chain of stages:
expert → demonstrations → BC → initial policy → GAIL → PPO fine-tuning → evaluation
Each stage would answer a different question:
BC: "start by copying what the expert does"
GAIL: "refine the behavior so it resembles the expert's"
PPO: "now optimize the task reward"
It’s a good way to organize the reasoning. But, as we saw at the end of part 4, it’s not what ML-Agents does when you put all three in one configuration. And the difference matters when designing the experiment.
3. What ML-Agents actually does: everything at once
In ML-Agents 1.1.0 there are no stages. There is a single PPO run with two extra sources of knowledge plugged into it:
- BC is a supervised loss applied after each PPO update, with a learning rate that fades out over
steps(part 3). - GAIL is one more reward signal, with its own head in the critic, that lasts the whole run (part 4).
- PPO is the only RL algorithm: the one that optimizes the reward.
Careful with one detail: BC isn’t just “information” for PPO. It has its own optimizer and takes its own gradient steps on the policy’s weights, after each PPO update. So while BC is active, two optimizers pull on the same network: PPO toward more reward, BC toward the expert’s actions. GAIL, on the other hand, doesn’t touch the policy: it only changes the reward PPO sees.
The configuration prepared for the eighth run was exactly that: environment reward at weight 1.0, GAIL at 0.15, and BC at 0.5 for the first 2 million steps. It’s not a chain: it’s a mix whose balance shifts over time, because BC fades out and GAIL doesn’t.
A note on BC. The ML-Agents documentation says BC “can be used on its own or as a pre-training step”,
but in the 1.1.0 code BC is always a module running inside an RL
trainer (PPO, SAC, or POCA): there is no trainer that does
only BC, without interacting with the environment. A “pure” BC like the one in part 3 would have to be
written separately (ML-Agents ships a reader for .demo files that would be a starting point).
4. The sequential version: two runs
If we want the real chain, with an imitation stage and then a fine-tuning stage, ML-Agents allows it with a command-line option:
mlagents-learn imitation.yaml --run-id=imit
mlagents-learn finetuning.yaml --run-id=ft --initialize-from=imit
The second run starts from the first one’s checkpoint. Reading the code that does that loading, there are three details worth knowing before interpreting anything:
- It loads more than the policy. It loads every registered module: the policy, the critic, PPO’s optimizer, and the weights of each reward provider, including the GAIL discriminator if the second run uses it too. (It doesn’t save or load BC’s own optimizer or the discriminator’s: those start over.)
- It tolerates differences. Loading isn’t strict: if the second configuration no longer has GAIL, the weights of its critic head are ignored; whatever is missing keeps its fresh initialization; and if a piece doesn’t fit, a warning goes to the log and it’s loaded as far as possible or reset. For example, removing GAIL changes the critic’s shape, and PPO’s optimizer can’t load its previous state: it starts over. That’s convenient, but it also means a configuration mistake won’t stop the run: you have to read the log.
- The step counter goes back to zero. And with it, every linear schedule from part 2: they go
back to the initial values in the second run’s configuration and decay over its own
max_steps. If the project’s configuration is copied, that means a learning rate of , of , and of 0.2. For fine-tuning that’s exactly the opposite of what we want: after learning to imitate, the second run starts with a learning rate and entropy pressure as high as at the start of the first one. Unless they’re lowered by hand in the second configuration, fine-tuning may start by undoing what imitation built.
5. Two architectures worth comparing
With that, the two concrete options are:
Architecture A — simultaneous, one run
demonstrations ─┬─► BC (fades out) ────────┐
└─► GAIL (reward) ─────────┤
environment reward ────────────────────────┴─► PPO ─► policy
This is the one prepared for Agentic Racing. Upside: a single run, and the environment reward is present from the start, so imitation is never optimized on its own. Risk: the balance between signals is set by each one’s scale, not just by the weight written in the configuration (section 7).
Architecture B — sequential, two runs
run 1: PPO + BC + GAIL ──checkpoint──► run 2: PPO + environment reward only
Upside: it cleanly separates what each stage contributes, because the intermediate checkpoint can be evaluated. Risk: once imitation is removed, nothing stops the policy from drifting away from the expert. That can be good (finding something better) or bad (forgetting useful skills). And as we saw in part 1, PPO alone with this reward already found an easy way out once: in the fourth run it learned to stop on purpose.
Neither is better in advance. They have to be tested.
6. What does it mean for a policy to be “better”? In this project, it’s not just lap time
Before comparing, we have to decide what we measure. The obvious choice would be driving: completed laps, lap time, off-track exits. But Agentic Racing has something more important, and it’s what changes the experiment’s design the most.
The demo’s pilot doesn’t exist to be fast. It exists to follow the strategist. The 42-number observation vector includes 6 directive channels (aggression, risk tolerance, and the directive type: attack, defend, conserve, push), and during training they are randomized every episode, precisely so the policy learns to drive differently depending on what the strategist says.
The expert does that. Measured on the fixed circuit:
| Forced directive | Lap time |
|---|---|
| conserve, aggression 0.15 | 112 s |
| random per episode | 100 s |
| attack, aggression 0.85 | 79 s |
The slowest lap takes about 33 s longer than the fastest, roughly 40% more: that’s the real lever the LLM strategist has over the race.
Now look at the task reward. Its terms are the ones from part 1: progress, target speed, racing line, edges, exits, completed lap. And the target speed it rewards depends on the track’s curvature, not on the directive. No term of the reward reads the directive channels.
That has an uncomfortable consequence:
- A policy that imitates the expert has a reason to respond to the directive: the expert does, and the directive is in its observations.
- A policy that only optimizes the reward has none. To it, the 6 channels are noise. The most profitable thing is to drive whichever way pays the most reward, whatever the directive.
In other words, in architecture B, fine-tuning on the task reward could improve lap time and, at the same time, erase exactly what the demo needs: that the car drives differently when the strategist says “attack” than when it says “conserve”. Nobody measured it, so it’s a question, not a finding. But it’s the most important question in the experiment.
That’s why, in this project, evaluation needs three families of metrics, not two:
Driving
- lap fraction, completed laps, and lap time;
- episode end reasons: off track, stalled car, no progress;
- recovery from noisy starts.
Imitation
- similarity of trajectories and of the action distribution to the expert’s, especially braking;
- speed per section and behavior in the four corners.
Directive response
- lap time under each forced directive;
- the gap between “attack” and “conserve”, compared with the expert’s.
The project’s evaluator (eval.exe) already covers much of the first family: lap fraction, episode end
reasons, and speeds, and it evaluates trained policies from noisy starts (the expert is always evaluated
from clean ones). It only reports lap time directly in its population mode, so that would have to be
added. It can also force one directive on every car, so the third family comes almost for free. The
second would have to be built.
7. The weight of each signal
In part 4 we did the math: with an undecided discriminator, GAIL at weight 0.15 pays about 0.10 per decision, and the designed reward, driving well, about 0.07. The 0.15 looked small and wasn’t.
The ML-Agents documentation points the same way. It recommends keeping the GAIL strength below about 0.1 when the demonstrations are suboptimal and there’s an environment reward, so the agent focuses on the reward instead of copying; its reference example that combines all three (PushBlock) uses 0.01. And it warns about the survivor bias we saw in part 4: an always-positive per-step reward pushes the agent to stretch the episode out, and can drown the task reward. If the agent seems to ignore the environment reward, the advice is to lower the GAIL strength.
So the GAIL weight isn’t a configuration detail: it’s one of the experiment’s variables. At minimum, you’d test the planned 0.15 and a value on the order the documentation recommends.
8. Designing a reproducible experiment
With all of the above, part 4’s experiment grows like this:
| Variant | What it trains on | In ML-Agents |
|---|---|---|
| A. PPO from scratch | environment reward | extrinsic, one run |
| B. Pure BC | demonstrations only | custom code: there’s no BC trainer without the environment |
| C. GAIL + environment | imitation as a reward | gail + extrinsic, one run |
| D. Simultaneous | BC + GAIL + environment | the configuration prepared for the eighth run |
| E. Sequential | D, then fine-tuning | a second run with --initialize-from, extrinsic only |
And, for each one:
| Metric | A | B | C | D | E |
|---|---|---|---|---|---|
| Completed laps | ? | ? | ? | ? | ? |
| Lap time | ? | ? | ? | ? | ? |
| Off-track exits | ? | ? | ? | ? | ? |
| Similarity to the expert | ? | ? | ? | ? | ? |
| Attack / conserve gap | ? | ? | ? | ? | ? |
| Recovery from noisy starts | ? | ? | ? | ? | ? |
For that table to mean anything, each variant has to compete under comparable conditions:
- The same interaction budget. E gets D’s steps plus those of its second run; compare at equal total steps, or at least report them. And demonstrations cost something too: the recorder runs 300 s of simulation with the expert by default, in every arena at once, and those steps don’t show up in PPO’s step counter.
- Several seeds.
mlagents-learnaccepts--seed; with a single run per variant there’s no way to tell an effect from chance. Report the mean and the spread, not the best run. - The same environment and the same demonstrations. The same build, the same physics, the same set
of
.demofiles for every variant that uses them. - The same evaluation rules. The same number of episodes, the same starts, the same rule for picking the checkpoint (the last one, or the best by a metric fixed in advance, never the one that “looks best”).
- An honest generalization test. All demonstrations would come from the fixed circuit, which is the default track. The procedural track generator is still in the code; evaluating there, without having trained there, would show whether the policy learned to drive or memorized one circuit.
The project already has a precedent for this kind of rigor. In the experiment pitting the LLM strategist against the heuristic one, it ran series of 18 races with the grid rotated, reported the mean with its standard deviation, and the first result went against expectations. It was published that way. The same discipline would apply here.
9. How to read the results (once they exist)
Without numbers, what we can write in advance is what each pattern would mean:
- If A already completes laps on the fixed circuit, the original motivation for imitation goes away. What’s left to measure is whether imitation teaches something PPO doesn’t: directive response.
- If B drives well from clean starts but fails from noisy ones, that’s part 3’s compounding-errors problem, exactly as expected.
- If C or D lengthen episodes but don’t increase completed laps, suspect GAIL’s survivor bias before celebrating.
- If E improves lap time but the attack/conserve gap shrinks, it won on the wrong metric for this demo.
- If the differences between variants are smaller than the spread across seeds, the honest result is “no detectable effect”.
A useful experiment isn’t the one that proves the architecture works. It’s the one designed so it could reveal that it doesn’t.
10. What if the expert isn’t perfect?
It isn’t. The expert brakes by a fixed rule, a target speed computed from the upcoming corner, and has recovery behaviors that depend on timers the policy can’t see (part 3). No imitation technique, on its own, is designed to beat it: BC learns its decisions, GAIL learns to resemble it.
Beating it takes a signal that says “this is better”, and that signal is the reward. That’s the appeal of architecture B. But “better according to the reward” is only “better” if the reward measures what we want, and section 6 showed that in this project it doesn’t measure the most important thing. Before any fine-tuning, you’d have to decide whether the reward needs a term that rewards following the directive. That’s no longer about combining algorithms: it’s redesigning the objective, which is part 1’s problem.
11. What if a single piece is already enough?
We have to take that seriously. If BC produces a reliable policy, GAIL may add complexity without improving anything. If D works, E may make it worse. Every stage has to justify its existence with evidence:
What concrete improvement does this stage bring, and how much does it cost to get it?
Agentic Racing took that question further than any variant in the table. After the eighth run, the decision was not to train any pilot at all: the heuristic expert, with the directives wired into its pace, braking, and racing line, became the demo’s pilot. The project’s goal was to show the loop between pilot and strategist, and that didn’t need a neural network: it needed a pilot whose behavior changed with the directive, and the expert already did that.
That’s not a conclusion about BC, GAIL, or PPO. It’s an engineering conclusion: sometimes the simplest piece that meets the goal is the right one, and the more sophisticated pipeline stays an experiment for another day.
12. Implemented, executed, evaluated
It’s worth separating three levels of maturity, because they’re easy to blur:
- Implemented: the code exists.
- Executed: it was run and there’s evidence.
- Evaluated: it was measured and compared against a baseline.
Applied to Agentic Racing:
| Piece | Implemented | Executed | Evaluated |
|---|---|---|---|
| Heuristic expert | yes | yes: it’s the demo’s pilot | yes: full laps on the fixed circuit and lap times per directive |
Demonstration recording (eval.exe -record) | yes | no record of it ever being used | — |
| PPO | yes | yes: eight runs on procedural tracks | yes: around 10% of a lap (8–13%); on the fixed circuit, never |
| BC | configured | no: removed before the eighth run | — |
| GAIL | configured | no: removed together with BC | — |
| Simultaneous BC + GAIL + PPO | configured | no | — |
| Sequential fine-tuning | no | no | — |
A configuration existing doesn’t prove it works. A run having been executed doesn’t prove it beats a baseline either. This table is what a reader should be able to ask any RL project for before believing one of its curves.
Conclusion
BC, GAIL, and PPO attack different parts of the problem:
- BC uses demonstrations to learn the expert’s actions directly.
- GAIL turns the comparison with the expert into a reward.
- PPO optimizes a policy with whatever reward it gets.
The BC → GAIL → PPO chain is a good way to think about them together, but in ML-Agents the natural
combination is simultaneous: a single PPO run, with BC pulling the policy toward the expert and GAIL
adding its reward. The sequential version
exists, through --initialize-from, and comes with its own traps: the schedules restart, and removing
imitation leaves the policy at the mercy of the reward.
And in this project that reward has a concrete blind spot: it knows nothing about the strategist’s directives, which are the pilot’s whole reason to exist. That’s the kind of finding that only shows up when you design the experiment with the real system in hand, not a generic architecture in your head.
The interesting question isn’t whether we can chain three algorithms, but whether each one contributes something measurable that the others don’t.
Agentic Racing has almost everything it needs to answer it: the expert, the recorder, the evaluator, and the configuration. What it’s missing is the most important part: running it.
If you landed directly on this article, the four earlier parts are linked at the top, and Agentic Racing: A Pilot That Never Learned to Drive tells the project’s full story.
References
- Ho, J. and Ermon, S. (2016). Generative Adversarial Imitation Learning. NeurIPS.
- Schulman, J. et al. (2017). Proximal Policy Optimization Algorithms.
- Ross, S., Gordon, G., and Bagnell, J. A. (2011). A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning. AISTATS.
- Rajeswaran, A. et al. (2018). Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations. RSS. (A precedent for initializing with BC and fine-tuning with RL.)
- Kostrikov, I. et al. (2019). Discriminator-Actor-Critic: Addressing Sample Inefficiency and Reward Bias in Adversarial Imitation Learning. ICLR.
- Unity Technologies. Unity ML-Agents Toolkit,
mlagents1.1.0 release: the imitation learning section ofdocs/ML-Agents-Overview.md, thegailandbehavioral_cloningsettings indocs/Training-Configuration-File.md,--initialize-fromindocs/Training-ML-Agents.md, and checkpoint loading inml-agents/mlagents/trainers/model_saver/torch_model_saver.py. - The project’s training configuration, expert, evaluator, and devlog are at github.com/alulema/agentic-racing.