Series: Reinforcement Learning from scratch, with Agentic Racing as the lab

  1. Part 1 · How Does an Agent Learn? Teaching a Car to Drive
  2. Part 2 · How Do We Improve a Policy Without Breaking It? PPO Explained from Scratch
  3. Part 3 · Can We Teach by Imitation? Behavioral Cloning Explained from Scratch
  4. Part 4 · Can We Learn to Behave Like an Expert? GAIL Explained from Scratch
  5. Part 5 · Can We Combine Imitation and RL? BC, GAIL, and PPO as One System (this article)

Where we left off

We have asked four questions so far:

PartQuestionShort answer
1Can an agent learn by trial and error?yes, from a reward; and it will find any shortcut that reward leaves open
2How do we improve the policy without breaking it?PPO: clipped updates that stay close to the previous policy
3Can we teach it by showing what an expert does?BC: supervised learning on (s,a)(s, a) pairs; fragile outside the demonstrations
4Can we make it behave like the expert?GAIL: a discriminator turns imitation into a learned reward

Part 4 ended with the next one:

Can we combine all of these ideas?

The most concrete version of that question is this: can we use an expert’s demonstrations to start from a useful policy, and then let reinforcement learning improve it?

The usual warning, bigger this time

No combination was ever run in Agentic Racing. BC and GAIL were configured for the eighth run and removed before it launched; no demonstration was ever recorded; there is no policy trained with imitation and no later fine-tuning. So this article doesn’t report a pipeline: it designs the experiment the project never got to run, using what the repository does have (the expert, the recorder, the evaluator, the configuration that was written) and what the ML-Agents code says about how the pieces would combine.

As in the earlier parts, every results table is a template full of ”?”, not a result.


1. Why combine: the Agentic Racing case

The motivation isn’t abstract. In parts 1 and 2 we saw what happened with PPO alone: eight runs on procedural tracks, and the policy stayed around 10% of a lap (between 8 and 13%). In the eighth, the most careful one, 91% of episodes ended with the car off the track and only 4% with a completed lap.

Meanwhile the expert, the heuristic pilot from part 3, does drive. When the project moved to a fixed circuit (a rounded rectangle about 2 km long), the expert drove it without leaving the track or getting stuck, and once the directives were wired into its driving, every evaluation car completed full laps (always from clean starts, which is how the evaluator runs the expert).

That asymmetry is exactly the case that justifies combining:

PPO alone: learns from the reward, but can't find how to complete a lap
Expert:    completes the lap, but learns nothing

The hypothesis is that the demonstrations spare PPO the hardest part of exploration (getting to the point of driving a lap) and that the reward then does the rest. It’s a reasonable hypothesis. But there’s also a competing explanation nobody ruled out: the fixed circuit made the problem much easier, and PPO alone was never tried on it. It may not need help anymore. We’ll come back to that.


2. The intuitive chain: BC → GAIL → PPO

The natural way to picture the combination is a chain of stages:

expert → demonstrations → BC → initial policy → GAIL → PPO fine-tuning → evaluation

Each stage would answer a different question:

BC:   "start by copying what the expert does"
GAIL: "refine the behavior so it resembles the expert's"
PPO:  "now optimize the task reward"

It’s a good way to organize the reasoning. But, as we saw at the end of part 4, it’s not what ML-Agents does when you put all three in one configuration. And the difference matters when designing the experiment.


3. What ML-Agents actually does: everything at once

In ML-Agents 1.1.0 there are no stages. There is a single PPO run with two extra sources of knowledge plugged into it:

  • BC is a supervised loss applied after each PPO update, with a learning rate that fades out over steps (part 3).
  • GAIL is one more reward signal, with its own head in the critic, that lasts the whole run (part 4).
  • PPO is the only RL algorithm: the one that optimizes the reward.

Careful with one detail: BC isn’t just “information” for PPO. It has its own optimizer and takes its own gradient steps on the policy’s weights, after each PPO update. So while BC is active, two optimizers pull on the same network: PPO toward more reward, BC toward the expert’s actions. GAIL, on the other hand, doesn’t touch the policy: it only changes the reward PPO sees.

The configuration prepared for the eighth run was exactly that: environment reward at weight 1.0, GAIL at 0.15, and BC at 0.5 for the first 2 million steps. It’s not a chain: it’s a mix whose balance shifts over time, because BC fades out and GAIL doesn’t.

A note on BC. The ML-Agents documentation says BC “can be used on its own or as a pre-training step”, but in the 1.1.0 code BC is always a module running inside an RL trainer (PPO, SAC, or POCA): there is no trainer that does only BC, without interacting with the environment. A “pure” BC like the one in part 3 would have to be written separately (ML-Agents ships a reader for .demo files that would be a starting point).


4. The sequential version: two runs

If we want the real chain, with an imitation stage and then a fine-tuning stage, ML-Agents allows it with a command-line option:

mlagents-learn imitation.yaml   --run-id=imit
mlagents-learn finetuning.yaml  --run-id=ft --initialize-from=imit

The second run starts from the first one’s checkpoint. Reading the code that does that loading, there are three details worth knowing before interpreting anything:

  1. It loads more than the policy. It loads every registered module: the policy, the critic, PPO’s optimizer, and the weights of each reward provider, including the GAIL discriminator if the second run uses it too. (It doesn’t save or load BC’s own optimizer or the discriminator’s: those start over.)
  2. It tolerates differences. Loading isn’t strict: if the second configuration no longer has GAIL, the weights of its critic head are ignored; whatever is missing keeps its fresh initialization; and if a piece doesn’t fit, a warning goes to the log and it’s loaded as far as possible or reset. For example, removing GAIL changes the critic’s shape, and PPO’s optimizer can’t load its previous state: it starts over. That’s convenient, but it also means a configuration mistake won’t stop the run: you have to read the log.
  3. The step counter goes back to zero. And with it, every linear schedule from part 2: they go back to the initial values in the second run’s configuration and decay over its own max_steps. If the project’s configuration is copied, that means a learning rate of 3×10−43 \times 10^{-4}, β\beta of 10−210^{-2}, and ε\varepsilon of 0.2. For fine-tuning that’s exactly the opposite of what we want: after learning to imitate, the second run starts with a learning rate and entropy pressure as high as at the start of the first one. Unless they’re lowered by hand in the second configuration, fine-tuning may start by undoing what imitation built.

Two ways to combine BC, GAIL, and PPO in ML-Agents. Top, the simultaneous one: a single run where the demonstrations feed the BC loss and the GAIL discriminator, PPO optimizes the reward and BC also adjusts the policy with its own optimizer. Bottom, the sequential one: an imitation run saves a checkpoint, and a second run starts from it with initialize-from and trains only on the environment reward. A note recalls that the step counter goes back to zero and the schedules start over.


5. Two architectures worth comparing

With that, the two concrete options are:

Architecture A — simultaneous, one run

demonstrations ─┬─► BC (fades out) ────────┐
                └─► GAIL (reward) ─────────┤
environment reward ────────────────────────┴─► PPO ─► policy

This is the one prepared for Agentic Racing. Upside: a single run, and the environment reward is present from the start, so imitation is never optimized on its own. Risk: the balance between signals is set by each one’s scale, not just by the weight written in the configuration (section 7).

Architecture B — sequential, two runs

run 1: PPO + BC + GAIL  ──checkpoint──►  run 2: PPO + environment reward only

Upside: it cleanly separates what each stage contributes, because the intermediate checkpoint can be evaluated. Risk: once imitation is removed, nothing stops the policy from drifting away from the expert. That can be good (finding something better) or bad (forgetting useful skills). And as we saw in part 1, PPO alone with this reward already found an easy way out once: in the fourth run it learned to stop on purpose.

Neither is better in advance. They have to be tested.


6. What does it mean for a policy to be “better”? In this project, it’s not just lap time

Before comparing, we have to decide what we measure. The obvious choice would be driving: completed laps, lap time, off-track exits. But Agentic Racing has something more important, and it’s what changes the experiment’s design the most.

The demo’s pilot doesn’t exist to be fast. It exists to follow the strategist. The 42-number observation vector includes 6 directive channels (aggression, risk tolerance, and the directive type: attack, defend, conserve, push), and during training they are randomized every episode, precisely so the policy learns to drive differently depending on what the strategist says.

The expert does that. Measured on the fixed circuit:

Forced directiveLap time
conserve, aggression 0.15112 s
random per episode100 s
attack, aggression 0.8579 s

The slowest lap takes about 33 s longer than the fastest, roughly 40% more: that’s the real lever the LLM strategist has over the race.

Now look at the task reward. Its terms are the ones from part 1: progress, target speed, racing line, edges, exits, completed lap. And the target speed it rewards depends on the track’s curvature, not on the directive. No term of the reward reads the directive channels.

That has an uncomfortable consequence:

  • A policy that imitates the expert has a reason to respond to the directive: the expert does, and the directive is in its observations.
  • A policy that only optimizes the reward has none. To it, the 6 channels are noise. The most profitable thing is to drive whichever way pays the most reward, whatever the directive.

In other words, in architecture B, fine-tuning on the task reward could improve lap time and, at the same time, erase exactly what the demo needs: that the car drives differently when the strategist says “attack” than when it says “conserve”. Nobody measured it, so it’s a question, not a finding. But it’s the most important question in the experiment.

The expert's lap time on the fixed circuit by directive: 112 seconds conserving, 100 with a random directive, 79 attacking: the slowest lap takes roughly 40% longer. Next to it, two columns with empty bars and question marks: a policy that imitates the expert, which should keep that spread if it learned it, and a policy fine-tuned only on the task reward, which has nothing forcing it to keep it because the reward doesn't read the directive.

That’s why, in this project, evaluation needs three families of metrics, not two:

Driving

  • lap fraction, completed laps, and lap time;
  • episode end reasons: off track, stalled car, no progress;
  • recovery from noisy starts.

Imitation

  • similarity of trajectories and of the action distribution to the expert’s, especially braking;
  • speed per section and behavior in the four corners.

Directive response

  • lap time under each forced directive;
  • the gap between “attack” and “conserve”, compared with the expert’s.

The project’s evaluator (eval.exe) already covers much of the first family: lap fraction, episode end reasons, and speeds, and it evaluates trained policies from noisy starts (the expert is always evaluated from clean ones). It only reports lap time directly in its population mode, so that would have to be added. It can also force one directive on every car, so the third family comes almost for free. The second would have to be built.


7. The weight of each signal

In part 4 we did the math: with an undecided discriminator, GAIL at weight 0.15 pays about 0.10 per decision, and the designed reward, driving well, about 0.07. The 0.15 looked small and wasn’t.

The ML-Agents documentation points the same way. It recommends keeping the GAIL strength below about 0.1 when the demonstrations are suboptimal and there’s an environment reward, so the agent focuses on the reward instead of copying; its reference example that combines all three (PushBlock) uses 0.01. And it warns about the survivor bias we saw in part 4: an always-positive per-step reward pushes the agent to stretch the episode out, and can drown the task reward. If the agent seems to ignore the environment reward, the advice is to lower the GAIL strength.

So the GAIL weight isn’t a configuration detail: it’s one of the experiment’s variables. At minimum, you’d test the planned 0.15 and a value on the order the documentation recommends.


8. Designing a reproducible experiment

With all of the above, part 4’s experiment grows like this:

VariantWhat it trains onIn ML-Agents
A. PPO from scratchenvironment rewardextrinsic, one run
B. Pure BCdemonstrations onlycustom code: there’s no BC trainer without the environment
C. GAIL + environmentimitation as a rewardgail + extrinsic, one run
D. SimultaneousBC + GAIL + environmentthe configuration prepared for the eighth run
E. SequentialD, then fine-tuninga second run with --initialize-from, extrinsic only

And, for each one:

MetricABCDE
Completed laps?????
Lap time?????
Off-track exits?????
Similarity to the expert?????
Attack / conserve gap?????
Recovery from noisy starts?????

For that table to mean anything, each variant has to compete under comparable conditions:

  • The same interaction budget. E gets D’s steps plus those of its second run; compare at equal total steps, or at least report them. And demonstrations cost something too: the recorder runs 300 s of simulation with the expert by default, in every arena at once, and those steps don’t show up in PPO’s step counter.
  • Several seeds. mlagents-learn accepts --seed; with a single run per variant there’s no way to tell an effect from chance. Report the mean and the spread, not the best run.
  • The same environment and the same demonstrations. The same build, the same physics, the same set of .demo files for every variant that uses them.
  • The same evaluation rules. The same number of episodes, the same starts, the same rule for picking the checkpoint (the last one, or the best by a metric fixed in advance, never the one that “looks best”).
  • An honest generalization test. All demonstrations would come from the fixed circuit, which is the default track. The procedural track generator is still in the code; evaluating there, without having trained there, would show whether the policy learned to drive or memorized one circuit.

The project already has a precedent for this kind of rigor. In the experiment pitting the LLM strategist against the heuristic one, it ran series of 18 races with the grid rotated, reported the mean with its standard deviation, and the first result went against expectations. It was published that way. The same discipline would apply here.


9. How to read the results (once they exist)

Without numbers, what we can write in advance is what each pattern would mean:

  • If A already completes laps on the fixed circuit, the original motivation for imitation goes away. What’s left to measure is whether imitation teaches something PPO doesn’t: directive response.
  • If B drives well from clean starts but fails from noisy ones, that’s part 3’s compounding-errors problem, exactly as expected.
  • If C or D lengthen episodes but don’t increase completed laps, suspect GAIL’s survivor bias before celebrating.
  • If E improves lap time but the attack/conserve gap shrinks, it won on the wrong metric for this demo.
  • If the differences between variants are smaller than the spread across seeds, the honest result is “no detectable effect”.

A useful experiment isn’t the one that proves the architecture works. It’s the one designed so it could reveal that it doesn’t.


10. What if the expert isn’t perfect?

It isn’t. The expert brakes by a fixed rule, a target speed computed from the upcoming corner, and has recovery behaviors that depend on timers the policy can’t see (part 3). No imitation technique, on its own, is designed to beat it: BC learns its decisions, GAIL learns to resemble it.

Beating it takes a signal that says “this is better”, and that signal is the reward. That’s the appeal of architecture B. But “better according to the reward” is only “better” if the reward measures what we want, and section 6 showed that in this project it doesn’t measure the most important thing. Before any fine-tuning, you’d have to decide whether the reward needs a term that rewards following the directive. That’s no longer about combining algorithms: it’s redesigning the objective, which is part 1’s problem.


11. What if a single piece is already enough?

We have to take that seriously. If BC produces a reliable policy, GAIL may add complexity without improving anything. If D works, E may make it worse. Every stage has to justify its existence with evidence:

What concrete improvement does this stage bring, and how much does it cost to get it?

Agentic Racing took that question further than any variant in the table. After the eighth run, the decision was not to train any pilot at all: the heuristic expert, with the directives wired into its pace, braking, and racing line, became the demo’s pilot. The project’s goal was to show the loop between pilot and strategist, and that didn’t need a neural network: it needed a pilot whose behavior changed with the directive, and the expert already did that.

That’s not a conclusion about BC, GAIL, or PPO. It’s an engineering conclusion: sometimes the simplest piece that meets the goal is the right one, and the more sophisticated pipeline stays an experiment for another day.


12. Implemented, executed, evaluated

It’s worth separating three levels of maturity, because they’re easy to blur:

  • Implemented: the code exists.
  • Executed: it was run and there’s evidence.
  • Evaluated: it was measured and compared against a baseline.

Applied to Agentic Racing:

PieceImplementedExecutedEvaluated
Heuristic expertyesyes: it’s the demo’s pilotyes: full laps on the fixed circuit and lap times per directive
Demonstration recording (eval.exe -record)yesno record of it ever being used—
PPOyesyes: eight runs on procedural tracksyes: around 10% of a lap (8–13%); on the fixed circuit, never
BCconfiguredno: removed before the eighth run—
GAILconfiguredno: removed together with BC—
Simultaneous BC + GAIL + PPOconfiguredno—
Sequential fine-tuningnono—

A configuration existing doesn’t prove it works. A run having been executed doesn’t prove it beats a baseline either. This table is what a reader should be able to ask any RL project for before believing one of its curves.


Conclusion

BC, GAIL, and PPO attack different parts of the problem:

  • BC uses demonstrations to learn the expert’s actions directly.
  • GAIL turns the comparison with the expert into a reward.
  • PPO optimizes a policy with whatever reward it gets.

The BC → GAIL → PPO chain is a good way to think about them together, but in ML-Agents the natural combination is simultaneous: a single PPO run, with BC pulling the policy toward the expert and GAIL adding its reward. The sequential version exists, through --initialize-from, and comes with its own traps: the schedules restart, and removing imitation leaves the policy at the mercy of the reward.

And in this project that reward has a concrete blind spot: it knows nothing about the strategist’s directives, which are the pilot’s whole reason to exist. That’s the kind of finding that only shows up when you design the experiment with the real system in hand, not a generic architecture in your head.

The interesting question isn’t whether we can chain three algorithms, but whether each one contributes something measurable that the others don’t.

Agentic Racing has almost everything it needs to answer it: the expert, the recorder, the evaluator, and the configuration. What it’s missing is the most important part: running it.

If you landed directly on this article, the four earlier parts are linked at the top, and Agentic Racing: A Pilot That Never Learned to Drive tells the project’s full story.


References