Series: Reinforcement Learning from scratch, with Agentic Racing as the lab

  1. Part 1 · How Does an Agent Learn? Teaching a Car to Drive
  2. Part 2 · How Do We Improve a Policy Without Breaking It? PPO Explained from Scratch
  3. Part 3 · Can We Teach by Imitation? Behavioral Cloning Explained from Scratch
  4. Part 4 · Can We Learn to Behave Like an Expert? GAIL Explained from Scratch (this article)
  5. Part 5 · Can We Combine Imitation and RL? BC, GAIL, and PPO as One System

Where we left off

Part 3 asked a simple question: can we teach an agent by imitating an expert? The answer was yes. With Behavioral Cloning (BC) we collect examples (si,ai)(s_i, a_i) and train a policy so that:

πθ(si)≈ai\pi_\theta(s_i) \approx a_i

But a problem showed up: if the agent reaches a state that wasn’t in the demonstrations, it may not know what to do, and each mistake takes it to states where it makes more mistakes. And in Agentic Racing another one showed up: the expert, the project’s heuristic pilot, decides with memory and timers the policy doesn’t see, so there are situations where no memoryless network can copy its exact action.

That opens a more interesting question:

Do we really need to tell the agent which action to take in every state?

Maybe we can teach it something different:

“I want your behavior to look like the expert’s.”

That’s the idea behind Generative Adversarial Imitation Learning (GAIL), which part 3 announced as the series’ next stop.

An honest warning, as in the previous parts

In Agentic Racing, GAIL met the same fate as BC: it was configured but never run. The configuration prepared for the eighth run had a GAIL block next to the BC one; it was removed before the run was launched, and that run was done with plain PPO. The log doesn’t record any demonstration being recorded, so there’s no trained discriminator, no GAIL metrics, and no BC-vs-GAIL comparison to show either.

What does exist is enough to make this article concrete: the expert, the demonstration recorder, the configuration that was written, and the ML-Agents code that would have run it. So, as in part 3, the theory comes with what the project actually had and with what would have happened with it. The result tables in this article are proposed experiments, not results.


1. From copying actions to imitating behavior

Behavioral Cloning can be pictured like this:

       observation
            │
            ▼
  what did the expert do?
            │
            ▼
      expert's action
            │
            ▼
         train π

For example:

observation:
right-hand corner
high speed
car slightly off line

             ↓

expert:
steer    = 0.45
throttle = 0.30
brake    = 0.10

BC tries to learn that mapping. GAIL asks a different question:

Does the behavior our agent produces look like the expert’s?

And to answer it, it introduces a new component: a discriminator.


2. The idea of the discriminator

Imagine two drivers, the expert and the agent, and a judge watching what each one does:

         EXPERT                  AGENT
           │                       │
        behavior                behavior
           │                       │
           └──────────┬────────────┘
                      ▼
              ┌───────────────┐
              │ Discriminator │
              └───────┬───────┘
                      ▼
          from the expert or the agent?

The discriminator tries to tell the expert’s behavior apart from the agent’s. If it gets an experience from the expert, it should answer “expert”; if it gets one from the agent, “agent”.

The interesting part is that the agent is learning too, and its goal is to make that classification harder and harder. We have a game:

Discriminator: "I want to tell them apart".
Agent:         "I want you to be unable to tell me apart from the expert".

3. Why does this resemble a GAN?

The name isn’t a coincidence. GAIL is inspired by Generative Adversarial Networks (GANs). In a GAN, simplifying:

generator ──► generated data ──► discriminator ──► real or generated?

The generator tries to produce data that looks real, and the discriminator tries to tell real from generated.

GAIL does something similar, with one fundamental difference:

In GAIL, the “generator” doesn’t generate images or vectors directly. The policy generates behavior by interacting with the environment.

The policy doesn’t produce an isolated sample. It produces a sequence:

(s0,a0,s1,a1,s2,a2,…)(s_0, a_0, s_1, a_1, s_2, a_2, \ldots)

and the states in that sequence aren’t chosen by the policy: the environment’s physics produces them from its actions. That’s why GAIL can’t be trained like a normal GAN, by differentiating through the generator. It needs reinforcement learning.


4. Behavior matters, not an isolated action

BC works with the mapping st→ats_t \rightarrow a_t. GAIL works with (st,at)(s_t, a_t) pairs, but looked at as a distribution: how often the policy visits each state and which actions it takes there.

Imagine a stretch of track:

expert: brake 0.30 → turn 0.42 → accelerate 0.70 → correct −0.05
agent:  brake 0.28 → turn 0.45 → accelerate 0.66 → correct −0.08

None of the agent’s actions match the expert’s, but the behavior is almost the same: it brakes, turns, accelerates and corrects in the same places and with similar magnitudes. BC would penalize each of those small differences. A discriminator, if they’re differences the size of the expert’s own normal variation, would have no way of using them to tell the two apart.

That’s what GAIL tries to capture.


5. An intuition with Agentic Racing

Agentic Racing’s circuit today is a rounded rectangle about 2 km long, with four identical corners. The expert, part 3’s heuristic pilot, drives it lap after lap. Each lap leaves a trajectory:

τexpert=((s0,a0), (s1,a1), …, (sT,aT))\tau_{\text{expert}} = \big((s_0, a_0),\ (s_1, a_1),\ \ldots,\ (s_T, a_T)\big)

where each sts_t is the same 42 observation numbers the policy receives (the ones from part 1) and each ata_t is the same three continuous actions: steer, throttle and brake.

Now we let the agent drive, and it produces its own trajectories τagent\tau_{\text{agent}}. GAIL trains a discriminator to tell the pairs from one and the other apart.

At first it’s easy: the freshly initialized agent zigzags, brakes on the straights or accelerates in the corners, and the discriminator spots it immediately. As the policy improves, its pairs look more and more like the expert’s, and the discriminator starts to hesitate. The goal is to reach a point where the discriminator struggles to tell them apart.


6. The discriminator

Formally, the discriminator is a function:

Dϕ(s,a)∈(0,1)D_\phi(s, a) \in (0, 1)

where ss is the state, aa the action, and ϕ\phi the parameters of a neural network. Its output is read as the probability that the pair comes from the expert.

For example, early in training:

D(expert pair) = 0.95
D(agent pair)  = 0.08

The discriminator is very sure. Later on:

D(expert pair) = 0.60
D(agent pair)  = 0.45

Telling them apart gets harder.

In ML-Agents 1.1.0, the discriminator is a dense network that receives the observation and, if the use_actions option is on, also the action and a flag saying whether the episode ended at that step. By default it has 2 layers of 128 neurons and a sigmoid output, which is the probability above. The configuration prepared for Agentic Racing turned use_actions on, so the discriminator would have compared full observation–action pairs.


7. The discriminator’s objective

The discriminator is trained as a binary classifier:

max⁡D  EπE[log⁡D(s,a)]+Eπ[log⁡(1−D(s,a))]\max_{D}\ \ \mathbb{E}_{\pi_E}\big[\log D(s,a)\big] + \mathbb{E}_{\pi}\big[\log\big(1 - D(s,a)\big)\big]

where πE\pi_E is the expert’s behavior and π\pi the agent’s:

expert pairs ──► D should give high values
agent pairs  ──► D should give low values

In ML-Agents, the discriminator’s loss is that expression with the sign flipped (apart from a small numerical margin inside the logarithms). It also adds a gradient penalty with weight 10: it evaluates the discriminator at points between expert and agent pairs, and penalizes the norm of its gradient with respect to the input for drifting away from 1. It’s a technique borrowed from GANs (Gulrajani et al., 2017) to keep the discriminator from becoming so sharp that it stops giving a useful signal. The code itself explains why it’s there: it “adds stability esp. for off-policy” (PPO is on-policy, so here it’s a safeguard more than a necessity).

But the important thing is that the policy is learning too.


8. What does the policy want?

The policy wants to generate behavior that looks like the expert’s:

discriminator ──► "this looks like the agent" ──► signal ──► policy
      ▲                                                        │
      │                                                        ▼
      └──────────────────── new behavior ◄────────────────────┘

The policy changes, the behavior changes, the discriminator is retrained on that new behavior, and the process continues. It’s an adversarial game.


9. The learning loop

The full loop is:

  1. the agent interacts with the environment;
  2. it generates trajectories;
  3. the discriminator scores those pairs;
  4. its output is turned into a reward;
  5. the policy is updated with that reward;
  6. the discriminator is updated to tell them apart better;
  7. the agent interacts again.

The GAIL loop in ML-Agents: the agent drives and fills the buffer; the discriminator scores every observation–action pair and that score becomes the GAIL reward, added to the environment's reward; PPO updates the policy, and on every minibatch the discriminator takes a step to better tell the expert's demonstrations apart from the agent's experiences.

In ML-Agents, steps 5 and 6 are interleaved. On every minibatch of the PPO update (the 30 from part 2), the policy is updated first, and then the discriminator takes a step with that same minibatch of agent experiences and a same-size sample of the demonstrations. Policy and discriminator move forward together.


10. GAIL is still reinforcement learning

This distinction matters: GAIL doesn’t replace reinforcement learning. What changes is the source of the learning signal.

In traditional RL:

environment ──► reward ──► policy

In GAIL:

expert + agent behavior ──► discriminator ──► imitation reward ──► policy

The policy still learns by interacting with the environment, with an RL algorithm. In ML-Agents that algorithm is PPO, the same one from part 2: to PPO, GAIL is simply one more reward signal. The difference is that we don’t design this reward: the discriminator builds it.


11. A learned reward

This is one of GAIL’s most powerful ideas. In part 1 we hand-designed a reward with nine components (progress, target speed, edge proximity, lap completed…) and saw how many versions it took and how many times the agent found shortcuts. GAIL proposes something else: let the reward come from comparing the agent with the expert.

ML-Agents 1.1.0 computes it like this, for every step:

rGAIL(s,a)=−log⁡(1−D(s,a))r_{\text{GAIL}}(s, a) = -\log\big(1 - D(s, a)\big)

(with a small numerical margin so it’s never infinite).

The reading is direct:

  • if the discriminator thinks the pair is the agent’s, D≈0D \approx 0 and the reward is almost 0;
  • if it thinks it’s the expert’s, D→1D \rightarrow 1 and the reward grows;
  • at D=0.5D = 0.5, when it can’t decide, the reward is log⁡2≈0.69\log 2 \approx 0.69.

One detail is worth noticing: this reward is never negative. We’ll come back to that in section 18.

the agent acts
       ↓
 discriminator
       ↓
"does it look like the expert?"
       ↓
    reward
       ↓
    policy

12. So does GAIL remove the need to design rewards?

Not entirely. That’s an oversimplified reading.

GAIL learns an imitation signal from demonstrations, but the full system still has an environment, observations, actions, an RL algorithm, hyperparameters and, almost always, other rewards. And learning from an expert doesn’t guarantee the learned behavior is optimal:

If the expert drives badly, GAIL can learn to drive like that expert.

In practice, GAIL is almost never used alone. The configuration prepared for Agentic Racing kept the whole designed reward from part 1, with weight 1.0, and added GAIL’s with a weight of 0.15. GAIL didn’t replace the reward: it complemented it.


13. How is it different from Behavioral Cloning?

CharacteristicBehavioral CloningGAIL
Type of learningsupervisedimitation through RL
Learns actions directly?yesno: learns to produce a similar distribution
Uses demonstrations?yesyes
Uses a discriminator?noyes
Needs to interact with the environment?noyes
Learns in the states the policy itself visits?noyes
Uses a reward?noyes, a learned one

In one sentence each:

  • Behavioral Cloning: “do what the expert did in this situation”.
  • GAIL: “produce behavior that looks like the expert’s”.

The difference looks small, but conceptually it’s huge.


14. The BC problem GAIL tackles

Recall part 3’s compounding errors. BC learns from the states the expert visited, but afterward the policy visits its own:

expert trajectory  ─────────────────────►

agent trajectory   ───────────╮
                              ╰────────► out of distribution

GAIL has a conceptual advantage: the policy generates its own states by interacting with the environment, and the discriminator scores those too. If the policy drifts into a state the expert never visited, the discriminator recognizes it as “the agent’s” and the GAIL reward drops. The way to earn reward again is to get back to zones that look like the expert’s. In other words, GAIL doesn’t just teach the agent what to do where the expert was: it gives it a reason to return there.

In Agentic Racing that would have been especially relevant. As we saw in part 3, the demonstrations would always be recorded from clean spawns (centered, aligned, at 14 m/s), while training starts with heading and position noise. With BC, those first states would fall outside what the demonstrations teach. With GAIL, the reward would push the policy to reach expert-like states as soon as possible.

That doesn’t mean GAIL magically removes the problem. It means training takes into account the distribution of states the policy itself generates, which is exactly what BC didn’t do.


15. Occupancy measures: the math behind GAIL

An elegant way to understand GAIL is through occupancy measures. Simplifying, a policy induces a distribution over the state–action pairs it visits:

ρπ(s,a)\rho_\pi(s, a)

and the expert induces its own:

ρE(s,a)\rho_E(s, a)

GAIL looks for a policy whose distribution resembles the expert’s:

ρπ≈ρE\rho_\pi \approx \rho_E

There’s an elegant result behind this (Ho and Ermon, 2016): when the state is fully observable, each policy has a unique distribution of this kind, so matching the distributions is, at bottom, matching the expert’s policy:

π(a∣s)=πE(a∣s)\pi(a \mid s) = \pi_E(a \mid s)

in the states the expert visits. So what changes compared to BC? Where the errors are measured. BC measures them in the expert’s states; GAIL, in the states the policy actually goes through. A mistake that takes the policy out of the expert’s zone costs BC little and GAIL a lot.

This connects to a limit we saw in part 3. The expert having memory isn’t, by itself, an obstacle: in theory there’s always a memoryless policy that reproduces the same distribution of state–action pairs. What is an obstacle in Agentic Racing is something else. First, the 42 observations aren’t the full state: they don’t include, for example, the previous steering command, which the expert’s filter depends on. Second, the policy that would match the expert would have to be multimodal (for the same observation, sometimes accelerate and sometimes reverse), and a Gaussian policy like part 2’s can only have one mode per action. GAIL can’t invent the missing information or give the policy a shape it doesn’t have.


16. GAIL as an adversarial problem

Putting both parts together, GAIL can be written as:

min⁡π max⁡D  EπE[log⁡D(s,a)]+Eπ[log⁡(1−D(s,a))]\min_{\pi}\ \max_{D}\ \ \mathbb{E}_{\pi_E}\big[\log D(s,a)\big] + \mathbb{E}_{\pi}\big[\log\big(1 - D(s,a)\big)\big]

The discriminator tries to maximize its ability to tell the expert from the agent, and the policy tries to make that distinction hard. (Ho and Ermon’s original formulation adds a policy entropy term; in ML-Agents, the PPO entropy term we saw in part 2 plays a similar, though not identical, role. Ho and Ermon also define DD the other way around, as the probability that the pair comes from the agent; the resulting reward is equivalent.)

It isn’t exactly a GAN applied to trajectories: as we saw in section 3, the environment sits between the policy and what the discriminator sees, which is why the policy learns with RL.


17. Why not just copy the actions?

Imagine that in some situation the expert steers with steer = 0.35, and our agent, in a very similar situation, gets practically the same result with steer = 0.32.

BC would directly penalize that difference. GAIL asks something else: does the resulting behavior look like the expert’s? If 0.03 differences are normal within what the expert itself does, the discriminator has no way of using them.

In ML-Agents this has an explicit knob. With use_actions off, the discriminator doesn’t even see the actions: it only compares the states each one goes through, and it’s enough for the agent to reach the same places in the same way. With use_actions on, as in Agentic Racing’s configuration, it compares full pairs, but it’s still comparing distributions, not actions one by one.

That’s why GAIL can capture something closer to a style of behavior.


18. But GAIL isn’t magic either

GAIL can run into trouble:

  • the discriminator can learn too fast and leave the policy without a useful signal;
  • or too slowly, and then its reward doesn’t distinguish anything;
  • the demonstrations can be insufficient or cover few situations;
  • the expert can be suboptimal;
  • the observations may not contain what the expert uses to decide;
  • adversarial training can be unstable.

And there’s a less obvious one, which in Agentic Racing would have been concrete. As we saw in section 11, the GAIL reward is never negative. In a task where the episode ends when the agent fails, a reward that’s always positive per step means that surviving longer pays more, regardless of whether the agent looks like the expert. Kostrikov et al. (2019) describe this bias in detail.

In part 1 we saw what happens in Agentic Racing when the reward has an easy way out: in the fourth run, the agent learned to stop on purpose because that paid more. With GAIL, the temptation would go the other way: stretching the episode out.

And that effect wouldn’t have been small. The configuration’s 0.15 looks like a low weight next to the environment reward’s 1.0, but you have to look at the scale. With an undecided discriminator (D=0.5D = 0.5), GAIL pays 0.15×log⁡2≈0.100.15 \times \log 2 \approx 0.10 per decision, on every decision. The designed reward, driving well at 21 m/s, pays at most about 0.07 per decision (0.04 for progress, 0.025 for target speed and almost nothing for the racing line). In other words: per step, GAIL would have weighed as much as or more than the whole designed reward. And since any episode end cuts off that future reward, finishing the lap sooner would also have cost the agent something; the lap bonus, of 12 to 20, would have had to make up for it. The no-progress cutoff (less than 8 meters in 5 seconds ends the episode) limits the worst version of the problem, which is standing still, but doesn’t remove it. It’s exactly the kind of thing you’d have to measure before trusting the result.

Which brings us to the important question:

How do we know whether the agent is really learning to imitate the expert?

Watching the discriminator’s loss change isn’t enough. You have to look at the behavior.


19. How do we evaluate GAIL?

ML-Agents reports several GAIL-specific metrics in TensorBoard: the discriminator loss (Losses/GAIL Loss), the gradient penalty, and the two most useful ones for watching the game play out: Policy/GAIL Expert Estimate and Policy/GAIL Policy Estimate, the discriminator’s mean output on the demonstrations and on the agent’s experiences. If those two curves get closer, the agent is becoming hard to tell apart. If the agent’s estimate stays stuck at 0, the discriminator has won and the policy isn’t getting any signal.

But, as in part 3, the algorithm’s metrics aren’t enough. In Agentic Racing you’d also have to look at:

Imitation metrics:

  • how similar the agent’s trajectories are to the expert’s;
  • the action distribution, especially the brake, which part 2’s policy barely used;
  • speed on each stretch and behavior in the corners.

Driving metrics:

  • fraction of the lap completed and full laps;
  • lap time;
  • why each episode ends: off track, stopped, no progress;
  • ability to recover from noisy spawns.

The project’s evaluator already measures almost all the driving metrics. The imitation ones would have to be built.


20. The expert’s role is fundamental again

It’s tempting to think “GAIL learns by itself”. Not exactly. The expert is still the source of knowledge:

expert ──► bad demonstrations ──► GAIL ──► bad behavior

GAIL doesn’t know the expert is wrong. That’s why everything we saw in part 3 about demonstrations still applies, and in Agentic Racing it applies with the same details: the recordings would always start from clean spawns, would cover a single circuit, and the expert has recovery behaviors that depend on timers. For GAIL, coverage also matters for another reason: the discriminator can only recognize as “the expert’s” what appears in the demonstrations.

The quality of the demonstrations matters, for GAIL too.


21. A heuristic expert can be enough

One of the most interesting parts of the project is that you don’t need a human driver:

heuristic controller ──► demonstrations ──► GAIL ──► neural policy

This lets you use a traditional controller as a teacher. And an idea shows up that I find especially valuable from an engineering point of view:

We can use traditional software to generate the data that teaches a machine learning model.

There’s no need to replace the traditional system right away. It can be the source of knowledge. In Agentic Racing, in fact, that controller ended up being the demo’s pilot, with an LLM strategist telling it how to drive through directives. The network that would have learned from it never came to exist, but the pattern still holds.


22. BC + GAIL

Now we can combine the two imitation methods:

  • BC: “copy the expert”.
  • GAIL: “make your behavior look like the expert’s”.

The classic intuition is sequential: BC provides a reasonable initial policy, and GAIL keeps refining it while it interacts with the environment, especially in the states where BC fails because it left the demonstrations’ distribution.

In ML-Agents the combination isn’t sequential but simultaneous. As we saw in part 3, BC runs inside training with a weight that fades over time; GAIL, on the other hand, is a reward that lasts the whole run. The configuration prepared for Agentic Racing combined them like this:

PieceConfigurationWhat it does
Environment rewardstrength: 1.0, gamma: 0.995the designed reward from part 1
GAILstrength: 0.15, use_actions: trueimitation reward for the whole run
BCstrength: 0.5, steps: 2000000initial push toward the expert: its learning rate is half of PPO’s and fades out at 2 million steps
PPOthe one from part 2optimizes the sum of the rewards

The demonstrations for BC and for GAIL would have been the same: the ones recorded by eval.exe -record.


23. And where does PPO come in?

The series starts connecting all its pieces. The most common intuition is a chain: BC to get started, GAIL to refine the imitation, and PPO to optimize the reward. In ML-Agents, the real structure is different: PPO is the only RL algorithm, and BC and GAIL are two sources of knowledge plugged into it. GAIL only changes the reward PPO sees; BC also takes its own gradient steps on the policy’s weights, with its own optimizer, after each PPO update.

How BC, GAIL and the environment reward combine in ML-Agents with the configuration prepared for Agentic Racing: the demonstrations feed both the BC update and the GAIL discriminator; the environment reward (weight 1.0) and GAIL's (weight 0.15) each have their own value estimate in the critic; PPO combines their advantages and updates the policy, while BC pulls toward the expert with a weight that fades out at 2 million steps.

There’s an internal ML-Agents detail worth mentioning. Each reward has its own “head” in the critic, its own discount factor and its own GAE computation. Then the advantages of all the rewards are averaged to form the advantage PPO uses. So GAIL’s weight isn’t just the configuration’s 0.15: it also depends on the scale of its reward compared to the environment’s.

This leads to a more interesting question than “which algorithm is better?”:

How do we combine different sources of knowledge to build a better policy?


24. A possible experiment with Agentic Racing

We could compare four strategies, all on the fixed circuit:

ExperimentLearning signalsIn ML-Agents
A. PPO from scratchenvironment rewardextrinsic
B. BC + PPOenvironment reward + BCextrinsic + behavioral_cloning
C. GAILimitation reward onlygail
D. BC + GAIL + PPOeverythingthe configuration prepared for the eighth run, adapted to the fixed circuit

And measure, for each one:

MetricABCD
Lap fraction????
Laps completed????
Lap time????
Track exits????
Similarity to the expert????
Recovery from noisy spawns????

This is a proposed experiment. None of those cells has a value, because none of this was run. We do know something about experiment A from the previous parts: on the procedural tracks, the lap fraction stayed around 10%. It was never tried on the fixed circuit.

The point wouldn’t be to prove that one technique “wins”, but to understand why each strategy works or fails.


25. The real challenge: what does “looking like the expert” mean?

Suppose two policies that complete the track. A follows the expert’s trajectory exactly. B takes different lines, but gets better times. Which is better?

It depends on the goal. If we want to imitate the expert, A. If we want to win the race, B.

That’s the fundamental difference between imitation learning and reinforcement learning. GAIL brings us closer to the expert’s behavior. PPO, with the environment reward, can take us away from it if doing so pays more. And that may be exactly what we want.

In Agentic Racing, that tension would take a very concrete form: the heuristic pilot brakes according to a fixed rule, the target speed computed from the upcoming corner. A policy that looked too much like the expert would inherit that rule. One that optimized only the reward could find a better one, or, as in part 1, an easy way out.


26. Imitating doesn’t mean surpassing

This is probably the most important idea in the article. If the expert has a certain level, GAIL tries to learn behavior similar to that level. GAIL’s goal is imitation, not improvement.

To surpass the expert you need another signal saying “this behavior is better”. That signal is the environment reward, and the relative weight of each one decides which way the policy pulls. With 1.0 for the environment and 0.15 for GAIL, as in Agentic Racing’s configuration, the intent was to use imitation as a guide, not as the goal. But, as we saw in section 18, intent and effect don’t have to match: given the scale of each reward, GAIL could have ended up in charge.

That’s why combining imitation learning and reinforcement learning can be much more powerful than either one alone.


27. A way to understand the whole series

So far we’ve asked four questions:

PartQuestionCore idea
1. Reinforcement LearningCan it learn by trial and error?environment → reward → policy
2. PPOHow do we improve the policy without destroying what it learned?controlled updates
3. Behavioral CloningCan we teach it by watching an expert?expert → examples → policy
4. GAILCan we make its behavior look like the expert’s?discriminator → imitation reward → policy

And now we can ask the next one:

Can we combine all of these ideas?

That is the question of part 5.


Conclusion

GAIL introduces a powerful idea:

We don’t need to teach the agent every correct action: we can teach it what it means to behave like an expert.

Behavioral Cloning learns directly from (s,a)(s, a) pairs. GAIL trains a discriminator that tries to tell ρE(s,a)\rho_E(s, a) apart from ρπ(s,a)\rho_\pi(s, a) and turns that competition into a reward that pulls the policy toward the expert’s behavior. The difference comes down to two questions:

  • Behavioral Cloning: what did the expert do here?
  • GAIL: does my behavior look like the expert’s?

But GAIL doesn’t remove the challenges. It still needs good demonstrations, good observations, a suitable environment, stable training and careful evaluation. And, like any reward, its own can have shortcuts.

In Agentic Racing, GAIL stayed in a configuration that was never launched. But that configuration says a lot about the intent: the designed reward as the main signal, imitation as a guide with weight 0.15, and BC as an initial push. It was a bet that the expert would show the agent the way and the reward would take it further. And it also leaves a lesson: a small weight in the configuration doesn’t guarantee a small influence, because what counts is the scale of each signal.

Imitating the expert can teach us to behave like it. Learning from rewards could let us go beyond what the expert knows how to do.

That’s where Behavioral Cloning, GAIL and PPO start to become pieces of the same system.

If you landed directly on this article, the three previous parts are linked at the top, and Agentic Racing: A Pilot That Never Learned to Drive tells the project’s full story.


References