Series: Reinforcement Learning from scratch, with Agentic Racing as the lab
- Part 1 · How Does an Agent Learn? Teaching a Car to Drive
- Part 2 · How Do We Improve a Policy Without Breaking It? PPO Explained from Scratch
- Part 3 · Can We Teach by Imitation? Behavioral Cloning Explained from Scratch
- Part 4 · Can We Learn to Behave Like an Expert? GAIL Explained from Scratch
- Part 5 · Can We Combine Imitation and RL? BC, GAIL, and PPO as One System
- Part 6 · How Do We Tell an Agent What We Want? Reward Engineering Explained from Scratch
- Part 7 · What Should an Agent Be Able to See? Observation Design Explained from Scratch (this article)
Where we left off
Part 6 was about how we tell the agent what we want. This part goes one step earlier:
What information should the agent get so it can decide well?
Part 1 already introduced the distinction between state and observation, and the table of the 42 numbers Agentic Racing’s pilot receives. Here we go back to that table more carefully: what each number means, how it’s scaled, what it enables, what’s missing, and what happened the times the project changed it.
As in part 6, this time there’s real material: the observations are in the code and have a history. The usual warning still applies: no pilot trained on these observations completed laps consistently, and the proposals at the end were never run.
1. The agent doesn’t see the world: it gets an observation
At each step, the environment is in a state , but the agent receives an observation that describes it only in part:
is the function that turns the state into an observation, and we design it. Agentic Racing’s simulator knows everything: every car’s exact position, the full circuit geometry, the physics forces. The pilot receives 42 numbers, 10 times a second.
The important consequence is this: the agent can only tell apart the situations its observation tells apart. If two situations that call for different actions produce the same observation, no policy that depends on that observation alone can handle both well.
2. The 42 numbers, one by one
This is how the pilot’s observation is built in the code:
| Group | Values | What it holds | How it’s scaled |
|---|---|---|---|
| Rays | 27 | 9 rays fanned over ±75°, up to 70 m, that recognize the edge walls; per ray: whether what it hit is a wall, whether it hit nothing, and the distance | two binary values and the distance as a fraction of the ray length (1 if it hit nothing) |
| Speed | 2 | forward and lateral | divided by top speed (55 m/s) |
| Track | 4 | heading error relative to the racing line; the car’s lateral offset; the racing line’s offset; lap progress | heading ÷ 180°; offsets ÷ half the track width, clipped to ±2; progress from 0 to 1 |
| Curvature | 3 | how much the centreline turns 0–22, 18–45, and 40–75 m ahead | ÷ 90°, clipped to ±1 |
| Directive | 6 | aggression, risk tolerance, and the directive type as one-hot | levels 0.15, 0.5, or 0.85; a 4-value one-hot |
Seen against the questions a driver needs to answer, the design covers almost everything:
| Question | What answers it |
|---|---|
| Where am I relative to the track? | lateral offset, rays |
| Where am I pointing, and where does the track go? | heading error |
| How fast am I going? Am I sliding? | forward and lateral speed |
| How tight is the next corner? | curvature ahead, rays |
| What does the strategist want? | directive channels |
| Is there a car ahead? | nothing |
We’ll get to the last row in section 8.
Two design details are worth noticing. The directive type is one-hot (four values, one on), which is what the ML-Agents documentation recommends for categorical variables: “attack” isn’t “twice defend”. And aggression and risk aren’t continuous: they’re drawn from 0.15, 0.5, and 0.85 because those are the three levels the LLM strategist can pick. The code comment explains it: that way the policy sees in training the same possible values it will see once the LLM writes those channels (not with the same frequency: training draws them at random, and the LLM doesn’t choose at random). It’s a way of preventing a training/deployment mismatch before it exists.
3. Scales: two layers of normalization
The ML-Agents documentation recommends normalizing every observation to or , and most of the pilot’s numbers are by construction: speed divided by top speed, heading by 180°, curvature by 90° with a clip.
Two details are worth noting. The lateral offsets are clipped at ±2 half-widths, not ±1. The code doesn’t say why, and in practice the clip almost never kicks in: the walls sit at the edge, and the episode ends 2 m beyond it, at about 1.33 half-widths. And lap progress, which is in , jumps from 1 to 0 when crossing the finish line: two neighbouring points on the track get values at opposite ends of the scale.
On top of that, the training configuration turns on normalize: true: ML-Agents accumulates a mean
and variance of each input over the whole training run (the 27 ray values included), rescales with them
before the network, and clips the result to ±5. Those statistics keep changing while training and are
frozen into the exported model, so inference uses the final ones.
A warning the project confirms: good normalization doesn’t make up for missing information. The first six runs had well-scaled observations and still couldn’t see corners coming.
4. Reacting or anticipating
Some observations describe the present: speed, lateral offset, distance to the walls. Others help anticipate: the track’s curvature ahead.
The first six runs could only anticipate through the rays, which back then reached 40 m. At 15 m/s that’s under three seconds to see a corner, brake, and turn. The seventh run added the three curvature observations and lengthened the rays to 70 m:
The critic improved visibly: its error stopped growing and settled. But two honest readings are in order:
- It can’t be credited to the observations alone. The same run replaced the speed reward with part 6’s target speed. Two changes at once, a single result: it’s the experimental-design mistake section 9 proposes avoiding.
- The metric that mattered barely moved. Completed laps went from 2% to 3%. A better observation was necessary, but not sufficient; as we explained in part 6, a good part of the problem was in the procedural tracks, and in the first seven runs lateral grip also braked the car in every turn.
A fair question about anticipation is whether the information really “exists”. Curvature is computed by reading the circuit geometry ahead, which a real car wouldn’t have without a map. Here that’s not a problem: the policy would run in the same simulator it trains in, and a racing driver knows the circuit. But the right question is always the same: will this information be available where the policy is going to be used?
5. Vector or camera?
ML-Agents supports vector observations (numbers) and visual ones (images from a virtual camera). Agentic Racing uses only vectors and rays: no cameras, no rendered textures.
The project’s plan started from vector observations and, from there, chose to train on CPU, without a GPU: with vectors, a small network, and PPO, the bottleneck is simulating Unity, not computing gradients. The plan itself notes that switching to cameras would change that calculation. The ML-Agents documentation points the same way: visual observations are typically less efficient and slower to train, and are best used only when the problem can’t be solved with vectors or rays. Here, the rays and curvature already handed the pilot, in summary, what a camera would have forced it to learn from pixels.
| Goal | Reasonable starting point |
|---|---|
| Validate that basic control works | a compact vector, like Agentic Racing’s |
| Study the effect of each signal | vectors with controlled ablations |
| Learn to drive from a camera | images, with enough variation in scenarios |
| Transfer to real sensors | only what those sensors can measure |
6. Does the agent need memory?
The project’s policy is purely reactive: it decides from the current observation alone,
. It doesn’t stack previous observations (NumStackedVectorObservations = 1) and
doesn’t use a recurrent network (the configuration has no memory block).
That has a consequence we already saw in parts 3 and 4. The expert, the heuristic pilot, does have memory: timers that count how long it has been slow against a wall, a reverse window to get unstuck, and a filtered steering command that depends on the previous one. When the expert reverses because its timer counts more than a second slow against a wall, the agent’s observation at that instant can be almost the same as at another moment when the expert accelerates. For a memoryless policy, they’re the same situation with two different answers.
Formally, the pilot’s problem is partially observable (a POMDP): a single observation doesn’t identify everything that matters. There are three ways to attack it, from least to most complex:
- Add the missing variable, if it exists. For example, the previous steering command.
- Stack recent observations. ML-Agents allows it with one setting for the vector and another for the ray sensor; the documentation describes it as “limited memory” without the complexity of a recurrent network.
- A recurrent policy (LSTM), which learns what to remember.
None was tried. And memory has a limit: if a signal is never observed and can’t be inferred from the history, no recurrent network can make it up.
7. Designing to generalize, not to memorize
The classic risk is an observation that lets the policy learn “at this position in the world, turn left” instead of “when the track turns left, turn left”. The ML-Agents documentation sums it up in a rule: encode positions in relative coordinates whenever possible.
The project’s pilot follows it almost everywhere: it doesn’t get its absolute position, and its relation to the track is measured against the centreline and the racing line. With one interesting exception: lap progress, a number from 0 to 1.
Progress is the fraction of the lap covered from the finish line, so it was always a position along the track. While training used nine different procedural tracks, one per arena, “37% of the lap” was a different spot on each track. But on the fixed circuit, where the project ended up (and which is now the default track for all nine arenas), that number is a single position code: “37% of the lap” is always the same spot on the same circuit. A policy trained there could learn to turn at 37% instead of turning when the track turns. It would work on that circuit and fail on any other.
Nobody trained on the fixed circuit, so it’s a risk, not a finding. But it shows something general: whether an observation is relative or absolute also depends on the training environment. The same variable can be harmless with nine tracks and dangerous with one.
8. What the pilot doesn’t see: the rivals
The demo runs six cars on the same track. The pilot sees none of them.
The rays only detect the edge walls, and none of the other 15 observations describes another car. It’s not a one-line oversight: in training, each arena has a single car, so there were no rivals to observe. The project’s plan included a multi-agent phase with observations of the nearest rivals; it never happened, because the RL pilot was replaced by the heuristic one first.
And the heuristic doesn’t see rivals either. There’s a revealing detail there. One of the directive channels is called “risk tolerance”, and the original design defined it as tolerance to proximity: accepting tighter gaps, running wheel to wheel. In the current code, that channel only adjusts how close the car stays to the centreline versus the racing line. It can’t mean “proximity” because the pilot has no proximity observation at all. A directive can only modulate what the pilot can perceive.
The overtakes you see in the demo, then, aren’t maneuvers: they come from pace differences between cars that don’t know the others exist. Contacts are left to the physics, and the race director logs them as incidents. (In the race scene, in fact, the cars don’t even have the ray sensor: the heuristic pilot is driving, and it doesn’t use it.)
9. Three deciders, three observations
Agentic Racing has three “agents” making decisions, and each one deliberately sees different things:
The asymmetry between the pilot and the strategist is the demo’s central idea, and it’s written into the plan from the start. The strategist sees more context (the standings, gaps in seconds, lap times, numbered corners, its own notes) but less immediacy: it doesn’t see speed, angle, or rays, and can’t react to what happens in the next few seconds. The pilot is the opposite. The code respects that boundary: the telemetry sent to the LLM describes rivals only by what’s observable from outside (position, gap, last lap time, trend).
That’s observation design too, at another level: deciding what each part of a system sees is deciding what problem each part solves.
The expert, for its part, has privileged information: it reads the track geometry directly and remembers things the policy doesn’t see. Privileged information is fine for a teacher generating demonstrations, but you have to know it’s there: what the teacher uses and the student can’t see is exactly what the student won’t be able to copy (part 3).
10. Observation bugs: when the agent gets zeros
Before the first training run, the project went through two bugs that had nothing to do with the observation design and everything to do with its plumbing:
- A shape mismatch. ML-Agents negotiates the size of each observation with Python when it connects. Because of the order in which the car’s components were assembled, the agreed size only included the 12-value vector, but at runtime the 27 ray values arrived too. Result: an exception along the lines of “expected (12,), got (27,)”.
- Empty observations. With that fixed, a whole test run filled the log with warnings of “fewer observations than expected, padding with zeros”, about two million of them: the 12-value vector arrived empty at almost every step, and the car didn’t even act. The empty observation was a symptom that the whole decision pipeline was wired wrong.
Both were fixed by changing how and in what order the car’s components are assembled, and the first run started with a clean log. The lesson is practical: verify that the observation reaching the network is the one you think you’re sending. Zero padding doesn’t stop training; it just ruins it silently.
11. Observations and imitation
Observations also shape BC and GAIL (parts 3 and 4). The project’s recorder saves demonstrations with exactly the same 42 observations the policy receives, so there’s no mismatch of units or preprocessing between expert and student. The problem is a different one, and we’ve seen it: the expert decides with information that isn’t in those 42 observations. BC can’t copy a decision that depends on something it can’t see, and GAIL’s discriminator can’t demand it.
12. A reproducible experiment
If the RL pilot is picked up again on the fixed circuit, the natural experiment changes one thing at a time on top of the current observation:
| Variant | Change | Question |
|---|---|---|
| A. Baseline | the current 42 observations | the reference: never trained on the fixed circuit |
| B. No progress | remove lap progress | was the policy using it as a position code? |
| C. No curvature | remove the 3 curvature values | how much do they contribute, now without the reward mixed in? |
| D. With the previous steering | add the last steering command | does it narrow the gap with the expert? |
| E. Stacked | 3 stacked observations | are there situations that can only be told apart over time? |
| F. With rivals | relative position and speed of the nearest cars | do overtaking maneuvers appear? (requires training with several cars per arena) |
Requirements for the comparison to mean anything:
- Same build, same reward, same step budget. The seventh run showed what happens otherwise.
- Several seeds per variant, with mean and spread.
- Evaluate with
eval.exe, not with the reward (part 6): episode end reasons, lap fraction, mean brake. - A test off the fixed circuit for variant B: the procedural tracks are still in the code, behind a setting you have to switch off and rebuild, and they’re the way to know whether the policy learned to drive or learned the circuit.
- Check the plumbing on every change: that the negotiated size matches and that the log has no zero padding.
| Metric | A | B | C | D | E | F |
|---|---|---|---|---|---|---|
| Completed laps | ? | ? | ? | ? | ? | ? |
| Lap fraction | ? | ? | ? | ? | ? | ? |
| Mean brake | ? | ? | ? | ? | ? | ? |
| Results on unseen tracks | ? | ? | ? | ? | ? | ? |
None of this was run.
13. How do we know an observation is good?
There’s no universal score, but there are five questions, each with its answer in the project:
| Test | Question | In Agentic Racing |
|---|---|---|
| Sufficiency | are there different situations that look the same? | yes: the ones that depend on the expert’s memory, and any situation with rivals |
| Usefulness | does adding the signal help, repeatably? | unknown: the only change was mixed with another |
| Robustness | does it hold up on other tracks and starts? | not tested; lap progress is a risk on the fixed circuit |
| Availability | does the signal exist where the policy will run? | yes: everything is computed in the same simulator |
| Cost | is it worth what it costs to produce? | cheap: 42 numbers and 9 rays, no cameras |
Conclusion
Designing the observation is deciding, in part, what problem the agent solves. Agentic Racing’s pilot solves a specific one: driving while seeing walls 70 metres away, its own speed, how the track bends up to 75 metres ahead, and what the strategist asks for. It doesn’t solve the problem of racing against other cars, because it can’t see them.
The lessons the project leaves:
- scales matter, but normalizing doesn’t replace missing information;
- anticipating requires observing the near future, and you have to ask whether that information will exist where the policy runs;
- a memoryless policy can’t copy an expert that has memory;
- the same variable can be relative with many tracks and a position code with just one;
- a directive can only modulate what the pilot can perceive;
- changing the observation and the reward at the same time makes it impossible to know what worked;
- and you have to verify that what reaches the network is what you think you’re sending.
It’s not about the agent seeing everything. It’s about it seeing what’s needed for the behavior we want, with information it will actually have.
If you landed directly on this article, the six earlier parts are linked at the top, and Agentic Racing: A Pilot That Never Learned to Drive tells the project’s full story.
References
- Sutton, R. S. and Barto, A. G. (2018). Reinforcement Learning: An Introduction, 2nd edition. MIT Press.
- Kaelbling, L. P., Littman, M. L., and Cassandra, A. R. (1998). Planning and Acting in Partially Observable Stochastic Domains. Artificial Intelligence, 101(1–2), 99–134.
- Schulman, J. et al. (2017). Proximal Policy Optimization Algorithms.
- Ho, J. and Ermon, S. (2016). Generative Adversarial Imitation Learning. NeurIPS.
- Unity Technologies. Unity ML-Agents Toolkit,
mlagents1.1.0 release: the observations section ofdocs/Learning-Environment-Design-Agents.md(normalization, stacking, one-hot, relative coordinates, rays) andnetwork_settings(normalize,memory) indocs/Training-Configuration-File.md. - The observations (
RaceAgent.CollectObservationsinunity/Assets/Scripts/Agents/RaceAgent.cs), the ray sensor (TrainingArena.cs), the strategist’s telemetry (Strategy/RaceTelemetry.cs), and the devlog are at github.com/alulema/agentic-racing.