
Cognitive Maps in RL
Does an AI that already remembers a place still need a map?
Psychologist Edward C. Tolman showed that rats build an internal “cognitive map” of a maze even without a reward. Recent work (Wijmans et al., 2023) found that the same kind of spatial representation emerges inside the memory (LSTM) of an AI agent trained to navigate. That raised the question I wanted to answer: if an agent already forms a map in its memory, does handing it an external map still help — the way people use written notes to support what they remember?
I built a maze environment in Unity (ML-Agents) divided into 36 areas, with a switch, a goal, and seven guide boxes. Agents sense their surroundings with rays and are rewarded only when they reach the goal after hitting the switch first — so the task rewards genuine exploration. I trained four agents for three million steps each with PPO, varying two factors independently: whether the agent received a map-like image marking the areas it had already visited, and whether it had LSTM memory.
Without memory, the external map helped: reward rose from −0.8924 to −0.8659 (p = 0.038), and those agents covered more of the maze (31.5 areas per five minutes, versus 30.5). With memory, the same map changed nothing — reward was statistically indistinguishable (−0.9604 versus −0.9703, p = 0.100). The map also cost time: adding it roughly doubled training, from 1h25m to 2h45m, because the image had to be read through a camera sensor on every step.
The most honest reading is that memory-based agents may already hold the visit information the external map provides — consistent with the cognitive-map finding. But the policy loss of the LSTM agents stayed high throughout, which means three million steps was probably not enough for them to finish learning. Erik Wijmans, whose paper prompted this work, kindly advised me that recurrent agents take longer to train and are unusually sensitive to hyperparameters. I could not test a longer run with the compute available to me — and stating that limit plainly, rather than claiming a cleaner result, is what I took away from the project.
I wrote the paper in English and presented it as a poster during the SSH overseas research program in the United Kingdom.

Next project
Space Balloon