01 / Investigate · Undergraduate dissertation
Benchmarking embodied navigation in multimodal LLMs
Can an LLM find its way through a dynamic 3D world like a child?
For my Cambridge dissertation, I evaluated a multimodal LLM agent in Animal-AI, using children aged 4–7 as a developmental benchmark.
- 7 experimental conditions
- 168 Gemini episodes
- Children: 89.6% success · Gemini agent: 45.8% success
Results are specific to the model, prompt and scaffold configuration, tasks and samples used in this study.
The question
The tested multimodal LLM agent could describe a plausible route in text. I examined what happened when it had to turn that plan into movement in a physics-based 3D world.
The experiment varied detour demands, temporary reward occlusion and direct versus indirect landmarks. Success depended on a full perception-action loop: see the environment, decide what to do, move, then use the next observation to adjust.
My role
I designed and ran the comparative evaluation, integrated the agents through an API using an existing Python LLM-AAI scaffold, and piloted both Claude and Gemini before selecting Gemini for the main study. I collected the LLM trials, analysed the results in SPSS, and reviewed replays and reasoning traces to understand how failures unfolded.
The experimental arenas and child datasets were pre-existing research assets. I independently analysed the de-identified child datasets alongside the newly collected LLM trials.
The system
See → Think → Go / Turn → Observe again
Each episode began with a first-person image and environment state. The model could use Think to state its plan, then move using Go or Turn. Its actions and outcomes were logged for analysis.
What I found
Across the pooled comparison, children succeeded in 89.6% of trials; the tested Gemini agent succeeded in 45.8%. The children's odds of success were 10.15 times those of the tested agent.
The gap was not identical in every condition. The reduction in success under detour conditions did not remain significant after correction, and extra occlusion or landmark type did not reliably change agent success in this setup.
The replays were often more revealing than the headline score. The agent could describe a plausible route, then turn at the wrong angle, confuse a landmark with a ramp, or keep pushing against an obstacle it could not move.
planning ≠ navigating
What I took from it
The agent's behaviour depended on more than the reasoning it reported in text. Perception, prompts, motor control and environmental feedback all shaped the outcome. A stronger follow-up study would align observation, practice and action controls more closely across children and agents.