Selected work

Testing an LLM agent, leading a peer-support programme, and developing a product concept of my own.

01 / Investigate · Undergraduate dissertation

Benchmarking embodied navigation in multimodal LLMs

Can an LLM find its way through a dynamic 3D world like a child?

For my Cambridge dissertation, I evaluated a multimodal LLM agent in Animal-AI, using children aged 4–7 as a developmental benchmark.

  • 7 experimental conditions
  • 168 Gemini episodes
  • Children: 89.6% success · Gemini agent: 45.8% success
  • AI research
  • Experimental design
  • Python
  • Statistics

Results are specific to the model, prompt and scaffold configuration, tasks and samples used in this study.

The question

The tested multimodal LLM agent could describe a plausible route in text. I examined what happened when it had to turn that plan into movement in a physics-based 3D world.

The experiment varied detour demands, temporary reward occlusion and direct versus indirect landmarks. Success depended on a full perception-action loop: see the environment, decide what to do, move, then use the next observation to adjust.

My role

I designed and ran the comparative evaluation, integrated the agents through an API using an existing Python LLM-AAI scaffold, and piloted both Claude and Gemini before selecting Gemini for the main study. I collected the LLM trials, analysed the results in SPSS, and reviewed replays and reasoning traces to understand how failures unfolded.

The experimental arenas and child datasets were pre-existing research assets. I independently analysed the de-identified child datasets alongside the newly collected LLM trials.

The system

See → Think → Go / Turn → Observe again

Each episode began with a first-person image and environment state. The model could use Think to state its plan, then move using Go or Turn. Its actions and outcomes were logged for analysis.

What I found

Across the pooled comparison, children succeeded in 89.6% of trials; the tested Gemini agent succeeded in 45.8%. The children's odds of success were 10.15 times those of the tested agent.

The gap was not identical in every condition. The reduction in success under detour conditions did not remain significant after correction, and extra occlusion or landmark type did not reliably change agent success in this setup.

The replays were often more revealing than the headline score. The agent could describe a plausible route, then turn at the wrong angle, confuse a landmark with a ramp, or keep pushing against an obstacle it could not move.

planning ≠ navigating

What I took from it

The agent's behaviour depended on more than the reasoning it reported in text. Perception, prompts, motor control and environmental feedback all shaped the outcome. A stronger follow-up study would align observation, practice and action controls more closely across children and agents.

02 / Build to last · Peer support

Leading a youth peer-support programme

How do you keep a student-led peer-support service safe and consistent as it grows?

As Head of the Peer-counselling Division at Heartbreak Vaccine, I led the team delivering its peer-support service. The wider HVC team conducted the original research behind the programme.

  • 150+ hours of peer support
  • 69 recorded peer-support cases
  • 4.5/5 average satisfaction rating among respondents
  • Leadership
  • Service operations
  • Team coordination
  • Community

The challenge

Teenagers often turn to friends when they are struggling, but informal support can vary widely. Heartbreak Vaccine set out to make peer support more consistent while making clear that it was not a substitute for professional therapy.

The wider team used surveys and interviews to understand the need. My work began with service delivery: leading the division and helping put the peer-support model into practice.

How the programme worked

Recruit and train → Intake and consent → Match → Support → Record → Follow up

Peer mentors completed four weeks of training and assessment. Participants entered through a structured intake process, gave consent, and were matched with a mentor. The programme kept internal records and used feedback and follow-up to support continuity.

My role

As Head of the Peer-counselling Division, I coordinated the team and ongoing cases, helping the programme maintain continuity as it developed.

The initial survey and interview research belonged to the wider HVC team. The figures below also reflect the wider team's work.

What the team delivered

  • More than 150 hours of peer support
  • 69 recorded peer-support cases
  • 4.5/5 average satisfaction rating among respondents

69 refers to recorded support cases, not necessarily 69 unique participants.

03 / Create · Personal product · In progress

A personal map for memorable meals

How might we make it easier to record, rate and revisit the restaurants worth remembering?

A planned restaurant journal that will bring places, photos, menus, Michelin stars, ratings and reviews together on a world map.

Currently in the kitchen

  • Product concept
  • Early-stage exploration
  • Vibe Coding

The starting point

I remember trips through meals, but the details rarely stay together. The address sits in a maps app, photos disappear into the camera roll, and menus or notes end up somewhere else.

I wanted one place to remember both where I ate and what made the experience worth keeping.

The planned first version

The first version is planned as a personal restaurant journal built around a world map. Each entry could include:

  • Restaurant address and map location
  • Personal rating
  • Restaurant and dish photos
  • Written review
  • Menu
  • Michelin-star information

The aim is simple: make it easy to record a meal, look back on it, and find the restaurant again later.

My role

This is a personal product concept. I am still exploring what the first version could be. Information architecture, experience design, prototyping and build all come later.

Current status

Concept ↑ You are here → Structure → Prototype → Build → Test

The project is in progress. This section will grow as I add real screens, tested flows and, eventually, a live product. Until then, I would rather show an honest work in progress than invent results too early.