Continual Agent Learning for Lifelong Adaptation and Improvement in Complex Environments
Supervisor: Dr Diana Benavides-Prado
Project Description
Over the past few years, artificial intelligence has undergone a massive shift from static, single-turn models to semi-autonomous AI agents capable of planning, using tools, and interacting with the environment. Despite these advancements, current agent architectures remain largely and fundamentally static. Once trained, they struggle to adapt to new tasks, environments, or modalities without incurring expensive retraining from scratch or suffering from catastrophic forgetting, in which learning new skills wipes out previously acquired knowledge.
On the other hand, research in continual machine learning has expanded significantly over the past decade as a direct response to this limitation. Continual machine learning focuses on systems that accumulate knowledge across a sequence of tasks, minimising the need for repeated retraining while actively mitigating forgetting. While several robust methods have been proposed for continual learning in standard supervised and reinforcement learning settings, continual agent learning remains scarcely explored.
To achieve true autonomy, next-generation AI systems must become true continual learning agents. They need to operate in dynamic, open-ended worlds over long horizons, tackling tasks requiring hundreds or thousands of sequential steps, while continuously acquiring new skills and retaining or even improving past ones. Furthermore, they must seamlessly integrate and reason across multiple modalities (such as text, vision, and action spaces). Building these lifelong learning agents requires addressing core algorithmic hurdles at the intersection of reinforcement learning, large language models (LLMs), and machine learning, both experimentally and theoretically. This project aims to pioneer novel methods and architectures that allow agents to learn indefinitely, improve via knowledge transfer, and maintain robust long-term memories.
Research Objectives
The project will investigate how agents can effectively perform tasks continually whilst improving their own knowledge. The successful candidate will develop novel frameworks that allow agents to autonomously navigate distribution shifts, manage long-term changing distributions, and transfer knowledge fluidly across tasks.
Research directions may include, but are not limited to:
- Language-Based Continual Agents: Address the critical bottlenecks of language-based agents operating in open-ended setups over time. This includes studying major limitations, such as maintaining agent memories across long-horizon task streams, efficiently combining parametric memory and non-parametric retrieval, and designing advanced strategies, such as anticipatory learning, to scale effectively to a massive number of sequential tasks.
- Multimodal Continual Agents: Extend lifelong learning beyond text by investigating how multimodal agents adapt to online, sequential task execution. This direction focuses on the unique challenges of cross-modal alignment and grounding, specifically exploring how an agent can continually update its internal world models when confronted with entirely new environments.
- Mitigating Catastrophic Forgetting in Active Agents: Develop novel regularisation, modular/parametric isolation, experience replay or other types of techniques specifically tailored for continual agents. The goal is to ensure that updating an agent's knowledge and/or behaviour for Task B does not degrade its performance on Task A.
- Long-Horizon Planning and Exploration: Investigate how continual learning agents handle complex tasks characterised by sparse rewards and massive temporal dependencies. This could include investigating how agents can robustly recall, adapt, and hierarchically compose skills learned long ago to solve entirely novel challenges.
- Forward and Backward Knowledge Transfer: Design mechanism-driven approaches where learning a new task is accelerated by past knowledge (forward transfer), and where learning a new task actively optimises or refines performance on older tasks (backward transfer).
- Evaluation Benchmarks for Lifelong Agents: Create or extend simulation environments and robust evaluation protocols that accurately measure an agent's adaptability, sample efficiency, and robustness against non-stationarity over long deployment cycles.
Your research direction will be shaped by the intersection of your interests and expertise, which you will refine into a detailed proposal during the early months of the PhD. We welcome candidates with backgrounds in Computer Science, Artificial Intelligence, Data Science, Mathematics, or Statistics. Strong programming skills (e.g., Python, PyTorch) and a solid foundation in machine learning are requirements, and some experience with foundation models and familiarity with agent frameworks are highly desirable. Demonstrated interest in research (e.g., MSc thesis, preprints/publications, or open-source ML projects) will be an advantage.
Expected Outcomes:
- Novel algorithmic frameworks for mitigating catastrophic forgetting and achieving knowledge transfer during sequential agent learning tasks.
- Advanced language and/or multimodal agent architectures optimised for long-horizon task learning and execution.
- Theoretically grounded mechanisms for achieving positive forward and backward knowledge transfer in deep reinforcement learning and/or LLM agents.
- Publications in leading AI conferences and journals (e.g., NeurIPS, ICLR, ICML, JMLR).
- Open-source codebases, benchmarks, or toolkits to facilitate further community research in continual agent learning.
This PhD provides the opportunity to pioneer the foundations of lifelong AI, moving beyond static models toward truly autonomous, evolving agents. If you are passionate about pushing the boundaries of machine learning and shaping the future of autonomous systems, this project offers a unique platform to make a lasting impact on research.
Tuition fees and stipend
The PhD student will receive tuition fees at the home rate and a London stipend at QMUL stipend rates (currently in 2026/27 of £22,618 per year, to be confirmed for subsequent years) annually during the PhD period, which can span for 3 years.
For more information about the project, please contact Dr. Diana Benavides-Prado (d.benavidesprado@qmul.ac.uk).
Supervisor
Dr Diana Benavides-Prado (she/her) – d.benavidesprado@qmul.ac.uk
Personal Homepage: https://dianabenavidesprado.github.io/
Google Scholar: https://scholar.google.com/citations?user=ayeIzIgAAAAJ&hl=en&oi=ao
Centre for Multimodal AI page:
https://www.seresearch.qmul.ac.uk/cmai/people/dbenavidesprado/
How to apply
Queen Mary is interested in developing the next generation of outstanding researchers and decided to invest in specific research areas. Applicants should submit their application following the instructions at: https://www.qmul.ac.uk/eecs/phd/how-to-apply/
The application should include the following:
- CV (max 2 pages)
- Cover letter (max 4,500 characters) stating clearly in the first page whether you are eligible for a scholarship as a UK resident (https://epsrc.ukri.org/skills/students/guidance-on-epsrc-studentships/eligibility)
- Research proposal (max 500 words)
- 2 References
- Certificate of English Language (for students whose first language is not English)
- Other Certificates
Please note that to qualify as a home student for the purpose of the scholarships, a student must have no restrictions on how long they can stay in the UK and have been ordinarily resident in the UK for at least 3 years prior to the start of the studentship. For more information, please see: (https://epsrc.ukri.org/skills/students/guidance-on-epsrc-studentships/eligibility)
Application Deadline
The deadline for applications is Wednesday, 30th September 2026.
Interviews will be held in early October. The successful candidate will start in January 2027.
For specific enquiries, contact Dr Diana Benavides-Prado at d.benavidesprado@qmul.ac.uk.
For general enquiries contact Mrs Melissa Yeo at m.yeo@qmul.ac.uk (administrative enquiries) or Dr Arkaitz Zubiaga at a.zubiaga@qmul.ac.uk (academic enquiries) with the subject “EECS 2026 PhD scholarships enquiry”.