Home
Projects
Publications
People
Join the Lab
Contact
9
Redistribution-based Cost Inference Improves Sparse Safe Offline RL
Safe offline RL typically assumes access to dense per-step cost annotations, but in practice supervisors provide only trajectory-level …
Ebenezer Gelo
,
Geraud Nangue Tasse
,
Steven James
,
Benjamin Rosman
PDF
Cite
The Goal-Directed Frame for General Agents
Reinforcement learning is often framed around episodic, discounted, or average scalar rewards. While useful, these views miss a core …
Geraud Nangue Tasse
,
Steven James
,
Benjamin Rosman
PDF
Cite
CORDA: A Benchmark for Hierarchical Harm-Centric Moral Reasoning in Large Language Models
The key question in moral judgement is not simply whether someone chooses the “right” answer, but how they decide what …
Siddarth Singh
,
Victoria Williams
,
Simon Rosen
,
Ebenezer Gelo
,
Helen Sarah Robertson
,
Ibrahim Suder
,
Benjamin Rosman
,
Geraud Nangue Tasse
,
Steven James
PDF
Cite
AI Agent Safety is a Reinforcement Learning Problem
With the rapid advancement and deployment of Agentic AI, our scientific understanding of capabilities and limitations has not kept …
Reginald McLean
,
Tabitha Edith Lee
,
Montaser Mohammedalamen
,
Kevin Roice
,
Glen Berseth
,
Patrick Pilarski
,
Marlos C. Machado
,
Alyssa Lefaivre Škopac
,
Benjamin Rosman
PDF
Cite
Schema-Based Understandability in Human-Robot Interactions: A Cognitive Framework
Robots are entering hospitals, airports, classrooms and homes, yet a persistent barrier to effective human–robot interaction is often …
Victoria Williams
,
Benjamin Rosman
PDF
Cite
The Five Senses: Assessing Non-verbal Communication in Multicultural Human–Robot Interaction
As social robots move from research laboratories into everyday settings, they increasingly encounter users whose sensory expectations …
Victoria Williams
,
Benjamin Rosman
PDF
Cite
Finding the FrameStack: Learning What to Remember for Non-Markovian Reinforcement Learning
Recent success in developing increasingly general purpose agents based on sequence models has led to increased focus on the problem of …
Geraud Nangue Tasse
,
Matthew Riemer
,
Benjamin Rosman
,
Tim Klinger
PDF
Cite
Procedural Generation of Semantically Correct Levels in Video Games using Reward Shaping
The generation of video game levels traditionally relies on manual efforts from skilled professionals, resulting in significant …
Luke Kerker
,
Branden Ingram
,
Pravesh Ranchod
PDF
Cite
Project
A Linear Network Theory of Iterated Learning
Language provides one of the primary examples of human’s ability to systematically generalize — reasoning about new …
Devon Jarvis
,
Richard Klein
,
Benjamin Rosman
,
Andrew Saxe
PDF
Cite
Optimal Task Generalisation in Cooperative Multi-Agent Reinforcement Learning
While task generalisation is widely studied in the context of single-agent reinforcement learning (RL), little research exists in the …
Simon Rosen
,
Abdel Mfougouon Njupoun
,
Geraud Nangue Tasse
,
Steven James
,
Benjamin Rosman
PDF
Cite
Project
»
Cite
×