Weighted Composition for Entropy-regularised Reinforcement Learning

Abstract

One avenue for creating a generally intelligent agent is to equip it with the ability to combine its previously learned behaviours to generate new ones. Currently, most reinforcement learning approaches focus on learning task-specific skills with little consideration for learning skills that can later be composed to solve new problems immediately. In this work, we consider the setting where an agent is required to compose its existing skills to attain goals with different preferences. In particular, we show how this problem can be formulated as a contextual bandit and demonstrate how relevant algorithms can learn the optimal weighted composition of skills given a context. Experiments in a high-dimensional environment indicate that an agent operating in our framework is capable of learning these weights directly from data autonomously.

Publication
ORiON
Caston Nyabadza
Caston Nyabadza

I love to learn more and more about the new technologies of this modern world but with a strong bias to Reinforcement Learning

Benjamin Rosman
Benjamin Rosman
Lab Director

I am a Professor in the School of Computer Science and Applied Mathematics at the University of the Witwatersrand in Johannesburg. I work in robotics, artificial intelligence, decision theory and machine learning.

Steven James
Steven James
Lab Director

My research interests include reinforcement learning and planning.

Geraud Nangue Tasse
Geraud Nangue Tasse
Lecturer

I am interested in reinforcement learning (RL) since it is the subfield of machine learning with the most potential for achieving AGI.