← all papers · overview

Coachable agents for interactive gameplay

Roberto Capobianco (Sony AI·Zurich·Switzerland)·Harm van Seijen (Sony AI·North America·various locations)·Nolan D. Bard (Sony AI·North America·various locations)·Neil Burch (Sony AI·North America·various locations)·Fatima Davelouis (Sony AI·North America·various locations)·Josh Davidson (Sony AI·North America·various locations)·Alisa Devlic (Sony AI·Zurich·Switzerland)·Yunshu Du (Sony AI·North America·various locations)·Ishan Durugkar (Sony AI·North America·various locations)·Siddhant Gangapurwala (Sony AI·North America·various locations)·Daniel Hernandez (Sony AI·North America·various locations)·G. Zacharias Holland (Sony AI·North America·various locations)·Sahil Jain (Sony AI·North America·various locations)·Kenta Kawamoto (Sony AI·Tokyo·Japan)·Raksha Kumaraswamy (Sony AI·North America·various locations)·Patrick MacAlpine (Sony AI·North America·various locations)·Dustin R. Morrill (Sony AI·North America·various locations)·Declan Oller (Sony AI·North America·various locations)·Francesco Riccio (Sony AI·Zurich·Switzerland)·Akanksha Saran (Sony AI·North America·various locations)·Craig Sherstan (Sony AI·Tokyo·Japan)·Kaushik Subramanian (Sony AI·Zurich·Switzerland)·Thomas J. Walsh (Sony AI·North America·various locations)·Samuel Barrett (Sony AI·North America·various locations)·Kizza N. Frisbee (Sony AI·North America·various locations)·Mady Govil (Sony AI·North America·various locations)·Johannes Günther (Sony AI·North America·various locations)·Varun R. Kompella (Sony AI·North America·various locations)·James A. MacGlashan (Sony AI·North America·various locations)·Maxwell Svetlik (Sony AI·North America·various locations)·Michael D. Thomure (Sony AI·North America·various locations)·Jaden B. Travnik (Sony AI·North America·various locations)·Kevin Waugh (Sony AI·North America·various locations)·Elahe Aghapour (Sony AI·North America·various locations)·Florian Fuchs (Sony AI·Zurich·Switzerland)·Andreanne Lemay (Sony AI·North America·various locations)·Shruti Mishra (Sony AI·Zurich·Switzerland)·Takuma Seno (Sony AI·Tokyo·Japan)·Peter Stone (Sony AI·North America·various locations)·Michael Spranger (Sony AI·Tokyo·Japan)·Peter R. Wurman (Sony AI·North America·various locations)·2026

Abstract

Reinforcement learning has proven to be a valuable tool in the creation of advanced AI and robotic systems, contributing to everything from game playing to robotics to foundation models. Through trial-and-error, these AI systems typically learn one, near-optimal behavior to solve their tasks. However, there are many use cases in which one would like to assert some level of control, preferably in real time, over how the task is solved. We refer to these modifications of a core task as styles. We combine universal value function approximators (UVFAs) with carefully selected training scenarios, learning algorithms, and data augmentation to create a framework for coaching agents that exhibit styles in complex domains. We demonstrate the framework's application in the AAA video games Horizon Forbidden West and Gran Turismo, and in an open-source humanoid test domain. Despite the different nature of the domains -- car racing, stylized game combat, and humanoid walking -- each agent shows strong coherence to the style requests while still satisfying the main task in its domain. Importantly, the techniques outlined in this paper allow an end user to choose the final behavior at run time, giving them flexible control over the final executed performance.

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).