Start of funding 01.01.2019

Instrinsically-motivated Intelligent Systems - Unsupervised Learning through Physical Interaction and Observation

Prof. Dr. Nils Thuerey
Technische Universität München
Lehrstuhl für Informatik 15 - Computer Graphik und Visualisierung

Damian Mrowca
Stanford University
Stanford Neuroscience and Artificial Intelligence Laboratory



Infants explore their world in a seemingly arbitrary but intrinsically structured way, commonly described as curiosity. Attracted by coherences they haven’t yet understood, babies seem to intentionally create events that are new, informative and exciting to them. An artificial system that could mimic such abilities would be of great use for applications in computer vision, robotics, reinforcement learning, and many other areas. This research aims at further investigation of possibilities for an autonomous intelligent system to learn without supervision and self-supervise in realistic physical environment. In the scope of this project, we are combining computer science techniques, e.g deep learning, with the psychology successes to make yet another step towards better artificial intelligence. The results are expected to be initial steps toward creating flexible self-supervised autonomous agents and further develop efficient learning methods. This project benefits from partners that come from various areas of research, Computer Science, Neuroscience and Psychology.

Final report:
In this BaCaTec project we focused on a joint research with Stanford university. Recently, neural-network based forward dynamics models have been proposed that attempt to learn the dynamics of physical systems in a deterministic way. While near-term motion can be predicted accurately, long-term predictions suffer from accumulating input and prediction errors which can lead to plausible but different trajectories that diverge from the ground truth. A system that predicts distributions of the future physical states for long time horizons based on its uncertainty is thus a promising solution. In this project, we introduce a novel robust Monte Carlo sampling based graph-convolutional dropout method that allows us to sample multiple plausible trajectories for an initial state given a neural-network based forward dynamics predictor. By introducing a new shape preservation loss and training our dynamics model recurrently, we stabilize long-term predictions. We show that our model's long-term forward dynamics prediction errors on complicated physical interactions of rigid and deformable objects of various shapes are significantly lower than existing strong baselines. Lastly, we demonstrate how generating multiple trajectories with our Monte Carlo dropout method can be used to train model-free reinforcement learning agents faster and to better solutions on simple manipulation tasks.