Start of funding 01.07.2016

Webly-Supervised Deep Learning for Large-Scale Human Activity Recognition in Videos

Prof. Dr. Nassir Navab
Technische Universität München
Lehrstuhl für Informatik 16 - Informatikanwendungen in der Medizin

Prof. Dr. Silvio Savarese
Stanford University
Department of Computer Science



There are an estimated 3.5 trillion photographs in the world. Facebook alone reports 6 billion photo uploads per month. Every minute, 72 hours of video are uploaded to YouTube. Cisco estimates that in the next few years, visual data (images and video) will account for over 85% of total internet traffic. Yet, we still lack effective computational methods to analyze big visual data. In the joint project, we focus on developing a new approach for human activity recognition in large-scale video collections. Our main idea is to take advantage of the large amount of data already available on the web in a weakly supervised scenario, thus it does not require manual labelling. To this end, we will design a deep learning architecture that incrementally learns the model by starting from the images retrieved from a search engine (e.g. Google). Then, these images will be enriched with “confident” frames extracted videos, and so the initial weak classifier is incrementally adapted to the video domain.