We propose an unsupervised variational model for disentangling video into
independent factors, i.e. each factor's future can be predicted from its past
without considering the others. We show that our approach often learns factors
which are interpretable as objects in a scene.
Related papers
Ranked by semantic similarity β how closely each paper's abstract matches this one (100% = near-identical topic).