Segmenting Invisible Moving Objects

Abstract: Biological visual systems are exceptionally good at perceiving objects that undergo changes in appearance, pose, and position. In this paper, we aim to train a computational model with similar functionality to segment the moving objects in videos. We target the challenging cases when objects are ``invisible'' in the RGB video sequence, for example, breaking camouflage, where visual appearance from a static scene can barely provide informative cues, or locating the objects as a whole even under partial occlusion. To this end, we make the following contributions: (i) In order to train a motion segmentation model, we propose a scalable pipeline for generating synthetic training data, significantly reducing the requirements for labour-intensive annotations; (ii) We introduce a dual-head architecture (hybrid of ConvNets and Transformer) that takes a sequence of optical flows as input, and learns to segment the moving objects even when they are partially occluded or stop moving at certain points in videos; (iii) We conduct thorough ablation studies to analyse the critical components in data simulation, and validate the necessity of Transformer layers for aggregating temporal information and for developing object permanence. When evaluating on the MoCA camouflage dataset, the model trained only on synthetic data demonstrates state-of-the-art segmentation performance, even outperforming strong supervised approaches. In addition, we also evaluate on the popular benchmarks DAVIS2016 and SegTrackv2, and show competitive performance despite only processing optical flow.

14/06/2020

Yu Liu, Lianghua Huang, Pan Pan and
Bin Wang, Yinghui Xu, Rong Jin

high dynamic range, inverse tone mapping, image sensor, dynamic range, camera response function, quantization, computational photography, deep learning, convolutional neural network, computer vision

1:01

05/01/2021

soft color segmentation, layer decomposition, image editing, video editing, color, segmentation, neural network, generative

1:01

19/08/2021

nice-gan, reusing discriminators for encoding, unsupervised image-to-image translation, decoupled training, multi-scale discriminators, adversarial loss, no independent component for encoding, shared layers, residual attention, cyclegan

1:01

16/11/2020

action synthesis, video synthesis, joint generative model, human action generation, end-to-end learning, conditional video generation

3:02

16/11/2020

computational imaging, 3d vision, optical computing, optimal coding, depth imaging, structured light, hardware-in-the-loop, differentiable systems, 3d reconstruction, time-of-flight