Look and Think Twice: Capturing Top-Down Visual Attention with Feedback Convolutional Neural Networks

Abstract
While feedforward deep convolutional neural networks
(CNNs) have been a great success in computer vision, it
is important to note that the human visual cortex generally
contains more feedback than feedforward connections. In
this paper, we will briefly introduce the background of feedbacks
in the human visual cortex, which motivates us to develop
a computational feedback mechanism in deep neural
networks. In addition to the feedforward inference in traditional
neural networks, a feedback loop is introduced to
infer the activation status of hidden layer neurons according
to the “goal” of the network, e.g., high-level semantic
labels. We analogize this mechanism as “Look and Think
Twice.” The feedback networks help better visualize and
understand how deep neural networks work, and capture
visual attention on expected objects, even in images with
cluttered background and multiple objects. Experiments on
ImageNet dataset demonstrate its effectiveness in solving
tasks such as image classification and object localization.
Cite
@inproceedings{look-and-think-twice-capturing-top-down-visual-2015,
author = {Chunshui Cao and Xianming Liu and Yi Yang and Yinan Yu and Jiang Wang and Zilei Wang and Yongzhen Huang and Liang Wang and Chang Huang and Wei Xu and Deva Ramanan and Thomas S. Huang},
title = {Look and Think Twice: Capturing Top-Down Visual Attention with Feedback Convolutional Neural Networks},
booktitle = {IEEE International Conference on Computer Vision},
year = {2015},
doi = {10.1109/iccv.2015.338},
url = {https://doi.org/10.1109/iccv.2015.338},
}