Microsoft COCO: Common objects in context

Abstract
We present a new dataset with the goal of advancing the state-of-the-art in object recognition by placing the question of
object recognition in the context of the broader question of scene understanding. This is achieved by gathering images of complex
everyday scenes containing common objects in their natural context. Objects are labeled using per-instance segmentations to aid in
precise object localization. Our dataset contains photos of 91 objects types that would be easily recognizable by a 4 year old. With a
total of 2.5 million labeled instances in 328k images, the creation of our dataset drew upon extensive crowd worker involvement via
novel user interfaces for category detection, instance spotting and instance segmentation. We present a detailed statistical analysis of
the dataset in comparison to PASCAL, ImageNet, and SUN. Finally, we provide baseline performance analysis for bounding box and
segmentation detection results using a Deformable Parts Model.
Cite
@incollection{microsoft-coco-common-objects-in-context-2014,
author = {Tsung-Yi Lin and Michael Maire and Serge Belongie and James Hays and Pietro Perona and Deva Ramanan and Piotr Dollár and C Lawrence Zitnick},
title = {Microsoft COCO: Common objects in context},
journal = {Computer Vision–ECCV 2014},
pages = {740--755},
year = {2014},
doi = {10.1007/978-3-319-10602-1_48},
url = {https://doi.org/10.1007/978-3-319-10602-1_48},
}