The data includes 200 images of natural scenes (plaque, packaging instructions, small advertisements, menus, posters, etc.), Internet images (magazine covers, comic covers, etc.), Document images (text documents, etc.).
Dataset containing open-ended questions about images. These questions require an understanding of vision, language and commonsense knowledge to answer.
Cityscapes is a large-scale urban street-scene dataset with stereo video and high-quality pixel-level annotations, built for benchmarking semantic segmentation, instance segmentation, and panoptic scene understanding for autonomous driving and smart-city computer vision.