<- Back to Datasets

The BreakingNews Dataset

Multimodal dataset for news article analysis

The BreakingNews Dataset

To foster research on multi-modal news article analysis, we propose the BreakingNews dataset, that includes images, captions, geo-location information and comments. This dataset includes approximately 100,000 news articles from several major newspapers and media agencies, collected between the 1st of January and the 31st of December of 2014. All articles include at least one image, and cover a wide variety of topics, including sports, politics, arts, healthcare or local news. The copyright of all text and images resides with the original owners.

View this Dataset
->
Institut de Robòtica i Informàtica Industrial, CSIC-UPC.
https://www.iri.upc.edu
Task
Image Captioning
Annotation Types
Bounding Boxes
100000
Items
3
Classes
100000
Labels
Models using this dataset
Last updated on 
January 20, 2022
Licensed under 
Research Only
Label your own datasets on V7
Try our trial or talk to one of our experts.