OmniCorpus  by OpenGVLab

10 Billion-level multimodal dataset for advanced AI training

Created 2 years ago
428 stars

Top 69.7% on SourcePulse

GitHubView on GitHub
Project Summary

OmniCorpus addresses the need for massive, diverse multimodal datasets by providing a 10 billion-level collection of interleaved images and text. Aimed at researchers and engineers developing multimodal large language models (MLLMs) and advanced retrieval systems, it offers unprecedented scale and quality, significantly boosting performance on downstream tasks.

How It Works

This dataset comprises 8.6 billion images and 1,696 billion text tokens across 2.2 billion documents, sourced from Common Crawl, Chinese internet resources, and YouTube. Its key advantages are its scale (1.7x images, 12.5x text vs. LAION-5B) and diversity, including bilingual content and varied document types. A rigorous five-stage data pipeline ensures high quality through extraction, filtering, and deduplication. The flexible streaming format supports various data structures.

Quick Start & Requirements

Processed documents are available on the Hugging Face and OpenDataLab platforms. The project also provides code for interleaved image-text pre-training and few-shot evaluation scripts. Specific hardware or software prerequisites are not detailed, but MLLM training typically demands substantial GPU resources. The project paper is available for further details.

Highlighted Details

  • Unprecedented Scale: Features 8.6 billion images and 1,696 billion text tokens, making it the largest multimodal dataset to date.
  • Enhanced MLLM Training: Proven to improve performance for models like Flamingo, EMU, and InternVL on multimodal in-context learning and fine-tuning.
  • Long Text-Image Retrieval: Facilitates retrieval models capable of handling longer text queries.
  • Rigorous Data Curation: Utilizes a five-stage pipeline for main body extraction, text/image filtering, and deduplication.

Maintenance & Community

The project has seen recent releases of processed data, models, and code (Oct/Aug 2024) and was accepted to ICLR 2025. Contact information for the development team is provided. No specific community channels are listed.

Licensing & Compatibility

The dataset is distributed under CC BY 4.0, while the accompanying code uses the Apache License 2.0. Use is restricted for sensitive content, harmful outcomes, and rights violations. Derived works must acknowledge OmniCorpus, and users must comply with the Terms of Use of original data sources (Common Crawl, Chinese regulations, YouTube).

Limitations & Caveats

Usage is strictly prohibited for applications involving sensitive content, harmful outcomes, or rights violations. Users must also adhere to the terms of the original data sources. The immense scale necessitates significant computational resources for effective utilization.

Health Check
Last Commit

1 year ago

Responsiveness

Inactive

Pull Requests (30d)
0
Issues (30d)
0
Star History
0 stars in the last 30 days

Explore Similar Projects

Feedback? Help us improve.