Discover and explore top open-source AI tools and projects—updated daily.
OpenGVLab10 Billion-level multimodal dataset for advanced AI training
Top 69.7% on SourcePulse
OmniCorpus addresses the need for massive, diverse multimodal datasets by providing a 10 billion-level collection of interleaved images and text. Aimed at researchers and engineers developing multimodal large language models (MLLMs) and advanced retrieval systems, it offers unprecedented scale and quality, significantly boosting performance on downstream tasks.
How It Works
This dataset comprises 8.6 billion images and 1,696 billion text tokens across 2.2 billion documents, sourced from Common Crawl, Chinese internet resources, and YouTube. Its key advantages are its scale (1.7x images, 12.5x text vs. LAION-5B) and diversity, including bilingual content and varied document types. A rigorous five-stage data pipeline ensures high quality through extraction, filtering, and deduplication. The flexible streaming format supports various data structures.
Quick Start & Requirements
Processed documents are available on the Hugging Face and OpenDataLab platforms. The project also provides code for interleaved image-text pre-training and few-shot evaluation scripts. Specific hardware or software prerequisites are not detailed, but MLLM training typically demands substantial GPU resources. The project paper is available for further details.
Highlighted Details
Maintenance & Community
The project has seen recent releases of processed data, models, and code (Oct/Aug 2024) and was accepted to ICLR 2025. Contact information for the development team is provided. No specific community channels are listed.
Licensing & Compatibility
The dataset is distributed under CC BY 4.0, while the accompanying code uses the Apache License 2.0. Use is restricted for sensitive content, harmful outcomes, and rights violations. Derived works must acknowledge OmniCorpus, and users must comply with the Terms of Use of original data sources (Common Crawl, Chinese regulations, YouTube).
Limitations & Caveats
Usage is strictly prohibited for applications involving sensitive content, harmful outcomes, or rights violations. Users must also adhere to the terms of the original data sources. The immense scale necessitates significant computational resources for effective utilization.
1 year ago
Inactive
veekaybee