monet jasperai
Dataset Card for MONET MONET (Massive, Open, Non-redundant and Enriched Text-to-image dataset) is a large-scale, curated image-text dataset designed for training text-to-image (T2I) systems. It contains 103.8 million high-quality image-text pairs distilled from 2.9 billion raw pairs across nine heterogeneous open sources (6 real and 3 synthetic) through successive stages of safety filtering, domain-based filtering, exact and near-duplicate removal, and re-captioning with… See the full description on the dataset page: https://huggingface.co/datasets/jasperai/monet.
- 유형
- dataset
- 라이선스
- apache-2.0
- 언어
- en
- 다운로드
- 184,360
- 좋아요
- 149
- 접근
- public
- 파일
- 0
README
--- license: apache-2.0 pretty_name: MONET task_categories: - text-to-image - image-feature-extraction - zero-shot-image-classification language: - en size_categories: - 100M<n<1B tags: - text-to-image - image-text - multimodal - captioning - synthetic-data configs: - config_name: parquet data_file…