Conceptual 12M is a dataset containing (image-URL, caption) pairs collected for vision-and-language pre-training.
-
Updated
Jul 14, 2025
Conceptual 12M is a dataset containing (image-URL, caption) pairs collected for vision-and-language pre-training.
Dataset and Evaluation Scripts for Obstacle Detection via Semantic Segmentation in a Marine Environment
The official code of "Beyond Walking: A Large-Scale Image-Text Benchmark for Text-based Person Anomaly Search"
Wearanize+ is a multimodal sleep dataset containing overnight sleep data from 130 young, healthy participants using PSG and three wearables
[ACM MM 2025] Beyond Snapshots: A Multimodal User-Level Dataset for Depression Detection in Dynamic Social Media Streams
The MM5 dataset is a comprehensive multimodal dataset capturing RGB, Depth, Thermal (LWIR), Ultraviolet (UV), and Near-Infrared (NIR) images.
This study introduces MultiBanFakeDetect, a novel multimodal dataset for Bangla fake news detection, combining textual and visual information. It features TextFakeNet for text analysis and MultiFusionFake for integrating multimodal data.
Accessory code and meta information to the HODOR dataset
AI-powered TaxaMask workbench for morphology mask annotation, taxonomic evidence review, 2D/STL datasets, and TIF/CT micro-CT volume workflows with SAM/VLM.
MSVD-Indonesian: A Benchmark for Multimodal Video-Text Tasks in Indonesian (Bahasa Indonesia).
🚀 Easy Dataset - 强大的LLM微调数据集创建工具 | 支持文档解析、智能分割、数据清洗、问题生成、多轮对话、Arena盲测评估 | 可视化界面 + Docker部署 + LLaMA Factory集成
三维二进制立方体与周易64卦同构模型 · 先后天八卦本体几何结构
Add a description, image, and links to the multimodal-dataset topic page so that developers can more easily learn about it.
To associate your repository with the multimodal-dataset topic, visit your repo's landing page and select "manage topics."