CI-VID: A Coherent Interleaved Text-Video Dataset
Fuente:
arXiv
Salvato in:
| Autori principali: | Ju, Yiming, Hu, Jijin, Luo, Zhengxiong, Deng, Haoge, Zhao, hanyu, Du, Li, Wu, Chengwei, Hao, Donglin, Wang, Xinlong, Pan, Tengfei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DataCube: A Video Retrieval Platform via Natural Language Semantic Profiling
di: Ju, Yiming, et al.
Pubblicazione: (2026)
di: Ju, Yiming, et al.
Pubblicazione: (2026)
You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale
di: Ma, Baorui, et al.
Pubblicazione: (2024)
di: Ma, Baorui, et al.
Pubblicazione: (2024)
Autoregressive Video Generation without Vector Quantization
di: Deng, Haoge, et al.
Pubblicazione: (2024)
di: Deng, Haoge, et al.
Pubblicazione: (2024)
Beyond IID: Optimizing Instruction Learning from the Perspective of Instruction Interaction and Dependency
di: Zhao, Hanyu, et al.
Pubblicazione: (2024)
di: Zhao, Hanyu, et al.
Pubblicazione: (2024)
Scaling Towards the Information Boundary of Instruction Sets: The Infinity Instruct Subject Technical Report
di: Du, Li, et al.
Pubblicazione: (2025)
di: Du, Li, et al.
Pubblicazione: (2025)
LINA: Linear Autoregressive Image Generative Models with Continuous Tokens
di: Wang, Jiahao, et al.
Pubblicazione: (2026)
di: Wang, Jiahao, et al.
Pubblicazione: (2026)
XS-VID: An Extremely Small Video Object Detection Dataset
di: Guo, Jiahao, et al.
Pubblicazione: (2024)
di: Guo, Jiahao, et al.
Pubblicazione: (2024)
SpatialVID: A Large-Scale Video Dataset with Spatial Annotations
di: Wang, Jiahao, et al.
Pubblicazione: (2025)
di: Wang, Jiahao, et al.
Pubblicazione: (2025)
CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation
di: Chen, Wei, et al.
Pubblicazione: (2024)
di: Chen, Wei, et al.
Pubblicazione: (2024)
Mitigating Training Imbalance in LLM Fine-Tuning via Selective Parameter Merging
di: Ju, Yiming, et al.
Pubblicazione: (2024)
di: Ju, Yiming, et al.
Pubblicazione: (2024)
Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning
di: Wang, Chendong, et al.
Pubblicazione: (2025)
di: Wang, Chendong, et al.
Pubblicazione: (2025)
Uniform Discrete Diffusion with Metric Path for Video Generation
di: Deng, Haoge, et al.
Pubblicazione: (2025)
di: Deng, Haoge, et al.
Pubblicazione: (2025)
Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation
di: Liao, Chao, et al.
Pubblicazione: (2025)
di: Liao, Chao, et al.
Pubblicazione: (2025)
MozzaVID: Mozzarella Volumetric Image Dataset
di: Pieta, Pawel Tomasz, et al.
Pubblicazione: (2024)
di: Pieta, Pawel Tomasz, et al.
Pubblicazione: (2024)
Interleaved Scene Graphs for Interleaved Text-and-Image Generation Assessment
di: Chen, Dongping, et al.
Pubblicazione: (2024)
di: Chen, Dongping, et al.
Pubblicazione: (2024)
Design-based Estimation Theory for Complex Experiments
di: Chang, Haoge
Pubblicazione: (2023)
di: Chang, Haoge
Pubblicazione: (2023)
UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models
di: Wang, Jiaqi, et al.
Pubblicazione: (2026)
di: Wang, Jiaqi, et al.
Pubblicazione: (2026)
FastVID: Dynamic Density Pruning for Fast Video Large Language Models
di: Shen, Leqi, et al.
Pubblicazione: (2025)
di: Shen, Leqi, et al.
Pubblicazione: (2025)
Accelerate Scaling of LLM Finetuning via Quantifying the Coverage and Depth of Instruction Set
di: Wu, Chengwei, et al.
Pubblicazione: (2025)
di: Wu, Chengwei, et al.
Pubblicazione: (2025)
ClimateVID -- Social Media Videos Analysis and Challenges Involved
di: Xu, Shiqi, et al.
Pubblicazione: (2026)
di: Xu, Shiqi, et al.
Pubblicazione: (2026)
Semantic-E2VID: a Semantic-Enriched Paradigm for Event-to-Video Reconstruction
di: Wu, Jingqian, et al.
Pubblicazione: (2025)
di: Wu, Jingqian, et al.
Pubblicazione: (2025)
KeyVID: Keyframe-Aware Video Diffusion for Audio-Synchronized Visual Animation
di: Wang, Xingrui, et al.
Pubblicazione: (2025)
di: Wang, Xingrui, et al.
Pubblicazione: (2025)
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models
di: Diao, Haiwen, et al.
Pubblicazione: (2025)
di: Diao, Haiwen, et al.
Pubblicazione: (2025)
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models
di: Zhang, Yongheng, et al.
Pubblicazione: (2025)
di: Zhang, Yongheng, et al.
Pubblicazione: (2025)
A Multiphase Interleaved SLH Converter With Feedforward Decoupling and Mode‐Switching for Hi‐Fi Audio
di: Chengguo Qian, et al.
Pubblicazione: (2026)
di: Chengguo Qian, et al.
Pubblicazione: (2026)
Interleave-VLA: Enhancing Robot Manipulation with Interleaved Image-Text Instructions
di: Fan, Cunxin, et al.
Pubblicazione: (2025)
di: Fan, Cunxin, et al.
Pubblicazione: (2025)
Randomization Inference For the Always-Reporter Average Treatment Effect
di: Chang, Haoge, et al.
Pubblicazione: (2026)
di: Chang, Haoge, et al.
Pubblicazione: (2026)
Generation of Smoke Dataset for Power Equipment and Study of Image Semantic Segmentation
di: Rong Chang, et al.
Pubblicazione: (2024)
di: Rong Chang, et al.
Pubblicazione: (2024)
A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation
di: Feng, Yukang, et al.
Pubblicazione: (2025)
di: Feng, Yukang, et al.
Pubblicazione: (2025)
Interleaving Reasoning for Better Text-to-Image Generation
di: Huang, Wenxuan, et al.
Pubblicazione: (2025)
di: Huang, Wenxuan, et al.
Pubblicazione: (2025)
Fine-grained spatial-temporal perception for gas leak segmentation
di: Zhao, Xinlong, et al.
Pubblicazione: (2025)
di: Zhao, Xinlong, et al.
Pubblicazione: (2025)
ATIR: Towards Audio-Text Interleaved Contextual Retrieval
di: Zhao, Tong, et al.
Pubblicazione: (2026)
di: Zhao, Tong, et al.
Pubblicazione: (2026)
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis
di: Yang, Yifan, et al.
Pubblicazione: (2024)
di: Yang, Yifan, et al.
Pubblicazione: (2024)
HyperE2VID: Improving Event-Based Video Reconstruction via Hypernetworks
di: Ercan, Burak, et al.
Pubblicazione: (2023)
di: Ercan, Burak, et al.
Pubblicazione: (2023)
SG2VID: Scene Graphs Enable Fine-Grained Control for Video Synthesis
di: Sivakumar, Ssharvien Kumar, et al.
Pubblicazione: (2025)
di: Sivakumar, Ssharvien Kumar, et al.
Pubblicazione: (2025)
TVIR: Building Deep Research Agents Towards Text--Visual Interleaved Report Generation
di: Ma, Xinkai, et al.
Pubblicazione: (2026)
di: Ma, Xinkai, et al.
Pubblicazione: (2026)
Scaling Speech-Text Pre-training with Synthetic Interleaved Data
di: Zeng, Aohan, et al.
Pubblicazione: (2024)
di: Zeng, Aohan, et al.
Pubblicazione: (2024)
Boundary-Protection W8A8 HiFloat8 Quantization for Large-Scale Text-to-Video Diffusion Transformers
di: Zhao, Yiming
Pubblicazione: (2026)
di: Zhao, Yiming
Pubblicazione: (2026)
MoSA: Motion-Coherent Human Video Generation via Structure-Appearance Decoupling
di: Wang, Haoyu, et al.
Pubblicazione: (2025)
di: Wang, Haoyu, et al.
Pubblicazione: (2025)
Global Context Compression with Interleaved Vision-Text Transformation
di: Jiao, Dian, et al.
Pubblicazione: (2026)
di: Jiao, Dian, et al.
Pubblicazione: (2026)
Documenti analoghi
-
DataCube: A Video Retrieval Platform via Natural Language Semantic Profiling
di: Ju, Yiming, et al.
Pubblicazione: (2026) -
You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale
di: Ma, Baorui, et al.
Pubblicazione: (2024) -
Autoregressive Video Generation without Vector Quantization
di: Deng, Haoge, et al.
Pubblicazione: (2024) -
Beyond IID: Optimizing Instruction Learning from the Perspective of Instruction Interaction and Dependency
di: Zhao, Hanyu, et al.
Pubblicazione: (2024) -
Scaling Towards the Information Boundary of Instruction Sets: The Infinity Instruct Subject Technical Report
di: Du, Li, et al.
Pubblicazione: (2025)