Towards Multimodal Video Paragraph Captioning Models Robust to Missing Modality
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Sishuo, Li, Lei, Ren, Shuhuai, Gao, Rundong, Liu, Yuanxin, Bi, Xiaohan, Sun, Xu, Hou, Lu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TempCompass: Do Video LLMs Really Understand Videos?
by: Liu, Yuanxin, et al.
Published: (2024)
by: Liu, Yuanxin, et al.
Published: (2024)
VITATECS: A Diagnostic Dataset for Temporal Concept Understanding of Video-Language Models
by: Li, Shicheng, et al.
Published: (2023)
by: Li, Shicheng, et al.
Published: (2023)
DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models
by: Yao, Linli, et al.
Published: (2024)
by: Yao, Linli, et al.
Published: (2024)
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
by: Ren, Shuhuai, et al.
Published: (2023)
by: Ren, Shuhuai, et al.
Published: (2023)
UVE: Are MLLMs Unified Evaluators for AI-Generated Videos?
by: Liu, Yuanxin, et al.
Published: (2025)
by: Liu, Yuanxin, et al.
Published: (2025)
GEM-VPC: A dual Graph-Enhanced Multimodal integration for Video Paragraph Captioning
by: Wang, Eileen, et al.
Published: (2024)
by: Wang, Eileen, et al.
Published: (2024)
Next Block Prediction: Video Generation via Semi-Autoregressive Modeling
by: Ren, Shuhuai, et al.
Published: (2025)
by: Ren, Shuhuai, et al.
Published: (2025)
TEMPLE: Incentivizing Temporal Understanding of Video Large Language Models via Progressive Pre-SFT Alignment
by: Li, Shicheng, et al.
Published: (2025)
by: Li, Shicheng, et al.
Published: (2025)
PRA-PoE: Robust Multimodal Alzheimer's Diagnosis with Arbitrary Missing Modalities
by: Yang, Guangqian, et al.
Published: (2026)
by: Yang, Guangqian, et al.
Published: (2026)
LaDiC: Are Diffusion Models Really Inferior to Autoregressive Counterparts for Image-to-Text Generation?
by: Wang, Yuchi, et al.
Published: (2024)
by: Wang, Yuchi, et al.
Published: (2024)
Purify-then-Align: Towards Robust Human Sensing under Modality Missing with Knowledge Distillation from Noisy Multimodal Teacher
by: Weng, Pengcheng, et al.
Published: (2026)
by: Weng, Pengcheng, et al.
Published: (2026)
TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions
by: Yao, Linli, et al.
Published: (2026)
by: Yao, Linli, et al.
Published: (2026)
Reconstruct before Query: Continual Missing Modality Learning with Decomposed Prompt Collaboration
by: Zhao, Shu, et al.
Published: (2024)
by: Zhao, Shu, et al.
Published: (2024)
Multimodal Knowledge Distillation for Egocentric Action Recognition Robust to Missing Modalities
by: Santos-Villafranca, Maria, et al.
Published: (2025)
by: Santos-Villafranca, Maria, et al.
Published: (2025)
Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
by: Tang, Yunlong, et al.
Published: (2025)
by: Tang, Yunlong, et al.
Published: (2025)
TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos
by: Yao, Linli, et al.
Published: (2025)
by: Yao, Linli, et al.
Published: (2025)
Chameleon: Images Are What You Need For Multimodal Learning Robust To Missing Modalities
by: Liaqat, Muhammad Irzam, et al.
Published: (2024)
by: Liaqat, Muhammad Irzam, et al.
Published: (2024)
Paragraph-to-Image Generation with Information-Enriched Diffusion Model
by: Wu, Weijia, et al.
Published: (2023)
by: Wu, Weijia, et al.
Published: (2023)
LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models
by: Wei, Hongchen, et al.
Published: (2025)
by: Wei, Hongchen, et al.
Published: (2025)
Towards Stable Cross-Domain Depression Recognition under Missing Modalities
by: Chen, Jiuyi, et al.
Published: (2025)
by: Chen, Jiuyi, et al.
Published: (2025)
Dealing with All-stage Missing Modality: Towards A Universal Model with Robust Reconstruction and Personalization
by: Zhao, Yunpeng, et al.
Published: (2024)
by: Zhao, Yunpeng, et al.
Published: (2024)
Exploring Missing Modality in Multimodal Egocentric Datasets
by: Ramazanova, Merey, et al.
Published: (2024)
by: Ramazanova, Merey, et al.
Published: (2024)
Towards Robust Multimodal Emotion Recognition under Missing Modalities and Distribution Shifts
by: Zhong, Guowei, et al.
Published: (2025)
by: Zhong, Guowei, et al.
Published: (2025)
Retrieval-Augmented Egocentric Video Captioning
by: Xu, Jilan, et al.
Published: (2024)
by: Xu, Jilan, et al.
Published: (2024)
Robust Multimodal Learning with Missing Modalities via Parameter-Efficient Adaptation
by: Reza, Md Kaykobad, et al.
Published: (2023)
by: Reza, Md Kaykobad, et al.
Published: (2023)
MissBench: Benchmarking Multimodal Affective Analysis under Imbalanced Missing Modalities
by: Pham, Tien Anh, et al.
Published: (2026)
by: Pham, Tien Anh, et al.
Published: (2026)
Gradient-Guided Modality Decoupling for Missing-Modality Robustness
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
Towards Robust Optical-SAR Object Detection under Missing Modalities: A Dynamic Quality-Aware Fusion Framework
by: Zhao, Zhicheng, et al.
Published: (2025)
by: Zhao, Zhicheng, et al.
Published: (2025)
Robust Divergence Learning for Missing-Modality Segmentation
by: Cheng, Runze, et al.
Published: (2024)
by: Cheng, Runze, et al.
Published: (2024)
Dual-task Mutual Reinforcing Embedded Joint Video Paragraph Retrieval and Grounding
by: Wang, Mengzhao, et al.
Published: (2024)
by: Wang, Mengzhao, et al.
Published: (2024)
Siamese Learning with Joint Alignment and Regression for Weakly-Supervised Video Paragraph Grounding
by: Tan, Chaolei, et al.
Published: (2024)
by: Tan, Chaolei, et al.
Published: (2024)
Instruction-guided Multi-Granularity Segmentation and Captioning with Large Multimodal Model
by: Zhou, Li, et al.
Published: (2024)
by: Zhou, Li, et al.
Published: (2024)
Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers
by: Chen, Tsai-Shien, et al.
Published: (2024)
by: Chen, Tsai-Shien, et al.
Published: (2024)
Synergistic Prompting for Robust Visual Recognition with Missing Modalities
by: Zhang, Zhihui, et al.
Published: (2025)
by: Zhang, Zhihui, et al.
Published: (2025)
MCMoE: Completing Missing Modalities with Mixture of Experts for Incomplete Multimodal Action Quality Assessment
by: Xu, Huangbiao, et al.
Published: (2025)
by: Xu, Huangbiao, et al.
Published: (2025)
Explicit Temporal-Semantic Modeling for Dense Video Captioning via Context-Aware Cross-Modal Interaction
by: Jia, Mingda, et al.
Published: (2025)
by: Jia, Mingda, et al.
Published: (2025)
Robust Multi-Modal Face Anti-Spoofing with Domain Adaptation: Tackling Missing Modalities, Noisy Pseudo-Labels, and Model Degradation
by: Hsu, Ming-Tsung, et al.
Published: (2025)
by: Hsu, Ming-Tsung, et al.
Published: (2025)
Benchmarking Multi-modal Semantic Segmentation under Sensor Failures: Missing and Noisy Modality Robustness
by: Liao, Chenfei, et al.
Published: (2025)
by: Liao, Chenfei, et al.
Published: (2025)
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search
by: Yu, Linhao, et al.
Published: (2025)
by: Yu, Linhao, et al.
Published: (2025)
Controlling Decision Drift in Multimodal Sentiment Analysis with Missing Modalities
by: Chen, Chenglizhao, et al.
Published: (2026)
by: Chen, Chenglizhao, et al.
Published: (2026)
Similar Items
-
TempCompass: Do Video LLMs Really Understand Videos?
by: Liu, Yuanxin, et al.
Published: (2024) -
VITATECS: A Diagnostic Dataset for Temporal Concept Understanding of Video-Language Models
by: Li, Shicheng, et al.
Published: (2023) -
DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models
by: Yao, Linli, et al.
Published: (2024) -
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
by: Ren, Shuhuai, et al.
Published: (2023) -
UVE: Are MLLMs Unified Evaluators for AI-Generated Videos?
by: Liu, Yuanxin, et al.
Published: (2025)