Leveraging multimodal explanatory annotations for video interpretation with Modality Specific Dataset
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ancarani, Elisa, Tores, Julie, Sassatelli, Lucile, Sun, Rémy, Wu, Hui-Yin, Precioso, Frédéric |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SynCB: A Synergy Concept-Based Model with Dynamic Routing Between Concepts and Complementary Neural Branches
von: Julie, Tores, et al.
Veröffentlicht: (2026)
von: Julie, Tores, et al.
Veröffentlicht: (2026)
MObyGaze: a film dataset of multimodal objectification densely annotated by experts
von: Tores, Julie, et al.
Veröffentlicht: (2025)
von: Tores, Julie, et al.
Veröffentlicht: (2025)
DiVR: incorporating context from diverse VR scenes for human trajectory prediction
von: Gallo, Franz Franco, et al.
Veröffentlicht: (2024)
von: Gallo, Franz Franco, et al.
Veröffentlicht: (2024)
Visual Objectification in Films: Towards a New AI Task for Video Interpretation
von: Tores, Julie, et al.
Veröffentlicht: (2024)
von: Tores, Julie, et al.
Veröffentlicht: (2024)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
GTPBD-MM: A Global Terraced Parcel and Boundary Dataset with Multi-Modality
von: Zhang, Zhiwei, et al.
Veröffentlicht: (2026)
von: Zhang, Zhiwei, et al.
Veröffentlicht: (2026)
Leveraging Automatic Personalised Nutrition: Food Image Recognition Benchmark and Dataset based on Nutrition Taxonomy
von: Romero-Tapiador, Sergio, et al.
Veröffentlicht: (2022)
von: Romero-Tapiador, Sergio, et al.
Veröffentlicht: (2022)
Anisotropic Modality Align
von: Yu, Xiaomin, et al.
Veröffentlicht: (2026)
von: Yu, Xiaomin, et al.
Veröffentlicht: (2026)
Robust Modality-incomplete Anomaly Detection: A Modality-instructive Framework with Benchmark
von: Miao, Bingchen, et al.
Veröffentlicht: (2024)
von: Miao, Bingchen, et al.
Veröffentlicht: (2024)
EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing
von: Li, Huilai, et al.
Veröffentlicht: (2026)
von: Li, Huilai, et al.
Veröffentlicht: (2026)
Gradient-Guided Modality Decoupling for Missing-Modality Robustness
von: Wang, Hao, et al.
Veröffentlicht: (2024)
von: Wang, Hao, et al.
Veröffentlicht: (2024)
Towards Emotion Analysis in Short-form Videos: A Large-Scale Dataset and Baseline
von: Wu, Xuecheng, et al.
Veröffentlicht: (2023)
von: Wu, Xuecheng, et al.
Veröffentlicht: (2023)
Similarity over Factuality: Are we making progress on multimodal out-of-context misinformation detection?
von: Papadopoulos, Stefanos-Iordanis, et al.
Veröffentlicht: (2024)
von: Papadopoulos, Stefanos-Iordanis, et al.
Veröffentlicht: (2024)
Robust Self-Paced Hashing for Cross-Modal Retrieval with Noisy Labels
von: Pu, Ruitao, et al.
Veröffentlicht: (2025)
von: Pu, Ruitao, et al.
Veröffentlicht: (2025)
Interactive Spatial-Frequency Fusion Mamba for Multi-Modal Image Fusion
von: Zhu, Yixin, et al.
Veröffentlicht: (2026)
von: Zhu, Yixin, et al.
Veröffentlicht: (2026)
Cross-Modal and Uni-Modal Soft-Label Alignment for Image-Text Retrieval
von: Huang, Hailang, et al.
Veröffentlicht: (2024)
von: Huang, Hailang, et al.
Veröffentlicht: (2024)
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing
von: Zheng, Junjie, et al.
Veröffentlicht: (2025)
von: Zheng, Junjie, et al.
Veröffentlicht: (2025)
Towards Holistic Language-video Representation: the language model-enhanced MSR-Video to Text Dataset
von: Yang, Yuchen, et al.
Veröffentlicht: (2024)
von: Yang, Yuchen, et al.
Veröffentlicht: (2024)
Consistent and Invariant Generalization Learning for Short-video Misinformation Detection
von: Guo, Hanghui, et al.
Veröffentlicht: (2025)
von: Guo, Hanghui, et al.
Veröffentlicht: (2025)
AdaptaGen: Domain-Specific Image Generation through Hierarchical Semantic Optimization Framework
von: Zhang, Suoxiang, et al.
Veröffentlicht: (2025)
von: Zhang, Suoxiang, et al.
Veröffentlicht: (2025)
VC-Bench: Pioneering the Video Connecting Benchmark with a Dataset and Evaluation Metrics
von: Yin, Zhiyu, et al.
Veröffentlicht: (2026)
von: Yin, Zhiyu, et al.
Veröffentlicht: (2026)
Towards Identity-Aware Cross-Modal Retrieval: a Dataset and a Baseline
von: Messina, Nicola, et al.
Veröffentlicht: (2024)
von: Messina, Nicola, et al.
Veröffentlicht: (2024)
SSTFB: Leveraging self-supervised pretext learning and temporal self-attention with feature branching for real-time video polyp segmentation
von: Xu, Ziang, et al.
Veröffentlicht: (2024)
von: Xu, Ziang, et al.
Veröffentlicht: (2024)
Joint Explicit and Implicit Cross-Modal Interaction Network for Anterior Chamber Inflammation Diagnosis
von: Shao, Qian, et al.
Veröffentlicht: (2023)
von: Shao, Qian, et al.
Veröffentlicht: (2023)
Does SpatioTemporal information benefit Two video summarization benchmarks?
von: Ganesh, Aashutosh, et al.
Veröffentlicht: (2024)
von: Ganesh, Aashutosh, et al.
Veröffentlicht: (2024)
Reviewing Intelligent Cinematography: AI research for camera-based video production
von: Azzarelli, Adrian, et al.
Veröffentlicht: (2024)
von: Azzarelli, Adrian, et al.
Veröffentlicht: (2024)
Subjective evaluation of UHD video coded using VVC with LCEVC and ML-VVC
von: Ramzan, Naeem, et al.
Veröffentlicht: (2026)
von: Ramzan, Naeem, et al.
Veröffentlicht: (2026)
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
von: Chen, Liyang, et al.
Veröffentlicht: (2025)
von: Chen, Liyang, et al.
Veröffentlicht: (2025)
CMTA: Leveraging Cross-Modal Temporal Artifacts for Generalizable AI-Generated Video Detection
von: Wang, Hang, et al.
Veröffentlicht: (2026)
von: Wang, Hang, et al.
Veröffentlicht: (2026)
Quantifying and Enhancing Multi-modal Robustness with Modality Preference
von: Yang, Zequn, et al.
Veröffentlicht: (2024)
von: Yang, Zequn, et al.
Veröffentlicht: (2024)
MSCT: Differential Cross-Modal Attention for Deepfake Detection
von: Wei, Fangda, et al.
Veröffentlicht: (2026)
von: Wei, Fangda, et al.
Veröffentlicht: (2026)
Find the Cliffhanger: Multi-Modal Trailerness in Soap Operas
von: Bretti, Carlo, et al.
Veröffentlicht: (2024)
von: Bretti, Carlo, et al.
Veröffentlicht: (2024)
3D2M Dataset: A 3-Dimension diverse Mesh Dataset
von: Dasgupta, Sankarshan
Veröffentlicht: (2024)
von: Dasgupta, Sankarshan
Veröffentlicht: (2024)
Towards Robust and Realible Multimodal Misinformation Recognition with Incomplete Modality
von: Zhou, Hengyang, et al.
Veröffentlicht: (2025)
von: Zhou, Hengyang, et al.
Veröffentlicht: (2025)
Modality-Aware Shot Relating and Comparing for Video Scene Detection
von: Tan, Jiawei, et al.
Veröffentlicht: (2024)
von: Tan, Jiawei, et al.
Veröffentlicht: (2024)
CoNav: Collaborative Cross-Modal Reasoning for Embodied Navigation
von: Hao, Haihong, et al.
Veröffentlicht: (2025)
von: Hao, Haihong, et al.
Veröffentlicht: (2025)
Neighbor-aware Instance Refining with Noisy Labels for Cross-Modal Retrieval
von: Liu, Yizhi, et al.
Veröffentlicht: (2025)
von: Liu, Yizhi, et al.
Veröffentlicht: (2025)
EALD-MLLM: Emotion Analysis in Long-sequential and De-identity videos with Multi-modal Large Language Model
von: Li, Deng, et al.
Veröffentlicht: (2024)
von: Li, Deng, et al.
Veröffentlicht: (2024)
Multi-Modal Image Fusion via Intervention-Stable Feature Learning
von: Wang, Xue, et al.
Veröffentlicht: (2026)
von: Wang, Xue, et al.
Veröffentlicht: (2026)
Towards Universal Modal Tracking with Online Dense Temporal Token Learning
von: Zheng, Yaozong, et al.
Veröffentlicht: (2025)
von: Zheng, Yaozong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SynCB: A Synergy Concept-Based Model with Dynamic Routing Between Concepts and Complementary Neural Branches
von: Julie, Tores, et al.
Veröffentlicht: (2026) -
MObyGaze: a film dataset of multimodal objectification densely annotated by experts
von: Tores, Julie, et al.
Veröffentlicht: (2025) -
DiVR: incorporating context from diverse VR scenes for human trajectory prediction
von: Gallo, Franz Franco, et al.
Veröffentlicht: (2024) -
Visual Objectification in Films: Towards a New AI Task for Video Interpretation
von: Tores, Julie, et al.
Veröffentlicht: (2024) -
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)