SAVEn-Vid: Synergistic Audio-Visual Integration for Enhanced Understanding in Long Video Context
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Jungang, Tao, Sicheng, Yan, Yibo, Gu, Xiaojie, Xu, Haodong, Zheng, Xu, Lyu, Yuanhuiyi, Zhang, Linfeng, Hu, Xuming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
von: Zou, Xin, et al.
Veröffentlicht: (2025)
von: Zou, Xin, et al.
Veröffentlicht: (2025)
VSI: Visual Subtitle Integration for Keyframe Selection to enhance Long Video Understanding
von: He, Jianxiang, et al.
Veröffentlicht: (2025)
von: He, Jianxiang, et al.
Veröffentlicht: (2025)
RealRAG: Retrieval-augmented Realistic Image Generation via Self-reflective Contrastive Learning
von: Lyu, Yuanhuiyi, et al.
Veröffentlicht: (2025)
von: Lyu, Yuanhuiyi, et al.
Veröffentlicht: (2025)
Understanding-in-Generation: Reinforcing Generative Capability of Unified Model via Infusing Understanding into Generation
von: Lyu, Yuanhuiyi, et al.
Veröffentlicht: (2025)
von: Lyu, Yuanhuiyi, et al.
Veröffentlicht: (2025)
Learning Robust Anymodal Segmentor with Unimodal and Cross-modal Distillation
von: Zheng, Xu, et al.
Veröffentlicht: (2024)
von: Zheng, Xu, et al.
Veröffentlicht: (2024)
MMUnlearner: Reformulating Multimodal Machine Unlearning in the Era of Multimodal Large Language Models
von: Huo, Jiahao, et al.
Veröffentlicht: (2025)
von: Huo, Jiahao, et al.
Veröffentlicht: (2025)
OmniSAM: Omnidirectional Segment Anything Model for UDA in Panoramic Semantic Segmentation
von: Zhong, Ding, et al.
Veröffentlicht: (2025)
von: Zhong, Ding, et al.
Veröffentlicht: (2025)
MAGIC++: Efficient and Resilient Modality-Agnostic Semantic Segmentation via Hierarchical Modality Selection
von: Zheng, Xu, et al.
Veröffentlicht: (2024)
von: Zheng, Xu, et al.
Veröffentlicht: (2024)
RTV-Bench: Benchmarking MLLM Continuous Perception, Understanding and Reasoning through Real-Time Video
von: Xun, Shuhang, et al.
Veröffentlicht: (2025)
von: Xun, Shuhang, et al.
Veröffentlicht: (2025)
StruVis: Enhancing Reasoning-based Text-to-Image Generation via Thinking with Structured Vision
von: Lyu, Yuanhuiyi, et al.
Veröffentlicht: (2026)
von: Lyu, Yuanhuiyi, et al.
Veröffentlicht: (2026)
Unlocking Multimodal Document Intelligence: From Current Triumphs to Future Frontiers of Visual Document Retrieval
von: Yan, Yibo, et al.
Veröffentlicht: (2026)
von: Yan, Yibo, et al.
Veröffentlicht: (2026)
CoheMark: A Novel Sentence-Level Watermark for Enhanced Text Quality
von: Zhang, Junyan, et al.
Veröffentlicht: (2025)
von: Zhang, Junyan, et al.
Veröffentlicht: (2025)
EventBind: Learning a Unified Representation to Bind Them All for Event-based Open-world Understanding
von: Zhou, Jiazhou, et al.
Veröffentlicht: (2023)
von: Zhou, Jiazhou, et al.
Veröffentlicht: (2023)
Long-Video Audio Synthesis with Multi-Agent Collaboration
von: Zhang, Yehang, et al.
Veröffentlicht: (2025)
von: Zhang, Yehang, et al.
Veröffentlicht: (2025)
Learning Modality-agnostic Representation for Semantic Segmentation from Any Modalities
von: Zheng, Xu, et al.
Veröffentlicht: (2024)
von: Zheng, Xu, et al.
Veröffentlicht: (2024)
Image Anything: Towards Reasoning-coherent and Training-free Multi-modal Image Generation
von: Lyu, Yuanhuiyi, et al.
Veröffentlicht: (2024)
von: Lyu, Yuanhuiyi, et al.
Veröffentlicht: (2024)
PyraVid: Hierarchical Multimodal Memory for Long-Horizon Video Reasoning
von: Yan, Sikuan, et al.
Veröffentlicht: (2026)
von: Yan, Sikuan, et al.
Veröffentlicht: (2026)
Reducing Unimodal Bias in Multi-Modal Semantic Segmentation with Multi-Scale Functional Entropy Regularization
von: Zheng, Xu, et al.
Veröffentlicht: (2025)
von: Zheng, Xu, et al.
Veröffentlicht: (2025)
EgoIntent: An Egocentric Step-level Benchmark for Understanding What, Why, and Next
von: Pan, Ye, et al.
Veröffentlicht: (2026)
von: Pan, Ye, et al.
Veröffentlicht: (2026)
MOSS-ChatV: Reinforcement Learning with Process Reasoning Reward for Video Temporal Reasoning
von: Tao, Sicheng, et al.
Veröffentlicht: (2025)
von: Tao, Sicheng, et al.
Veröffentlicht: (2025)
StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding
von: Wang, Junxi, et al.
Veröffentlicht: (2026)
von: Wang, Junxi, et al.
Veröffentlicht: (2026)
A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges
von: Yan, Yibo, et al.
Veröffentlicht: (2024)
von: Yan, Yibo, et al.
Veröffentlicht: (2024)
Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression Methods
von: Liao, Chenfei, et al.
Veröffentlicht: (2025)
von: Liao, Chenfei, et al.
Veröffentlicht: (2025)
Retrieval Augmented Generation and Understanding in Vision: A Survey and New Outlook
von: Zheng, Xu, et al.
Veröffentlicht: (2025)
von: Zheng, Xu, et al.
Veröffentlicht: (2025)
AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficiency in Audio LLMs
von: He, Peize, et al.
Veröffentlicht: (2025)
von: He, Peize, et al.
Veröffentlicht: (2025)
Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models
von: Zou, Xin, et al.
Veröffentlicht: (2024)
von: Zou, Xin, et al.
Veröffentlicht: (2024)
MemorySAM: Memorize Modalities and Semantics with Segment Anything Model 2 for Multi-modal Semantic Segmentation
von: Liao, Chenfei, et al.
Veröffentlicht: (2025)
von: Liao, Chenfei, et al.
Veröffentlicht: (2025)
VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI
von: Cheng, Sijie, et al.
Veröffentlicht: (2024)
von: Cheng, Sijie, et al.
Veröffentlicht: (2024)
Sharp Eyes and Memory for VideoLLMs: Information-Aware Visual Token Pruning for Efficient and Reliable VideoLLM Reasoning
von: Qin, Jialong, et al.
Veröffentlicht: (2025)
von: Qin, Jialong, et al.
Veröffentlicht: (2025)
Centering the Value of Every Modality: Towards Efficient and Resilient Modality-agnostic Semantic Segmentation
von: Zheng, Xu, et al.
Veröffentlicht: (2024)
von: Zheng, Xu, et al.
Veröffentlicht: (2024)
OmniBind: Teach to Build Unequal-Scale Modality Interaction for Omni-Bind of All
von: Lyu, Yuanhuiyi, et al.
Veröffentlicht: (2024)
von: Lyu, Yuanhuiyi, et al.
Veröffentlicht: (2024)
UniBind: LLM-Augmented Unified and Balanced Representation Space to Bind Them All
von: Lyu, Yuanhuiyi, et al.
Veröffentlicht: (2024)
von: Lyu, Yuanhuiyi, et al.
Veröffentlicht: (2024)
ExACT: Language-guided Conceptual Reasoning and Uncertainty Estimation for Event-based Action Recognition and More
von: Zhou, Jiazhou, et al.
Veröffentlicht: (2024)
von: Zhou, Jiazhou, et al.
Veröffentlicht: (2024)
VideoCompressa: Data-Efficient Video Understanding via Joint Temporal Compression and Spatial Reconstruction
von: Wang, Shaobo, et al.
Veröffentlicht: (2025)
von: Wang, Shaobo, et al.
Veröffentlicht: (2025)
TextVidBench: A Benchmark for Long Video Scene Text Understanding
von: Zhong, Yangyang, et al.
Veröffentlicht: (2025)
von: Zhong, Yangyang, et al.
Veröffentlicht: (2025)
EIT-1M: One Million EEG-Image-Text Pairs for Human Visual-textual Recognition and More
von: Zheng, Xu, et al.
Veröffentlicht: (2024)
von: Zheng, Xu, et al.
Veröffentlicht: (2024)
Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?
von: Dongfang, Zihao, et al.
Veröffentlicht: (2025)
von: Dongfang, Zihao, et al.
Veröffentlicht: (2025)
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
von: Dai, Song, et al.
Veröffentlicht: (2025)
von: Dai, Song, et al.
Veröffentlicht: (2025)
Unveiling Instruction-Specific Neurons & Experts: An Analytical Framework for LLM's Instruction-Following Capabilities
von: Zhang, Junyan, et al.
Veröffentlicht: (2025)
von: Zhang, Junyan, et al.
Veröffentlicht: (2025)
ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
von: Xu, Yicheng, et al.
Veröffentlicht: (2025)
von: Xu, Yicheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
von: Zou, Xin, et al.
Veröffentlicht: (2025) -
VSI: Visual Subtitle Integration for Keyframe Selection to enhance Long Video Understanding
von: He, Jianxiang, et al.
Veröffentlicht: (2025) -
RealRAG: Retrieval-augmented Realistic Image Generation via Self-reflective Contrastive Learning
von: Lyu, Yuanhuiyi, et al.
Veröffentlicht: (2025) -
Understanding-in-Generation: Reinforcing Generative Capability of Unified Model via Infusing Understanding into Generation
von: Lyu, Yuanhuiyi, et al.
Veröffentlicht: (2025) -
Learning Robust Anymodal Segmentor with Unimodal and Cross-modal Distillation
von: Zheng, Xu, et al.
Veröffentlicht: (2024)