video-SALMONN S: Memory-Enhanced Streaming Audio-Visual LLM
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Guangzhi, Li, Yixuan, Wu, Xiaodong, Yang, Yudong, Li, Wei, Ma, Zejun, Zhang, Chao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models
by: Tang, Changli, et al.
Published: (2025)
by: Tang, Changli, et al.
Published: (2025)
video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language Model
by: Sun, Guangzhi, et al.
Published: (2025)
by: Sun, Guangzhi, et al.
Published: (2025)
video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
by: Sun, Guangzhi, et al.
Published: (2024)
by: Sun, Guangzhi, et al.
Published: (2024)
Audio-centric Video Understanding Benchmark without Text Shortcut
by: Yang, Yudong, et al.
Published: (2025)
by: Yang, Yudong, et al.
Published: (2025)
Improving LLM Video Understanding with 16 Frames Per Second
by: Li, Yixuan, et al.
Published: (2025)
by: Li, Yixuan, et al.
Published: (2025)
Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding
by: Wu, Hang, et al.
Published: (2026)
by: Wu, Hang, et al.
Published: (2026)
Speech-Audio Compositional Attacks on Multimodal LLMs and Their Mitigation with SALMONN-Guard
by: Yang, Yudong, et al.
Published: (2025)
by: Yang, Yudong, et al.
Published: (2025)
Enhancing Multimodal LLM for Detailed and Accurate Video Captioning using Multi-Round Preference Optimization
by: Tang, Changli, et al.
Published: (2024)
by: Tang, Changli, et al.
Published: (2024)
Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning
by: Li, Zejun, et al.
Published: (2025)
by: Li, Zejun, et al.
Published: (2025)
FreshMem: Brain-Inspired Frequency-Space Hybrid Memory for Streaming Video Understanding
by: Li, Kangcong, et al.
Published: (2026)
by: Li, Kangcong, et al.
Published: (2026)
Audio-Sync Video Generation with Multi-Stream Temporal Control
by: Weng, Shuchen, et al.
Published: (2025)
by: Weng, Shuchen, et al.
Published: (2025)
STNet: Deep Audio-Visual Fusion Network for Robust Speaker Tracking
by: Li, Yidi, et al.
Published: (2024)
by: Li, Yidi, et al.
Published: (2024)
Stepping Stones: A Progressive Training Strategy for Audio-Visual Semantic Segmentation
by: Ma, Juncheng, et al.
Published: (2024)
by: Ma, Juncheng, et al.
Published: (2024)
Ref-AVS: Refer and Segment Objects in Audio-Visual Scenes
by: Wang, Yaoting, et al.
Published: (2024)
by: Wang, Yaoting, et al.
Published: (2024)
StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
by: Yang, Yanlai, et al.
Published: (2025)
by: Yang, Yanlai, et al.
Published: (2025)
Decoupled Audio-Visual Dataset Distillation
by: Li, Wenyuan, et al.
Published: (2025)
by: Li, Wenyuan, et al.
Published: (2025)
Audio-Guided Visual Perception for Audio-Visual Navigation
by: Wang, Yi, et al.
Published: (2025)
by: Wang, Yi, et al.
Published: (2025)
Towards Visual-Prompt Temporal Answering Grounding in Medical Instructional Video
by: Li, Bin, et al.
Published: (2022)
by: Li, Bin, et al.
Published: (2022)
Scaling the Long Video Understanding of Multimodal Large Language Models via Visual Memory Mechanism
by: Chen, Tao, et al.
Published: (2026)
by: Chen, Tao, et al.
Published: (2026)
StreamingAssistant: Efficient Visual Token Pruning for Accelerating Online Video Understanding
by: Jin, Xinqi, et al.
Published: (2025)
by: Jin, Xinqi, et al.
Published: (2025)
FluxMem: Adaptive Hierarchical Memory for Streaming Video Understanding
by: Xie, Yiweng, et al.
Published: (2026)
by: Xie, Yiweng, et al.
Published: (2026)
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs
by: Chowdhury, Sanjoy, et al.
Published: (2025)
by: Chowdhury, Sanjoy, et al.
Published: (2025)
CSMCIR: CoT-Enhanced Symmetric Alignment with Memory Bank for Composed Image Retrieval
by: Qian, Zhipeng, et al.
Published: (2026)
by: Qian, Zhipeng, et al.
Published: (2026)
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
by: Lin, Junming, et al.
Published: (2024)
by: Lin, Junming, et al.
Published: (2024)
Unsupervised Audio-Visual Segmentation with Modality Alignment
by: Bhosale, Swapnil, et al.
Published: (2024)
by: Bhosale, Swapnil, et al.
Published: (2024)
VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models
by: Li, Zejun, et al.
Published: (2024)
by: Li, Zejun, et al.
Published: (2024)
Visual Position Prompt for MLLM based Visual Grounding
by: Tang, Wei, et al.
Published: (2025)
by: Tang, Wei, et al.
Published: (2025)
QAVA: Query-Agnostic Visual Attack to Large Vision-Language Models
by: Zhang, Yudong, et al.
Published: (2025)
by: Zhang, Yudong, et al.
Published: (2025)
Flow caching for autoregressive video generation
by: Ma, Yuexiao, et al.
Published: (2026)
by: Ma, Yuexiao, et al.
Published: (2026)
Beyond Visual Memory: Mechanistic Diagnostics of Latent Visual Reasoning
by: Guo, Garvin, et al.
Published: (2026)
by: Guo, Garvin, et al.
Published: (2026)
BlobCtrl: Taming Controllable Blob for Element-level Image Editing
by: Li, Yaowei, et al.
Published: (2025)
by: Li, Yaowei, et al.
Published: (2025)
Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge
by: Xiong, Haomiao, et al.
Published: (2025)
by: Xiong, Haomiao, et al.
Published: (2025)
MemoNav: Working Memory Model for Visual Navigation
by: Li, Hongxin, et al.
Published: (2024)
by: Li, Hongxin, et al.
Published: (2024)
StreamTinyNet: video streaming analysis with spatial-temporal TinyML
by: Shalby, Hazem Hesham Yousef, et al.
Published: (2024)
by: Shalby, Hazem Hesham Yousef, et al.
Published: (2024)
SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation
by: Zhang, Jiwen, et al.
Published: (2026)
by: Zhang, Jiwen, et al.
Published: (2026)
Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs
by: Huang, Siyuan, et al.
Published: (2026)
by: Huang, Siyuan, et al.
Published: (2026)
Phantom: Subject-consistent video generation via cross-modal alignment
by: Liu, Lijie, et al.
Published: (2025)
by: Liu, Lijie, et al.
Published: (2025)
CRD: Collaborative Representation Distance for Practical Anomaly Detection
by: Han, Chao, et al.
Published: (2023)
by: Han, Chao, et al.
Published: (2023)
StrLoRA: Towards Streaming Continual Visual Instruction Tuning for MLLMs
by: Che, Chang, et al.
Published: (2026)
by: Che, Chang, et al.
Published: (2026)
DiffAttn: Diffusion-Based Drivers' Visual Attention Prediction with LLM-Enhanced Semantic Reasoning
by: Liu, Weimin, et al.
Published: (2026)
by: Liu, Weimin, et al.
Published: (2026)
Similar Items
-
video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models
by: Tang, Changli, et al.
Published: (2025) -
video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language Model
by: Sun, Guangzhi, et al.
Published: (2025) -
video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
by: Sun, Guangzhi, et al.
Published: (2024) -
Audio-centric Video Understanding Benchmark without Text Shortcut
by: Yang, Yudong, et al.
Published: (2025) -
Improving LLM Video Understanding with 16 Frames Per Second
by: Li, Yixuan, et al.
Published: (2025)