From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents
Fuente:
arXiv
Salvato in:
| Autori principali: | Lian, Niu, Wang, Yuting, Yao, Hanshu, Wang, Jinpeng, Chen, Bin, Wang, Yaowei, Zhang, Min, Xia, Shu-Tao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AutoSSVH: Exploring Automated Frame Sampling for Efficient Self-Supervised Video Hashing
di: Lian, Niu, et al.
Pubblicazione: (2025)
di: Lian, Niu, et al.
Pubblicazione: (2025)
Imagine Before Concentration: Diffusion-Guided Registers Enhance Partially Relevant Video Retrieval
di: Li, Jun, et al.
Pubblicazione: (2026)
di: Li, Jun, et al.
Pubblicazione: (2026)
HLFormer: Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning
di: Li, Jun, et al.
Pubblicazione: (2025)
di: Li, Jun, et al.
Pubblicazione: (2025)
Efficient Self-Supervised Video Hashing with Selective State Spaces
di: Wang, Jinpeng, et al.
Pubblicazione: (2024)
di: Wang, Jinpeng, et al.
Pubblicazione: (2024)
GMMFormer: Gaussian-Mixture-Model Based Transformer for Efficient Partially Relevant Video Retrieval
di: Wang, Yuting, et al.
Pubblicazione: (2023)
di: Wang, Yuting, et al.
Pubblicazione: (2023)
Revisiting Uncertainty: On Evidential Learning for Partially Relevant Video Retrieval
di: Li, Jun, et al.
Pubblicazione: (2026)
di: Li, Jun, et al.
Pubblicazione: (2026)
Love Me, Love My Label: Rethinking the Role of Labels in Prompt Retrieval for Visual In-Context Learning
di: Luo, Tianci, et al.
Pubblicazione: (2026)
di: Luo, Tianci, et al.
Pubblicazione: (2026)
Rebalancing Contrastive Alignment with Bottlenecked Semantic Increments in Text-Video Retrieval
di: Xiao, Jian, et al.
Pubblicazione: (2025)
di: Xiao, Jian, et al.
Pubblicazione: (2025)
SemCORE: A Semantic-Enhanced Generative Cross-Modal Retrieval Framework with MLLMs
di: Li, Haoxuan, et al.
Pubblicazione: (2025)
di: Li, Haoxuan, et al.
Pubblicazione: (2025)
From Swath to Full-Disc: Advancing Precipitation Retrieval with Multimodal Knowledge Expansion
di: Wang, Zheng, et al.
Pubblicazione: (2025)
di: Wang, Zheng, et al.
Pubblicazione: (2025)
Beyond Static Collision Handling: Adaptive Semantic ID Learning for Multimodal Recommendation at Industrial Scale
di: Pan, Yongsen, et al.
Pubblicazione: (2026)
di: Pan, Yongsen, et al.
Pubblicazione: (2026)
Dynamic Multimodal Fusion via Meta-Learning Towards Micro-Video Recommendation
di: Liu, Han, et al.
Pubblicazione: (2025)
di: Liu, Han, et al.
Pubblicazione: (2025)
Balancing Semantic Relevance and Engagement in Related Video Recommendations
di: Jaspal, Amit, et al.
Pubblicazione: (2025)
di: Jaspal, Amit, et al.
Pubblicazione: (2025)
Multimodal Misinformation Detection using Large Vision-Language Models
di: Tahmasebi, Sahar, et al.
Pubblicazione: (2024)
di: Tahmasebi, Sahar, et al.
Pubblicazione: (2024)
Hierarchical Vision-Language Reasoning for Multimodal Multiple-Choice Question Answering
di: Zhou, Ao, et al.
Pubblicazione: (2025)
di: Zhou, Ao, et al.
Pubblicazione: (2025)
AgentDisCo: Towards Disentanglement and Collaboration in Open-ended Deep Research Agents
di: Jin, Jiarui, et al.
Pubblicazione: (2026)
di: Jin, Jiarui, et al.
Pubblicazione: (2026)
VisTopics: A Visual Semantic Unsupervised Approach to Topic Modeling of Video and Image Data
di: Lokmanoglu, Ayse D, et al.
Pubblicazione: (2025)
di: Lokmanoglu, Ayse D, et al.
Pubblicazione: (2025)
MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval
di: Wu, Qiyu, et al.
Pubblicazione: (2025)
di: Wu, Qiyu, et al.
Pubblicazione: (2025)
Why Multi-Interest Fairness Matters: Hypergraph Contrastive Multi-Interest Learning for Fair Conversational Recommender System
di: Zheng, Yongsen, et al.
Pubblicazione: (2025)
di: Zheng, Yongsen, et al.
Pubblicazione: (2025)
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
di: Kong, Fanheng, et al.
Pubblicazione: (2025)
di: Kong, Fanheng, et al.
Pubblicazione: (2025)
MMSRARec: Summarization and Retrieval Augumented Sequential Recommendation Based on Multimodal Large Language Model
di: Wang, Haoyu, et al.
Pubblicazione: (2025)
di: Wang, Haoyu, et al.
Pubblicazione: (2025)
Very Efficient Listwise Multimodal Reranking for Long Documents
di: Sun, Yiqun, et al.
Pubblicazione: (2026)
di: Sun, Yiqun, et al.
Pubblicazione: (2026)
CM$^3$: Calibrating Multimodal Recommendation
di: Zhou, Xin, et al.
Pubblicazione: (2025)
di: Zhou, Xin, et al.
Pubblicazione: (2025)
GenState-AI: State-Aware Dataset for Text-to-Video Retrieval on AI-Generated Videos
di: Li, Minghan, et al.
Pubblicazione: (2026)
di: Li, Minghan, et al.
Pubblicazione: (2026)
Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval
di: Ning, Hailong, et al.
Pubblicazione: (2025)
di: Ning, Hailong, et al.
Pubblicazione: (2025)
CoPRS: Learning Positional Prior from Chain-of-Thought for Reasoning Segmentation
di: Lu, Zhenyu, et al.
Pubblicazione: (2025)
di: Lu, Zhenyu, et al.
Pubblicazione: (2025)
Interactive Multi-Turn Retrieval for Health Videos
di: Wu, Chengzheng, et al.
Pubblicazione: (2026)
di: Wu, Chengzheng, et al.
Pubblicazione: (2026)
Any2Any: Incomplete Multimodal Retrieval with Conformal Prediction
di: Li, Po-han, et al.
Pubblicazione: (2024)
di: Li, Po-han, et al.
Pubblicazione: (2024)
VKIE: The Application of Key Information Extraction on Video Text
di: An, Siyu, et al.
Pubblicazione: (2023)
di: An, Siyu, et al.
Pubblicazione: (2023)
Benchmarking Multimodal Large Language Models for Missing Modality Completion in Product Catalogues
di: Fu, Junchen, et al.
Pubblicazione: (2026)
di: Fu, Junchen, et al.
Pubblicazione: (2026)
Automating Steering for Safe Multimodal Large Language Models
di: Wu, Lyucheng, et al.
Pubblicazione: (2025)
di: Wu, Lyucheng, et al.
Pubblicazione: (2025)
CLEAR: Null-Space Projection for Cross-Modal De-Redundancy in Multimodal Recommendation
di: Zhan, Hao, et al.
Pubblicazione: (2026)
di: Zhan, Hao, et al.
Pubblicazione: (2026)
Learning Partially-Decorrelated Common Spaces for Ad-hoc Video Search
di: Hu, Fan, et al.
Pubblicazione: (2025)
di: Hu, Fan, et al.
Pubblicazione: (2025)
OTCR: Optimal Transmission, Compression and Representation for Multimodal Information Extraction
di: Li, Yang, et al.
Pubblicazione: (2025)
di: Li, Yang, et al.
Pubblicazione: (2025)
CLOSP: A Unified Semantic Space for SAR, MSI, and Text in Remote Sensing
di: Cambrin, Daniele Rege, et al.
Pubblicazione: (2025)
di: Cambrin, Daniele Rege, et al.
Pubblicazione: (2025)
CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception
di: Li, Liupeng, et al.
Pubblicazione: (2026)
di: Li, Liupeng, et al.
Pubblicazione: (2026)
Mitigating Modality Bias in Multi-modal Entity Alignment from a Causal Perspective
di: Su, Taoyu, et al.
Pubblicazione: (2025)
di: Su, Taoyu, et al.
Pubblicazione: (2025)
Verifying Cross-modal Entity Consistency in News using Vision-language Models
di: Tahmasebi, Sahar, et al.
Pubblicazione: (2025)
di: Tahmasebi, Sahar, et al.
Pubblicazione: (2025)
PREMISE: Matching-based Prediction for Accurate Review Recommendation
di: Han, Wei, et al.
Pubblicazione: (2025)
di: Han, Wei, et al.
Pubblicazione: (2025)
Multimodal Neural Databases
di: Trappolini, Giovanni, et al.
Pubblicazione: (2023)
di: Trappolini, Giovanni, et al.
Pubblicazione: (2023)
Documenti analoghi
-
AutoSSVH: Exploring Automated Frame Sampling for Efficient Self-Supervised Video Hashing
di: Lian, Niu, et al.
Pubblicazione: (2025) -
Imagine Before Concentration: Diffusion-Guided Registers Enhance Partially Relevant Video Retrieval
di: Li, Jun, et al.
Pubblicazione: (2026) -
HLFormer: Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning
di: Li, Jun, et al.
Pubblicazione: (2025) -
Efficient Self-Supervised Video Hashing with Selective State Spaces
di: Wang, Jinpeng, et al.
Pubblicazione: (2024) -
GMMFormer: Gaussian-Mixture-Model Based Transformer for Efficient Partially Relevant Video Retrieval
di: Wang, Yuting, et al.
Pubblicazione: (2023)