Free-MoRef: Instantly Multiplexing Context Perception Capabilities of Video-MLLMs within Single Inference
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Kuo, Zheng, Quanlong, Xie, Junlin, Zhang, Yanhao, Luo, Jinguo, Lu, Haonan, Lin, Liang, Zhou, Fan, Li, Guanbin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
H2VU-Benchmark: A Comprehensive Benchmark for Hierarchical Holistic Video Understanding
di: Wu, Qi, et al.
Pubblicazione: (2025)
di: Wu, Qi, et al.
Pubblicazione: (2025)
OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning
di: Liang, Zhijia, et al.
Pubblicazione: (2026)
di: Liang, Zhijia, et al.
Pubblicazione: (2026)
Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM
di: Liu, Peng, et al.
Pubblicazione: (2025)
di: Liu, Peng, et al.
Pubblicazione: (2025)
Mobile-Agent-RAG: Driving Smart Multi-Agent Coordination with Contextual Knowledge Empowerment for Long-Horizon Mobile Automation
di: Zhou, Yuxiang, et al.
Pubblicazione: (2025)
di: Zhou, Yuxiang, et al.
Pubblicazione: (2025)
An Empirical Study on Configuring In-Context Learning Demonstrations for Unleashing MLLMs' Sentimental Perception Capability
di: Wu, Daiqing, et al.
Pubblicazione: (2025)
di: Wu, Daiqing, et al.
Pubblicazione: (2025)
LLMI3D: MLLM-based 3D Perception from a Single 2D Image
di: Yang, Fan, et al.
Pubblicazione: (2024)
di: Yang, Fan, et al.
Pubblicazione: (2024)
Achieving Net‐Zero Through AI ‐Enabled Dynamic Capability and Green Servitization in Chinese Furniture Manufacturing Industry
di: Jinguo Li
Pubblicazione: (2026)
di: Jinguo Li
Pubblicazione: (2026)
From Pixels to Feelings: Aligning MLLMs with Human Cognitive Perception of Images
di: Chen, Yiming, et al.
Pubblicazione: (2025)
di: Chen, Yiming, et al.
Pubblicazione: (2025)
Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
di: Xiao, Tong, et al.
Pubblicazione: (2025)
di: Xiao, Tong, et al.
Pubblicazione: (2025)
Large Multimodal Agents: A Survey
di: Xie, Junlin, et al.
Pubblicazione: (2024)
di: Xie, Junlin, et al.
Pubblicazione: (2024)
Collaborative Multi-Agent Scripts Generation for Enhancing Imperfect-Information Reasoning in Murder Mystery Games
di: Zhong, Keyang, et al.
Pubblicazione: (2026)
di: Zhong, Keyang, et al.
Pubblicazione: (2026)
RefTok: Reference-Based Tokenization for Video Generation
di: Fan, Xiang, et al.
Pubblicazione: (2025)
di: Fan, Xiang, et al.
Pubblicazione: (2025)
RefDecoder: Enhancing Visual Generation with Conditional Video Decoding
di: Fan, Xiang, et al.
Pubblicazione: (2026)
di: Fan, Xiang, et al.
Pubblicazione: (2026)
Credible Teacher for Semi-Supervised Object Detection in Open Scene
di: Zhuang, Jingyu, et al.
Pubblicazione: (2024)
di: Zhuang, Jingyu, et al.
Pubblicazione: (2024)
Video-MSR: Benchmarking Multi-hop Spatial Reasoning Capabilities of MLLMs
di: Zhu, Rui, et al.
Pubblicazione: (2026)
di: Zhu, Rui, et al.
Pubblicazione: (2026)
VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-Tuning
di: Wang, Qi, et al.
Pubblicazione: (2025)
di: Wang, Qi, et al.
Pubblicazione: (2025)
VLDrive: Vision-Augmented Lightweight MLLMs for Efficient Language-grounded Autonomous Driving
di: Zhang, Ruifei, et al.
Pubblicazione: (2025)
di: Zhang, Ruifei, et al.
Pubblicazione: (2025)
LocRef-Diffusion:Tuning-Free Layout and Appearance-Guided Generation
di: Deng, Fan, et al.
Pubblicazione: (2024)
di: Deng, Fan, et al.
Pubblicazione: (2024)
Exploring the Design Space of Visual Context Representation in Video MLLMs
di: Du, Yifan, et al.
Pubblicazione: (2024)
di: Du, Yifan, et al.
Pubblicazione: (2024)
X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding
di: Sun, Peiwen, et al.
Pubblicazione: (2026)
di: Sun, Peiwen, et al.
Pubblicazione: (2026)
RefDrop: Controllable Consistency in Image or Video Generation via Reference Feature Guidance
di: Fan, Jiaojiao, et al.
Pubblicazione: (2024)
di: Fan, Jiaojiao, et al.
Pubblicazione: (2024)
MEIA: Multimodal Embodied Perception and Interaction in Unknown Environments
di: Liu, Yang, et al.
Pubblicazione: (2024)
di: Liu, Yang, et al.
Pubblicazione: (2024)
CodePercept: Code-Grounded Visual STEM Perception for MLLMs
di: Guan, Tongkun, et al.
Pubblicazione: (2026)
di: Guan, Tongkun, et al.
Pubblicazione: (2026)
Multifaceted Evaluation of Audio-Visual Capability for MLLMs: Effectiveness, Efficiency, Generalizability and Robustness
di: Zhao, Yusheng, et al.
Pubblicazione: (2025)
di: Zhao, Yusheng, et al.
Pubblicazione: (2025)
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
di: Wang, Yi, et al.
Pubblicazione: (2025)
di: Wang, Yi, et al.
Pubblicazione: (2025)
ChartEdit: How Far Are MLLMs From Automating Chart Analysis? Evaluating MLLMs' Capability via Chart Editing
di: Zhao, Xuanle, et al.
Pubblicazione: (2025)
di: Zhao, Xuanle, et al.
Pubblicazione: (2025)
RefAlign: Representation Alignment for Reference-to-Video Generation
di: Wang, Lei, et al.
Pubblicazione: (2026)
di: Wang, Lei, et al.
Pubblicazione: (2026)
DSPNet: Dual-vision Scene Perception for Robust 3D Question Answering
di: Luo, Jingzhou, et al.
Pubblicazione: (2025)
di: Luo, Jingzhou, et al.
Pubblicazione: (2025)
Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
di: Han, Su Ho, et al.
Pubblicazione: (2025)
di: Han, Su Ho, et al.
Pubblicazione: (2025)
V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding
di: Ge, Junqi, et al.
Pubblicazione: (2024)
di: Ge, Junqi, et al.
Pubblicazione: (2024)
Ref-Long: Benchmarking the Long-context Referencing Capability of Long-context Language Models
di: Wu, Junjie, et al.
Pubblicazione: (2025)
di: Wu, Junjie, et al.
Pubblicazione: (2025)
Tuning-Free Long Video Generation via Global-Local Collaborative Diffusion
di: Ma, Yongjia, et al.
Pubblicazione: (2025)
di: Ma, Yongjia, et al.
Pubblicazione: (2025)
360° Image Perception with MLLMs: A Comprehensive Benchmark and a Training-Free Method
di: Tran, Huyen T. T., et al.
Pubblicazione: (2026)
di: Tran, Huyen T. T., et al.
Pubblicazione: (2026)
A Retrospect to Multi-prompt Learning across Vision and Language
di: Chen, Ziliang, et al.
Pubblicazione: (2025)
di: Chen, Ziliang, et al.
Pubblicazione: (2025)
MarvelOVD: Marrying Object Recognition and Vision-Language Models for Robust Open-Vocabulary Object Detection
di: Wang, Kuo, et al.
Pubblicazione: (2024)
di: Wang, Kuo, et al.
Pubblicazione: (2024)
Linking Perception, Confidence and Accuracy in MLLMs
di: Du, Yuetian, et al.
Pubblicazione: (2026)
di: Du, Yuetian, et al.
Pubblicazione: (2026)
Region-Adaptive Video Sharpening via Rate-Perception Optimization
di: Pang, Yingxue, et al.
Pubblicazione: (2025)
di: Pang, Yingxue, et al.
Pubblicazione: (2025)
HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses through Reasoning MLLMs
di: Qin, Zheng, et al.
Pubblicazione: (2025)
di: Qin, Zheng, et al.
Pubblicazione: (2025)
Ref-DGS: Reflective Dual Gaussian Splatting
di: Fan, Ningjing, et al.
Pubblicazione: (2026)
di: Fan, Ningjing, et al.
Pubblicazione: (2026)
Interpretable Embeddings for Segmentation-Free Single-Cell Analysis in Multiplex Imaging
di: Gutwein, Simon, et al.
Pubblicazione: (2024)
di: Gutwein, Simon, et al.
Pubblicazione: (2024)
Documenti analoghi
-
H2VU-Benchmark: A Comprehensive Benchmark for Hierarchical Holistic Video Understanding
di: Wu, Qi, et al.
Pubblicazione: (2025) -
OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning
di: Liang, Zhijia, et al.
Pubblicazione: (2026) -
Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM
di: Liu, Peng, et al.
Pubblicazione: (2025) -
Mobile-Agent-RAG: Driving Smart Multi-Agent Coordination with Contextual Knowledge Empowerment for Long-Horizon Mobile Automation
di: Zhou, Yuxiang, et al.
Pubblicazione: (2025) -
An Empirical Study on Configuring In-Context Learning Demonstrations for Unleashing MLLMs' Sentimental Perception Capability
di: Wu, Daiqing, et al.
Pubblicazione: (2025)