Training-Free Reasoning and Reflection in MLLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wei, Hongchen, Chen, Zhenzhong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LOP: Learning Optimal Pruning for Efficient On-Demand MLLMs Scaling
von: Zhang, Zhihan, et al.
Veröffentlicht: (2025)
von: Zhang, Zhihan, et al.
Veröffentlicht: (2025)
Visual Context Window Extension: A New Perspective for Long Video Understanding
von: Wei, Hongchen, et al.
Veröffentlicht: (2024)
von: Wei, Hongchen, et al.
Veröffentlicht: (2024)
ETC: Extreme Token Compression via Task-aware Visual Information Distillation in VLMs
von: Gao, Yiling, et al.
Veröffentlicht: (2026)
von: Gao, Yiling, et al.
Veröffentlicht: (2026)
LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models
von: Wei, Hongchen, et al.
Veröffentlicht: (2025)
von: Wei, Hongchen, et al.
Veröffentlicht: (2025)
Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
von: Han, Su Ho, et al.
Veröffentlicht: (2025)
von: Han, Su Ho, et al.
Veröffentlicht: (2025)
FreeRet: MLLMs as Training-Free Retrievers
von: Zhu, Yuhan, et al.
Veröffentlicht: (2025)
von: Zhu, Yuhan, et al.
Veröffentlicht: (2025)
RSFAKE-1M: A Large-Scale Dataset for Detecting Diffusion-Generated Remote Sensing Forgeries
von: Tan, Zhihong, et al.
Veröffentlicht: (2025)
von: Tan, Zhihong, et al.
Veröffentlicht: (2025)
Math Blind: Failures in Diagram Understanding Undermine Reasoning in MLLMs
von: Sun, Yanpeng, et al.
Veröffentlicht: (2025)
von: Sun, Yanpeng, et al.
Veröffentlicht: (2025)
From Training-Free to Adaptive: Empirical Insights into MLLMs' Understanding of Detection Information
von: Jiao, Qirui, et al.
Veröffentlicht: (2024)
von: Jiao, Qirui, et al.
Veröffentlicht: (2024)
Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
von: Tong, Jintao, et al.
Veröffentlicht: (2025)
von: Tong, Jintao, et al.
Veröffentlicht: (2025)
Visual Jigsaw Post-Training Improves MLLMs
von: Wu, Penghao, et al.
Veröffentlicht: (2025)
von: Wu, Penghao, et al.
Veröffentlicht: (2025)
EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs
von: Chen, Zhenghao, et al.
Veröffentlicht: (2026)
von: Chen, Zhenghao, et al.
Veröffentlicht: (2026)
VER-Bench: Evaluating MLLMs on Reasoning with Fine-Grained Visual Evidence
von: Qiang, Chenhui, et al.
Veröffentlicht: (2025)
von: Qiang, Chenhui, et al.
Veröffentlicht: (2025)
Can MLLMs Reason Beyond Language? VisReason: A Comprehensive Benchmark for Vision-Centric Reasoning
von: Guo, Longteng, et al.
Veröffentlicht: (2026)
von: Guo, Longteng, et al.
Veröffentlicht: (2026)
Reason-before-Retrieve: One-Stage Reflective Chain-of-Thoughts for Training-Free Zero-Shot Composed Image Retrieval
von: Tang, Yuanmin, et al.
Veröffentlicht: (2024)
von: Tang, Yuanmin, et al.
Veröffentlicht: (2024)
CameraBench: Benchmarking Visual Reasoning in MLLMs via Photography
von: Fang, I-Sheng, et al.
Veröffentlicht: (2025)
von: Fang, I-Sheng, et al.
Veröffentlicht: (2025)
[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs
von: Wang, Ao, et al.
Veröffentlicht: (2024)
von: Wang, Ao, et al.
Veröffentlicht: (2024)
Video-R1: Reinforcing Video Reasoning in MLLMs
von: Feng, Kaituo, et al.
Veröffentlicht: (2025)
von: Feng, Kaituo, et al.
Veröffentlicht: (2025)
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
von: Ouyang, Kun, et al.
Veröffentlicht: (2025)
von: Ouyang, Kun, et al.
Veröffentlicht: (2025)
Touch-R1: Reinforcing Touch Reasoning in MLLMs
von: Lai, Yingxin, et al.
Veröffentlicht: (2026)
von: Lai, Yingxin, et al.
Veröffentlicht: (2026)
Training-Free Multimodal Deepfake Detection via Graph Reasoning
von: Liu, Yuxin, et al.
Veröffentlicht: (2025)
von: Liu, Yuxin, et al.
Veröffentlicht: (2025)
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators
von: Zhou, Jiazhou, et al.
Veröffentlicht: (2026)
von: Zhou, Jiazhou, et al.
Veröffentlicht: (2026)
Can MLLMs Reason About Visual Persuasion? Evaluating the Efficacy and Faithfulness of Reasoning
von: Lee, Naeun, et al.
Veröffentlicht: (2026)
von: Lee, Naeun, et al.
Veröffentlicht: (2026)
Open Eyes, Then Reason: Fine-grained Visual Mathematical Understanding in MLLMs
von: Zhang, Shan, et al.
Veröffentlicht: (2025)
von: Zhang, Shan, et al.
Veröffentlicht: (2025)
DiffVC-RT: Towards Practical Real-Time Diffusion-based Perceptual Neural Video Compression
von: Ma, Wenzhuo, et al.
Veröffentlicht: (2026)
von: Ma, Wenzhuo, et al.
Veröffentlicht: (2026)
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
von: Lu, Lidong, et al.
Veröffentlicht: (2025)
von: Lu, Lidong, et al.
Veröffentlicht: (2025)
Think 360°: Evaluating the Width-centric Reasoning Capability of MLLMs Beyond Depth
von: Chen, Mingrui, et al.
Veröffentlicht: (2026)
von: Chen, Mingrui, et al.
Veröffentlicht: (2026)
Unlocking the Forgery Detection Potential of Vanilla MLLMs: A Novel Training-Free Pipeline
von: Zuo, Rui, et al.
Veröffentlicht: (2025)
von: Zuo, Rui, et al.
Veröffentlicht: (2025)
360° Image Perception with MLLMs: A Comprehensive Benchmark and a Training-Free Method
von: Tran, Huyen T. T., et al.
Veröffentlicht: (2026)
von: Tran, Huyen T. T., et al.
Veröffentlicht: (2026)
VisualQuest: A Benchmark for Abstract Visual Reasoning in MLLMs
von: Xiao, Kelaiti, et al.
Veröffentlicht: (2025)
von: Xiao, Kelaiti, et al.
Veröffentlicht: (2025)
Empowering Lightweight MLLMs with Reasoning via Long CoT SFT
von: Ou, Linyu, et al.
Veröffentlicht: (2025)
von: Ou, Linyu, et al.
Veröffentlicht: (2025)
AbductiveMLLM: Boosting Visual Abductive Reasoning Within MLLMs
von: Chang, Boyu, et al.
Veröffentlicht: (2026)
von: Chang, Boyu, et al.
Veröffentlicht: (2026)
POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs
von: Wang, Haicheng, et al.
Veröffentlicht: (2026)
von: Wang, Haicheng, et al.
Veröffentlicht: (2026)
Incentivizing Cardiologist-Like Reasoning in MLLMs for Interpretable Echocardiographic Diagnosis
von: Qin, Yi, et al.
Veröffentlicht: (2026)
von: Qin, Yi, et al.
Veröffentlicht: (2026)
Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
von: Kim, Sanghwan, et al.
Veröffentlicht: (2025)
von: Kim, Sanghwan, et al.
Veröffentlicht: (2025)
VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?
von: Liu, Yuanxin, et al.
Veröffentlicht: (2025)
von: Liu, Yuanxin, et al.
Veröffentlicht: (2025)
MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs
von: Yuan, Jiakang, et al.
Veröffentlicht: (2025)
von: Yuan, Jiakang, et al.
Veröffentlicht: (2025)
Connecting the Dots: Training-Free Visual Grounding via Agentic Reasoning
von: Luo, Liqin, et al.
Veröffentlicht: (2025)
von: Luo, Liqin, et al.
Veröffentlicht: (2025)
Intention-driven Ego-to-Exo Video Generation
von: Luo, Hongchen, et al.
Veröffentlicht: (2024)
von: Luo, Hongchen, et al.
Veröffentlicht: (2024)
Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning
von: Deng, Huilin, et al.
Veröffentlicht: (2025)
von: Deng, Huilin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LOP: Learning Optimal Pruning for Efficient On-Demand MLLMs Scaling
von: Zhang, Zhihan, et al.
Veröffentlicht: (2025) -
Visual Context Window Extension: A New Perspective for Long Video Understanding
von: Wei, Hongchen, et al.
Veröffentlicht: (2024) -
ETC: Extreme Token Compression via Task-aware Visual Information Distillation in VLMs
von: Gao, Yiling, et al.
Veröffentlicht: (2026) -
LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models
von: Wei, Hongchen, et al.
Veröffentlicht: (2025) -
Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
von: Han, Su Ho, et al.
Veröffentlicht: (2025)