POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Haicheng, Liu, Yuan, Liu, Yikun, Yu, Zhemeng, Zhao, Zhongyin, You, Yangxiu, Yu, Zilin, Tian, Le, Zhou, Xiao, Zhou, Jie, Xie, Weidi, Wang, Yanfeng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
POINTS-GUI-G: GUI-Grounding Journey
by: Zhao, Zhongyin, et al.
Published: (2026)
by: Zhao, Zhongyin, et al.
Published: (2026)
POINTS-Reader: Distillation-Free Adaptation of Vision-Language Models for Document Conversion
by: Liu, Yuan, et al.
Published: (2025)
by: Liu, Yuan, et al.
Published: (2025)
POINTS-Seeker: Towards Training a Multimodal Agentic Search Model from Scratch
by: Liu, Yikun, et al.
Published: (2026)
by: Liu, Yikun, et al.
Published: (2026)
VersaViT: Enhancing MLLM Vision Backbones via Task-Guided Optimization
by: Liu, Yikun, et al.
Published: (2026)
by: Liu, Yikun, et al.
Published: (2026)
POINTS: Improving Your Vision-language Model with Affordable Strategies
by: Liu, Yuan, et al.
Published: (2024)
by: Liu, Yuan, et al.
Published: (2024)
A Sanity Check on Composed Image Retrieval
by: Liu, Yikun, et al.
Published: (2026)
by: Liu, Yikun, et al.
Published: (2026)
POINTS1.5: Building a Vision-Language Model towards Real World Applications
by: Liu, Yuan, et al.
Published: (2024)
by: Liu, Yuan, et al.
Published: (2024)
Zero-shot Composed Text-Image Retrieval
by: Liu, Yikun, et al.
Published: (2023)
by: Liu, Yikun, et al.
Published: (2023)
When Seeing Is not Enough: Revealing the Limits of Active Reasoning in MLLMs
by: Liu, Hongcheng, et al.
Published: (2025)
by: Liu, Hongcheng, et al.
Published: (2025)
Knowledge-enhanced Visual-Language Pretraining for Computational Pathology
by: Zhou, Xiao, et al.
Published: (2024)
by: Zhou, Xiao, et al.
Published: (2024)
Intelligent Grimm -- Open-ended Visual Storytelling via Latent Diffusion Models
by: Liu, Chang, et al.
Published: (2023)
by: Liu, Chang, et al.
Published: (2023)
FOLDER: Accelerating Multi-modal Large Language Models with Enhanced Performance
by: Wang, Haicheng, et al.
Published: (2025)
by: Wang, Haicheng, et al.
Published: (2025)
When Eyes and Ears Disagree: Can MLLMs Discern Audio-Visual Confusion?
by: Ye, Qilang, et al.
Published: (2025)
by: Ye, Qilang, et al.
Published: (2025)
Drawing the Line: Enhancing Trustworthiness of MLLMs Through the Power of Refusal
by: Wang, Yuhao, et al.
Published: (2024)
by: Wang, Yuhao, et al.
Published: (2024)
LamRA: Large Multimodal Model as Your Advanced Retrieval Assistant
by: Liu, Yikun, et al.
Published: (2024)
by: Liu, Yikun, et al.
Published: (2024)
Towards Building Multilingual Language Model for Medicine
by: Qiu, Pengcheng, et al.
Published: (2024)
by: Qiu, Pengcheng, et al.
Published: (2024)
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
by: Ouyang, Kun, et al.
Published: (2025)
by: Ouyang, Kun, et al.
Published: (2025)
ChestX-Reasoner: Advancing Radiology Foundation Models with Reasoning through Step-by-Step Verification
by: Fan, Ziqing, et al.
Published: (2025)
by: Fan, Ziqing, et al.
Published: (2025)
MLLMs-Augmented Visual-Language Representation Learning
by: Liu, Yanqing, et al.
Published: (2023)
by: Liu, Yanqing, et al.
Published: (2023)
MatchTime: Towards Automatic Soccer Game Commentary Generation
by: Rao, Jiayuan, et al.
Published: (2024)
by: Rao, Jiayuan, et al.
Published: (2024)
Multi-Sentence Grounding for Long-term Instructional Video
by: Li, Zeqian, et al.
Published: (2023)
by: Li, Zeqian, et al.
Published: (2023)
AbductiveMLLM: Boosting Visual Abductive Reasoning Within MLLMs
by: Chang, Boyu, et al.
Published: (2026)
by: Chang, Boyu, et al.
Published: (2026)
Can Textual Reasoning Improve the Performance of MLLMs on Fine-grained Visual Classification?
by: Zhu, Jie, et al.
Published: (2026)
by: Zhu, Jie, et al.
Published: (2026)
Audio-Visual Segmentation via Unlabeled Frame Exploitation
by: Liu, Jinxiang, et al.
Published: (2024)
by: Liu, Jinxiang, et al.
Published: (2024)
Can Visual Foundation Models Achieve Long-term Point Tracking?
by: Aydemir, Görkay, et al.
Published: (2024)
by: Aydemir, Görkay, et al.
Published: (2024)
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
by: Zhang, Xiaoman, et al.
Published: (2023)
by: Zhang, Xiaoman, et al.
Published: (2023)
Agentic Active Omni-Modal Perception for Multi-Hop Audio-Visual Reasoning
by: Xu, Ke, et al.
Published: (2026)
by: Xu, Ke, et al.
Published: (2026)
SpotSound: Enhancing Large Audio-Language Models with Fine-Grained Temporal Grounding
by: Sun, Luoyi, et al.
Published: (2026)
by: Sun, Luoyi, et al.
Published: (2026)
Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs
by: Zhang, Xin, et al.
Published: (2026)
by: Zhang, Xin, et al.
Published: (2026)
Empowering Lightweight MLLMs with Reasoning via Long CoT SFT
by: Ou, Linyu, et al.
Published: (2025)
by: Ou, Linyu, et al.
Published: (2025)
Visual-RolePlay: Universal Jailbreak Attack on MultiModal Large Language Models via Role-playing Image Character
by: Ma, Siyuan, et al.
Published: (2024)
by: Ma, Siyuan, et al.
Published: (2024)
Dual‐Mode Thermo‐Responsive Janus Cellulosic Textiles for Visual Sensing, Adaptive Thermal Regulation, and Synergistic Bio‐Protection
by: Zhaochuan Yu, et al.
Published: (2026)
by: Zhaochuan Yu, et al.
Published: (2026)
GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization
by: Wang, Yikun, et al.
Published: (2025)
by: Wang, Yikun, et al.
Published: (2025)
A Robust Fault Detection Filter for Linear Time-Varying System with Non-Gaussian Noise
by: Zhang, Zhemeng, et al.
Published: (2025)
by: Zhang, Zhemeng, et al.
Published: (2025)
Grounded Question-Answering in Long Egocentric Videos
by: Di, Shangzhe, et al.
Published: (2023)
by: Di, Shangzhe, et al.
Published: (2023)
VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-Tuning
by: Wang, Qi, et al.
Published: (2025)
by: Wang, Qi, et al.
Published: (2025)
Chain of Mindset: Reasoning with Adaptive Cognitive Modes
by: Jiang, Tianyi, et al.
Published: (2026)
by: Jiang, Tianyi, et al.
Published: (2026)
M2K-VDG: Model-Adaptive Multimodal Knowledge Anchor Enhanced Video-grounded Dialogue Generation
by: Liu, Hongcheng, et al.
Published: (2024)
by: Liu, Hongcheng, et al.
Published: (2024)
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs
by: Li, Qi, et al.
Published: (2026)
by: Li, Qi, et al.
Published: (2026)
EchoSight: Advancing Visual-Language Models with Wiki Knowledge
by: Yan, Yibin, et al.
Published: (2024)
by: Yan, Yibin, et al.
Published: (2024)
Similar Items
-
POINTS-GUI-G: GUI-Grounding Journey
by: Zhao, Zhongyin, et al.
Published: (2026) -
POINTS-Reader: Distillation-Free Adaptation of Vision-Language Models for Document Conversion
by: Liu, Yuan, et al.
Published: (2025) -
POINTS-Seeker: Towards Training a Multimodal Agentic Search Model from Scratch
by: Liu, Yikun, et al.
Published: (2026) -
VersaViT: Enhancing MLLM Vision Backbones via Task-Guided Optimization
by: Liu, Yikun, et al.
Published: (2026) -
POINTS: Improving Your Vision-language Model with Affordable Strategies
by: Liu, Yuan, et al.
Published: (2024)