Enhancing Multimodal In-Context Learning via Inductive-Deductive Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Haoyu, Wang, Haonan, Chen, Yuyan, Chen, Jun, Liu, Gang, Wang, Qian, Yan, Jiahong, Xiao, Yanghua |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
XMeCap: Meme Caption Generation with Sub-Image Adaptability
by: Chen, Yuyan, et al.
Published: (2024)
by: Chen, Yuyan, et al.
Published: (2024)
Look Before You Decide: Prompting Active Deduction of MLLMs for Assumptive Reasoning
by: Li, Yian, et al.
Published: (2024)
by: Li, Yian, et al.
Published: (2024)
Fact-R1: Towards Explainable Video Misinformation Detection with Deep Reasoning
by: Zhang, Fanrui, et al.
Published: (2025)
by: Zhang, Fanrui, et al.
Published: (2025)
HOTVCOM: Generating Buzzworthy Comments for Videos
by: Chen, Yuyan, et al.
Published: (2024)
by: Chen, Yuyan, et al.
Published: (2024)
How Long Can Unified Multimodal Models Generate Images Reliably? Taming Long-Horizon Interleaved Image Generation via Context Curation
by: Chen, Haoyu, et al.
Published: (2026)
by: Chen, Haoyu, et al.
Published: (2026)
True Multimodal In-Context Learning Needs Attention to the Visual Context
by: Chen, Shuo, et al.
Published: (2025)
by: Chen, Shuo, et al.
Published: (2025)
Multimodal Video Emotion Recognition with Reliable Reasoning Priors
by: Wang, Zhepeng, et al.
Published: (2025)
by: Wang, Zhepeng, et al.
Published: (2025)
FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning
by: Ge, Haonan, et al.
Published: (2025)
by: Ge, Haonan, et al.
Published: (2025)
Enhancing Multimodal In-Context Learning for Image Classification through Coreset Optimization
by: Chen, Huiyi, et al.
Published: (2025)
by: Chen, Huiyi, et al.
Published: (2025)
A Neuro-Symbolic Framework Combining Inductive and Deductive Reasoning for Autonomous Driving Planning
by: Wei, Hongyan, et al.
Published: (2026)
by: Wei, Hongyan, et al.
Published: (2026)
LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts
by: Xiao, Yijia, et al.
Published: (2024)
by: Xiao, Yijia, et al.
Published: (2024)
Aligning Vision to Language: Annotation-Free Multimodal Knowledge Graph Construction for Enhanced LLMs Reasoning
by: Liu, Junming, et al.
Published: (2025)
by: Liu, Junming, et al.
Published: (2025)
VTPerception-R1: Enhancing Multimodal Reasoning via Explicit Visual and Textual Perceptual Grounding
by: Ding, Yizhuo, et al.
Published: (2025)
by: Ding, Yizhuo, et al.
Published: (2025)
TrackTeller: Temporal Multimodal 3D Grounding for Behavior-Dependent Object References
by: Yu, Jiahong, et al.
Published: (2025)
by: Yu, Jiahong, et al.
Published: (2025)
Socratic Questioning: Learn to Self-guide Multimodal Reasoning in the Wild
by: Hu, Wanpeng, et al.
Published: (2025)
by: Hu, Wanpeng, et al.
Published: (2025)
R-C2: Cycle-Consistent Reinforcement Learning Improves Multimodal Reasoning
by: Zhang, Zirui, et al.
Published: (2026)
by: Zhang, Zirui, et al.
Published: (2026)
M$^2$IV: Towards Efficient and Fine-grained Multimodal In-Context Learning via Representation Engineering
by: Li, Yanshu, et al.
Published: (2025)
by: Li, Yanshu, et al.
Published: (2025)
Fostering Video Reasoning via Next-Event Prediction
by: Wang, Haonan, et al.
Published: (2025)
by: Wang, Haonan, et al.
Published: (2025)
OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks
by: Wang, Jiayu, et al.
Published: (2025)
by: Wang, Jiayu, et al.
Published: (2025)
UniEval: Unified Holistic Evaluation for Unified Multimodal Understanding and Generation
by: Li, Yi, et al.
Published: (2025)
by: Li, Yi, et al.
Published: (2025)
CAPTCHA Solving for Native GUI Agents: Automated Reasoning-Action Data Generation and Self-Corrective Training
by: Chen, Yuxi, et al.
Published: (2026)
by: Chen, Yuxi, et al.
Published: (2026)
Discovering Pathology Rationale and Token Allocation for Efficient Multimodal Pathology Reasoning
by: Xu, Zhe, et al.
Published: (2025)
by: Xu, Zhe, et al.
Published: (2025)
GEMeX-RMCoT: An Enhanced Med-VQA Dataset for Region-Aware Multimodal Chain-of-Thought Reasoning
by: Liu, Bo, et al.
Published: (2025)
by: Liu, Bo, et al.
Published: (2025)
CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning
by: Wu, Hang, et al.
Published: (2026)
by: Wu, Hang, et al.
Published: (2026)
ParkingScenes: A Structured Dataset for End-to-End Autonomous Parking in Simulation Scenes
by: Chen, Haonan, et al.
Published: (2026)
by: Chen, Haonan, et al.
Published: (2026)
DiG: Differential Grounding for Enhancing Fine-Grained Perception in Multimodal Large Language Model
by: Tao, Zhou, et al.
Published: (2025)
by: Tao, Zhou, et al.
Published: (2025)
Visual-Noise Guided In-Context Distillation for Multimodal Large Language Model Unlearning
by: Chen, Junkai, et al.
Published: (2026)
by: Chen, Junkai, et al.
Published: (2026)
Perception Before Reasoning: Two-Stage Reinforcement Learning for Visual Reasoning in Vision-Language Models
by: Chen, Yan, et al.
Published: (2025)
by: Chen, Yan, et al.
Published: (2025)
Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start
by: Wei, Lai, et al.
Published: (2025)
by: Wei, Lai, et al.
Published: (2025)
Reasoning Can Hurt the Inductive Abilities of Large Language Models
by: Jin, Haibo, et al.
Published: (2025)
by: Jin, Haibo, et al.
Published: (2025)
When Models Judge Themselves: Unsupervised Self-Evolution for Multimodal Reasoning
by: Wu, Zhengxian, et al.
Published: (2026)
by: Wu, Zhengxian, et al.
Published: (2026)
CaptchaMind: Training CAPTCHA Solvers via Reinforcement Learning with Explicit Reasoning Supervision
by: Wang, Pengcheng, et al.
Published: (2026)
by: Wang, Pengcheng, et al.
Published: (2026)
Any-to-Any Learning in Computational Pathology via Triplet Multimodal Pretraining
by: Sun, Qichen, et al.
Published: (2025)
by: Sun, Qichen, et al.
Published: (2025)
PixelPrune: Pixel-Level Adaptive Visual Token Reduction via Predictive Coding
by: Wang, Nan, et al.
Published: (2026)
by: Wang, Nan, et al.
Published: (2026)
Self-Rewarded Multimodal Coherent Reasoning Across Diverse Visual Domains
by: Zhang, Jesen, et al.
Published: (2025)
by: Zhang, Jesen, et al.
Published: (2025)
CAD-GPT: Synthesising CAD Construction Sequence with Spatial Reasoning-Enhanced Multimodal LLMs
by: Wang, Siyu, et al.
Published: (2024)
by: Wang, Siyu, et al.
Published: (2024)
A Versatile Pathology Co-pilot via Reasoning Enhanced Multimodal Large Language Model
by: Xu, Zhe, et al.
Published: (2025)
by: Xu, Zhe, et al.
Published: (2025)
Trifuse: Enhancing Attention-Based GUI Grounding via Multimodal Fusion
by: Ma, Longhui, et al.
Published: (2026)
by: Ma, Longhui, et al.
Published: (2026)
From Illusion to Intention: Visual Rationale Learning for Vision-Language Reasoning
by: Wang, Changpeng, et al.
Published: (2025)
by: Wang, Changpeng, et al.
Published: (2025)
CoV: Chain-of-View Prompting for Spatial Reasoning
by: Zhao, Haoyu, et al.
Published: (2026)
by: Zhao, Haoyu, et al.
Published: (2026)
Similar Items
-
XMeCap: Meme Caption Generation with Sub-Image Adaptability
by: Chen, Yuyan, et al.
Published: (2024) -
Look Before You Decide: Prompting Active Deduction of MLLMs for Assumptive Reasoning
by: Li, Yian, et al.
Published: (2024) -
Fact-R1: Towards Explainable Video Misinformation Detection with Deep Reasoning
by: Zhang, Fanrui, et al.
Published: (2025) -
HOTVCOM: Generating Buzzworthy Comments for Videos
by: Chen, Yuyan, et al.
Published: (2024) -
How Long Can Unified Multimodal Models Generate Images Reliably? Taming Long-Horizon Interleaved Image Generation via Context Curation
by: Chen, Haoyu, et al.
Published: (2026)