Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Haoran, Wang, Hongyu, Li, Jiaze, Chen, Shunpeng, Tong, Zizhao, Ju, Jianzhong, Luo, Zhenbo, Luan, Jian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Federated Joint Learning for Domain and Class Generalization
von: Xu, Haoran, et al.
Veröffentlicht: (2026)
von: Xu, Haoran, et al.
Veröffentlicht: (2026)
TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding
von: Xu, Boshen, et al.
Veröffentlicht: (2025)
von: Xu, Boshen, et al.
Veröffentlicht: (2025)
Think-Clip-Sample: Slow-Fast Frame Selection for Video Understanding
von: Tan, Wenhui, et al.
Veröffentlicht: (2026)
von: Tan, Wenhui, et al.
Veröffentlicht: (2026)
LLaVA-SG: Leveraging Scene Graphs as Visual Semantic Expression in Vision-Language Models
von: Wang, Jingyi, et al.
Veröffentlicht: (2024)
von: Wang, Jingyi, et al.
Veröffentlicht: (2024)
Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation
von: Li, Jiaze, et al.
Veröffentlicht: (2026)
von: Li, Jiaze, et al.
Veröffentlicht: (2026)
REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding
von: Li, Jiaze, et al.
Veröffentlicht: (2025)
von: Li, Jiaze, et al.
Veröffentlicht: (2025)
MSJoE: Jointly Evolving MLLM and Sampler for Efficient Long-Form Video Understanding
von: Tan, Wenhui, et al.
Veröffentlicht: (2026)
von: Tan, Wenhui, et al.
Veröffentlicht: (2026)
Divide-and-Conquer Inference for Large-Scale Visual Recognition with Multimodal Large Language Models
von: Ye, Zhipeng, et al.
Veröffentlicht: (2026)
von: Ye, Zhipeng, et al.
Veröffentlicht: (2026)
Direction-Aware Diagonal Autoregressive Image Generation
von: Xu, Yijia, et al.
Veröffentlicht: (2025)
von: Xu, Yijia, et al.
Veröffentlicht: (2025)
Divide and Conquer: Heterogeneous Noise Integration for Diffusion-based Adversarial Purification
von: Pei, Gaozheng, et al.
Veröffentlicht: (2025)
von: Pei, Gaozheng, et al.
Veröffentlicht: (2025)
ProxyThinker: Test-Time Guidance through Small Visual Reasoners
von: Xiao, Zilin, et al.
Veröffentlicht: (2025)
von: Xiao, Zilin, et al.
Veröffentlicht: (2025)
StreamPro: From Reactive Perception to Proactive Decision-Making in Streaming Video
von: Li, Ao, et al.
Veröffentlicht: (2026)
von: Li, Ao, et al.
Veröffentlicht: (2026)
Xiaomi MiMo-VL-Miloco Technical Report
von: Li, Jiaze, et al.
Veröffentlicht: (2025)
von: Li, Jiaze, et al.
Veröffentlicht: (2025)
SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
Federated Balanced Learning
von: Li, Jiaze, et al.
Veröffentlicht: (2026)
von: Li, Jiaze, et al.
Veröffentlicht: (2026)
ThinkOmni: Lifting Textual Reasoning to Omni-modal Scenarios via Guidance Decoding
von: Guan, Yiran, et al.
Veröffentlicht: (2026)
von: Guan, Yiran, et al.
Veröffentlicht: (2026)
MosaicThinker: On-Device Visual Spatial Reasoning for Embodied AI via Iterative Construction of Space Representation
von: Wang, Haoming, et al.
Veröffentlicht: (2026)
von: Wang, Haoming, et al.
Veröffentlicht: (2026)
PatchCue: Enhancing Vision-Language Model Reasoning with Patch-Based Visual Cues
von: Qi, Yukun, et al.
Veröffentlicht: (2026)
von: Qi, Yukun, et al.
Veröffentlicht: (2026)
Divide and Conquer: Rethinking the Training Paradigm of Neural Radiance Fields
von: Ma, Rongkai, et al.
Veröffentlicht: (2024)
von: Ma, Rongkai, et al.
Veröffentlicht: (2024)
ControlThinker: Unveiling Latent Semantics for Controllable Image Generation through Visual Reasoning
von: Han, Feng, et al.
Veröffentlicht: (2025)
von: Han, Feng, et al.
Veröffentlicht: (2025)
Divide and Conquer: Object Co-occurrence Helps Mitigate Simplicity Bias in OOD Detection
von: Dai, Boyang, et al.
Veröffentlicht: (2026)
von: Dai, Boyang, et al.
Veröffentlicht: (2026)
EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models
von: Fang, Yiyang, et al.
Veröffentlicht: (2026)
von: Fang, Yiyang, et al.
Veröffentlicht: (2026)
DCText: Scheduled Attention Masking for Visual Text Generation via Divide-and-Conquer Strategy
von: Song, Jaewoo, et al.
Veröffentlicht: (2025)
von: Song, Jaewoo, et al.
Veröffentlicht: (2025)
Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
von: Guan, Yiran, et al.
Veröffentlicht: (2026)
von: Guan, Yiran, et al.
Veröffentlicht: (2026)
Divide and Conquer: Static-Dynamic Collaboration for Few-Shot Class-Incremental Learning
von: Bao, Kexin, et al.
Veröffentlicht: (2026)
von: Bao, Kexin, et al.
Veröffentlicht: (2026)
MERG3R: A Divide-and-Conquer Approach to Large-Scale Neural Visual Geometry
von: Cheng, Leo Kaixuan, et al.
Veröffentlicht: (2026)
von: Cheng, Leo Kaixuan, et al.
Veröffentlicht: (2026)
Divide-and-Conquer: Tree-structured Strategy with Answer Distribution Estimator for Goal-Oriented Visual Dialogue
von: Cai, Shuo, et al.
Veröffentlicht: (2025)
von: Cai, Shuo, et al.
Veröffentlicht: (2025)
Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
von: Xiao, Tong, et al.
Veröffentlicht: (2025)
von: Xiao, Tong, et al.
Veröffentlicht: (2025)
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes
von: Wang, Yuji, et al.
Veröffentlicht: (2025)
von: Wang, Yuji, et al.
Veröffentlicht: (2025)
ViSA-Enhanced Aerial VLN: A Visual-Spatial Reasoning Enhanced Framework for Aerial Vision-Language Navigation
von: Tong, Haoyu, et al.
Veröffentlicht: (2026)
von: Tong, Haoyu, et al.
Veröffentlicht: (2026)
Towards Better Visualizing the Decision Basis of Networks via Unfold and Conquer Attribution Guidance
von: Hong, Jung-Ho, et al.
Veröffentlicht: (2023)
von: Hong, Jung-Ho, et al.
Veröffentlicht: (2023)
Cook and Clean Together: Teaching Embodied Agents for Parallel Task Execution
von: Liang, Dingkang, et al.
Veröffentlicht: (2025)
von: Liang, Dingkang, et al.
Veröffentlicht: (2025)
BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent
von: Zhang, Shaojie, et al.
Veröffentlicht: (2025)
von: Zhang, Shaojie, et al.
Veröffentlicht: (2025)
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning
von: Li, Chenglin, et al.
Veröffentlicht: (2026)
von: Li, Chenglin, et al.
Veröffentlicht: (2026)
Knot So Simple: A Minimalistic Environment for Spatial Reasoning
von: Chen, Zizhao, et al.
Veröffentlicht: (2025)
von: Chen, Zizhao, et al.
Veröffentlicht: (2025)
Unified Thinker: A General Reasoning Modular Core for Image Generation
von: Zhou, Sashuai, et al.
Veröffentlicht: (2026)
von: Zhou, Sashuai, et al.
Veröffentlicht: (2026)
VITAL: Visual-Semantic Dual Supervision for Enhanced and Interpretable Latent Reasoning in Medical MLLMs
von: Li, Qiaoru, et al.
Veröffentlicht: (2026)
von: Li, Qiaoru, et al.
Veröffentlicht: (2026)
Divide and Conquer Self-Supervised Learning for High-Content Imaging
von: Farndale, Lucas, et al.
Veröffentlicht: (2025)
von: Farndale, Lucas, et al.
Veröffentlicht: (2025)
Beyond Visual Memory: Mechanistic Diagnostics of Latent Visual Reasoning
von: Guo, Garvin, et al.
Veröffentlicht: (2026)
von: Guo, Garvin, et al.
Veröffentlicht: (2026)
DivCon: Divide and Conquer for Complex Numerical and Spatial Reasoning in Text-to-Image Generation
von: Jia, Yuhao, et al.
Veröffentlicht: (2024)
von: Jia, Yuhao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Federated Joint Learning for Domain and Class Generalization
von: Xu, Haoran, et al.
Veröffentlicht: (2026) -
TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding
von: Xu, Boshen, et al.
Veröffentlicht: (2025) -
Think-Clip-Sample: Slow-Fast Frame Selection for Video Understanding
von: Tan, Wenhui, et al.
Veröffentlicht: (2026) -
LLaVA-SG: Leveraging Scene Graphs as Visual Semantic Expression in Vision-Language Models
von: Wang, Jingyi, et al.
Veröffentlicht: (2024) -
Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation
von: Li, Jiaze, et al.
Veröffentlicht: (2026)