CoTZero: Annotation-Free Human-Like Vision Reasoning via Hierarchical Synthetic CoT
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Du, Chengyi, Niu, Yazhe, Shen, Dazhong, Xu, Luxin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Let Androids Dream of Electric Sheep: A Human-Inspired Image Implication Understanding and Reasoning Framework
von: Zhang, Chenhao, et al.
Veröffentlicht: (2025)
von: Zhang, Chenhao, et al.
Veröffentlicht: (2025)
Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT
von: Asfour, Alaa, et al.
Veröffentlicht: (2026)
von: Asfour, Alaa, et al.
Veröffentlicht: (2026)
VG-CoT: Towards Trustworthy Visual Reasoning via Grounded Chain-of-Thought
von: Lim, Byeonggeuk, et al.
Veröffentlicht: (2026)
von: Lim, Byeonggeuk, et al.
Veröffentlicht: (2026)
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
von: Zhao, Qingqing, et al.
Veröffentlicht: (2025)
von: Zhao, Qingqing, et al.
Veröffentlicht: (2025)
Pretrained Reversible Generation as Unsupervised Visual Representation Learning
von: Xue, Rongkun, et al.
Veröffentlicht: (2024)
von: Xue, Rongkun, et al.
Veröffentlicht: (2024)
$\texttt{Complex-Edit}$: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark
von: Yang, Siwei, et al.
Veröffentlicht: (2025)
von: Yang, Siwei, et al.
Veröffentlicht: (2025)
CodePlot-CoT: Mathematical Visual Reasoning by Thinking with Code-Driven Images
von: Duan, Chengqi, et al.
Veröffentlicht: (2025)
von: Duan, Chengqi, et al.
Veröffentlicht: (2025)
LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
von: Xu, Guowei, et al.
Veröffentlicht: (2024)
von: Xu, Guowei, et al.
Veröffentlicht: (2024)
MedCLM: Learning to Localize and Reason via a CoT-Curriculum in Medical Vision-Language Models
von: Kim, Soo Yong, et al.
Veröffentlicht: (2025)
von: Kim, Soo Yong, et al.
Veröffentlicht: (2025)
Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT
von: Dong, Zhuobai, et al.
Veröffentlicht: (2025)
von: Dong, Zhuobai, et al.
Veröffentlicht: (2025)
AD^2-Bench: A Hierarchical CoT Benchmark for MLLM in Autonomous Driving under Adverse Conditions
von: Wei, Zhaoyang, et al.
Veröffentlicht: (2025)
von: Wei, Zhaoyang, et al.
Veröffentlicht: (2025)
MetaphorStar: Image Metaphor Understanding and Reasoning with End-to-End Visual Reinforcement Learning
von: Zhang, Chenhao, et al.
Veröffentlicht: (2026)
von: Zhang, Chenhao, et al.
Veröffentlicht: (2026)
Empowering Lightweight MLLMs with Reasoning via Long CoT SFT
von: Ou, Linyu, et al.
Veröffentlicht: (2025)
von: Ou, Linyu, et al.
Veröffentlicht: (2025)
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
Mitigating Visual Forgetting via Take-along Visual Conditioning for Multi-modal Long CoT Reasoning
von: Sun, Hai-Long, et al.
Veröffentlicht: (2025)
von: Sun, Hai-Long, et al.
Veröffentlicht: (2025)
EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
von: Pei, Baoqi, et al.
Veröffentlicht: (2025)
von: Pei, Baoqi, et al.
Veröffentlicht: (2025)
AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning
von: Li, Xiping, et al.
Veröffentlicht: (2025)
von: Li, Xiping, et al.
Veröffentlicht: (2025)
UML-CoT: Structured Reasoning and Planning with Unified Modeling Language for Robotic Room Cleaning
von: Chen, Hongyu, et al.
Veröffentlicht: (2025)
von: Chen, Hongyu, et al.
Veröffentlicht: (2025)
Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
von: Lee, Daeun, et al.
Veröffentlicht: (2025)
von: Lee, Daeun, et al.
Veröffentlicht: (2025)
CardioCoT: Hierarchical Reasoning for Multimodal Survival Analysis
von: Rui, Shaohao, et al.
Veröffentlicht: (2025)
von: Rui, Shaohao, et al.
Veröffentlicht: (2025)
Multi-Object Grounding via Hierarchical Contrastive Siamese Transformers
von: Du, Chengyi, et al.
Veröffentlicht: (2025)
von: Du, Chengyi, et al.
Veröffentlicht: (2025)
Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning
von: Li, Ang, et al.
Veröffentlicht: (2025)
von: Li, Ang, et al.
Veröffentlicht: (2025)
MC-CoT: A Modular Collaborative CoT Framework for Zero-shot Medical-VQA with LLM and MLLM Integration
von: Wei, Lai, et al.
Veröffentlicht: (2024)
von: Wei, Lai, et al.
Veröffentlicht: (2024)
Meta-CoT: Enhancing Granularity and Generalization in Image Editing
von: Zhang, Shiyi, et al.
Veröffentlicht: (2026)
von: Zhang, Shiyi, et al.
Veröffentlicht: (2026)
CSMCIR: CoT-Enhanced Symmetric Alignment with Memory Bank for Composed Image Retrieval
von: Qian, Zhipeng, et al.
Veröffentlicht: (2026)
von: Qian, Zhipeng, et al.
Veröffentlicht: (2026)
FakeVLM-R1: Internalizing Physical Laws via CoT for Synthetic Image Detection
von: Zhu, Leqi, et al.
Veröffentlicht: (2026)
von: Zhu, Leqi, et al.
Veröffentlicht: (2026)
X-Ray-CoT: Interpretable Chest X-ray Diagnosis with Vision-Language Models via Chain-of-Thought Reasoning
von: Ng, Chee, et al.
Veröffentlicht: (2025)
von: Ng, Chee, et al.
Veröffentlicht: (2025)
CoT-Drive: Efficient Motion Forecasting for Autonomous Driving with LLMs and Chain-of-Thought Prompting
von: Liao, Haicheng, et al.
Veröffentlicht: (2025)
von: Liao, Haicheng, et al.
Veröffentlicht: (2025)
Rethinking the Spatial Inconsistency in Classifier-Free Diffusion Guidance
von: Shen, Dazhong, et al.
Veröffentlicht: (2024)
von: Shen, Dazhong, et al.
Veröffentlicht: (2024)
Beyond Textual CoT: Interleaved Text-Image Chains with Deep Confidence Reasoning for Image Editing
von: Zou, Zhentao, et al.
Veröffentlicht: (2025)
von: Zou, Zhentao, et al.
Veröffentlicht: (2025)
NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation
von: Qiu, Longtian, et al.
Veröffentlicht: (2025)
von: Qiu, Longtian, et al.
Veröffentlicht: (2025)
MedVR: Annotation-Free Medical Visual Reasoning via Agentic Reinforcement Learning
von: Jiang, Zheng, et al.
Veröffentlicht: (2026)
von: Jiang, Zheng, et al.
Veröffentlicht: (2026)
C-CoT: Counterfactual Chain-of-Thought with Vision-Language Models for Safe Autonomous Driving
von: Tian, Kefei, et al.
Veröffentlicht: (2026)
von: Tian, Kefei, et al.
Veröffentlicht: (2026)
ESTR-CoT: Towards Explainable and Accurate Event Stream based Scene Text Recognition with Chain-of-Thought Reasoning
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
CoT-RVS: Zero-Shot Chain-of-Thought Reasoning Segmentation for Videos
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2025)
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2025)
CoT-Seg: Rethinking Segmentation with Chain-of-Thought Reasoning and Self-Correction
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2026)
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2026)
ViRC: Enhancing Visual Interleaved Mathematical CoT with Reason Chunking
von: Wang, Lihong, et al.
Veröffentlicht: (2025)
von: Wang, Lihong, et al.
Veröffentlicht: (2025)
DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
CoT4AD: A Vision-Language-Action Model with Explicit Chain-of-Thought Reasoning for Autonomous Driving
von: Wang, Zhaohui, et al.
Veröffentlicht: (2025)
von: Wang, Zhaohui, et al.
Veröffentlicht: (2025)
Aligning Vision to Language: Annotation-Free Multimodal Knowledge Graph Construction for Enhanced LLMs Reasoning
von: Liu, Junming, et al.
Veröffentlicht: (2025)
von: Liu, Junming, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Let Androids Dream of Electric Sheep: A Human-Inspired Image Implication Understanding and Reasoning Framework
von: Zhang, Chenhao, et al.
Veröffentlicht: (2025) -
Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT
von: Asfour, Alaa, et al.
Veröffentlicht: (2026) -
VG-CoT: Towards Trustworthy Visual Reasoning via Grounded Chain-of-Thought
von: Lim, Byeonggeuk, et al.
Veröffentlicht: (2026) -
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
von: Zhao, Qingqing, et al.
Veröffentlicht: (2025) -
Pretrained Reversible Generation as Unsupervised Visual Representation Learning
von: Xue, Rongkun, et al.
Veröffentlicht: (2024)