CoT-Seg: Rethinking Segmentation with Chain-of-Thought Reasoning and Self-Correction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kao, Shiu-hong, Huang, Chak Ho, Liu, Huaiqian, Tai, Yu-Wing, Tang, Chi-Keung |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CoT-RVS: Zero-Shot Chain-of-Thought Reasoning Segmentation for Videos
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2025)
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2025)
Think Before You Segment: High-Quality Reasoning Segmentation with GPT Chain of Thoughts
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2025)
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2025)
InceptionHuman: Controllable Prompt-to-NeRF for Photorealistic 3D Human Generation
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2023)
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2023)
Deceptive-NeRF/3DGS: Diffusion-Generated Pseudo-Observations for High-Quality Sparse-View Reconstruction
von: Liu, Xinhang, et al.
Veröffentlicht: (2023)
von: Liu, Xinhang, et al.
Veröffentlicht: (2023)
UVRM: A Scalable 3D Reconstruction Model from Unposed Videos
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2025)
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2025)
ReasonNavi: Human-Inspired Global Map Reasoning for Zero-Shot Embodied Navigation
von: Ao, Yuzhuo, et al.
Veröffentlicht: (2026)
von: Ao, Yuzhuo, et al.
Veröffentlicht: (2026)
SANeRF-HQ: Segment Anything for NeRF in High Quality
von: Liu, Yichen, et al.
Veröffentlicht: (2023)
von: Liu, Yichen, et al.
Veröffentlicht: (2023)
FED-NeRF: Achieve High 3D Consistency and Temporal Coherence for Face Video Editing on Dynamic NeRF
von: Zhang, Hao, et al.
Veröffentlicht: (2024)
von: Zhang, Hao, et al.
Veröffentlicht: (2024)
Multimodal Generation of Animatable 3D Human Models with AvatarForge
von: Liu, Xinhang, et al.
Veröffentlicht: (2025)
von: Liu, Xinhang, et al.
Veröffentlicht: (2025)
ChatCam: Empowering Camera Control through Conversational AI
von: Liu, Xinhang, et al.
Veröffentlicht: (2024)
von: Liu, Xinhang, et al.
Veröffentlicht: (2024)
CoT-Segmenter: Enhancing OOD Detection in Dense Road Scenes via Chain-of-Thought Reasoning
von: Song, Jeonghyo, et al.
Veröffentlicht: (2025)
von: Song, Jeonghyo, et al.
Veröffentlicht: (2025)
RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation
von: Wen, Junwei, et al.
Veröffentlicht: (2026)
von: Wen, Junwei, et al.
Veröffentlicht: (2026)
Agentic 3D Scene Generation with Spatially Contextualized VLMs
von: Liu, Xinhang, et al.
Veröffentlicht: (2025)
von: Liu, Xinhang, et al.
Veröffentlicht: (2025)
ReelWave: Multi-Agentic Movie Sound Generation through Multimodal LLM Conversation
von: Wang, Zixuan, et al.
Veröffentlicht: (2025)
von: Wang, Zixuan, et al.
Veröffentlicht: (2025)
WorldCraft: Photo-Realistic 3D World Creation and Customization via LLM Agents
von: Liu, Xinhang, et al.
Veröffentlicht: (2025)
von: Liu, Xinhang, et al.
Veröffentlicht: (2025)
ImageGen-CoT: Enhancing Text-to-Image In-context Learning with Chain-of-Thought Reasoning
von: Liao, Jiaqi, et al.
Veröffentlicht: (2025)
von: Liao, Jiaqi, et al.
Veröffentlicht: (2025)
MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
von: Chen, Xinyan, et al.
Veröffentlicht: (2025)
von: Chen, Xinyan, et al.
Veröffentlicht: (2025)
Audio-Agent: Leveraging LLMs For Audio Generation, Editing and Composition
von: Wang, Zixuan, et al.
Veröffentlicht: (2024)
von: Wang, Zixuan, et al.
Veröffentlicht: (2024)
Navigating Motion Agents in Dynamic and Cluttered Environments through LLM Reasoning
von: Zhao, Yubo, et al.
Veröffentlicht: (2025)
von: Zhao, Yubo, et al.
Veröffentlicht: (2025)
AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning
von: Li, Xiping, et al.
Veröffentlicht: (2025)
von: Li, Xiping, et al.
Veröffentlicht: (2025)
VG-CoT: Towards Trustworthy Visual Reasoning via Grounded Chain-of-Thought
von: Lim, Byeonggeuk, et al.
Veröffentlicht: (2026)
von: Lim, Byeonggeuk, et al.
Veröffentlicht: (2026)
X-Ray-CoT: Interpretable Chest X-ray Diagnosis with Vision-Language Models via Chain-of-Thought Reasoning
von: Ng, Chee, et al.
Veröffentlicht: (2025)
von: Ng, Chee, et al.
Veröffentlicht: (2025)
Beyond and Free from Diffusion: Invertible Guided Consistency Training
von: Hsu, Chia-Hong, et al.
Veröffentlicht: (2025)
von: Hsu, Chia-Hong, et al.
Veröffentlicht: (2025)
StableKD: Breaking Inter-block Optimization Entanglement for Stable Knowledge Distillation
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2023)
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2023)
X-CoT: Explainable Text-to-Video Retrieval via LLM-based Chain-of-Thought Reasoning
von: Pulakurthi, Prasanna Reddy, et al.
Veröffentlicht: (2025)
von: Pulakurthi, Prasanna Reddy, et al.
Veröffentlicht: (2025)
CoT-Pose: Chain-of-Thought Reasoning for 3D Pose Generation from Abstract Prompts
von: Cha, Junuk, et al.
Veröffentlicht: (2025)
von: Cha, Junuk, et al.
Veröffentlicht: (2025)
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
von: Zhao, Qingqing, et al.
Veröffentlicht: (2025)
von: Zhao, Qingqing, et al.
Veröffentlicht: (2025)
Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
von: Shao, Hao, et al.
Veröffentlicht: (2024)
von: Shao, Hao, et al.
Veröffentlicht: (2024)
CoT4AD: A Vision-Language-Action Model with Explicit Chain-of-Thought Reasoning for Autonomous Driving
von: Wang, Zhaohui, et al.
Veröffentlicht: (2025)
von: Wang, Zhaohui, et al.
Veröffentlicht: (2025)
Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought
von: Zhang, Shuyi, et al.
Veröffentlicht: (2025)
von: Zhang, Shuyi, et al.
Veröffentlicht: (2025)
CIR-CoT: Towards Interpretable Composed Image Retrieval via End-to-End Chain-of-Thought Reasoning
von: Lin, Weihuang, et al.
Veröffentlicht: (2025)
von: Lin, Weihuang, et al.
Veröffentlicht: (2025)
Inpaint4DNeRF: Promptable Spatio-Temporal NeRF Inpainting with Generative Diffusion Models
von: Jiang, Han, et al.
Veröffentlicht: (2023)
von: Jiang, Han, et al.
Veröffentlicht: (2023)
Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
von: Lee, Daeun, et al.
Veröffentlicht: (2025)
von: Lee, Daeun, et al.
Veröffentlicht: (2025)
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
CoT-PL: Chain-of-Thought Pseudo-Labeling for Open-Vocabulary Object Detection
von: Choi, Hojun, et al.
Veröffentlicht: (2025)
von: Choi, Hojun, et al.
Veröffentlicht: (2025)
DragVideo: Interactive Drag-style Video Editing
von: Deng, Yufan, et al.
Veröffentlicht: (2023)
von: Deng, Yufan, et al.
Veröffentlicht: (2023)
ESTR-CoT: Towards Explainable and Accurate Event Stream based Scene Text Recognition with Chain-of-Thought Reasoning
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks
von: Qi, Yu, et al.
Veröffentlicht: (2025)
von: Qi, Yu, et al.
Veröffentlicht: (2025)
SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents
von: Huang-Menders, Alexander, et al.
Veröffentlicht: (2025)
von: Huang-Menders, Alexander, et al.
Veröffentlicht: (2025)
CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding
von: Abdelrahman, Eslam, et al.
Veröffentlicht: (2023)
von: Abdelrahman, Eslam, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
CoT-RVS: Zero-Shot Chain-of-Thought Reasoning Segmentation for Videos
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2025) -
Think Before You Segment: High-Quality Reasoning Segmentation with GPT Chain of Thoughts
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2025) -
InceptionHuman: Controllable Prompt-to-NeRF for Photorealistic 3D Human Generation
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2023) -
Deceptive-NeRF/3DGS: Diffusion-Generated Pseudo-Observations for High-Quality Sparse-View Reconstruction
von: Liu, Xinhang, et al.
Veröffentlicht: (2023) -
UVRM: A Scalable 3D Reconstruction Model from Unposed Videos
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2025)