CoT-RVS: Zero-Shot Chain-of-Thought Reasoning Segmentation for Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Kao, Shiu-hong, Tai, Yu-Wing, Tang, Chi-Keung |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CoT-Seg: Rethinking Segmentation with Chain-of-Thought Reasoning and Self-Correction
by: Kao, Shiu-hong, et al.
Published: (2026)
by: Kao, Shiu-hong, et al.
Published: (2026)
Think Before You Segment: High-Quality Reasoning Segmentation with GPT Chain of Thoughts
by: Kao, Shiu-hong, et al.
Published: (2025)
by: Kao, Shiu-hong, et al.
Published: (2025)
InceptionHuman: Controllable Prompt-to-NeRF for Photorealistic 3D Human Generation
by: Kao, Shiu-hong, et al.
Published: (2023)
by: Kao, Shiu-hong, et al.
Published: (2023)
Deceptive-NeRF/3DGS: Diffusion-Generated Pseudo-Observations for High-Quality Sparse-View Reconstruction
by: Liu, Xinhang, et al.
Published: (2023)
by: Liu, Xinhang, et al.
Published: (2023)
UVRM: A Scalable 3D Reconstruction Model from Unposed Videos
by: Kao, Shiu-hong, et al.
Published: (2025)
by: Kao, Shiu-hong, et al.
Published: (2025)
ReasonNavi: Human-Inspired Global Map Reasoning for Zero-Shot Embodied Navigation
by: Ao, Yuzhuo, et al.
Published: (2026)
by: Ao, Yuzhuo, et al.
Published: (2026)
FED-NeRF: Achieve High 3D Consistency and Temporal Coherence for Face Video Editing on Dynamic NeRF
by: Zhang, Hao, et al.
Published: (2024)
by: Zhang, Hao, et al.
Published: (2024)
SANeRF-HQ: Segment Anything for NeRF in High Quality
by: Liu, Yichen, et al.
Published: (2023)
by: Liu, Yichen, et al.
Published: (2023)
Multimodal Generation of Animatable 3D Human Models with AvatarForge
by: Liu, Xinhang, et al.
Published: (2025)
by: Liu, Xinhang, et al.
Published: (2025)
ChatCam: Empowering Camera Control through Conversational AI
by: Liu, Xinhang, et al.
Published: (2024)
by: Liu, Xinhang, et al.
Published: (2024)
Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought
by: Zhang, Shuyi, et al.
Published: (2025)
by: Zhang, Shuyi, et al.
Published: (2025)
CoT-Segmenter: Enhancing OOD Detection in Dense Road Scenes via Chain-of-Thought Reasoning
by: Song, Jeonghyo, et al.
Published: (2025)
by: Song, Jeonghyo, et al.
Published: (2025)
X-CoT: Explainable Text-to-Video Retrieval via LLM-based Chain-of-Thought Reasoning
by: Pulakurthi, Prasanna Reddy, et al.
Published: (2025)
by: Pulakurthi, Prasanna Reddy, et al.
Published: (2025)
Agentic 3D Scene Generation with Spatially Contextualized VLMs
by: Liu, Xinhang, et al.
Published: (2025)
by: Liu, Xinhang, et al.
Published: (2025)
ReelWave: Multi-Agentic Movie Sound Generation through Multimodal LLM Conversation
by: Wang, Zixuan, et al.
Published: (2025)
by: Wang, Zixuan, et al.
Published: (2025)
DragVideo: Interactive Drag-style Video Editing
by: Deng, Yufan, et al.
Published: (2023)
by: Deng, Yufan, et al.
Published: (2023)
Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
by: Lee, Daeun, et al.
Published: (2025)
by: Lee, Daeun, et al.
Published: (2025)
ImageGen-CoT: Enhancing Text-to-Image In-context Learning with Chain-of-Thought Reasoning
by: Liao, Jiaqi, et al.
Published: (2025)
by: Liao, Jiaqi, et al.
Published: (2025)
MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
by: Chen, Xinyan, et al.
Published: (2025)
by: Chen, Xinyan, et al.
Published: (2025)
WorldCraft: Photo-Realistic 3D World Creation and Customization via LLM Agents
by: Liu, Xinhang, et al.
Published: (2025)
by: Liu, Xinhang, et al.
Published: (2025)
AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning
by: Li, Xiping, et al.
Published: (2025)
by: Li, Xiping, et al.
Published: (2025)
X-Ray-CoT: Interpretable Chest X-ray Diagnosis with Vision-Language Models via Chain-of-Thought Reasoning
by: Ng, Chee, et al.
Published: (2025)
by: Ng, Chee, et al.
Published: (2025)
Beyond and Free from Diffusion: Invertible Guided Consistency Training
by: Hsu, Chia-Hong, et al.
Published: (2025)
by: Hsu, Chia-Hong, et al.
Published: (2025)
StableKD: Breaking Inter-block Optimization Entanglement for Stable Knowledge Distillation
by: Kao, Shiu-hong, et al.
Published: (2023)
by: Kao, Shiu-hong, et al.
Published: (2023)
CoT-Pose: Chain-of-Thought Reasoning for 3D Pose Generation from Abstract Prompts
by: Cha, Junuk, et al.
Published: (2025)
by: Cha, Junuk, et al.
Published: (2025)
VG-CoT: Towards Trustworthy Visual Reasoning via Grounded Chain-of-Thought
by: Lim, Byeonggeuk, et al.
Published: (2026)
by: Lim, Byeonggeuk, et al.
Published: (2026)
Navigating Motion Agents in Dynamic and Cluttered Environments through LLM Reasoning
by: Zhao, Yubo, et al.
Published: (2025)
by: Zhao, Yubo, et al.
Published: (2025)
Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
by: Shao, Hao, et al.
Published: (2024)
by: Shao, Hao, et al.
Published: (2024)
CIR-CoT: Towards Interpretable Composed Image Retrieval via End-to-End Chain-of-Thought Reasoning
by: Lin, Weihuang, et al.
Published: (2025)
by: Lin, Weihuang, et al.
Published: (2025)
Inpaint4DNeRF: Promptable Spatio-Temporal NeRF Inpainting with Generative Diffusion Models
by: Jiang, Han, et al.
Published: (2023)
by: Jiang, Han, et al.
Published: (2023)
Audio-Agent: Leveraging LLMs For Audio Generation, Editing and Composition
by: Wang, Zixuan, et al.
Published: (2024)
by: Wang, Zixuan, et al.
Published: (2024)
CoT-PL: Chain-of-Thought Pseudo-Labeling for Open-Vocabulary Object Detection
by: Choi, Hojun, et al.
Published: (2025)
by: Choi, Hojun, et al.
Published: (2025)
RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation
by: Wen, Junwei, et al.
Published: (2026)
by: Wen, Junwei, et al.
Published: (2026)
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
by: Jiang, Dongzhi, et al.
Published: (2025)
by: Jiang, Dongzhi, et al.
Published: (2025)
Trace Anything: Representing Any Video in 4D via Trajectory Fields
by: Liu, Xinhang, et al.
Published: (2025)
by: Liu, Xinhang, et al.
Published: (2025)
CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks
by: Qi, Yu, et al.
Published: (2025)
by: Qi, Yu, et al.
Published: (2025)
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
by: Zhao, Qingqing, et al.
Published: (2025)
by: Zhao, Qingqing, et al.
Published: (2025)
ESTR-CoT: Towards Explainable and Accurate Event Stream based Scene Text Recognition with Chain-of-Thought Reasoning
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding
by: Abdelrahman, Eslam, et al.
Published: (2023)
by: Abdelrahman, Eslam, et al.
Published: (2023)
Reason-before-Retrieve: One-Stage Reflective Chain-of-Thoughts for Training-Free Zero-Shot Composed Image Retrieval
by: Tang, Yuanmin, et al.
Published: (2024)
by: Tang, Yuanmin, et al.
Published: (2024)
Similar Items
-
CoT-Seg: Rethinking Segmentation with Chain-of-Thought Reasoning and Self-Correction
by: Kao, Shiu-hong, et al.
Published: (2026) -
Think Before You Segment: High-Quality Reasoning Segmentation with GPT Chain of Thoughts
by: Kao, Shiu-hong, et al.
Published: (2025) -
InceptionHuman: Controllable Prompt-to-NeRF for Photorealistic 3D Human Generation
by: Kao, Shiu-hong, et al.
Published: (2023) -
Deceptive-NeRF/3DGS: Diffusion-Generated Pseudo-Observations for High-Quality Sparse-View Reconstruction
by: Liu, Xinhang, et al.
Published: (2023) -
UVRM: A Scalable 3D Reconstruction Model from Unposed Videos
by: Kao, Shiu-hong, et al.
Published: (2025)