CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Abdelrahman, Eslam, Ayman, Mohamed, Ahmed, Mahmoud, Slim, Habib, Elhoseiny, Mohamed |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Kestrel: 3D Multimodal LLM for Part-Aware Grounded Description
von: Ahmed, Mahmoud, et al.
Veröffentlicht: (2024)
von: Ahmed, Mahmoud, et al.
Veröffentlicht: (2024)
3DCoMPaT200: Language-Grounded Compositional Understanding of Parts and Materials of 3D Shapes
von: Ahmed, Mahmoud, et al.
Veröffentlicht: (2025)
von: Ahmed, Mahmoud, et al.
Veröffentlicht: (2025)
3DCoMPaT$^{++}$: An improved Large-scale 3D Vision Dataset for Compositional Recognition
von: Slim, Habib, et al.
Veröffentlicht: (2023)
von: Slim, Habib, et al.
Veröffentlicht: (2023)
InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Shows
von: Ataallah, Kirolos, et al.
Veröffentlicht: (2024)
von: Ataallah, Kirolos, et al.
Veröffentlicht: (2024)
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
von: Ataallah, Kirolos, et al.
Veröffentlicht: (2024)
von: Ataallah, Kirolos, et al.
Veröffentlicht: (2024)
VG-CoT: Towards Trustworthy Visual Reasoning via Grounded Chain-of-Thought
von: Lim, Byeonggeuk, et al.
Veröffentlicht: (2026)
von: Lim, Byeonggeuk, et al.
Veröffentlicht: (2026)
iMotion-LLM: Instruction-Conditioned Trajectory Generation
von: Felemban, Abdulwahab, et al.
Veröffentlicht: (2024)
von: Felemban, Abdulwahab, et al.
Veröffentlicht: (2024)
ToddlerDiffusion: Interactive Structured Image Generation with Cascaded Schrödinger Bridge
von: Abdelrahman, Eslam, et al.
Veröffentlicht: (2023)
von: Abdelrahman, Eslam, et al.
Veröffentlicht: (2023)
CoT-Pose: Chain-of-Thought Reasoning for 3D Pose Generation from Abstract Prompts
von: Cha, Junuk, et al.
Veröffentlicht: (2025)
von: Cha, Junuk, et al.
Veröffentlicht: (2025)
MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
von: Chen, Xinyan, et al.
Veröffentlicht: (2025)
von: Chen, Xinyan, et al.
Veröffentlicht: (2025)
StoryGPT-V: Large Language Models as Consistent Story Visualizers
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2023)
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2023)
CoT-RVS: Zero-Shot Chain-of-Thought Reasoning Segmentation for Videos
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2025)
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2025)
CoT-Seg: Rethinking Segmentation with Chain-of-Thought Reasoning and Self-Correction
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2026)
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2026)
3DRef: 3D Dataset and Benchmark for Reflection Detection in RGB and Lidar Data
von: Zhao, Xiting, et al.
Veröffentlicht: (2024)
von: Zhao, Xiting, et al.
Veröffentlicht: (2024)
CoT-PL: Chain-of-Thought Pseudo-Labeling for Open-Vocabulary Object Detection
von: Choi, Hojun, et al.
Veröffentlicht: (2025)
von: Choi, Hojun, et al.
Veröffentlicht: (2025)
Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
von: Shao, Hao, et al.
Veröffentlicht: (2024)
von: Shao, Hao, et al.
Veröffentlicht: (2024)
Pathology-CoT: Learning Visual Chain-of-Thought Agent from Expert Whole Slide Image Diagnosis Behavior
von: Wang, Sheng, et al.
Veröffentlicht: (2025)
von: Wang, Sheng, et al.
Veröffentlicht: (2025)
Goldfish: Vision-Language Understanding of Arbitrarily Long Videos
von: Ataallah, Kirolos, et al.
Veröffentlicht: (2024)
von: Ataallah, Kirolos, et al.
Veröffentlicht: (2024)
ImageGen-CoT: Enhancing Text-to-Image In-context Learning with Chain-of-Thought Reasoning
von: Liao, Jiaqi, et al.
Veröffentlicht: (2025)
von: Liao, Jiaqi, et al.
Veröffentlicht: (2025)
Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought
von: Zhang, Shuyi, et al.
Veröffentlicht: (2025)
von: Zhang, Shuyi, et al.
Veröffentlicht: (2025)
AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning
von: Li, Xiping, et al.
Veröffentlicht: (2025)
von: Li, Xiping, et al.
Veröffentlicht: (2025)
CoT-Drive: Efficient Motion Forecasting for Autonomous Driving with LLMs and Chain-of-Thought Prompting
von: Liao, Haicheng, et al.
Veröffentlicht: (2025)
von: Liao, Haicheng, et al.
Veröffentlicht: (2025)
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
von: Zhao, Qingqing, et al.
Veröffentlicht: (2025)
von: Zhao, Qingqing, et al.
Veröffentlicht: (2025)
CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks
von: Qi, Yu, et al.
Veröffentlicht: (2025)
von: Qi, Yu, et al.
Veröffentlicht: (2025)
X-CoT: Explainable Text-to-Video Retrieval via LLM-based Chain-of-Thought Reasoning
von: Pulakurthi, Prasanna Reddy, et al.
Veröffentlicht: (2025)
von: Pulakurthi, Prasanna Reddy, et al.
Veröffentlicht: (2025)
CoT-Segmenter: Enhancing OOD Detection in Dense Road Scenes via Chain-of-Thought Reasoning
von: Song, Jeonghyo, et al.
Veröffentlicht: (2025)
von: Song, Jeonghyo, et al.
Veröffentlicht: (2025)
C-CoT: Counterfactual Chain-of-Thought with Vision-Language Models for Safe Autonomous Driving
von: Tian, Kefei, et al.
Veröffentlicht: (2026)
von: Tian, Kefei, et al.
Veröffentlicht: (2026)
Efficient Self-supervised Vision Pretraining with Local Masked Reconstruction
von: Chen, Jun, et al.
Veröffentlicht: (2022)
von: Chen, Jun, et al.
Veröffentlicht: (2022)
CIR-CoT: Towards Interpretable Composed Image Retrieval via End-to-End Chain-of-Thought Reasoning
von: Lin, Weihuang, et al.
Veröffentlicht: (2025)
von: Lin, Weihuang, et al.
Veröffentlicht: (2025)
A Dynamic Programming Framework for Discovering Count and Values of Multilevel Image Thresholding
von: Hegazy, Eslam, et al.
Veröffentlicht: (2026)
von: Hegazy, Eslam, et al.
Veröffentlicht: (2026)
MMIS: Multimodal Dataset for Interior Scene Visual Generation and Recognition
von: Kassab, Hozaifa, et al.
Veröffentlicht: (2024)
von: Kassab, Hozaifa, et al.
Veröffentlicht: (2024)
INSTA-YOLO: Real-Time Instance Segmentation
von: Mohamed, Eslam, et al.
Veröffentlicht: (2021)
von: Mohamed, Eslam, et al.
Veröffentlicht: (2021)
M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
von: Chen, Qiguang, et al.
Veröffentlicht: (2024)
von: Chen, Qiguang, et al.
Veröffentlicht: (2024)
Bootstrapping Grounded Chain-of-Thought in Multimodal LLMs for Data-Efficient Model Adaptation
von: Xia, Jiaer, et al.
Veröffentlicht: (2025)
von: Xia, Jiaer, et al.
Veröffentlicht: (2025)
VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
CoRGI: Verified Chain-of-Thought Reasoning with Post-hoc Visual Grounding
von: Yi, Shixin, et al.
Veröffentlicht: (2025)
von: Yi, Shixin, et al.
Veröffentlicht: (2025)
FishNet++: Analyzing the capabilities of Multimodal Large Language Models in marine biology
von: Khan, Faizan Farooq, et al.
Veröffentlicht: (2025)
von: Khan, Faizan Farooq, et al.
Veröffentlicht: (2025)
EfficientVITON: An Efficient Virtual Try-On Model using Optimized Diffusion Process
von: Atef, Mostafa, et al.
Veröffentlicht: (2025)
von: Atef, Mostafa, et al.
Veröffentlicht: (2025)
Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
von: Lee, Daeun, et al.
Veröffentlicht: (2025)
von: Lee, Daeun, et al.
Veröffentlicht: (2025)
X-Ray-CoT: Interpretable Chest X-ray Diagnosis with Vision-Language Models via Chain-of-Thought Reasoning
von: Ng, Chee, et al.
Veröffentlicht: (2025)
von: Ng, Chee, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Kestrel: 3D Multimodal LLM for Part-Aware Grounded Description
von: Ahmed, Mahmoud, et al.
Veröffentlicht: (2024) -
3DCoMPaT200: Language-Grounded Compositional Understanding of Parts and Materials of 3D Shapes
von: Ahmed, Mahmoud, et al.
Veröffentlicht: (2025) -
3DCoMPaT$^{++}$: An improved Large-scale 3D Vision Dataset for Compositional Recognition
von: Slim, Habib, et al.
Veröffentlicht: (2023) -
InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Shows
von: Ataallah, Kirolos, et al.
Veröffentlicht: (2024) -
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
von: Ataallah, Kirolos, et al.
Veröffentlicht: (2024)