TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes
Fuente:
arXiv
Saved in:
| Main Authors: | Jin, Bu, Zheng, Yupeng, Li, Pengfei, Li, Weize, Zheng, Yuhang, Hu, Sujie, Liu, Xinyu, Zhu, Jinwei, Yan, Zhijie, Sun, Haiyang, Zhan, Kun, Jia, Peng, Long, Xiaoxiao, Chen, Yilun, Zhao, Hao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UniArt: Unified 3D Representation for Generating 3D Articulated Objects with Open-Set Articulation
by: Jin, Bu, et al.
Published: (2025)
by: Jin, Bu, et al.
Published: (2025)
VistaBot: View-Robust Robot Manipulation via Spatiotemporal-Aware View Synthesis
by: Gu, Songen, et al.
Published: (2026)
by: Gu, Songen, et al.
Published: (2026)
CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning
by: Xing, Long, et al.
Published: (2025)
by: Xing, Long, et al.
Published: (2025)
MonoOcc: Digging into Monocular Semantic Occupancy Prediction
by: Zheng, Yupeng, et al.
Published: (2024)
by: Zheng, Yupeng, et al.
Published: (2024)
Semi-Supervised Vision-Centric 3D Occupancy World Model for Autonomous Driving
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
ChartCap: Mitigating Hallucination of Dense Chart Captioning
by: Lim, Junyoung, et al.
Published: (2025)
by: Lim, Junyoung, et al.
Published: (2025)
PlanAgent: A Multi-modal Large Language Agent for Closed-loop Vehicle Motion Planning
by: Zheng, Yupeng, et al.
Published: (2024)
by: Zheng, Yupeng, et al.
Published: (2024)
GaussianGrasper: 3D Language Gaussian Splatting for Open-vocabulary Robotic Grasping
by: Zheng, Yuhang, et al.
Published: (2024)
by: Zheng, Yuhang, et al.
Published: (2024)
Controllable 3D Outdoor Scene Generation via Scene Graphs
by: Liu, Yuheng, et al.
Published: (2025)
by: Liu, Yuheng, et al.
Published: (2025)
OccTENS: 3D Occupancy World Model via Temporal Next-Scale Prediction
by: Jin, Bu, et al.
Published: (2025)
by: Jin, Bu, et al.
Published: (2025)
RubiCap: Rubric-Guided Reinforcement Learning for Dense Image Captioning
by: Huang, Tzu-Heng, et al.
Published: (2026)
by: Huang, Tzu-Heng, et al.
Published: (2026)
CapGeo: A Caption-Assisted Approach to Geometric Reasoning
by: Li, Yuying, et al.
Published: (2025)
by: Li, Yuying, et al.
Published: (2025)
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing
by: Xing, Long, et al.
Published: (2025)
by: Xing, Long, et al.
Published: (2025)
PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance
by: Zheng, Yupeng, et al.
Published: (2026)
by: Zheng, Yupeng, et al.
Published: (2026)
ExCap3D: Expressive 3D Scene Understanding via Object Captioning with Varying Detail
by: Yeshwanth, Chandan, et al.
Published: (2025)
by: Yeshwanth, Chandan, et al.
Published: (2025)
Enhancing Indoor Occupancy Prediction via Sparse Query-Based Multi-Level Consistent Knowledge Distillation
by: Li, Xiang, et al.
Published: (2026)
by: Li, Xiang, et al.
Published: (2026)
DenseAnnotate: Enabling Scalable Dense Caption Collection for Images and 3D Scenes via Spoken Descriptions
by: Lin, Xiaoyu, et al.
Published: (2025)
by: Lin, Xiaoyu, et al.
Published: (2025)
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning
by: Li, Yizhe, et al.
Published: (2025)
by: Li, Yizhe, et al.
Published: (2025)
LiloDriver: A Lifelong Learning Framework for Closed-loop Motion Planning in Long-tail Autonomous Driving Scenarios
by: Yao, Huaiyuan, et al.
Published: (2025)
by: Yao, Huaiyuan, et al.
Published: (2025)
A Comprehensive Survey of 3D Dense Captioning: Localizing and Describing Objects in 3D Scenes
by: Yu, Ting, et al.
Published: (2024)
by: Yu, Ting, et al.
Published: (2024)
GaRe: Relightable 3D Gaussian Splatting for Outdoor Scenes from Unconstrained Photo Collections
by: Bai, Haiyang, et al.
Published: (2025)
by: Bai, Haiyang, et al.
Published: (2025)
CodecCap: High-Fidelity Codec-Inspired Residual Modeling for Dense Video Captioning
by: Lin, Zihan, et al.
Published: (2026)
by: Lin, Zihan, et al.
Published: (2026)
DualCap: Enhancing Lightweight Image Captioning via Dual Retrieval with Similar Scenes Visual Prompts
by: Li, Binbin, et al.
Published: (2025)
by: Li, Binbin, et al.
Published: (2025)
OmniVTA: Visuo-Tactile World Modeling for Contact-Rich Robotic Manipulation
by: Zheng, Yuhang, et al.
Published: (2026)
by: Zheng, Yuhang, et al.
Published: (2026)
TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions
by: Yao, Linli, et al.
Published: (2026)
by: Yao, Linli, et al.
Published: (2026)
SceneCompleter: Dense 3D Scene Completion for Generative Novel View Synthesis
by: Chen, Weiliang, et al.
Published: (2025)
by: Chen, Weiliang, et al.
Published: (2025)
Technical Report for Soccernet 2023 -- Dense Video Captioning
by: Ruan, Zheng, et al.
Published: (2024)
by: Ruan, Zheng, et al.
Published: (2024)
FingerCap: Fine-grained Finger-level Hand Motion Captioning
by: Shen, Xin, et al.
Published: (2025)
by: Shen, Xin, et al.
Published: (2025)
DecoRec: Decomposed 3D Scene Reconstruction from Single-View Images via Object-Level Diffusion
by: Ping, Yuhan, et al.
Published: (2026)
by: Ping, Yuhan, et al.
Published: (2026)
Stereo 3D Gaussian Splatting SLAM for Outdoor Urban Scenes
by: Li, Xiaohan, et al.
Published: (2025)
by: Li, Xiaohan, et al.
Published: (2025)
Adaptive Surface Normal Constraint for Geometric Estimation from Monocular Images
by: Long, Xiaoxiao, et al.
Published: (2024)
by: Long, Xiaoxiao, et al.
Published: (2024)
AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
by: Chai, Wenhao, et al.
Published: (2024)
by: Chai, Wenhao, et al.
Published: (2024)
PaveCap: The First Multimodal Framework for Comprehensive Pavement Condition Assessment with Dense Captioning and PCI Estimation
by: Kyem, Blessing Agyei, et al.
Published: (2024)
by: Kyem, Blessing Agyei, et al.
Published: (2024)
You Only Gaussian Once: Controllable 3D Gaussian Splatting for Ultra-Densely Sampled Scenes
by: Jia, Jinrang, et al.
Published: (2026)
by: Jia, Jinrang, et al.
Published: (2026)
The coherence peak of unconventional superconductors in the charge channel
by: Li, Pengfei, et al.
Published: (2025)
by: Li, Pengfei, et al.
Published: (2025)
CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era
by: Cheng, Kanzhi, et al.
Published: (2025)
by: Cheng, Kanzhi, et al.
Published: (2025)
EmoTransCap: Dataset and Pipeline for Emotion Transition-Aware Speech Captioning in Discourses
by: Xu, Shuhao, et al.
Published: (2026)
by: Xu, Shuhao, et al.
Published: (2026)
LVD-2M: A Long-take Video Dataset with Temporally Dense Captions
by: Xiong, Tianwei, et al.
Published: (2024)
by: Xiong, Tianwei, et al.
Published: (2024)
XMeCap: Meme Caption Generation with Sub-Image Adaptability
by: Chen, Yuyan, et al.
Published: (2024)
by: Chen, Yuyan, et al.
Published: (2024)
PosePilot: Steering Camera Pose for Generative World Models with Self-supervised Depth
by: Jin, Bu, et al.
Published: (2025)
by: Jin, Bu, et al.
Published: (2025)
Similar Items
-
UniArt: Unified 3D Representation for Generating 3D Articulated Objects with Open-Set Articulation
by: Jin, Bu, et al.
Published: (2025) -
VistaBot: View-Robust Robot Manipulation via Spatiotemporal-Aware View Synthesis
by: Gu, Songen, et al.
Published: (2026) -
CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning
by: Xing, Long, et al.
Published: (2025) -
MonoOcc: Digging into Monocular Semantic Occupancy Prediction
by: Zheng, Yupeng, et al.
Published: (2024) -
Semi-Supervised Vision-Centric 3D Occupancy World Model for Autonomous Driving
by: Li, Xiang, et al.
Published: (2025)