Leveraging VLM-Based Pipelines to Annotate 3D Objects
Fuente:
arXiv
Saved in:
| Main Authors: | Kabra, Rishabh, Matthey, Loic, Lerchner, Alexander, Mitra, Niloy J. |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How to Spin an Object: First, Get the Shape Right
by: Kabra, Rishabh, et al.
Published: (2024)
by: Kabra, Rishabh, et al.
Published: (2024)
Animal Avatars: Reconstructing Animatable 3D Animals from Casual Videos
by: Sabathier, Remy, et al.
Published: (2024)
by: Sabathier, Remy, et al.
Published: (2024)
ActionMesh: Animated 3D Mesh Generation with Temporal 3D Diffusion
by: Sabathier, Remy, et al.
Published: (2026)
by: Sabathier, Remy, et al.
Published: (2026)
CHOIR: Contact-aware 4D Hand-Object Interaction Reconstruction
by: Xu, Hao, et al.
Published: (2026)
by: Xu, Hao, et al.
Published: (2026)
A Mixed Diet Makes DINO An Omnivorous Vision Encoder
by: Kabra, Rishabh, et al.
Published: (2026)
by: Kabra, Rishabh, et al.
Published: (2026)
Diffusion 3D Features (Diff3F): Decorating Untextured Shapes with Distilled Semantic Features
by: Dutt, Niladri Shekhar, et al.
Published: (2023)
by: Dutt, Niladri Shekhar, et al.
Published: (2023)
ProteusNeRF: Fast Lightweight NeRF Editing using 3D-Aware Image Context
by: Wang, Binglun, et al.
Published: (2023)
by: Wang, Binglun, et al.
Published: (2023)
SMILe: Leveraging Submodular Mutual Information For Robust Few-Shot Object Detection
by: Majee, Anay, et al.
Published: (2024)
by: Majee, Anay, et al.
Published: (2024)
SAGE: Structure-Aware Generative Video Transitions between Diverse Clips
by: Kan, Mia, et al.
Published: (2025)
by: Kan, Mia, et al.
Published: (2025)
3D Annotation Of Arbitrary Objects In The Wild
by: Blomqvist, Kenneth, et al.
Published: (2021)
by: Blomqvist, Kenneth, et al.
Published: (2021)
GANFusion: Feed-Forward Text-to-3D with Diffusion in GAN Space
by: Attaiki, Souhaib, et al.
Published: (2024)
by: Attaiki, Souhaib, et al.
Published: (2024)
GOEmbed: Gradient Origin Embeddings for Representation Agnostic 3D Feature Learning
by: Karnewar, Animesh, et al.
Published: (2023)
by: Karnewar, Animesh, et al.
Published: (2023)
ART3mis: Ray-Based Textual Annotation on 3D Cultural Objects
by: Arampatzakis, Vasileios, et al.
Published: (2026)
by: Arampatzakis, Vasileios, et al.
Published: (2026)
Beyond the Visible: Disocclusion-Aware Editing via Proxy Dynamic Graphs
by: Qi, Anran, et al.
Published: (2025)
by: Qi, Anran, et al.
Published: (2025)
SMF: Template-free and Rig-free Animation Transfer using Kinetic Codes
by: Muralikrishnan, Sanjeev, et al.
Published: (2025)
by: Muralikrishnan, Sanjeev, et al.
Published: (2025)
Leveraging Multi-Rater Annotations to Calibrate Object Detectors in Microscopy Imaging
by: Campi, Francesco, et al.
Published: (2026)
by: Campi, Francesco, et al.
Published: (2026)
LIM: Large Interpolator Model for Dynamic Reconstruction
by: Sabathier, Remy, et al.
Published: (2025)
by: Sabathier, Remy, et al.
Published: (2025)
Leveraging 2D-VLM for Label-Free 3D Segmentation in Large-Scale Outdoor Scene Understanding
by: Nishimura, Toshihiko, et al.
Published: (2026)
by: Nishimura, Toshihiko, et al.
Published: (2026)
Betsu-Betsu: Multi-View Separable 3D Reconstruction of Two Interacting Objects
by: Gopal, Suhas, et al.
Published: (2025)
by: Gopal, Suhas, et al.
Published: (2025)
Self-Improving VLM Judges Without Human Annotations
by: Lin, Inna Wanyin, et al.
Published: (2025)
by: Lin, Inna Wanyin, et al.
Published: (2025)
JOG3R: Towards 3D-Consistent Video Generators
by: Huang, Chun-Hao Paul, et al.
Published: (2025)
by: Huang, Chun-Hao Paul, et al.
Published: (2025)
STONE: A Submodular Optimization Framework for Active 3D Object Detection
by: Mao, Ruiyu, et al.
Published: (2024)
by: Mao, Ruiyu, et al.
Published: (2024)
OpenWorldSAM: Extending SAM2 for Universal Image Segmentation with Language Prompts
by: Xiao, Shiting, et al.
Published: (2025)
by: Xiao, Shiting, et al.
Published: (2025)
FlairGPT: Repurposing LLMs for Interior Designs
by: Littlefair, Gabrielle, et al.
Published: (2025)
by: Littlefair, Gabrielle, et al.
Published: (2025)
Leveraging Automatic CAD Annotations for Supervised Learning in 3D Scene Understanding
by: Rao, Yuchen, et al.
Published: (2025)
by: Rao, Yuchen, et al.
Published: (2025)
Multi-Stage VLM Pipeline for Zero-Shot Traffic Accident Understanding
by: Tatematsu, Fumiya, et al.
Published: (2026)
by: Tatematsu, Fumiya, et al.
Published: (2026)
WorldReel: 4D Video Generation with Consistent Geometry and Motion Modeling
by: Fang, Shaoheng, et al.
Published: (2025)
by: Fang, Shaoheng, et al.
Published: (2025)
A Modular Pipeline for 3D Object Tracking Using RGB Cameras
by: Bredereke, Lars, et al.
Published: (2025)
by: Bredereke, Lars, et al.
Published: (2025)
BLiSS: Bootstrapped Linear Shape Space
by: Muralikrishnan, Sanjeev, et al.
Published: (2023)
by: Muralikrishnan, Sanjeev, et al.
Published: (2023)
Neural Assets: 3D-Aware Multi-Object Scene Synthesis with Image Diffusion Models
by: Wu, Ziyi, et al.
Published: (2024)
by: Wu, Ziyi, et al.
Published: (2024)
Neural Geometry Processing via Spherical Neural Surfaces
by: Williamson, Romy, et al.
Published: (2024)
by: Williamson, Romy, et al.
Published: (2024)
EZ-SP: Fast and Lightweight Superpoint-Based 3D Segmentation
by: Geist, Louis, et al.
Published: (2025)
by: Geist, Louis, et al.
Published: (2025)
RenderDiffusion: Image Diffusion for 3D Reconstruction, Inpainting and Generation
by: Anciukevičius, Titas, et al.
Published: (2022)
by: Anciukevičius, Titas, et al.
Published: (2022)
SuperGaussian: Repurposing Video Models for 3D Super Resolution
by: Shen, Yuan, et al.
Published: (2024)
by: Shen, Yuan, et al.
Published: (2024)
VLM-in-the-Loop: A Plug-In Quality Assurance Module for ECG Digitization Pipelines
by: Li, Jiachen, et al.
Published: (2026)
by: Li, Jiachen, et al.
Published: (2026)
Neural Semantic Surface Maps
by: Morreale, Luca, et al.
Published: (2023)
by: Morreale, Luca, et al.
Published: (2023)
MonoSAOD: Monocular 3D Object Detection with Sparsely Annotated Label
by: Jung, Junyoung, et al.
Published: (2026)
by: Jung, Junyoung, et al.
Published: (2026)
From Programs to Poses: Factored Real-World Scene Generation via Learned Program Libraries
by: Hsu, Joy, et al.
Published: (2025)
by: Hsu, Joy, et al.
Published: (2025)
MonetGPT: Solving Puzzles Enhances MLLMs' Image Retouching Skills
by: Dutt, Niladri Shekhar, et al.
Published: (2025)
by: Dutt, Niladri Shekhar, et al.
Published: (2025)
Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents
by: Baik, Sangwon, et al.
Published: (2026)
by: Baik, Sangwon, et al.
Published: (2026)
Similar Items
-
How to Spin an Object: First, Get the Shape Right
by: Kabra, Rishabh, et al.
Published: (2024) -
Animal Avatars: Reconstructing Animatable 3D Animals from Casual Videos
by: Sabathier, Remy, et al.
Published: (2024) -
ActionMesh: Animated 3D Mesh Generation with Temporal 3D Diffusion
by: Sabathier, Remy, et al.
Published: (2026) -
CHOIR: Contact-aware 4D Hand-Object Interaction Reconstruction
by: Xu, Hao, et al.
Published: (2026) -
A Mixed Diet Makes DINO An Omnivorous Vision Encoder
by: Kabra, Rishabh, et al.
Published: (2026)