ODIN: A Single Model for 2D and 3D Segmentation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jain, Ayush, Katara, Pushkal, Gkanatsios, Nikolaos, Harley, Adam W., Sarch, Gabriel, Aggarwal, Kriti, Chaudhary, Vishrav, Fragkiadaki, Katerina |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
3D Diffuser Actor: Policy Diffusion with 3D Scene Representations
von: Ke, Tsung-Wei, et al.
Veröffentlicht: (2024)
von: Ke, Tsung-Wei, et al.
Veröffentlicht: (2024)
Energy-based Models are Zero-Shot Planners for Compositional Scene Rearrangement
von: Gkanatsios, Nikolaos, et al.
Veröffentlicht: (2023)
von: Gkanatsios, Nikolaos, et al.
Veröffentlicht: (2023)
3D FlowMatch Actor: Unified 3D Policy for Single- and Dual-Arm Manipulation
von: Gkanatsios, Nikolaos, et al.
Veröffentlicht: (2025)
von: Gkanatsios, Nikolaos, et al.
Veröffentlicht: (2025)
Unifying 2D and 3D Vision-Language Understanding
von: Jain, Ayush, et al.
Veröffentlicht: (2025)
von: Jain, Ayush, et al.
Veröffentlicht: (2025)
TAPIP3D: Tracking Any Point in Persistent 3D Geometry
von: Zhang, Bowei, et al.
Veröffentlicht: (2025)
von: Zhang, Bowei, et al.
Veröffentlicht: (2025)
Grounded Reinforcement Learning for Visual Reasoning
von: Sarch, Gabriel, et al.
Veröffentlicht: (2025)
von: Sarch, Gabriel, et al.
Veröffentlicht: (2025)
Dex4D: Task-Agnostic Point Track Policy for Sim-to-Real Dexterous Manipulation
von: Kuang, Yuxuan, et al.
Veröffentlicht: (2026)
von: Kuang, Yuxuan, et al.
Veröffentlicht: (2026)
Diffusion-ES: Gradient-free Planning with Diffusion for Autonomous Driving and Zero-Shot Instruction Following
von: Yang, Brian, et al.
Veröffentlicht: (2024)
von: Yang, Brian, et al.
Veröffentlicht: (2024)
RoboTAG: End-to-end Robot Configuration Estimation via Topological Alignment Graph
von: Liu, Yifan, et al.
Veröffentlicht: (2025)
von: Liu, Yifan, et al.
Veröffentlicht: (2025)
Aligning Text-to-Image Diffusion Models with Reward Backpropagation
von: Prabhudesai, Mihir, et al.
Veröffentlicht: (2023)
von: Prabhudesai, Mihir, et al.
Veröffentlicht: (2023)
HELPER-X: A Unified Instructable Embodied Agent to Tackle Four Interactive Vision-Language Domains with Memory-Augmented Language Models
von: Sarch, Gabriel, et al.
Veröffentlicht: (2024)
von: Sarch, Gabriel, et al.
Veröffentlicht: (2024)
Learning to Assist: Physics-Grounded Human-Human Control via Multi-Agent Reinforcement Learning
von: Shibata, Yuto, et al.
Veröffentlicht: (2026)
von: Shibata, Yuto, et al.
Veröffentlicht: (2026)
VLM Agents Generate Their Own Memories: Distilling Experience into Embodied Programs of Thought
von: Sarch, Gabriel, et al.
Veröffentlicht: (2024)
von: Sarch, Gabriel, et al.
Veröffentlicht: (2024)
Video Diffusion Alignment via Reward Gradients
von: Prabhudesai, Mihir, et al.
Veröffentlicht: (2024)
von: Prabhudesai, Mihir, et al.
Veröffentlicht: (2024)
Diffusion Beats Autoregressive in Data-Constrained Settings
von: Prabhudesai, Mihir, et al.
Veröffentlicht: (2025)
von: Prabhudesai, Mihir, et al.
Veröffentlicht: (2025)
Unified Multimodal Discrete Diffusion
von: Swerdlow, Alexander, et al.
Veröffentlicht: (2025)
von: Swerdlow, Alexander, et al.
Veröffentlicht: (2025)
Semantic Segmentation and Depth Estimation for Real-Time Lunar Surface Mapping Using 3D Gaussian Splatting
von: Vila, Guillem Casadesus, et al.
Veröffentlicht: (2026)
von: Vila, Guillem Casadesus, et al.
Veröffentlicht: (2026)
More than Segmentation: Benchmarking SAM 3 for Segmentation, 3D Perception, and Reconstruction in Robotic Surgery
von: Dong, Wenzhen, et al.
Veröffentlicht: (2025)
von: Dong, Wenzhen, et al.
Veröffentlicht: (2025)
SR3D: Unleashing Single-view 3D Reconstruction for Transparent and Specular Object Grasping
von: Zhang, Mingxu, et al.
Veröffentlicht: (2025)
von: Zhang, Mingxu, et al.
Veröffentlicht: (2025)
Zero-Shot Open-Vocabulary Tracking with Large Pre-Trained Models
von: Chu, Wen-Hsuan, et al.
Veröffentlicht: (2023)
von: Chu, Wen-Hsuan, et al.
Veröffentlicht: (2023)
Gradient-Driven 3D Segmentation and Affordance Transfer in Gaussian Splatting Using 2D Masks
von: Joseph, Joji, et al.
Veröffentlicht: (2024)
von: Joseph, Joji, et al.
Veröffentlicht: (2024)
DreamScene4D: Dynamic Multi-Object Scene Generation from Monocular Videos
von: Chu, Wen-Hsuan, et al.
Veröffentlicht: (2024)
von: Chu, Wen-Hsuan, et al.
Veröffentlicht: (2024)
Semantics-Guided Moving Object Segmentation with 3D LiDAR
von: Gu, Shuo, et al.
Veröffentlicht: (2022)
von: Gu, Shuo, et al.
Veröffentlicht: (2022)
WildScenes: A Benchmark for 2D and 3D Semantic Segmentation in Large-scale Natural Environments
von: Vidanapathirana, Kavisha, et al.
Veröffentlicht: (2023)
von: Vidanapathirana, Kavisha, et al.
Veröffentlicht: (2023)
Domain Adaptation-Based Crossmodal Knowledge Distillation for 3D Semantic Segmentation
von: Kang, Jialiang, et al.
Veröffentlicht: (2025)
von: Kang, Jialiang, et al.
Veröffentlicht: (2025)
EmbodiedSAM: Online Segment Any 3D Thing in Real Time
von: Xu, Xiuwei, et al.
Veröffentlicht: (2024)
von: Xu, Xiuwei, et al.
Veröffentlicht: (2024)
A Generalization of CLAP from 3D Localization to Image Processing, A Connection With RANSAC & Hough Transforms
von: Hou, Ruochen, et al.
Veröffentlicht: (2025)
von: Hou, Ruochen, et al.
Veröffentlicht: (2025)
Learning to Identify Out-of-Distribution Objects for 3D LiDAR Anomaly Segmentation
von: Mosco, Simone, et al.
Veröffentlicht: (2026)
von: Mosco, Simone, et al.
Veröffentlicht: (2026)
3D Hierarchical Panoptic Segmentation in Real Orchard Environments Across Different Sensors
von: Sodano, Matteo, et al.
Veröffentlicht: (2025)
von: Sodano, Matteo, et al.
Veröffentlicht: (2025)
QueryOcc: Query-based Self-Supervision for 3D Semantic Occupancy
von: Lilja, Adam, et al.
Veröffentlicht: (2025)
von: Lilja, Adam, et al.
Veröffentlicht: (2025)
SpaCeFormer: Fast Proposal-Free Open-Vocabulary 3D Instance Segmentation
von: Choy, Chris, et al.
Veröffentlicht: (2026)
von: Choy, Chris, et al.
Veröffentlicht: (2026)
Overlap-Aware Feature Learning for Robust Unsupervised Domain Adaptation for 3D Semantic Segmentation
von: Chen, Junjie, et al.
Veröffentlicht: (2025)
von: Chen, Junjie, et al.
Veröffentlicht: (2025)
PLVS: A SLAM System with Points, Lines, Volumetric Mapping, and 3D Incremental Segmentation
von: Freda, Luigi
Veröffentlicht: (2023)
von: Freda, Luigi
Veröffentlicht: (2023)
Iterative Refinement Improves Compositional Image Generation
von: Jaiswal, Shantanu, et al.
Veröffentlicht: (2026)
von: Jaiswal, Shantanu, et al.
Veröffentlicht: (2026)
Hamba: Single-view 3D Hand Reconstruction with Graph-guided Bi-Scanning Mamba
von: Dong, Haoye, et al.
Veröffentlicht: (2024)
von: Dong, Haoye, et al.
Veröffentlicht: (2024)
PhysX-Anything: Simulation-Ready Physical 3D Assets from Single Image
von: Cao, Ziang, et al.
Veröffentlicht: (2025)
von: Cao, Ziang, et al.
Veröffentlicht: (2025)
Single-View 3D Reconstruction via SO(2)-Equivariant Gaussian Sculpting Networks
von: Xu, Ruihan, et al.
Veröffentlicht: (2024)
von: Xu, Ruihan, et al.
Veröffentlicht: (2024)
Clutt3R-Seg: Sparse-view 3D Instance Segmentation for Language-grounded Grasping in Cluttered Scenes
von: Noh, Jeongho, et al.
Veröffentlicht: (2026)
von: Noh, Jeongho, et al.
Veröffentlicht: (2026)
Methods for the Segmentation of Reticular Structures Using 3D LiDAR Data: A Comparative Evaluation
von: Mora, Francisco J. Soler, et al.
Veröffentlicht: (2025)
von: Mora, Francisco J. Soler, et al.
Veröffentlicht: (2025)
Segment, Lift and Fit: Automatic 3D Shape Labeling from 2D Prompts
von: Li, Jianhao, et al.
Veröffentlicht: (2024)
von: Li, Jianhao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
3D Diffuser Actor: Policy Diffusion with 3D Scene Representations
von: Ke, Tsung-Wei, et al.
Veröffentlicht: (2024) -
Energy-based Models are Zero-Shot Planners for Compositional Scene Rearrangement
von: Gkanatsios, Nikolaos, et al.
Veröffentlicht: (2023) -
3D FlowMatch Actor: Unified 3D Policy for Single- and Dual-Arm Manipulation
von: Gkanatsios, Nikolaos, et al.
Veröffentlicht: (2025) -
Unifying 2D and 3D Vision-Language Understanding
von: Jain, Ayush, et al.
Veröffentlicht: (2025) -
TAPIP3D: Tracking Any Point in Persistent 3D Geometry
von: Zhang, Bowei, et al.
Veröffentlicht: (2025)