Learning Fine-Grained Correspondence with Cross-Perspective Perception for Open-Vocabulary 6D Object Pose Estimation
Fuente:
arXiv
Guardado en:
| Autores principales: | Qin, Yu, Fan, Shimeng, Yang, Fan, Xue, Zixuan, Mai, Zijie, Chen, Wenrui, Yang, Kailun, Li, Zhiyong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
NOVA: Next-step Open-Vocabulary Autoregression for 3D Multi-Object Tracking in Autonomous Driving
por: Luo, Kai, et al.
Publicado: (2026)
por: Luo, Kai, et al.
Publicado: (2026)
Learning Granularity-Aware Affordances from Human-Object Interaction for Tool-Based Functional Dexterous Grasping
por: Yang, Fan, et al.
Publicado: (2024)
por: Yang, Fan, et al.
Publicado: (2024)
One-Shot Affordance Grounding of Deformable Objects in Egocentric Organizing Scenes
por: Jia, Wanjun, et al.
Publicado: (2025)
por: Jia, Wanjun, et al.
Publicado: (2025)
O3N: Omnidirectional Open-Vocabulary Occupancy Prediction
por: Duan, Mengfei, et al.
Publicado: (2026)
por: Duan, Mengfei, et al.
Publicado: (2026)
Multi-Keypoint Affordance Representation for Functional Dexterous Grasping
por: Yang, Fan, et al.
Publicado: (2025)
por: Yang, Fan, et al.
Publicado: (2025)
Towards Precise 3D Human Pose Estimation with Multi-Perspective Spatial-Temporal Relational Transformers
por: Jiao, Jianbin, et al.
Publicado: (2024)
por: Jiao, Jianbin, et al.
Publicado: (2024)
Language-Driven Dual Style Mixing for Single-Domain Generalized Object Detection
por: Qin, Hongda, et al.
Publicado: (2025)
por: Qin, Hongda, et al.
Publicado: (2025)
Resource-Efficient Affordance Grounding with Complementary Depth and Semantic Prompts
por: Huang, Yizhou, et al.
Publicado: (2025)
por: Huang, Yizhou, et al.
Publicado: (2025)
EgoEV-HandPose: Egocentric 3D Hand Pose Estimation and Gesture Recognition with Stereo Event Cameras
por: Wang, Luming, et al.
Publicado: (2026)
por: Wang, Luming, et al.
Publicado: (2026)
DeProPose: Deficiency-Proof 3D Human Pose Estimation via Adaptive Multi-View Fusion
por: Jiao, Jianbin, et al.
Publicado: (2025)
por: Jiao, Jianbin, et al.
Publicado: (2025)
S$^3$-MonoDETR: Supervised Shape&Scale-perceptive Deformable Transformer for Monocular 3D Object Detection
por: He, Xuan, et al.
Publicado: (2023)
por: He, Xuan, et al.
Publicado: (2023)
CoBEVMoE: Heterogeneity-aware Feature Fusion with Dynamic Mixture-of-Experts for Collaborative Perception
por: Kong, Lingzhao, et al.
Publicado: (2025)
por: Kong, Lingzhao, et al.
Publicado: (2025)
GenMapping: Unleashing the Potential of Inverse Perspective Mapping for Robust Online HD Map Construction
por: Li, Siyu, et al.
Publicado: (2024)
por: Li, Siyu, et al.
Publicado: (2024)
Exploring Event-based Human Pose Estimation with 3D Event Representations
por: Yin, Xiaoting, et al.
Publicado: (2023)
por: Yin, Xiaoting, et al.
Publicado: (2023)
LF Tracy: A Unified Single-Pipeline Approach for Salient Object Detection in Light Field Cameras
por: Teng, Fei, et al.
Publicado: (2024)
por: Teng, Fei, et al.
Publicado: (2024)
HierDAMap: Towards Universal Domain Adaptive BEV Mapping via Hierarchical Perspective Priors
por: Li, Siyu, et al.
Publicado: (2025)
por: Li, Siyu, et al.
Publicado: (2025)
P2U-SLAM: A Monocular Wide-FoV SLAM System Based on Point Uncertainty and Pose Uncertainty
por: Zhang, Yufan, et al.
Publicado: (2024)
por: Zhang, Yufan, et al.
Publicado: (2024)
UniFucGrasp: Human-Hand-Inspired Unified Functional Grasp Annotation Strategy and Dataset for Diverse Dexterous Hands
por: Lin, Haoran, et al.
Publicado: (2025)
por: Lin, Haoran, et al.
Publicado: (2025)
Expression Prompt Collaboration Transformer for Universal Referring Video Object Segmentation
por: Chen, Jiajun, et al.
Publicado: (2023)
por: Chen, Jiajun, et al.
Publicado: (2023)
CFMW: Cross-modality Fusion Mamba for Robust Object Detection under Adverse Weather
por: Li, Haoyuan, et al.
Publicado: (2024)
por: Li, Haoyuan, et al.
Publicado: (2024)
Retrieval-Guided Photovoltaic Inventory Estimation from Satellite Imagery for Distribution Grid Planning
por: Guo, Muhao, et al.
Publicado: (2026)
por: Guo, Muhao, et al.
Publicado: (2026)
Beyond the Field-of-View: Enhancing Scene Visibility and Perception with Clip-Recurrent Transformer
por: Shi, Hao, et al.
Publicado: (2022)
por: Shi, Hao, et al.
Publicado: (2022)
EchoTrack: Auditory Referring Multi-Object Tracking for Autonomous Driving
por: Lin, Jiacheng, et al.
Publicado: (2024)
por: Lin, Jiacheng, et al.
Publicado: (2024)
CoCPF: Coordinate-based Continuous Projection Field for Ill-Posed Inverse Problem in Imaging
por: Chen, Zixuan, et al.
Publicado: (2024)
por: Chen, Zixuan, et al.
Publicado: (2024)
MambaMOS: LiDAR-based 3D Moving Object Segmentation with Motion-aware State Space Model
por: Zeng, Kang, et al.
Publicado: (2024)
por: Zeng, Kang, et al.
Publicado: (2024)
MarsSQE: Stereo Quality Enhancement for Martian Images Using Bi-level Cross-view Attention
por: Xu, Mai, et al.
Publicado: (2024)
por: Xu, Mai, et al.
Publicado: (2024)
Multi-Camera Self-Calibration in Sports Motion Capture: Leveraging Human and Stick Poses
por: Yang, Fan, et al.
Publicado: (2026)
por: Yang, Fan, et al.
Publicado: (2026)
DepTR-MOT: Unveiling the Potential of Depth-Informed Trajectory Refinement for Multi-Object Tracking
por: Deng, Buyin, et al.
Publicado: (2025)
por: Deng, Buyin, et al.
Publicado: (2025)
PVPUFormer: Probabilistic Visual Prompt Unified Transformer for Interactive Image Segmentation
por: Zhang, Xu, et al.
Publicado: (2023)
por: Zhang, Xu, et al.
Publicado: (2023)
PanoAffordanceNet: Towards Holistic Affordance Grounding in 360° Indoor Environments
por: Zhu, Guoliang, et al.
Publicado: (2026)
por: Zhu, Guoliang, et al.
Publicado: (2026)
NRSeg: Noise-Resilient Learning for BEV Semantic Segmentation via Driving World Models
por: Li, Siyu, et al.
Publicado: (2025)
por: Li, Siyu, et al.
Publicado: (2025)
Domain Generalization for In-Orbit 6D Pose Estimation
por: Legrand, Antoine, et al.
Publicado: (2024)
por: Legrand, Antoine, et al.
Publicado: (2024)
Towards Consistent Object Detection via LiDAR-Camera Synergy
por: Luo, Kai, et al.
Publicado: (2024)
por: Luo, Kai, et al.
Publicado: (2024)
Perception-Aware Video Semantic Communication
por: Huang, Yinhuan, et al.
Publicado: (2026)
por: Huang, Yinhuan, et al.
Publicado: (2026)
DARCS: Memory-Efficient Deep Compressed Sensing Reconstruction for Acceleration of 3D Whole-Heart Coronary MR Angiography
por: Xue, Zhihao, et al.
Publicado: (2024)
por: Xue, Zhihao, et al.
Publicado: (2024)
Breaking the Multi-Enhancement Bottleneck: Domain-Consistent Quality Enhancement for Compressed Images
por: Xing, Qunliang, et al.
Publicado: (2025)
por: Xing, Qunliang, et al.
Publicado: (2025)
FSAR-Cap: A Fine-Grained Two-Stage Annotated Dataset for SAR Image Captioning
por: Zhang, Jinqi, et al.
Publicado: (2025)
por: Zhang, Jinqi, et al.
Publicado: (2025)
Color Agnostic Cross-Spectral Disparity Estimation
por: Sippel, Frank, et al.
Publicado: (2023)
por: Sippel, Frank, et al.
Publicado: (2023)
Training-Free Robot Pose Estimation using Off-the-Shelf Foundational Models
por: Liang, Laurence
Publicado: (2025)
por: Liang, Laurence
Publicado: (2025)
Unveiling the Potential of Segment Anything Model 2 for RGB-Thermal Semantic Segmentation with Language Guidance
por: Zhao, Jiayi, et al.
Publicado: (2025)
por: Zhao, Jiayi, et al.
Publicado: (2025)
Ejemplares similares
-
NOVA: Next-step Open-Vocabulary Autoregression for 3D Multi-Object Tracking in Autonomous Driving
por: Luo, Kai, et al.
Publicado: (2026) -
Learning Granularity-Aware Affordances from Human-Object Interaction for Tool-Based Functional Dexterous Grasping
por: Yang, Fan, et al.
Publicado: (2024) -
One-Shot Affordance Grounding of Deformable Objects in Egocentric Organizing Scenes
por: Jia, Wanjun, et al.
Publicado: (2025) -
O3N: Omnidirectional Open-Vocabulary Occupancy Prediction
por: Duan, Mengfei, et al.
Publicado: (2026) -
Multi-Keypoint Affordance Representation for Functional Dexterous Grasping
por: Yang, Fan, et al.
Publicado: (2025)