WildRefer: 3D Object Localization in Large-scale Dynamic Scenes with Multi-modal Visual Data and Natural Language
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Zhenxiang, Peng, Xidong, Cong, Peishan, Zheng, Ge, Sun, Yujin, Hou, Yuenan, Zhu, Xinge, Yang, Sibei, Ma, Yuexin |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SemGeoMo: Dynamic Contextual Human Motion Generation with Semantic and Geometric Guidance
by: Cong, Peishan, et al.
Published: (2025)
by: Cong, Peishan, et al.
Published: (2025)
EvolvingGrasp: Evolutionary Grasp Generation via Efficient Preference Alignment
by: Zhu, Yufei, et al.
Published: (2025)
by: Zhu, Yufei, et al.
Published: (2025)
LaserHuman: Language-guided Scene-aware Human Motion Generation in Free Environment
by: Cong, Peishan, et al.
Published: (2024)
by: Cong, Peishan, et al.
Published: (2024)
MoE3D: Mixture of Experts meets Multi-Modal 3D Understanding
by: Li, Yu, et al.
Published: (2025)
by: Li, Yu, et al.
Published: (2025)
Gait Recognition in Large-scale Free Environment via Single LiDAR
by: Han, Xiao, et al.
Published: (2022)
by: Han, Xiao, et al.
Published: (2022)
Controllable Video Object Insertion via Multiview Priors
by: Qi, Xia, et al.
Published: (2026)
by: Qi, Xia, et al.
Published: (2026)
SocialMirror: Reconstructing 3D Human Interaction Behaviors from Monocular Videos with Semantic and Geometric Guidance
by: Xia, Qi, et al.
Published: (2026)
by: Xia, Qi, et al.
Published: (2026)
$L^3$:Scene-agnostic Visual Localization in the Wild
by: Zhang, Yu, et al.
Published: (2026)
by: Zhang, Yu, et al.
Published: (2026)
OccMamba: Semantic Occupancy Prediction with State Space Models
by: Li, Heng, et al.
Published: (2024)
by: Li, Heng, et al.
Published: (2024)
Part2Object: Hierarchical Unsupervised 3D Instance Segmentation
by: Shi, Cheng, et al.
Published: (2024)
by: Shi, Cheng, et al.
Published: (2024)
TASeg: Temporal Aggregation Network for LiDAR Semantic Segmentation
by: Wu, Xiaopei, et al.
Published: (2024)
by: Wu, Xiaopei, et al.
Published: (2024)
Learning to Adapt SAM for Segmenting Cross-domain Point Clouds
by: Peng, Xidong, et al.
Published: (2023)
by: Peng, Xidong, et al.
Published: (2023)
WildScenes: A Benchmark for 2D and 3D Semantic Segmentation in Large-scale Natural Environments
by: Vidanapathirana, Kavisha, et al.
Published: (2023)
by: Vidanapathirana, Kavisha, et al.
Published: (2023)
Curriculum Point Prompting for Weakly-Supervised Referring Image Segmentation
by: Dai, Qiyuan, et al.
Published: (2024)
by: Dai, Qiyuan, et al.
Published: (2024)
Ref-AVS: Refer and Segment Objects in Audio-Visual Scenes
by: Wang, Yaoting, et al.
Published: (2024)
by: Wang, Yaoting, et al.
Published: (2024)
HexPlane Representation for 3D Semantic Scene Understanding
by: Chen, Zeren, et al.
Published: (2025)
by: Chen, Zeren, et al.
Published: (2025)
SceneX: Procedural Controllable Large-scale Scene Generation
by: Zhou, Mengqi, et al.
Published: (2024)
by: Zhou, Mengqi, et al.
Published: (2024)
Exploring Textual Semantics Diversity for Image Transmission in Semantic Communication Systems using Visual Language Model
by: Huang, Peishan, et al.
Published: (2025)
by: Huang, Peishan, et al.
Published: (2025)
EasyHOI: Unleashing the Power of Large Models for Reconstructing Hand-Object Interactions in the Wild
by: Liu, Yumeng, et al.
Published: (2024)
by: Liu, Yumeng, et al.
Published: (2024)
Mpemba Effect in Large-Language Model Training Dynamics: A Minimal Analysis of the Valley-River model
by: Liu, Sibei, et al.
Published: (2025)
by: Liu, Sibei, et al.
Published: (2025)
Dynamical analysis and optimal control of an age‐structured epidemic model with asymptomatic infection and multiple transmission pathways
by: Yuenan Kang, et al.
Published: (2024)
by: Yuenan Kang, et al.
Published: (2024)
ReAL-AD: Towards Human-Like Reasoning in End-to-End Autonomous Driving
by: Lu, Yuhang, et al.
Published: (2025)
by: Lu, Yuhang, et al.
Published: (2025)
OctreeOcc: Efficient and Multi-Granularity Occupancy Prediction Using Octree Queries
by: Lu, Yuhang, et al.
Published: (2023)
by: Lu, Yuhang, et al.
Published: (2023)
Sufficient conditions for the variation of toughness under the distance spectral in graphs involving minimum degree
by: Li, Peishan
Published: (2025)
by: Li, Peishan
Published: (2025)
Commonsense Scene Graph-based Target Localization for Object Search
by: Ge, Wenqi, et al.
Published: (2024)
by: Ge, Wenqi, et al.
Published: (2024)
HUNTER: Unsupervised Human-centric 3D Detection via Transferring Knowledge from Synthetic Instances to Real Scenes
by: Yao, Yichen, et al.
Published: (2024)
by: Yao, Yichen, et al.
Published: (2024)
HUMOF: Human Motion Forecasting in Interactive Social Scenes
by: Sun, Caiyi, et al.
Published: (2025)
by: Sun, Caiyi, et al.
Published: (2025)
SceneVTG++: Controllable Multilingual Visual Text Generation in the Wild
by: Liu, Jiawei, et al.
Published: (2025)
by: Liu, Jiawei, et al.
Published: (2025)
Detection and Geographic Localization of Natural Objects in the Wild: A Case Study on Palms
by: Cui, Kangning, et al.
Published: (2025)
by: Cui, Kangning, et al.
Published: (2025)
TransXNet: Learning Both Global and Local Dynamics with a Dual Dynamic Token Mixer for Visual Recognition
by: Lou, Meng, et al.
Published: (2023)
by: Lou, Meng, et al.
Published: (2023)
STAGE: A Stream-Centric Generative World Model for Long-Horizon Driving-Scene Simulation
by: Wang, Jiamin, et al.
Published: (2025)
by: Wang, Jiamin, et al.
Published: (2025)
UniDemoiré: Towards Universal Image Demoiréing with Data Generation and Synthesis
by: Yang, Zemin, et al.
Published: (2025)
by: Yang, Zemin, et al.
Published: (2025)
EvoDefense: Co-Evolving Black-Box Defense with Large Language Models
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
Wild-Drive: Off-Road Scene Captioning and Path Planning via Robust Multi-modal Routing and Efficient Large Language Model
by: Wang, Zihang, et al.
Published: (2026)
by: Wang, Zihang, et al.
Published: (2026)
Visual Text Generation in the Wild
by: Zhu, Yuanzhi, et al.
Published: (2024)
by: Zhu, Yuanzhi, et al.
Published: (2024)
Show Me When and Where: Towards Referring Video Object Segmentation in the Wild
by: Gao, Mingqi, et al.
Published: (2026)
by: Gao, Mingqi, et al.
Published: (2026)
FreeTimeGS: Free Gaussian Primitives at Anytime and Anywhere for Dynamic Scene Reconstruction
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
Closed-Loop Transfer for Weakly-supervised Affordance Grounding
by: Tang, Jiajin, et al.
Published: (2025)
by: Tang, Jiajin, et al.
Published: (2025)
Why LVLMs Are More Prone to Hallucinations in Longer Responses: The Role of Context
by: Zheng, Ge, et al.
Published: (2025)
by: Zheng, Ge, et al.
Published: (2025)
Intervene-All-Paths: Unified Mitigation of LVLM Hallucinations across Alignment Formats
by: Qian, Jiaye, et al.
Published: (2025)
by: Qian, Jiaye, et al.
Published: (2025)
Similar Items
-
SemGeoMo: Dynamic Contextual Human Motion Generation with Semantic and Geometric Guidance
by: Cong, Peishan, et al.
Published: (2025) -
EvolvingGrasp: Evolutionary Grasp Generation via Efficient Preference Alignment
by: Zhu, Yufei, et al.
Published: (2025) -
LaserHuman: Language-guided Scene-aware Human Motion Generation in Free Environment
by: Cong, Peishan, et al.
Published: (2024) -
MoE3D: Mixture of Experts meets Multi-Modal 3D Understanding
by: Li, Yu, et al.
Published: (2025) -
Gait Recognition in Large-scale Free Environment via Single LiDAR
by: Han, Xiao, et al.
Published: (2022)