MiKASA: Multi-Key-Anchor & Scene-Aware Transformer for 3D Visual Grounding
Fuente:
arXiv
Guardado en:
| Autores principales: | Chang, Chun-Peng, Wang, Shaoxiang, Pagani, Alain, Stricker, Didier |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Uni-SLAM: Uncertainty-Aware Neural Implicit SLAM for Real-Time Dense Indoor Scene Reconstruction
por: Wang, Shaoxiang, et al.
Publicado: (2024)
por: Wang, Shaoxiang, et al.
Publicado: (2024)
3D Spatial Understanding in MLLMs: Disambiguation and Evaluation
por: Chang, Chun-Peng, et al.
Publicado: (2024)
por: Chang, Chun-Peng, et al.
Publicado: (2024)
Inpaint360GS: Efficient Object-Aware 3D Inpainting via Gaussian Splatting for 360° Scenes
por: Wang, Shaoxiang, et al.
Publicado: (2025)
por: Wang, Shaoxiang, et al.
Publicado: (2025)
SG-PGM: Partial Graph Matching Network with Semantic Geometric Fusion for 3D Scene Graph Alignment and Its Downstream Tasks
por: Xie, Yaxu, et al.
Publicado: (2024)
por: Xie, Yaxu, et al.
Publicado: (2024)
EgoForce: Forearm-Guided Camera-Space 3D Hand Pose from a Monocular Egocentric Camera
por: Millerdurai, Christen, et al.
Publicado: (2026)
por: Millerdurai, Christen, et al.
Publicado: (2026)
G3FA: Geometry-guided GAN for Face Animation
por: Javanmardi, Alireza, et al.
Publicado: (2024)
por: Javanmardi, Alireza, et al.
Publicado: (2024)
Invaria: Learning Scale and Density Invariance in Point Clouds via Next-Resolution Prediction
por: Chang, Chun-Peng, et al.
Publicado: (2026)
por: Chang, Chun-Peng, et al.
Publicado: (2026)
TalkingPose: Efficient Face and Gesture Animation with Feedback-guided Diffusion Model
por: Javanmardi, Alireza, et al.
Publicado: (2025)
por: Javanmardi, Alireza, et al.
Publicado: (2025)
Beyond Averages: Open-Vocabulary 3D Scene Understanding with Gaussian Splatting and Bag of Embeddings
por: Arafa, Abdalla, et al.
Publicado: (2025)
por: Arafa, Abdalla, et al.
Publicado: (2025)
ReLaGS: Relational Language Gaussian Splatting
por: Xie, Yaxu, et al.
Publicado: (2026)
por: Xie, Yaxu, et al.
Publicado: (2026)
Probing the Reliability of Driving VLMs: From Inconsistent Responses to Grounded Temporal Reasoning
por: Chang, Chun-Peng, et al.
Publicado: (2026)
por: Chang, Chun-Peng, et al.
Publicado: (2026)
EventEgo3D++: 3D Human Motion Capture from a Head-Mounted Event Camera
por: Millerdurai, Christen, et al.
Publicado: (2025)
por: Millerdurai, Christen, et al.
Publicado: (2025)
Object-Centric 2D Gaussian Splatting: Background Removal and Occlusion-Aware Pruning for Compact Object Models
por: Rogge, Marcel, et al.
Publicado: (2025)
por: Rogge, Marcel, et al.
Publicado: (2025)
SIMSPINE: A Biomechanics-Aware Simulation Framework for 3D Spine Motion Annotation and Benchmarking
por: Khan, Muhammad Saif Ullah, et al.
Publicado: (2026)
por: Khan, Muhammad Saif Ullah, et al.
Publicado: (2026)
RMS-FlowNet++: Efficient and Robust Multi-Scale Scene Flow Estimation for Large-Scale Point Clouds
por: Battrawy, Ramy, et al.
Publicado: (2024)
por: Battrawy, Ramy, et al.
Publicado: (2024)
SF3D-RGB: Scene Flow Estimation from Monocular Camera and Sparse LiDAR
por: Alhimdiat, Rajai, et al.
Publicado: (2026)
por: Alhimdiat, Rajai, et al.
Publicado: (2026)
EgoFlowNet: Non-Rigid Scene Flow from Point Clouds with Ego-Motion Support
por: Battrawy, Ramy, et al.
Publicado: (2024)
por: Battrawy, Ramy, et al.
Publicado: (2024)
ChangingGrounding: 3D Visual Grounding in Changing Scenes
por: Hu, Miao, et al.
Publicado: (2025)
por: Hu, Miao, et al.
Publicado: (2025)
TinyIceNet: Low-Power SAR Sea Ice Segmentation for On-Board FPGA Inference
por: Koutayni, Mhd Rashed Al, et al.
Publicado: (2026)
por: Koutayni, Mhd Rashed Al, et al.
Publicado: (2026)
Seeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving Simulation
por: Chang, Chun-Peng, et al.
Publicado: (2025)
por: Chang, Chun-Peng, et al.
Publicado: (2025)
Semi-Supervised Object Detection: A Survey on Progress from CNN to Transformer
por: Shehzadi, Tahira, et al.
Publicado: (2024)
por: Shehzadi, Tahira, et al.
Publicado: (2024)
D-MiSo: Editing Dynamic 3D Scenes using Multi-Gaussians Soup
por: Waczyńska, Joanna, et al.
Publicado: (2024)
por: Waczyńska, Joanna, et al.
Publicado: (2024)
From Objects to Anywhere: A Holistic Benchmark for Multi-level Visual Grounding in 3D Scenes
por: Wang, Tianxu, et al.
Publicado: (2025)
por: Wang, Tianxu, et al.
Publicado: (2025)
SemAttNet: Towards Attention-based Semantic Aware Guided Depth Completion
por: Nazir, Danish, et al.
Publicado: (2022)
por: Nazir, Danish, et al.
Publicado: (2022)
Towards Unconstrained 2D Pose Estimation of the Human Spine
por: Khan, Muhammad Saif Ullah, et al.
Publicado: (2025)
por: Khan, Muhammad Saif Ullah, et al.
Publicado: (2025)
Towards End-to-End Semi-Supervised Table Detection with Semantic Aligned Matching Transformer
por: Shehzadi, Tahira, et al.
Publicado: (2024)
por: Shehzadi, Tahira, et al.
Publicado: (2024)
UniGround: Universal 3D Visual Grounding via Training-Free Scene Parsing
por: Zhang, Jiaxi, et al.
Publicado: (2026)
por: Zhang, Jiaxi, et al.
Publicado: (2026)
LSVG: Language-Guided Scene Graphs with 2D-Assisted Multi-Modal Encoding for 3D Visual Grounding
por: Xiao, Feng, et al.
Publicado: (2025)
por: Xiao, Feng, et al.
Publicado: (2025)
ShapeAug: Occlusion Augmentation for Event Camera Data
por: Bendig, Katharina, et al.
Publicado: (2024)
por: Bendig, Katharina, et al.
Publicado: (2024)
ShapeAug++: More Realistic Shape Augmentation for Event Data
por: Bendig, Katharina, et al.
Publicado: (2024)
por: Bendig, Katharina, et al.
Publicado: (2024)
MILE: Mixture of Incremental LoRA Experts for Continual Semantic Segmentation across Domains and Modalities
por: Muralidhara, Shishir, et al.
Publicado: (2026)
por: Muralidhara, Shishir, et al.
Publicado: (2026)
Sensor Generalization for Adaptive Sensing in Event-based Object Detection via Joint Distribution Training
por: Saha, Aheli, et al.
Publicado: (2026)
por: Saha, Aheli, et al.
Publicado: (2026)
Domain-Incremental Semantic Segmentation for Autonomous Driving under Adverse Driving Conditions
por: Muralidhara, Shishir, et al.
Publicado: (2025)
por: Muralidhara, Shishir, et al.
Publicado: (2025)
PanoSAMic: Panoramic Image Segmentation from SAM Feature Encoding and Dual View Fusion
por: Chamseddine, Mahdi, et al.
Publicado: (2026)
por: Chamseddine, Mahdi, et al.
Publicado: (2026)
SAILS: Segment Anything with Incrementally Learned Semantics for Task-Invariant and Training-Free Continual Learning
por: Muralidhara, Shishir, et al.
Publicado: (2026)
por: Muralidhara, Shishir, et al.
Publicado: (2026)
DriverGaze360: OmniDirectional Driver Attention with Object-Level Guidance
por: Govil, Shreedhar, et al.
Publicado: (2025)
por: Govil, Shreedhar, et al.
Publicado: (2025)
PoseAdapt: Sustainable Human Pose Estimation via Continual Learning Benchmarks and Toolkit
por: Khan, Muhammad Saif Ullah, et al.
Publicado: (2024)
por: Khan, Muhammad Saif Ullah, et al.
Publicado: (2024)
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations
por: Yuan, Zhihao, et al.
Publicado: (2025)
por: Yuan, Zhihao, et al.
Publicado: (2025)
ReConText3D: Replay-based Continual Text-to-3D Generation
por: Khan, Muhammad Ahmed Ullah, et al.
Publicado: (2026)
por: Khan, Muhammad Ahmed Ullah, et al.
Publicado: (2026)
ProxyTransformation: Preshaping Point Cloud Manifold With Proxy Attention For 3D Visual Grounding
por: Peng, Qihang, et al.
Publicado: (2025)
por: Peng, Qihang, et al.
Publicado: (2025)
Ejemplares similares
-
Uni-SLAM: Uncertainty-Aware Neural Implicit SLAM for Real-Time Dense Indoor Scene Reconstruction
por: Wang, Shaoxiang, et al.
Publicado: (2024) -
3D Spatial Understanding in MLLMs: Disambiguation and Evaluation
por: Chang, Chun-Peng, et al.
Publicado: (2024) -
Inpaint360GS: Efficient Object-Aware 3D Inpainting via Gaussian Splatting for 360° Scenes
por: Wang, Shaoxiang, et al.
Publicado: (2025) -
SG-PGM: Partial Graph Matching Network with Semantic Geometric Fusion for 3D Scene Graph Alignment and Its Downstream Tasks
por: Xie, Yaxu, et al.
Publicado: (2024) -
EgoForce: Forearm-Guided Camera-Space 3D Hand Pose from a Monocular Egocentric Camera
por: Millerdurai, Christen, et al.
Publicado: (2026)