Salvato in:
| Autori principali: | Zheng, Meng, Marri, Samhita, Choudhuri, Anwesa, Planche, Benjamin, Gao, Zhongpai, Nguyen, Van Nguyen, Chen, Terrence, Chowdhary, Girish, Wu, Ziyan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2605.08434 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
PolypSegTrack: Unified Foundation Model for Colonoscopy Video Analysis
di: Choudhuri, Anwesa, et al.
Pubblicazione: (2025)
di: Choudhuri, Anwesa, et al.
Pubblicazione: (2025)
Render-FM: A Foundation Model for Real-time Photorealistic Volumetric Rendering
di: Gao, Zhongpai, et al.
Pubblicazione: (2025)
di: Gao, Zhongpai, et al.
Pubblicazione: (2025)
7DGS: Unified Spatial-Temporal-Angular Gaussian Splatting
di: Gao, Zhongpai, et al.
Pubblicazione: (2025)
di: Gao, Zhongpai, et al.
Pubblicazione: (2025)
6DGS: Enhanced Direction-Aware Gaussian Splatting for Volumetric Rendering
di: Gao, Zhongpai, et al.
Pubblicazione: (2024)
di: Gao, Zhongpai, et al.
Pubblicazione: (2024)
3D Vision-Language Gaussian Splatting
di: Peng, Qucheng, et al.
Pubblicazione: (2024)
di: Peng, Qucheng, et al.
Pubblicazione: (2024)
PlantTrack: Task-Driven Plant Keypoint Tracking with Zero-Shot Sim2Real Transfer
di: Marri, Samhita, et al.
Pubblicazione: (2024)
di: Marri, Samhita, et al.
Pubblicazione: (2024)
MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding
di: Su, Yuhao, et al.
Pubblicazione: (2025)
di: Su, Yuhao, et al.
Pubblicazione: (2025)
Universal Beta Splatting
di: Liu, Rong, et al.
Pubblicazione: (2025)
di: Liu, Rong, et al.
Pubblicazione: (2025)
HyReach: Vision-Guided Hybrid Manipulator Reaching in Unseen Cluttered Environments
di: Kamtikar, Shivani, et al.
Pubblicazione: (2026)
di: Kamtikar, Shivani, et al.
Pubblicazione: (2026)
Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal Grounding
di: Deng, Andong, et al.
Pubblicazione: (2024)
di: Deng, Andong, et al.
Pubblicazione: (2024)
Order-aware Interactive Segmentation
di: Wang, Bin, et al.
Pubblicazione: (2024)
di: Wang, Bin, et al.
Pubblicazione: (2024)
OW-VISCapTor: Abstractors for Open-World Video Instance Segmentation and Captioning
di: Choudhuri, Anwesa, et al.
Pubblicazione: (2024)
di: Choudhuri, Anwesa, et al.
Pubblicazione: (2024)
Consistent Instance Field for Dynamic Scene Understanding
di: Wu, Junyi, et al.
Pubblicazione: (2025)
di: Wu, Junyi, et al.
Pubblicazione: (2025)
Gazebo Plants: Simulating Plant-Robot Interaction with Cosserat Rods
di: Deng, Junchen, et al.
Pubblicazione: (2024)
di: Deng, Junchen, et al.
Pubblicazione: (2024)
Precision Harvesting in Cluttered Environments: Integrating End Effector Design with Dual Camera Perception
di: Koe, Kendall, et al.
Pubblicazione: (2025)
di: Koe, Kendall, et al.
Pubblicazione: (2025)
Automated Patient Positioning with Learned 3D Hand Gestures
di: Gao, Zhongpai, et al.
Pubblicazione: (2024)
di: Gao, Zhongpai, et al.
Pubblicazione: (2024)
Anatomy-Aware Conditional Image-Text Retrieval
di: Zheng, Meng, et al.
Pubblicazione: (2025)
di: Zheng, Meng, et al.
Pubblicazione: (2025)
DDGS-CT: Direction-Disentangled Gaussian Splatting for Realistic Volume Rendering
di: Gao, Zhongpai, et al.
Pubblicazione: (2024)
di: Gao, Zhongpai, et al.
Pubblicazione: (2024)
CHROME: Clothed Human Reconstruction with Occlusion-Resilience and Multiview-Consistency from a Single Image
di: Dutta, Arindam, et al.
Pubblicazione: (2025)
di: Dutta, Arindam, et al.
Pubblicazione: (2025)
Few-Shot 3D Volumetric Segmentation with Multi-Surrogate Fusion
di: Zheng, Meng, et al.
Pubblicazione: (2024)
di: Zheng, Meng, et al.
Pubblicazione: (2024)
From Particles to Fields: Reframing Photon Mapping with Continuous Gaussian Photon Fields
di: Tao, Jiachen, et al.
Pubblicazione: (2025)
di: Tao, Jiachen, et al.
Pubblicazione: (2025)
PBADet: A One-Stage Anchor-Free Approach for Part-Body Association
di: Gao, Zhongpai, et al.
Pubblicazione: (2024)
di: Gao, Zhongpai, et al.
Pubblicazione: (2024)
Exploring Cycle Consistency Learning in Interactive Volume Segmentation
di: Liu, Qin, et al.
Pubblicazione: (2023)
di: Liu, Qin, et al.
Pubblicazione: (2023)
FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models
di: Lin, Zijun, et al.
Pubblicazione: (2025)
di: Lin, Zijun, et al.
Pubblicazione: (2025)
CATNAV: Cached Vision-Language Traversability for Efficient Zero-Shot Robot Navigation
di: Potnis, Aditya, et al.
Pubblicazione: (2026)
di: Potnis, Aditya, et al.
Pubblicazione: (2026)
Active Semantic Mapping of Horticultural Environments Using Gaussian Splatting
di: Cuaran, Jose, et al.
Pubblicazione: (2026)
di: Cuaran, Jose, et al.
Pubblicazione: (2026)
Visual-Language-Guided Task Planning for Horticultural Robots
di: Cuaran, Jose, et al.
Pubblicazione: (2026)
di: Cuaran, Jose, et al.
Pubblicazione: (2026)
I-FailSense: Towards General Robotic Failure Detection with Vision-Language Models
di: Grislain, Clemence, et al.
Pubblicazione: (2025)
di: Grislain, Clemence, et al.
Pubblicazione: (2025)
Neural Finite-State Machines for Surgical Phase Recognition
di: Ding, Hao, et al.
Pubblicazione: (2024)
di: Ding, Hao, et al.
Pubblicazione: (2024)
DaRePlane: Direction-aware Representations for Dynamic Scene Reconstruction
di: Lou, Ange, et al.
Pubblicazione: (2024)
di: Lou, Ange, et al.
Pubblicazione: (2024)
DaReNeRF: Direction-aware Representation for Dynamic Scenes
di: Lou, Ange, et al.
Pubblicazione: (2024)
di: Lou, Ange, et al.
Pubblicazione: (2024)
IA-VLA: Input Augmentation for Vision-Language-Action models in settings with semantically complex tasks
di: Hannus, Eric, et al.
Pubblicazione: (2025)
di: Hannus, Eric, et al.
Pubblicazione: (2025)
Divide and Fuse: Body Part Mesh Recovery from Partially Visible Human Images
di: Luan, Tianyu, et al.
Pubblicazione: (2024)
di: Luan, Tianyu, et al.
Pubblicazione: (2024)
SAFE: Multitask Failure Detection for Vision-Language-Action Models
di: Gu, Qiao, et al.
Pubblicazione: (2025)
di: Gu, Qiao, et al.
Pubblicazione: (2025)
Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding
di: Vo, Khoa, et al.
Pubblicazione: (2025)
di: Vo, Khoa, et al.
Pubblicazione: (2025)
Adaptive Action Chunking at Inference-time for Vision-Language-Action Models
di: Liang, Yuanchang, et al.
Pubblicazione: (2026)
di: Liang, Yuanchang, et al.
Pubblicazione: (2026)
Self-supervised 3D Patient Modeling with Multi-modal Attentive Fusion
di: Zheng, Meng, et al.
Pubblicazione: (2024)
di: Zheng, Meng, et al.
Pubblicazione: (2024)
Automating Catheterization Labs with Real-Time Perception
di: Yang, Fan, et al.
Pubblicazione: (2024)
di: Yang, Fan, et al.
Pubblicazione: (2024)
Threading Optimization for Vision-Language-Action Model Inference in Low-Cost Smart Agricultural Manipulation
di: Truongcao, Keith, et al.
Pubblicazione: (2026)
di: Truongcao, Keith, et al.
Pubblicazione: (2026)
Active Semantic Mapping with Mobile Manipulator in Horticultural Environments
di: Cuaran, Jose, et al.
Pubblicazione: (2024)
di: Cuaran, Jose, et al.
Pubblicazione: (2024)
Documenti analoghi
-
PolypSegTrack: Unified Foundation Model for Colonoscopy Video Analysis
di: Choudhuri, Anwesa, et al.
Pubblicazione: (2025) -
Render-FM: A Foundation Model for Real-time Photorealistic Volumetric Rendering
di: Gao, Zhongpai, et al.
Pubblicazione: (2025) -
7DGS: Unified Spatial-Temporal-Angular Gaussian Splatting
di: Gao, Zhongpai, et al.
Pubblicazione: (2025) -
6DGS: Enhanced Direction-Aware Gaussian Splatting for Volumetric Rendering
di: Gao, Zhongpai, et al.
Pubblicazione: (2024) -
3D Vision-Language Gaussian Splatting
di: Peng, Qucheng, et al.
Pubblicazione: (2024)