YETI (YET to Intervene) Proactive Interventions by Multimodal AI Agents in Augmented Reality Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Bandyopadhyay, Saptarashmi, Bahirwani, Vikas, Aggarwal, Lavisha, Guda, Bhanu, Li, Lin, Colaco, Andrea |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
by: Raoufi, Behnam, et al.
Published: (2025)
by: Raoufi, Behnam, et al.
Published: (2025)
CCVA-FL: Cross-Client Variations Adaptive Federated Learning for Medical Imaging
by: Gupta, Sunny, et al.
Published: (2024)
by: Gupta, Sunny, et al.
Published: (2024)
Taming the Tail: Leveraging Asymmetric Loss and Pade Approximation to Overcome Medical Image Long-Tailed Class Imbalance
by: Kashyap, Pankhi, et al.
Published: (2024)
by: Kashyap, Pankhi, et al.
Published: (2024)
Ultra-Reduced-Impact-Encased-Logging (URIEL): propose a new method for selective sustainable logging and post-harvest silvicultural treatment in tropical forest using airborne robotics systems
by: Albiero, Daniel, et al.
Published: (2026)
by: Albiero, Daniel, et al.
Published: (2026)
Edge-Enabled Collaborative Object Detection for Real-Time Multi-Vehicle Perception
by: Richards, Everett, et al.
Published: (2025)
by: Richards, Everett, et al.
Published: (2025)
LLM-Guided Agentic Floor Plan Parsing for Accessible Indoor Navigation of Blind and Low-Vision People
by: Ayanzadeh, Aydin, et al.
Published: (2026)
by: Ayanzadeh, Aydin, et al.
Published: (2026)
Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention
by: Durrani, Hamza Ahmed, et al.
Published: (2026)
by: Durrani, Hamza Ahmed, et al.
Published: (2026)
FEDTAIL: Federated Long-Tailed Domain Generalization with Sharpness-Guided Gradient Matching
by: Gupta, Sunny, et al.
Published: (2025)
by: Gupta, Sunny, et al.
Published: (2025)
FedStein: Enhancing Multi-Domain Federated Learning Through James-Stein Estimator
by: Gupta, Sunny, et al.
Published: (2024)
by: Gupta, Sunny, et al.
Published: (2024)
UniVarFL: Uniformity and Variance Regularized Federated Learning for Heterogeneous Data
by: Gupta, Sunny, et al.
Published: (2025)
by: Gupta, Sunny, et al.
Published: (2025)
FedAlign: Federated Domain Generalization with Cross-Client Feature Alignment
by: Gupta, Sunny, et al.
Published: (2025)
by: Gupta, Sunny, et al.
Published: (2025)
OmniAcc: Personalized Accessibility Assistant Using Generative AI
by: Karki, Siddhant, et al.
Published: (2025)
by: Karki, Siddhant, et al.
Published: (2025)
WaveMix: A Resource-efficient Neural Network for Image Analysis
by: Jeevan, Pranav, et al.
Published: (2022)
by: Jeevan, Pranav, et al.
Published: (2022)
Which Backbone to Use: A Resource-efficient Domain Specific Comparison for Computer Vision
by: Jeevan, Pranav, et al.
Published: (2024)
by: Jeevan, Pranav, et al.
Published: (2024)
Making High-Level AI Design Decisions Explicit Using a Binary Stream System-Designation Approach
by: Mossbridge, Julia
Published: (2024)
by: Mossbridge, Julia
Published: (2024)
CoMoCAVs: Cohesive Decision-Guided Motion Planning for Connected and Autonomous Vehicles with Multi-Policy Reinforcement Learning
by: Hu, Pan
Published: (2025)
by: Hu, Pan
Published: (2025)
ChargingBoul: A Competitive Negotiating Agent with Novel Opponent Modeling
by: Shymanski, Joe
Published: (2025)
by: Shymanski, Joe
Published: (2025)
Context Engineering: From Prompts to Corporate Multi-Agent Architecture
by: Vishnyakova, Vera V.
Published: (2026)
by: Vishnyakova, Vera V.
Published: (2026)
Safe Road-Crossing by Autonomous Wheelchairs: a Novel Dataset and its Experimental Evaluation
by: Grigioni, Carlo, et al.
Published: (2024)
by: Grigioni, Carlo, et al.
Published: (2024)
FeedbackSTS-Det: Sparse Frames-Based Spatio-Temporal Semantic Feedback Network for Moving Infrared Small Target Detection
by: Huang, Yian, et al.
Published: (2026)
by: Huang, Yian, et al.
Published: (2026)
Habitat Classification from Ground-Level Imagery Using Deep Neural Networks
by: Shi, Hongrui, et al.
Published: (2025)
by: Shi, Hongrui, et al.
Published: (2025)
SERA-H: Beyond Native Sentinel Spatial Limits for High-Resolution Canopy Height Mapping
by: Boudras, Thomas, et al.
Published: (2025)
by: Boudras, Thomas, et al.
Published: (2025)
FlexDoc: Parameterized Sampling for Diverse Multilingual Synthetic Documents for Training Document Understanding Models
by: Dua, Karan, et al.
Published: (2025)
by: Dua, Karan, et al.
Published: (2025)
MaSC: A Masked Similarity Metric for Evaluating Concept-Driven Generation
by: Bartkowiak, Patryk, et al.
Published: (2026)
by: Bartkowiak, Patryk, et al.
Published: (2026)
Evaluating the Impact of Synthetic Data on Object Detection Tasks in Autonomous Driving
by: Özeren, Enes, et al.
Published: (2025)
by: Özeren, Enes, et al.
Published: (2025)
MARS: Multi-Agent Robotic System with Multimodal Large Language Models for Assistive Intelligence
by: Gao, Renjun
Published: (2025)
by: Gao, Renjun
Published: (2025)
Leum-VL Technical Report
by: He, Yuxuan, et al.
Published: (2026)
by: He, Yuxuan, et al.
Published: (2026)
FSFM: A Biologically-Inspired Framework for Selective Forgetting of Agent Memory
by: Gu, Yingjie, et al.
Published: (2026)
by: Gu, Yingjie, et al.
Published: (2026)
HATL: Hierarchical Adaptive-Transfer Learning Framework for Sign Language Machine Translation
by: Shahin, Nada, et al.
Published: (2026)
by: Shahin, Nada, et al.
Published: (2026)
FLD+: Data-efficient Evaluation Metric for Generative Models
by: Jeevan, Pranav, et al.
Published: (2024)
by: Jeevan, Pranav, et al.
Published: (2024)
WaveMixSR-V2: Enhancing Super-resolution with Higher Efficiency
by: Jeevan, Pranav, et al.
Published: (2024)
by: Jeevan, Pranav, et al.
Published: (2024)
Normalizing Flow-Based Metric for Image Generation
by: Jeevan, Pranav, et al.
Published: (2024)
by: Jeevan, Pranav, et al.
Published: (2024)
GLoT: A Novel Gated-Logarithmic Transformer for Efficient Sign Language Translation
by: Shahin, Nada, et al.
Published: (2025)
by: Shahin, Nada, et al.
Published: (2025)
Knowledge Equivalence in Digital Twins of Intelligent Systems
by: Zhang, Nan, et al.
Published: (2022)
by: Zhang, Nan, et al.
Published: (2022)
T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models
by: Chen, Yiteng, et al.
Published: (2025)
by: Chen, Yiteng, et al.
Published: (2025)
Efficient Temporally-Aware DeepFake Detection using H.264 Motion Vectors
by: Grönquist, Peter, et al.
Published: (2023)
by: Grönquist, Peter, et al.
Published: (2023)
Systematic Comparison of Projection Methods for Monocular 3D Human Pose Estimation on Fisheye Images
by: Käs, Stephanie, et al.
Published: (2025)
by: Käs, Stephanie, et al.
Published: (2025)
Distributed Intelligent System Architecture for UAV-Assisted Monitoring of Wind Energy Infrastructure
by: Svystun, Serhii, et al.
Published: (2024)
by: Svystun, Serhii, et al.
Published: (2024)
Is Single-View Mesh Reconstruction Ready for Robotics?
by: Nolte, Frederik, et al.
Published: (2025)
by: Nolte, Frederik, et al.
Published: (2025)
Semi supervised GAN for smart microscopy, fast and data efficient cell cycle classification
by: Manick, Rajeev, et al.
Published: (2026)
by: Manick, Rajeev, et al.
Published: (2026)
Similar Items
-
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
by: Raoufi, Behnam, et al.
Published: (2025) -
CCVA-FL: Cross-Client Variations Adaptive Federated Learning for Medical Imaging
by: Gupta, Sunny, et al.
Published: (2024) -
Taming the Tail: Leveraging Asymmetric Loss and Pade Approximation to Overcome Medical Image Long-Tailed Class Imbalance
by: Kashyap, Pankhi, et al.
Published: (2024) -
Ultra-Reduced-Impact-Encased-Logging (URIEL): propose a new method for selective sustainable logging and post-harvest silvicultural treatment in tropical forest using airborne robotics systems
by: Albiero, Daniel, et al.
Published: (2026) -
Edge-Enabled Collaborative Object Detection for Real-Time Multi-Vehicle Perception
by: Richards, Everett, et al.
Published: (2025)