MoXaRt: Audio-Visual Object-Guided Sound Interaction for XR
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Tianyu, Kim, Sieun, Zheng, Qianhui, Xu, Ruoyu, Ravi, Tejasvi, Kulkarni, Anuva, Passarella-Ward, Katrina, Zhu, Junyi, Kowdle, Adarsh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Enhancing XR Auditory Realism via Multimodal Scene-Aware Acoustic Rendering
von: Xu, Tianyu, et al.
Veröffentlicht: (2025)
von: Xu, Tianyu, et al.
Veröffentlicht: (2025)
Listen to the Unexpected: Self-Supervised Surprise Detection for Efficient Viewport Prediction
von: Khah, Arman Nik, et al.
Veröffentlicht: (2026)
von: Khah, Arman Nik, et al.
Veröffentlicht: (2026)
Widening the Role of Group Recommender Systems with CAJO
von: Ricci, Francesco, et al.
Veröffentlicht: (2025)
von: Ricci, Francesco, et al.
Veröffentlicht: (2025)
Masked Contrastive Pre-Training Improves Music Audio Key Detection
von: Yonay, Ori, et al.
Veröffentlicht: (2026)
von: Yonay, Ori, et al.
Veröffentlicht: (2026)
FaceValue: Exploring Real-Time Self-View Overlays to Prompt Meaning-Oriented Self-Awareness in Remote Meetings
von: Park, Gun Woo Warren, et al.
Veröffentlicht: (2026)
von: Park, Gun Woo Warren, et al.
Veröffentlicht: (2026)
GroundLink: Exploring How Contextual Meeting Snippets Can Close Common Ground Gaps in Editing 3D Scenes for Virtual Production
von: Woo, Gun, et al.
Veröffentlicht: (2026)
von: Woo, Gun, et al.
Veröffentlicht: (2026)
OOPrompt: Reifying Intents into Structured Artifacts for Modular and Iterative Prompting
von: Xu, Tengyou, et al.
Veröffentlicht: (2026)
von: Xu, Tengyou, et al.
Veröffentlicht: (2026)
Distorted Perspectives of LLM-Simulated Preferences: Can AI Mislead Design?
von: Kuric, Eduard, et al.
Veröffentlicht: (2026)
von: Kuric, Eduard, et al.
Veröffentlicht: (2026)
Exploring Diagnostic Prompting Approach for Multimodal LLM-based Visual Complexity Assessment: A Case Study of Amazon Search Result Pages
von: Murtadak, Divendar, et al.
Veröffentlicht: (2025)
von: Murtadak, Divendar, et al.
Veröffentlicht: (2025)
SonoHaptics: An Audio-Haptic Cursor for Gaze-Based Object Selection in XR
von: Cho, Hyunsung, et al.
Veröffentlicht: (2024)
von: Cho, Hyunsung, et al.
Veröffentlicht: (2024)
Conversational Forecasting Across Large Human Groups Using A Swarm of Surrogate AI Agents
von: Rosenberg, Louis, et al.
Veröffentlicht: (2026)
von: Rosenberg, Louis, et al.
Veröffentlicht: (2026)
Breaking Negative Cycles: A Reflection-To-Action System For Adaptive Change
von: Kim, Minsol Michelle, et al.
Veröffentlicht: (2026)
von: Kim, Minsol Michelle, et al.
Veröffentlicht: (2026)
AiGet: Transforming Everyday Moments into Hidden Knowledge Discovery with AI Assistance on Smart Glasses
von: Cai, Runze, et al.
Veröffentlicht: (2025)
von: Cai, Runze, et al.
Veröffentlicht: (2025)
Conversational Swarms of Humans and AI Agents enable Hybrid Collaborative Decision-making
von: Rosenberg, Louis, et al.
Veröffentlicht: (2024)
von: Rosenberg, Louis, et al.
Veröffentlicht: (2024)
Self-Improvement for Audio Large Language Model using Unlabeled Speech
von: Wang, Shaowen, et al.
Veröffentlicht: (2025)
von: Wang, Shaowen, et al.
Veröffentlicht: (2025)
Enhanced DareFightingICE Competitions: Sound Design and AI Competitions
von: Khan, Ibrahim, et al.
Veröffentlicht: (2024)
von: Khan, Ibrahim, et al.
Veröffentlicht: (2024)
ScaleMAP: Preserving Local Density and Neighborhood Structure in Low-Dimensional Embeddings
von: Poorna, Rajas, et al.
Veröffentlicht: (2026)
von: Poorna, Rajas, et al.
Veröffentlicht: (2026)
Step-Audio-R1 Technical Report
von: Tian, Fei, et al.
Veröffentlicht: (2025)
von: Tian, Fei, et al.
Veröffentlicht: (2025)
The Spatial and Temporal Resolution of Motor Intention in Multi-Target Prediction
von: Schmidt, Marie Dominique, et al.
Veröffentlicht: (2026)
von: Schmidt, Marie Dominique, et al.
Veröffentlicht: (2026)
AudioMiXR: Spatial Audio Object Manipulation with 6DoF for Sound Design in Augmented Reality
von: Woodard, Brandon, et al.
Veröffentlicht: (2025)
von: Woodard, Brandon, et al.
Veröffentlicht: (2025)
What Would GPT Click: Practical Effects of Human-AI Behavioral Misalignment and the Cost of Synthetic Participants in User Experience
von: Kuric, Eduard, et al.
Veröffentlicht: (2026)
von: Kuric, Eduard, et al.
Veröffentlicht: (2026)
Distilled HuBERT for Mobile Speech Emotion Recognition: A Cross-Corpus Validation Study
von: Ismail, Saifelden M.
Veröffentlicht: (2025)
von: Ismail, Saifelden M.
Veröffentlicht: (2025)
OrganicHAR: Towards Activity Discovery in Organic Settings for Privacy Preserving Sensors Using Efficient Video Analysis
von: Patidar, Prasoon, et al.
Veröffentlicht: (2026)
von: Patidar, Prasoon, et al.
Veröffentlicht: (2026)
Adaptive Background Music for a Fighting Game: A Multi-Instrument Volume Modulation Approach
von: Khan, Ibrahim, et al.
Veröffentlicht: (2023)
von: Khan, Ibrahim, et al.
Veröffentlicht: (2023)
Fighting Game Adaptive Background Music for Improved Gameplay
von: Khan, Ibrahim, et al.
Veröffentlicht: (2024)
von: Khan, Ibrahim, et al.
Veröffentlicht: (2024)
QoSGMAA: A Robust Multi-Order Graph Attention and Adversarial Framework for Sparse QoS Prediction
von: Du, Guanchen, et al.
Veröffentlicht: (2025)
von: Du, Guanchen, et al.
Veröffentlicht: (2025)
PromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching
von: Vosoughi, Ali, et al.
Veröffentlicht: (2025)
von: Vosoughi, Ali, et al.
Veröffentlicht: (2025)
Domain Specific Data Distillation and Multi-modal Embedding Generation
von: Peddiraju, Sharadind, et al.
Veröffentlicht: (2024)
von: Peddiraju, Sharadind, et al.
Veröffentlicht: (2024)
Human-Robot Creative Interactions (HRCI): Exploring Creativity in Artificial Agents Using a Story-Telling Game
von: Sandoval, Eduardo Benitez, et al.
Veröffentlicht: (2022)
von: Sandoval, Eduardo Benitez, et al.
Veröffentlicht: (2022)
GenComUI: Exploring Generative Visual Aids as Medium to Support Task-Oriented Human-Robot Communication
von: Ge, Yate, et al.
Veröffentlicht: (2025)
von: Ge, Yate, et al.
Veröffentlicht: (2025)
OBHS: An Optimized Block Huffman Scheme for Real-Time Audio Compression
von: Mahfi, Muntahi Safwan, et al.
Veröffentlicht: (2025)
von: Mahfi, Muntahi Safwan, et al.
Veröffentlicht: (2025)
Audio Foundation Models Outperform Symbolic Representations for Piano Performance Evaluation
von: Dhiman, Jai
Veröffentlicht: (2026)
von: Dhiman, Jai
Veröffentlicht: (2026)
From Framework to Reliable Practice: End-User Perspectives on Social Robots in Public Spaces
von: Oruma, Samson, et al.
Veröffentlicht: (2025)
von: Oruma, Samson, et al.
Veröffentlicht: (2025)
Domain Adaptation of the Pyannote Diarization Pipeline for Conversational Indonesian Audio
von: Prasetyo, Muhammad Daffa'i Rafi, et al.
Veröffentlicht: (2026)
von: Prasetyo, Muhammad Daffa'i Rafi, et al.
Veröffentlicht: (2026)
Dynamic Theater: Location-Based Immersive Dance Theater, Investigating User Guidance and Experience
von: Kim, You-Jin, et al.
Veröffentlicht: (2025)
von: Kim, You-Jin, et al.
Veröffentlicht: (2025)
Beyond Reality: Designing Personal Experiences and Interactive Narratives in AR Theater
von: Kim, You-Jin
Veröffentlicht: (2025)
von: Kim, You-Jin
Veröffentlicht: (2025)
Quantum-Enhanced Analysis and Grading of Vocal Performance
von: Agarwal, Rohan
Veröffentlicht: (2025)
von: Agarwal, Rohan
Veröffentlicht: (2025)
ParaNoise-SV: Integrated Approach for Noise-Robust Speaker Verification with Parallel Joint Learning of Speech Enhancement and Noise Extraction
von: Kim, Minu, et al.
Veröffentlicht: (2025)
von: Kim, Minu, et al.
Veröffentlicht: (2025)
Augmented Assembly: Object Recognition and Hand Tracking for Adaptive Assembly Instructions in Augmented Reality
von: Kyaw, Alexander Htet, et al.
Veröffentlicht: (2025)
von: Kyaw, Alexander Htet, et al.
Veröffentlicht: (2025)
Exploring Hierarchical Classification Performance for Time Series Data: Dissimilarity Measures and Classifier Comparisons
von: Alagoz, Celal
Veröffentlicht: (2024)
von: Alagoz, Celal
Veröffentlicht: (2024)
Ähnliche Einträge
-
Enhancing XR Auditory Realism via Multimodal Scene-Aware Acoustic Rendering
von: Xu, Tianyu, et al.
Veröffentlicht: (2025) -
Listen to the Unexpected: Self-Supervised Surprise Detection for Efficient Viewport Prediction
von: Khah, Arman Nik, et al.
Veröffentlicht: (2026) -
Widening the Role of Group Recommender Systems with CAJO
von: Ricci, Francesco, et al.
Veröffentlicht: (2025) -
Masked Contrastive Pre-Training Improves Music Audio Key Detection
von: Yonay, Ori, et al.
Veröffentlicht: (2026) -
FaceValue: Exploring Real-Time Self-View Overlays to Prompt Meaning-Oriented Self-Awareness in Remote Meetings
von: Park, Gun Woo Warren, et al.
Veröffentlicht: (2026)