ManiWAV: Learning Robot Manipulation from In-the-Wild Audio-Visual Data
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Zeyi, Chi, Cheng, Cousineau, Eric, Kuppuswamy, Naveen, Burchfiel, Benjamin, Song, Shuran |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Single-Microphone-Based Sound Source Localization for Mobile Robots in Reverberant Environments
by: Wang, Jiang, et al.
Published: (2025)
by: Wang, Jiang, et al.
Published: (2025)
Audio Array-Based 3D UAV Trajectory Estimation with LiDAR Pseudo-Labeling
by: Lei, Allen, et al.
Published: (2024)
by: Lei, Allen, et al.
Published: (2024)
Long-Term, Store-Front Robotics: Interactive Music for Robotic Arm, Caxixi and Frame Drums
by: Savery, Richard, et al.
Published: (2024)
by: Savery, Richard, et al.
Published: (2024)
Evaluating Speech-in-Speech Perception via a Humanoid Robot
by: Meyer, Luke, et al.
Published: (2023)
by: Meyer, Luke, et al.
Published: (2023)
Theoretical Framework for the Optimization of Microphone Array Configuration for Humanoid Robot Audition
by: Tourbabin, Vladimir, et al.
Published: (2024)
by: Tourbabin, Vladimir, et al.
Published: (2024)
Direction of Arrival Estimation Using Microphone Array Processing for Moving Humanoid Robots
by: Tourbabin, Vladimir, et al.
Published: (2024)
by: Tourbabin, Vladimir, et al.
Published: (2024)
Accelerating Audio Research with Robotic Dummy Heads
by: Lu, Austin, et al.
Published: (2025)
by: Lu, Austin, et al.
Published: (2025)
ADD 2023: Towards Audio Deepfake Detection and Analysis in the Wild
by: Yi, Jiangyan, et al.
Published: (2024)
by: Yi, Jiangyan, et al.
Published: (2024)
DialogueSidon: Recovering Full-Duplex Dialogue Tracks from In-the-Wild Dialogue Audio
by: Nakata, Wataru, et al.
Published: (2026)
by: Nakata, Wataru, et al.
Published: (2026)
Disentangled Acoustic Fields For Multimodal Physical Scene Understanding
by: Yin, Jie, et al.
Published: (2024)
by: Yin, Jie, et al.
Published: (2024)
Music to Dance as Language Translation using Sequence Models
by: Correia, André, et al.
Published: (2024)
by: Correia, André, et al.
Published: (2024)
SonicBoom: Contact Localization Using Array of Microphones
by: Lee, Moonyoung, et al.
Published: (2024)
by: Lee, Moonyoung, et al.
Published: (2024)
Active Listener: Continuous Generation of Listener's Head Motion Response in Dyadic Interactions
by: Ghosh, Bishal, et al.
Published: (2024)
by: Ghosh, Bishal, et al.
Published: (2024)
On Adversarial Attacks In Acoustic Drone Localization
by: Shor, Tamir, et al.
Published: (2025)
by: Shor, Tamir, et al.
Published: (2025)
An Efficient GPU-based Implementation for Noise Robust Sound Source Localization
by: Lin, Zirui, et al.
Published: (2025)
by: Lin, Zirui, et al.
Published: (2025)
Human-mimetic binaural ear design and sound source direction estimation for task realization of musculoskeletal humanoids
by: Omura, Yusuke, et al.
Published: (2024)
by: Omura, Yusuke, et al.
Published: (2024)
AIM: Acoustic Inertial Measurement for Indoor Drone Localization and Tracking
by: Sun, Yimiao, et al.
Published: (2025)
by: Sun, Yimiao, et al.
Published: (2025)
Soft Acoustic Curvature Sensor: Design and Development
by: Sofla, Mohammad Sheikh, et al.
Published: (2024)
by: Sofla, Mohammad Sheikh, et al.
Published: (2024)
Uncovering the Visual Contribution in Audio-Visual Speech Recognition
by: Lin, Zhaofeng, et al.
Published: (2024)
by: Lin, Zhaofeng, et al.
Published: (2024)
Pengi: An Audio Language Model for Audio Tasks
by: Deshmukh, Soham, et al.
Published: (2023)
by: Deshmukh, Soham, et al.
Published: (2023)
Leveraging Mamba with Full-Face Vision for Audio-Visual Speech Enhancement
by: Chao, Rong, et al.
Published: (2025)
by: Chao, Rong, et al.
Published: (2025)
Online Audio-Visual Autoregressive Speaker Extraction
by: Pan, Zexu, et al.
Published: (2025)
by: Pan, Zexu, et al.
Published: (2025)
High-Fidelity Generative Audio Compression at 0.275kbps
by: Ma, Hao, et al.
Published: (2026)
by: Ma, Hao, et al.
Published: (2026)
SmoothSync: Dual-Stream Diffusion Transformers for Jitter-Robust Beat-Synchronized Gesture Generation from Quantized Audio
by: Jiang, Yujiao, et al.
Published: (2026)
by: Jiang, Yujiao, et al.
Published: (2026)
ALDAS: Audio-Linguistic Data Augmentation for Spoofed Audio Detection
by: Khanjani, Zahra, et al.
Published: (2024)
by: Khanjani, Zahra, et al.
Published: (2024)
ActiveRIR: Active Audio-Visual Exploration for Acoustic Environment Modeling
by: Somayazulu, Arjun, et al.
Published: (2024)
by: Somayazulu, Arjun, et al.
Published: (2024)
Audio Entailment: Assessing Deductive Reasoning for Audio Understanding
by: Deshmukh, Soham, et al.
Published: (2024)
by: Deshmukh, Soham, et al.
Published: (2024)
PANDORA: Diffusion Policy Learning for Dexterous Robotic Piano Playing
by: Huang, Yanjia, et al.
Published: (2025)
by: Huang, Yanjia, et al.
Published: (2025)
Audio Texture Manipulation by Exemplar-Based Analogy
by: Cheng, Kan Jen, et al.
Published: (2025)
by: Cheng, Kan Jen, et al.
Published: (2025)
AudioEval: Automatic Dual-Perspective and Multi-Dimensional Evaluation of Text-to-Audio-Generation
by: Wang, Hui, et al.
Published: (2025)
by: Wang, Hui, et al.
Published: (2025)
PAM: Prompting Audio-Language Models for Audio Quality Assessment
by: Deshmukh, Soham, et al.
Published: (2024)
by: Deshmukh, Soham, et al.
Published: (2024)
A Review on Sound Source Localization in Robotics: Focusing on Deep Learning Methods
by: Jalayer, Reza, et al.
Published: (2025)
by: Jalayer, Reza, et al.
Published: (2025)
AVR: Synergizing Foundation Models for Audio-Visual Humor Detection
by: Sharma, Sarthak, et al.
Published: (2024)
by: Sharma, Sarthak, et al.
Published: (2024)
MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models
by: Gong, Yitian, et al.
Published: (2026)
by: Gong, Yitian, et al.
Published: (2026)
Natural Language Supervision for General-Purpose Audio Representations
by: Elizalde, Benjamin, et al.
Published: (2023)
by: Elizalde, Benjamin, et al.
Published: (2023)
AudioLCM: Text-to-Audio Generation with Latent Consistency Models
by: Liu, Huadai, et al.
Published: (2024)
by: Liu, Huadai, et al.
Published: (2024)
Sim2Real Transfer for Audio-Visual Navigation with Frequency-Adaptive Acoustic Field Prediction
by: Chen, Changan, et al.
Published: (2024)
by: Chen, Changan, et al.
Published: (2024)
A Fast and Lightweight Model for Causal Audio-Visual Speech Separation
by: Sang, Wendi, et al.
Published: (2025)
by: Sang, Wendi, et al.
Published: (2025)
Audio-Visual Target Speaker Extraction with Reverse Selective Auditory Attention
by: Tao, Ruijie, et al.
Published: (2024)
by: Tao, Ruijie, et al.
Published: (2024)
BickGraphing: Web-Based Application for Visual Inspection of Audio Recordings
by: Seow, Kayley, et al.
Published: (2026)
by: Seow, Kayley, et al.
Published: (2026)
Similar Items
-
Single-Microphone-Based Sound Source Localization for Mobile Robots in Reverberant Environments
by: Wang, Jiang, et al.
Published: (2025) -
Audio Array-Based 3D UAV Trajectory Estimation with LiDAR Pseudo-Labeling
by: Lei, Allen, et al.
Published: (2024) -
Long-Term, Store-Front Robotics: Interactive Music for Robotic Arm, Caxixi and Frame Drums
by: Savery, Richard, et al.
Published: (2024) -
Evaluating Speech-in-Speech Perception via a Humanoid Robot
by: Meyer, Luke, et al.
Published: (2023) -
Theoretical Framework for the Optimization of Microphone Array Configuration for Humanoid Robot Audition
by: Tourbabin, Vladimir, et al.
Published: (2024)