Action2Sound: Ambient-Aware Generation of Action Sounds from Egocentric Videos
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Changan, Peng, Puyuan, Baid, Ami, Xue, Zihui, Hsu, Wei-Ning, Harwath, David, Grauman, Kristen |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SoundingActions: Learning How Actions Sound from Narrated Egocentric Videos
di: Chen, Changan, et al.
Pubblicazione: (2024)
di: Chen, Changan, et al.
Pubblicazione: (2024)
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
di: Peng, Puyuan, et al.
Pubblicazione: (2025)
di: Peng, Puyuan, et al.
Pubblicazione: (2025)
BAT: Learning to Reason about Spatial Sounds with Large Language Models
di: Zheng, Zhisheng, et al.
Pubblicazione: (2024)
di: Zheng, Zhisheng, et al.
Pubblicazione: (2024)
Learning Spatial Features from Audio-Visual Correspondence in Egocentric Videos
di: Majumder, Sagnik, et al.
Pubblicazione: (2023)
di: Majumder, Sagnik, et al.
Pubblicazione: (2023)
Probing the Robustness Properties of Neural Speech Codecs
di: Tseng, Wei-Cheng, et al.
Pubblicazione: (2025)
di: Tseng, Wei-Cheng, et al.
Pubblicazione: (2025)
Sim2Real Transfer for Audio-Visual Navigation with Frequency-Adaptive Acoustic Field Prediction
di: Chen, Changan, et al.
Pubblicazione: (2024)
di: Chen, Changan, et al.
Pubblicazione: (2024)
ActiveRIR: Active Audio-Visual Exploration for Acoustic Environment Modeling
di: Somayazulu, Arjun, et al.
Pubblicazione: (2024)
di: Somayazulu, Arjun, et al.
Pubblicazione: (2024)
Segmenting Collision Sound Sources in Egocentric Videos
di: Parida, Kranti Kumar, et al.
Pubblicazione: (2025)
di: Parida, Kranti Kumar, et al.
Pubblicazione: (2025)
Sound Event Detection with Boundary-Aware Optimization and Inference
di: Schmid, Florian, et al.
Pubblicazione: (2026)
di: Schmid, Florian, et al.
Pubblicazione: (2026)
Epic-Sounds: A Large-scale Dataset of Actions That Sound
di: Huh, Jaesung, et al.
Pubblicazione: (2023)
di: Huh, Jaesung, et al.
Pubblicazione: (2023)
Leveraging Sound Source Trajectories for Universal Sound Separation
di: Wu, Donghang, et al.
Pubblicazione: (2024)
di: Wu, Donghang, et al.
Pubblicazione: (2024)
Sound Zone Control Robust To Sound Speed Change
di: Bhattacharjee, Sankha Subhra, et al.
Pubblicazione: (2024)
di: Bhattacharjee, Sankha Subhra, et al.
Pubblicazione: (2024)
SoundLoCD: An Efficient Conditional Discrete Contrastive Latent Diffusion Model for Text-to-Sound Generation
di: Niu, Xinlei, et al.
Pubblicazione: (2024)
di: Niu, Xinlei, et al.
Pubblicazione: (2024)
Evaluating Sound Similarity Metrics for Differentiable, Iterative Sound-Matching
di: Salimi, Amir, et al.
Pubblicazione: (2025)
di: Salimi, Amir, et al.
Pubblicazione: (2025)
Boundary-Informed Sound Field Reconstruction
di: Sundström, David, et al.
Pubblicazione: (2025)
di: Sundström, David, et al.
Pubblicazione: (2025)
VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
di: Zheng, Zhisheng, et al.
Pubblicazione: (2025)
di: Zheng, Zhisheng, et al.
Pubblicazione: (2025)
Exploring Self-Supervised Audio Models for Generalized Anomalous Sound Detection
di: Han, Bing, et al.
Pubblicazione: (2025)
di: Han, Bing, et al.
Pubblicazione: (2025)
Listen and Move: Improving GANs Coherency in Agnostic Sound-to-Video Generation
di: Redondo, Rafael
Pubblicazione: (2024)
di: Redondo, Rafael
Pubblicazione: (2024)
Sounding Out Reconstruction Error-Based Evaluation of Generative Models of Expressive Performance
di: Peter, Silvan David, et al.
Pubblicazione: (2023)
di: Peter, Silvan David, et al.
Pubblicazione: (2023)
Measuring Sound Symbolism in Audio-visual Models
di: Tseng, Wei-Cheng, et al.
Pubblicazione: (2024)
di: Tseng, Wei-Cheng, et al.
Pubblicazione: (2024)
Noise-Robust Sound Event Detection and Counting via Language-Queried Sound Separation
di: Chen, Yuanjian, et al.
Pubblicazione: (2025)
di: Chen, Yuanjian, et al.
Pubblicazione: (2025)
DiffSound: Differentiable Modal Sound Rendering and Inverse Rendering for Diverse Inference Tasks
di: Jin, Xutong, et al.
Pubblicazione: (2024)
di: Jin, Xutong, et al.
Pubblicazione: (2024)
Retaining Mixture Representations for Domain Generalized Anomalous Sound Detection
di: Saengthong, Phurich, et al.
Pubblicazione: (2025)
di: Saengthong, Phurich, et al.
Pubblicazione: (2025)
Fractional Fourier Sound Synthesis
di: Gutiérrez, Esteban, et al.
Pubblicazione: (2025)
di: Gutiérrez, Esteban, et al.
Pubblicazione: (2025)
Diffuse Sound Field Synthesis
di: Zotter, Franz, et al.
Pubblicazione: (2024)
di: Zotter, Franz, et al.
Pubblicazione: (2024)
Sound Event Bounding Boxes
di: Ebbers, Janek, et al.
Pubblicazione: (2024)
di: Ebbers, Janek, et al.
Pubblicazione: (2024)
Disentangling Dual-Encoder Masked Autoencoder for Respiratory Sound Classification
di: Wei, Peidong, et al.
Pubblicazione: (2025)
di: Wei, Peidong, et al.
Pubblicazione: (2025)
Context-Aware Query Refinement for Target Sound Extraction: Handling Partially Matched Queries
di: Sato, Ryo, et al.
Pubblicazione: (2025)
di: Sato, Ryo, et al.
Pubblicazione: (2025)
SoundBeam meets M2D: Target Sound Extraction with Audio Foundation Model
di: Hernandez-Olivan, Carlos, et al.
Pubblicazione: (2024)
di: Hernandez-Olivan, Carlos, et al.
Pubblicazione: (2024)
Hierarchical Pooling Structure for Weakly Labeled Sound Event Detection
di: He, Ke-Xin, et al.
Pubblicazione: (2019)
di: He, Ke-Xin, et al.
Pubblicazione: (2019)
Sound Field Synthesis with Acoustic Waves
di: Mansour, Mohamed F.
Pubblicazione: (2024)
di: Mansour, Mohamed F.
Pubblicazione: (2024)
Fast Algorithm for Moving Sound Source
di: Yang, Dong
Pubblicazione: (2025)
di: Yang, Dong
Pubblicazione: (2025)
Joint Analysis of Acoustic Scenes and Sound Events Based on Semi-Supervised Training of Sound Events With Partial Labels
di: Imoto, Keisuke
Pubblicazione: (2025)
di: Imoto, Keisuke
Pubblicazione: (2025)
Language-Queried Target Sound Extraction Without Parallel Training Data
di: Ma, Hao, et al.
Pubblicazione: (2024)
di: Ma, Hao, et al.
Pubblicazione: (2024)
Codec-SUPERB: An In-Depth Analysis of Sound Codec Models
di: Wu, Haibin, et al.
Pubblicazione: (2024)
di: Wu, Haibin, et al.
Pubblicazione: (2024)
DiveSound: LLM-Assisted Automatic Taxonomy Construction for Diverse Audio Generation
di: Li, Baihan, et al.
Pubblicazione: (2024)
di: Li, Baihan, et al.
Pubblicazione: (2024)
MambaFoley: Foley Sound Generation using Selective State-Space Models
di: Colombo, Marco Furio, et al.
Pubblicazione: (2024)
di: Colombo, Marco Furio, et al.
Pubblicazione: (2024)
Generating Diverse Audio-Visual 360 Soundscapes for Sound Event Localization and Detection
di: Roman, Adrian S., et al.
Pubblicazione: (2025)
di: Roman, Adrian S., et al.
Pubblicazione: (2025)
Listen through the Sound: Generative Speech Restoration Leveraging Acoustic Context Representation
di: Chung, Soo-Whan, et al.
Pubblicazione: (2025)
di: Chung, Soo-Whan, et al.
Pubblicazione: (2025)
SoundReactor: Frame-level Online Video-to-Audio Generation
di: Saito, Koichi, et al.
Pubblicazione: (2025)
di: Saito, Koichi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
SoundingActions: Learning How Actions Sound from Narrated Egocentric Videos
di: Chen, Changan, et al.
Pubblicazione: (2024) -
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
di: Peng, Puyuan, et al.
Pubblicazione: (2025) -
BAT: Learning to Reason about Spatial Sounds with Large Language Models
di: Zheng, Zhisheng, et al.
Pubblicazione: (2024) -
Learning Spatial Features from Audio-Visual Correspondence in Egocentric Videos
di: Majumder, Sagnik, et al.
Pubblicazione: (2023) -
Probing the Robustness Properties of Neural Speech Codecs
di: Tseng, Wei-Cheng, et al.
Pubblicazione: (2025)