Modeling and Driving Human Body Soundfields through Acoustic Primitives
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Chao, Markovic, Dejan, Xu, Chenliang, Richard, Alexander |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
High-Quality Visually-Guided Sound Separation from Diverse Categories
von: Huang, Chao, et al.
Veröffentlicht: (2023)
von: Huang, Chao, et al.
Veröffentlicht: (2023)
ZeroSep: Separate Anything in Audio with Zero Training
von: Huang, Chao, et al.
Veröffentlicht: (2025)
von: Huang, Chao, et al.
Veröffentlicht: (2025)
Learning to Highlight Audio by Watching Movies
von: Huang, Chao, et al.
Veröffentlicht: (2025)
von: Huang, Chao, et al.
Veröffentlicht: (2025)
SoundCam: A Dataset for Finding Humans Using Room Acoustics
von: Wang, Mason, et al.
Veröffentlicht: (2023)
von: Wang, Mason, et al.
Veröffentlicht: (2023)
Real Acoustic Fields: An Audio-Visual Room Acoustics Dataset and Benchmark
von: Chen, Ziyang, et al.
Veröffentlicht: (2024)
von: Chen, Ziyang, et al.
Veröffentlicht: (2024)
Decoding Emotions: Unveiling Facial Expressions through Acoustic Sensing with Contrastive Attention
von: Wang, Guangjing, et al.
Veröffentlicht: (2024)
von: Wang, Guangjing, et al.
Veröffentlicht: (2024)
Sonicmesh: Enhancing 3D Human Mesh Reconstruction in Vision-Impaired Environments With Acoustic Signals
von: Liang, Xiaoxuan, et al.
Veröffentlicht: (2024)
von: Liang, Xiaoxuan, et al.
Veröffentlicht: (2024)
Prosody-Enhanced Acoustic Pre-training and Acoustic-Disentangled Prosody Adapting for Movie Dubbing
von: Zhang, Zhedong, et al.
Veröffentlicht: (2025)
von: Zhang, Zhedong, et al.
Veröffentlicht: (2025)
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling
von: Pham, Long-Khanh, et al.
Veröffentlicht: (2025)
von: Pham, Long-Khanh, et al.
Veröffentlicht: (2025)
Improving Acoustic Scene Classification with City Features
von: Cai, Yiqiang, et al.
Veröffentlicht: (2025)
von: Cai, Yiqiang, et al.
Veröffentlicht: (2025)
SOAF: Scene Occlusion-aware Neural Acoustic Field
von: Gao, Huiyu, et al.
Veröffentlicht: (2024)
von: Gao, Huiyu, et al.
Veröffentlicht: (2024)
Few-shot Acoustic Synthesis with Multimodal Flow Matching
von: Brunetto, Amandine
Veröffentlicht: (2026)
von: Brunetto, Amandine
Veröffentlicht: (2026)
NBM: an Open Dataset for the Acoustic Monitoring of Nocturnal Migratory Birds in Europe
von: Airale, Louis, et al.
Veröffentlicht: (2024)
von: Airale, Louis, et al.
Veröffentlicht: (2024)
Novel-View Acoustic Synthesis from 3D Reconstructed Rooms
von: Ahn, Byeongjoo, et al.
Veröffentlicht: (2023)
von: Ahn, Byeongjoo, et al.
Veröffentlicht: (2023)
NeRAF: 3D Scene Infused Neural Radiance and Acoustic Fields
von: Brunetto, Amandine, et al.
Veröffentlicht: (2024)
von: Brunetto, Amandine, et al.
Veröffentlicht: (2024)
How Would It Sound? Material-Controlled Multimodal Acoustic Profile Generation for Indoor Scenes
von: Saad, Mahnoor Fatima, et al.
Veröffentlicht: (2025)
von: Saad, Mahnoor Fatima, et al.
Veröffentlicht: (2025)
RapVerse: Coherent Vocals and Whole-Body Motions Generations from Text
von: Chen, Jiaben, et al.
Veröffentlicht: (2024)
von: Chen, Jiaben, et al.
Veröffentlicht: (2024)
ActiveRIR: Active Audio-Visual Exploration for Acoustic Environment Modeling
von: Somayazulu, Arjun, et al.
Veröffentlicht: (2024)
von: Somayazulu, Arjun, et al.
Veröffentlicht: (2024)
Lodge: A Coarse to Fine Diffusion Network for Long Dance Generation Guided by the Characteristic Dance Primitives
von: Li, Ronghui, et al.
Veröffentlicht: (2024)
von: Li, Ronghui, et al.
Veröffentlicht: (2024)
Separate to Collaborate: Dual-Stream Diffusion Model for Coordinated Piano Hand Motion Synthesis
von: Liu, Zihao, et al.
Veröffentlicht: (2025)
von: Liu, Zihao, et al.
Veröffentlicht: (2025)
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction
von: Yang, Shu-wen, et al.
Veröffentlicht: (2025)
von: Yang, Shu-wen, et al.
Veröffentlicht: (2025)
Synchronized Video-to-Audio Generation via Mel Quantization-Continuum Decomposition
von: Wang, Juncheng, et al.
Veröffentlicht: (2025)
von: Wang, Juncheng, et al.
Veröffentlicht: (2025)
Spiking Structured State Space Model for Monaural Speech Enhancement
von: Du, Yu, et al.
Veröffentlicht: (2023)
von: Du, Yu, et al.
Veröffentlicht: (2023)
GCDance: Genre-Controlled Music-Driven 3D Full Body Dance Generation
von: Liu, Xinran, et al.
Veröffentlicht: (2025)
von: Liu, Xinran, et al.
Veröffentlicht: (2025)
AV-Surf: Surface-Enhanced Geometry-Aware Novel-View Acoustic Synthesis
von: Baek, Hadam, et al.
Veröffentlicht: (2025)
von: Baek, Hadam, et al.
Veröffentlicht: (2025)
ExpGest: Expressive Speaker Generation Using Diffusion Model and Hybrid Audio-Text Guidance
von: Cheng, Yongkang, et al.
Veröffentlicht: (2024)
von: Cheng, Yongkang, et al.
Veröffentlicht: (2024)
Accuracy enhancement method for speech emotion recognition from spectrogram using temporal frequency correlation and positional information learning through knowledge transfer
von: Kim, Jeong-Yoon, et al.
Veröffentlicht: (2024)
von: Kim, Jeong-Yoon, et al.
Veröffentlicht: (2024)
A-JEPA: Joint-Embedding Predictive Architecture Can Listen
von: Fei, Zhengcong, et al.
Veröffentlicht: (2023)
von: Fei, Zhengcong, et al.
Veröffentlicht: (2023)
FLUX that Plays Music
von: Fei, Zhengcong, et al.
Veröffentlicht: (2024)
von: Fei, Zhengcong, et al.
Veröffentlicht: (2024)
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing
von: Liu, Zehua, et al.
Veröffentlicht: (2025)
von: Liu, Zehua, et al.
Veröffentlicht: (2025)
Bidirectional Autoregressive Diffusion Model for Dance Generation
von: Zhang, Canyu, et al.
Veröffentlicht: (2024)
von: Zhang, Canyu, et al.
Veröffentlicht: (2024)
MMAudioReverbs: Video-Guided Acoustic Modeling for Dereverberation and Room Impulse Response Estimation
von: Takahashi, Akira, et al.
Veröffentlicht: (2026)
von: Takahashi, Akira, et al.
Veröffentlicht: (2026)
SoundLoc3D: Invisible 3D Sound Source Localization and Classification Using a Multimodal RGB-D Acoustic Camera
von: He, Yuhang, et al.
Veröffentlicht: (2024)
von: He, Yuhang, et al.
Veröffentlicht: (2024)
Music Genre Classification using Large Language Models
von: Meguenani, Mohamed El Amine, et al.
Veröffentlicht: (2024)
von: Meguenani, Mohamed El Amine, et al.
Veröffentlicht: (2024)
Input Conditioned Layer Dropping in Speech Foundation Models
von: Hannan, Abdul, et al.
Veröffentlicht: (2025)
von: Hannan, Abdul, et al.
Veröffentlicht: (2025)
Acoustic Scene Classification: A Competition Review
von: Gharib, Shayan, et al.
Veröffentlicht: (2018)
von: Gharib, Shayan, et al.
Veröffentlicht: (2018)
EmoTalker: Emotionally Editable Talking Face Generation via Diffusion Model
von: Zhang, Bingyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Bingyuan, et al.
Veröffentlicht: (2024)
Enhancing Dance-to-Music Generation via Negative Conditioning Latent Diffusion Model
von: Sun, Changchang, et al.
Veröffentlicht: (2025)
von: Sun, Changchang, et al.
Veröffentlicht: (2025)
Omni-AVSR: Towards Unified Multimodal Speech Recognition with Large Language Models
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
SSAVSV: Towards Unified Model for Self-Supervised Audio-Visual Speaker Verification
von: Rajasekhar, Gnana Praveen, et al.
Veröffentlicht: (2025)
von: Rajasekhar, Gnana Praveen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
High-Quality Visually-Guided Sound Separation from Diverse Categories
von: Huang, Chao, et al.
Veröffentlicht: (2023) -
ZeroSep: Separate Anything in Audio with Zero Training
von: Huang, Chao, et al.
Veröffentlicht: (2025) -
Learning to Highlight Audio by Watching Movies
von: Huang, Chao, et al.
Veröffentlicht: (2025) -
SoundCam: A Dataset for Finding Humans Using Room Acoustics
von: Wang, Mason, et al.
Veröffentlicht: (2023) -
Real Acoustic Fields: An Audio-Visual Room Acoustics Dataset and Benchmark
von: Chen, Ziyang, et al.
Veröffentlicht: (2024)