BGM2Pose: Active 3D Human Pose Estimation with Non-Stationary Sounds
Fuente:
arXiv
Saved in:
| Main Authors: | Shibata, Yuto, Oumi, Yusuke, Irie, Go, Kimura, Akisato, Aoki, Yoshimitsu, Isogawa, Mariko |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Acoustic-based 3D Human Pose Estimation Robust to Human Position
by: Oumi, Yusuke, et al.
Published: (2024)
by: Oumi, Yusuke, et al.
Published: (2024)
Estimating Indoor Scene Depth Maps from Ultrasonic Echoes
by: Honma, Junpei, et al.
Published: (2024)
by: Honma, Junpei, et al.
Published: (2024)
DnR-nonverbal: Cinematic Audio Source Separation Dataset Containing Non-Verbal Sounds
by: Hasumi, Takuya, et al.
Published: (2025)
by: Hasumi, Takuya, et al.
Published: (2025)
Physics-Informed Machine Learning For Sound Field Estimation
by: Koyama, Shoichi, et al.
Published: (2024)
by: Koyama, Shoichi, et al.
Published: (2024)
Stream-based Active Learning for Anomalous Sound Detection in Machine Condition Monitoring
by: Ho, Tuan Vu, et al.
Published: (2024)
by: Ho, Tuan Vu, et al.
Published: (2024)
AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences
by: Kishi, Minoru, et al.
Published: (2025)
by: Kishi, Minoru, et al.
Published: (2025)
SoundBeam meets M2D: Target Sound Extraction with Audio Foundation Model
by: Hernandez-Olivan, Carlos, et al.
Published: (2024)
by: Hernandez-Olivan, Carlos, et al.
Published: (2024)
Audible Networks: Deconstructing and Manipulating Sounds with Deep Non-Negative Autoencoders
by: Burred, Juan José, et al.
Published: (2025)
by: Burred, Juan José, et al.
Published: (2025)
Real-time Speech Extraction Using Spatially Regularized Independent Low-rank Matrix Analysis and Rank-constrained Spatial Covariance Matrix Estimation
by: Ishikawa, Yuto, et al.
Published: (2024)
by: Ishikawa, Yuto, et al.
Published: (2024)
Human-CLAP: Human-perception-based contrastive language-audio pretraining
by: Takano, Taisei, et al.
Published: (2025)
by: Takano, Taisei, et al.
Published: (2025)
Active Learning of Non-semantic Speech Tasks with Pretrained Models
by: Lee, Harlin, et al.
Published: (2022)
by: Lee, Harlin, et al.
Published: (2022)
IPDnet: A Universal Direct-Path IPD Estimation Network for Sound Source Localization
by: Wang, Yabo, et al.
Published: (2024)
by: Wang, Yabo, et al.
Published: (2024)
Sound-Based Spin Estimation in Table Tennis: Dataset and Real-Time Classification Pipeline
by: Gossard, Thomas, et al.
Published: (2024)
by: Gossard, Thomas, et al.
Published: (2024)
A Few-Shot Learning Approach for Sound Source Distance Estimation Using Relation Networks
by: Sobhdel, Amirreza, et al.
Published: (2021)
by: Sobhdel, Amirreza, et al.
Published: (2021)
First-Shot Unsupervised Anomalous Sound Detection With Unknown Anomalies Estimated by Metadata-Assisted Audio Generation
by: Zhang, Hejing, et al.
Published: (2023)
by: Zhang, Hejing, et al.
Published: (2023)
Online Single-Channel Audio-Based Sound Speed Estimation for Robust Multi-Channel Audio Control
by: Fuglsig, Andreas Jonas, et al.
Published: (2026)
by: Fuglsig, Andreas Jonas, et al.
Published: (2026)
Leveraging Sound Source Trajectories for Universal Sound Separation
by: Wu, Donghang, et al.
Published: (2024)
by: Wu, Donghang, et al.
Published: (2024)
Sound Zone Control Robust To Sound Speed Change
by: Bhattacharjee, Sankha Subhra, et al.
Published: (2024)
by: Bhattacharjee, Sankha Subhra, et al.
Published: (2024)
Evaluating Sound Similarity Metrics for Differentiable, Iterative Sound-Matching
by: Salimi, Amir, et al.
Published: (2025)
by: Salimi, Amir, et al.
Published: (2025)
LatentVoiceGrad: Nonparallel Voice Conversion with Latent Diffusion/Flow-Matching Models
by: Kameoka, Hirokazu, et al.
Published: (2025)
by: Kameoka, Hirokazu, et al.
Published: (2025)
Rethinking Mean Opinion Scores in Speech Quality Assessment: Aggregation through Quantized Distribution Fitting
by: Kondo, Yuto, et al.
Published: (2025)
by: Kondo, Yuto, et al.
Published: (2025)
Selecting N-lowest scores for training MOS prediction models
by: Kondo, Yuto, et al.
Published: (2025)
by: Kondo, Yuto, et al.
Published: (2025)
JIS: A Speech Corpus of Japanese Idol Speakers with Various Speaking Styles
by: Kondo, Yuto, et al.
Published: (2025)
by: Kondo, Yuto, et al.
Published: (2025)
Noise-Robust Sound Event Detection and Counting via Language-Queried Sound Separation
by: Chen, Yuanjian, et al.
Published: (2025)
by: Chen, Yuanjian, et al.
Published: (2025)
DiffSound: Differentiable Modal Sound Rendering and Inverse Rendering for Diverse Inference Tasks
by: Jin, Xutong, et al.
Published: (2024)
by: Jin, Xutong, et al.
Published: (2024)
Fractional Fourier Sound Synthesis
by: Gutiérrez, Esteban, et al.
Published: (2025)
by: Gutiérrez, Esteban, et al.
Published: (2025)
Diffuse Sound Field Synthesis
by: Zotter, Franz, et al.
Published: (2024)
by: Zotter, Franz, et al.
Published: (2024)
Sound Event Bounding Boxes
by: Ebbers, Janek, et al.
Published: (2024)
by: Ebbers, Janek, et al.
Published: (2024)
SoundLoCD: An Efficient Conditional Discrete Contrastive Latent Diffusion Model for Text-to-Sound Generation
by: Niu, Xinlei, et al.
Published: (2024)
by: Niu, Xinlei, et al.
Published: (2024)
Sound Field Synthesis with Acoustic Waves
by: Mansour, Mohamed F.
Published: (2024)
by: Mansour, Mohamed F.
Published: (2024)
Fast Algorithm for Moving Sound Source
by: Yang, Dong
Published: (2025)
by: Yang, Dong
Published: (2025)
Boundary-Informed Sound Field Reconstruction
by: Sundström, David, et al.
Published: (2025)
by: Sundström, David, et al.
Published: (2025)
Joint Analysis of Acoustic Scenes and Sound Events Based on Semi-Supervised Training of Sound Events With Partial Labels
by: Imoto, Keisuke
Published: (2025)
by: Imoto, Keisuke
Published: (2025)
Reverberation-based Features for Sound Event Localization and Detection with Distance Estimation
by: Berghi, Davide, et al.
Published: (2025)
by: Berghi, Davide, et al.
Published: (2025)
SHEET: A Multi-purpose Open-source Speech Human Evaluation Estimation Toolkit
by: Huang, Wen-Chin, et al.
Published: (2025)
by: Huang, Wen-Chin, et al.
Published: (2025)
Automatic Inspection Based on Switch Sounds of Electric Point Machines
by: Shibata, Ayano, et al.
Published: (2025)
by: Shibata, Ayano, et al.
Published: (2025)
HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis
by: Nishimura, Yuto, et al.
Published: (2024)
by: Nishimura, Yuto, et al.
Published: (2024)
Audio Fingerprinting with Holographic Reduced Representations
by: Fujita, Yusuke, et al.
Published: (2024)
by: Fujita, Yusuke, et al.
Published: (2024)
Domain-Invariant Representation Learning of Bird Sounds
by: Moummad, Ilyass, et al.
Published: (2024)
by: Moummad, Ilyass, et al.
Published: (2024)
AudioSpa: Spatializing Sound Events with Text
by: Feng, Linfeng, et al.
Published: (2025)
by: Feng, Linfeng, et al.
Published: (2025)
Similar Items
-
Acoustic-based 3D Human Pose Estimation Robust to Human Position
by: Oumi, Yusuke, et al.
Published: (2024) -
Estimating Indoor Scene Depth Maps from Ultrasonic Echoes
by: Honma, Junpei, et al.
Published: (2024) -
DnR-nonverbal: Cinematic Audio Source Separation Dataset Containing Non-Verbal Sounds
by: Hasumi, Takuya, et al.
Published: (2025) -
Physics-Informed Machine Learning For Sound Field Estimation
by: Koyama, Shoichi, et al.
Published: (2024) -
Stream-based Active Learning for Anomalous Sound Detection in Machine Condition Monitoring
by: Ho, Tuan Vu, et al.
Published: (2024)