Physics-Aware Novel-View Acoustic Synthesis with Vision-Language Priors and 3D Acoustic Environment Modeling
Fuente:
arXiv
Guardado en:
| Autores principales: | Fan, Congyi, Guan, Jian, Lin, Youtian, Xu, Dongli, Ye, Tong, Zhu, Qiaoxi, Feng, Pengming, Wang, Wenwu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Align Your Rhythm: Generating Highly Aligned Dance Poses with Gating-Enhanced Rhythm-Aware Feature Representation
por: Fan, Congyi, et al.
Publicado: (2025)
por: Fan, Congyi, et al.
Publicado: (2025)
AV-Surf: Surface-Enhanced Geometry-Aware Novel-View Acoustic Synthesis
por: Baek, Hadam, et al.
Publicado: (2025)
por: Baek, Hadam, et al.
Publicado: (2025)
DualMark: Identifying Model and Training Data Origins in Generated Audio
por: Yang, Xuefeng, et al.
Publicado: (2025)
por: Yang, Xuefeng, et al.
Publicado: (2025)
SoundVista: Novel-View Ambient Sound Synthesis via Visual-Acoustic Binding
por: Chen, Mingfei, et al.
Publicado: (2025)
por: Chen, Mingfei, et al.
Publicado: (2025)
SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations
por: Wang, Chunshi, et al.
Publicado: (2025)
por: Wang, Chunshi, et al.
Publicado: (2025)
Uncertainty-Aware 3D Emotional Talking Face Synthesis with Emotion Prior Distillation
por: Shen, Nanhan, et al.
Publicado: (2026)
por: Shen, Nanhan, et al.
Publicado: (2026)
SonicSense: Object Perception from In-Hand Acoustic Vibration
por: Liu, Jiaxun, et al.
Publicado: (2024)
por: Liu, Jiaxun, et al.
Publicado: (2024)
Real Acoustic Fields: An Audio-Visual Room Acoustics Dataset and Benchmark
por: Chen, Ziyang, et al.
Publicado: (2024)
por: Chen, Ziyang, et al.
Publicado: (2024)
Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music Generation
por: Tong, Xinyi, et al.
Publicado: (2025)
por: Tong, Xinyi, et al.
Publicado: (2025)
HAFM: Hierarchical Autoregressive Foundation Model for Music Accompaniment Generation
por: Zhu, Jian, et al.
Publicado: (2026)
por: Zhu, Jian, et al.
Publicado: (2026)
Memo2496: Expert-Annotated Dataset and Dual-View Adaptive Framework for Music Emotion Recognition
por: Li, Qilin, et al.
Publicado: (2025)
por: Li, Qilin, et al.
Publicado: (2025)
PAVAS: Physics-Aware Video-to-Audio Synthesis
por: Hyun-Bin, Oh, et al.
Publicado: (2025)
por: Hyun-Bin, Oh, et al.
Publicado: (2025)
Deciphering GunType Hierarchy through Acoustic Analysis of Gunshot Recordings
por: Shah, Ankit, et al.
Publicado: (2025)
por: Shah, Ankit, et al.
Publicado: (2025)
EigeNet: Geometry-Informed Multi-Modal Learning for Few-shot Novel View RIR Prediction
por: Jing, Chong, et al.
Publicado: (2026)
por: Jing, Chong, et al.
Publicado: (2026)
WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
por: Ji, Shengpeng, et al.
Publicado: (2024)
por: Ji, Shengpeng, et al.
Publicado: (2024)
Quality-Aware End-to-End Audio-Visual Neural Speaker Diarization
por: He, Mao-Kui, et al.
Publicado: (2024)
por: He, Mao-Kui, et al.
Publicado: (2024)
AV-GS: Learning Material and Geometry Aware Priors for Novel View Acoustic Synthesis
por: Bhosale, Swapnil, et al.
Publicado: (2024)
por: Bhosale, Swapnil, et al.
Publicado: (2024)
Can We Hear from Events? Generating Speech from Event Camera
por: Fang, Jingping, et al.
Publicado: (2026)
por: Fang, Jingping, et al.
Publicado: (2026)
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues
por: Zhou, Dingkun, et al.
Publicado: (2025)
por: Zhou, Dingkun, et al.
Publicado: (2025)
It's Never Too Late: Fusing Acoustic Information into Large Language Models for Automatic Speech Recognition
por: Chen, Chen, et al.
Publicado: (2024)
por: Chen, Chen, et al.
Publicado: (2024)
SonicVisionLM: Playing Sound with Vision Language Models
por: Xie, Zhifeng, et al.
Publicado: (2024)
por: Xie, Zhifeng, et al.
Publicado: (2024)
MMED: A Multimodal Micro-Expression Dataset based on Audio-Visual Fusion
por: Wang, Junbo, et al.
Publicado: (2025)
por: Wang, Junbo, et al.
Publicado: (2025)
STCTS: Generative Semantic Compression for Ultra-Low Bitrate Speech via Explicit Text-Prosody-Timbre Decomposition
por: Wang, Siyu, et al.
Publicado: (2025)
por: Wang, Siyu, et al.
Publicado: (2025)
AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis with Acoustic Transfer
por: Fang, Pengjun, et al.
Publicado: (2026)
por: Fang, Pengjun, et al.
Publicado: (2026)
SpotSound: Enhancing Large Audio-Language Models with Fine-Grained Temporal Grounding
por: Sun, Luoyi, et al.
Publicado: (2026)
por: Sun, Luoyi, et al.
Publicado: (2026)
CLAIP-Emo: Parameter-Efficient Adaptation of Language-supervised models for In-the-Wild Audiovisual Emotion Recognition
por: Chen, Yin, et al.
Publicado: (2025)
por: Chen, Yin, et al.
Publicado: (2025)
Beyond Video-to-SFX: Video to Audio Synthesis with Environmentally Aware Speech
por: Niu, Xinlei, et al.
Publicado: (2025)
por: Niu, Xinlei, et al.
Publicado: (2025)
Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation
por: Xie, Zhifei, et al.
Publicado: (2026)
por: Xie, Zhifei, et al.
Publicado: (2026)
Enkidu: Universal Frequential Perturbation for Real-Time Audio Privacy Protection against Voice Deepfakes
por: Feng, Zhou, et al.
Publicado: (2025)
por: Feng, Zhou, et al.
Publicado: (2025)
Semantically consistent Video-to-Audio Generation using Multimodal Language Large Model
por: Chen, Gehui, et al.
Publicado: (2024)
por: Chen, Gehui, et al.
Publicado: (2024)
Multi-Domain Audio Question Answering Benchmark Toward Acoustic Content Reasoning
por: Yang, Chao-Han Huck, et al.
Publicado: (2025)
por: Yang, Chao-Han Huck, et al.
Publicado: (2025)
Exploring Acoustic Similarity in Emotional Speech and Music via Self-Supervised Representations
por: Sun, Yujia, et al.
Publicado: (2024)
por: Sun, Yujia, et al.
Publicado: (2024)
MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video Generation
por: Zhou, Yang-Hao, et al.
Publicado: (2026)
por: Zhou, Yang-Hao, et al.
Publicado: (2026)
Structured Multi-Track Accompaniment Arrangement via Style Prior Modelling
por: Zhao, Jingwei, et al.
Publicado: (2023)
por: Zhao, Jingwei, et al.
Publicado: (2023)
SLAM-LLM: A Modular, Open-Source Multimodal Large Language Model Framework and Best Practice for Speech, Language, Audio and Music Processing
por: Ma, Ziyang, et al.
Publicado: (2026)
por: Ma, Ziyang, et al.
Publicado: (2026)
Multimodal Fish Feeding Intensity Assessment in Aquaculture
por: Cui, Meng, et al.
Publicado: (2023)
por: Cui, Meng, et al.
Publicado: (2023)
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model
por: Ren, Yong, et al.
Publicado: (2025)
por: Ren, Yong, et al.
Publicado: (2025)
HDA-SELD: Hierarchical Cross-Modal Distillation with Multi-Level Data Augmentation for Low-Resource Audio-Visual Sound Event Localization and Detection
por: Wang, Qing, et al.
Publicado: (2025)
por: Wang, Qing, et al.
Publicado: (2025)
MG-Former: A Transformer-Based Framework for Music-Driven 3D Conducting Gesture Generation
por: Qiu, Ke, et al.
Publicado: (2026)
por: Qiu, Ke, et al.
Publicado: (2026)
Audio-Visual Separation with Hierarchical Fusion and Representation Alignment
por: Hu, Han, et al.
Publicado: (2025)
por: Hu, Han, et al.
Publicado: (2025)
Ejemplares similares
-
Align Your Rhythm: Generating Highly Aligned Dance Poses with Gating-Enhanced Rhythm-Aware Feature Representation
por: Fan, Congyi, et al.
Publicado: (2025) -
AV-Surf: Surface-Enhanced Geometry-Aware Novel-View Acoustic Synthesis
por: Baek, Hadam, et al.
Publicado: (2025) -
DualMark: Identifying Model and Training Data Origins in Generated Audio
por: Yang, Xuefeng, et al.
Publicado: (2025) -
SoundVista: Novel-View Ambient Sound Synthesis via Visual-Acoustic Binding
por: Chen, Mingfei, et al.
Publicado: (2025) -
SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations
por: Wang, Chunshi, et al.
Publicado: (2025)