Listen2Scene: Interactive material-aware binaural sound propagation for reconstructed 3D scenes
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ratnarajah, Anton, Manocha, Dinesh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Visual and audio scene classification for detecting discrepancies in video: a baseline method and experimental protocol
von: Apostolidis, Konstantinos, et al.
Veröffentlicht: (2024)
von: Apostolidis, Konstantinos, et al.
Veröffentlicht: (2024)
Aligned Better, Listen Better for Audio-Visual Large Language Models
von: Guo, Yuxin, et al.
Veröffentlicht: (2025)
von: Guo, Yuxin, et al.
Veröffentlicht: (2025)
FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2025)
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2025)
Statistics-aware Audio-visual Deepfake Detector
von: Astrid, Marcella, et al.
Veröffentlicht: (2024)
von: Astrid, Marcella, et al.
Veröffentlicht: (2024)
Estimating Indoor Scene Depth Maps from Ultrasonic Echoes
von: Honma, Junpei, et al.
Veröffentlicht: (2024)
von: Honma, Junpei, et al.
Veröffentlicht: (2024)
SoundLoc3D: Invisible 3D Sound Source Localization and Classification Using a Multimodal RGB-D Acoustic Camera
von: He, Yuhang, et al.
Veröffentlicht: (2024)
von: He, Yuhang, et al.
Veröffentlicht: (2024)
3D Audio-Visual Segmentation
von: Sokolov, Artem, et al.
Veröffentlicht: (2024)
von: Sokolov, Artem, et al.
Veröffentlicht: (2024)
Reverse the auditory processing pathway: Coarse-to-fine audio reconstruction from fMRI
von: Liu, Che, et al.
Veröffentlicht: (2024)
von: Liu, Che, et al.
Veröffentlicht: (2024)
InterDance:Reactive 3D Dance Generation with Realistic Duet Interactions
von: Li, Ronghui, et al.
Veröffentlicht: (2024)
von: Li, Ronghui, et al.
Veröffentlicht: (2024)
Nexus: An Omni-Perceptive And -Interactive Model for Language, Audio, And Vision
von: Liu, Che, et al.
Veröffentlicht: (2025)
von: Liu, Che, et al.
Veröffentlicht: (2025)
Efficient learning-based sound propagation for virtual and real-world audio processing applications
von: Ratnarajah, Anton Jeran
Veröffentlicht: (2024)
von: Ratnarajah, Anton Jeran
Veröffentlicht: (2024)
SAV-SE: Scene-aware Audio-Visual Speech Enhancement with Selective State Space Model
von: Qian, Xinyuan, et al.
Veröffentlicht: (2024)
von: Qian, Xinyuan, et al.
Veröffentlicht: (2024)
Art2Mus: Bridging Visual Arts and Music through Cross-Modal Generation
von: Rinaldi, Ivan, et al.
Veröffentlicht: (2024)
von: Rinaldi, Ivan, et al.
Veröffentlicht: (2024)
Sound2Vision: Generating Diverse Visuals from Audio through Cross-Modal Latent Alignment
von: Sung-Bin, Kim, et al.
Veröffentlicht: (2024)
von: Sung-Bin, Kim, et al.
Veröffentlicht: (2024)
Face2VoiceSync: Lightweight Face-Voice Consistency for Text-Driven Talking Face Generation
von: Kang, Fang, et al.
Veröffentlicht: (2025)
von: Kang, Fang, et al.
Veröffentlicht: (2025)
NeRF-3DTalker: Neural Radiance Field with 3D Prior Aided Audio Disentanglement for Talking Head Synthesis
von: Liu, Xiaoxing, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoxing, et al.
Veröffentlicht: (2025)
UniAV: Unified Audio-Visual Perception for Multi-Task Video Event Localization
von: Geng, Tiantian, et al.
Veröffentlicht: (2024)
von: Geng, Tiantian, et al.
Veröffentlicht: (2024)
AV-Surf: Surface-Enhanced Geometry-Aware Novel-View Acoustic Synthesis
von: Baek, Hadam, et al.
Veröffentlicht: (2025)
von: Baek, Hadam, et al.
Veröffentlicht: (2025)
Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation
von: Low, Chetwin, et al.
Veröffentlicht: (2025)
von: Low, Chetwin, et al.
Veröffentlicht: (2025)
AISHELL6-whisper: A Chinese Mandarin Audio-visual Whisper Speech Dataset with Speech Recognition Baselines
von: Li, Cancan, et al.
Veröffentlicht: (2025)
von: Li, Cancan, et al.
Veröffentlicht: (2025)
IIANet: An Intra- and Inter-Modality Attention Network for Audio-Visual Speech Separation
von: Li, Kai, et al.
Veröffentlicht: (2023)
von: Li, Kai, et al.
Veröffentlicht: (2023)
Spiking Tucker Fusion Transformer for Audio-Visual Zero-Shot Learning
von: Li, Wenrui, et al.
Veröffentlicht: (2024)
von: Li, Wenrui, et al.
Veröffentlicht: (2024)
SoundingActions: Learning How Actions Sound from Narrated Egocentric Videos
von: Chen, Changan, et al.
Veröffentlicht: (2024)
von: Chen, Changan, et al.
Veröffentlicht: (2024)
Pilot-guided Multimodal Semantic Communication for Audio-Visual Event Localization
von: Yu, Fei, et al.
Veröffentlicht: (2024)
von: Yu, Fei, et al.
Veröffentlicht: (2024)
Learning Self-Supervised Audio-Visual Representations for Sound Recommendations
von: Krishnamurthy, Sudha
Veröffentlicht: (2024)
von: Krishnamurthy, Sudha
Veröffentlicht: (2024)
VMAS: Video-to-Music Generation via Semantic Alignment in Web Music Videos
von: Lin, Yan-Bo, et al.
Veröffentlicht: (2024)
von: Lin, Yan-Bo, et al.
Veröffentlicht: (2024)
Comparative Analysis of Modality Fusion Approaches for Audio-Visual Person Identification and Verification
von: Farhadipour, Aref, et al.
Veröffentlicht: (2024)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2024)
Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper
von: Chen, Gehui, et al.
Veröffentlicht: (2025)
von: Chen, Gehui, et al.
Veröffentlicht: (2025)
AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment
von: Cao, Yuqin, et al.
Veröffentlicht: (2025)
von: Cao, Yuqin, et al.
Veröffentlicht: (2025)
Audio-Assisted Face Video Restoration with Temporal and Identity Complementary Learning
von: Cao, Yuqin, et al.
Veröffentlicht: (2025)
von: Cao, Yuqin, et al.
Veröffentlicht: (2025)
Diffusion Models as Masked Audio-Video Learners
von: Nunez, Elvis, et al.
Veröffentlicht: (2023)
von: Nunez, Elvis, et al.
Veröffentlicht: (2023)
Learning to Visually Localize Sound Sources from Mixtures without Prior Source Knowledge
von: Kim, Dongjin, et al.
Veröffentlicht: (2024)
von: Kim, Dongjin, et al.
Veröffentlicht: (2024)
Learning Sparsity for Effective and Efficient Music Performance Question Answering
von: Diao, Xingjian, et al.
Veröffentlicht: (2025)
von: Diao, Xingjian, et al.
Veröffentlicht: (2025)
Multifaceted Evaluation of Audio-Visual Capability for MLLMs: Effectiveness, Efficiency, Generalizability and Robustness
von: Zhao, Yusheng, et al.
Veröffentlicht: (2025)
von: Zhao, Yusheng, et al.
Veröffentlicht: (2025)
Temporal Working Memory: Query-Guided Segment Refinement for Enhanced Multimodal Understanding
von: Diao, Xingjian, et al.
Veröffentlicht: (2025)
von: Diao, Xingjian, et al.
Veröffentlicht: (2025)
Large Language Models are Strong Audio-Visual Speech Recognition Learners
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2024)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2024)
M&M: Multimodal-Multitask Model Integrating Audiovisual Cues in Cognitive Load Assessment
von: Nguyen-Phuoc, Long, et al.
Veröffentlicht: (2024)
von: Nguyen-Phuoc, Long, et al.
Veröffentlicht: (2024)
Mechanisms of Multimodal Synchronization: Insights from Decoder-Based Video-Text-to-Speech Synthesis
von: Gupta, Akshita, et al.
Veröffentlicht: (2024)
von: Gupta, Akshita, et al.
Veröffentlicht: (2024)
Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching
von: Wang, Yongqi, et al.
Veröffentlicht: (2024)
von: Wang, Yongqi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Visual and audio scene classification for detecting discrepancies in video: a baseline method and experimental protocol
von: Apostolidis, Konstantinos, et al.
Veröffentlicht: (2024) -
Aligned Better, Listen Better for Audio-Visual Large Language Models
von: Guo, Yuxin, et al.
Veröffentlicht: (2025) -
FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2025) -
Statistics-aware Audio-visual Deepfake Detector
von: Astrid, Marcella, et al.
Veröffentlicht: (2024) -
Estimating Indoor Scene Depth Maps from Ultrasonic Echoes
von: Honma, Junpei, et al.
Veröffentlicht: (2024)