Object-aware Sound Source Localization via Audio-Visual Scene Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Um, Sung Jin, Kim, Dongjin, Lee, Sangmin, Kim, Jung Uk |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning to Visually Localize Sound Sources from Mixtures without Prior Source Knowledge
by: Kim, Dongjin, et al.
Published: (2024)
by: Kim, Dongjin, et al.
Published: (2024)
Watch Video, Catch Keyword: Context-aware Keyword Attention for Moment Retrieval and Highlight Detection
by: Um, Sung Jin, et al.
Published: (2025)
by: Um, Sung Jin, et al.
Published: (2025)
Generate, Analyze, and Refine: Training-Free Sound Source Localization via MLLM Meta-Reasoning
by: Park, Subin, et al.
Published: (2026)
by: Park, Subin, et al.
Published: (2026)
See, Rank, and Filter: Important Word-Aware Clip Filtering via Scene Understanding for Moment Retrieval and Highlight Detection
by: Lee, YuEun, et al.
Published: (2025)
by: Lee, YuEun, et al.
Published: (2025)
From Adaptation to Generalization: Adaptive Visual Prompting for Medical Image Segmentation
by: Çetinkaya, Evren, et al.
Published: (2026)
by: Çetinkaya, Evren, et al.
Published: (2026)
Question-Aware Gaussian Experts for Audio-Visual Question Answering
by: Kim, Hongyeob, et al.
Published: (2025)
by: Kim, Hongyeob, et al.
Published: (2025)
MonoSAOD: Monocular 3D Object Detection with Sparsely Annotated Label
by: Jung, Junyoung, et al.
Published: (2026)
by: Jung, Junyoung, et al.
Published: (2026)
Learning Trimodal Relation for Audio-Visual Question Answering with Missing Modality
by: Park, Kyu Ri, et al.
Published: (2024)
by: Park, Kyu Ri, et al.
Published: (2024)
Improving Sound Source Localization with Joint Slot Attention on Image and Audio
by: Kim, Inho, et al.
Published: (2025)
by: Kim, Inho, et al.
Published: (2025)
MemoryTalker: Personalized Speech-Driven 3D Facial Animation via Audio-Guided Stylization
by: Kim, Hyung Kyu, et al.
Published: (2025)
by: Kim, Hyung Kyu, et al.
Published: (2025)
Informative Object-centric Next Best View for Object-aware 3D Gaussian Splatting in Cluttered Scenes
by: Jeong, Seunghoon, et al.
Published: (2026)
by: Jeong, Seunghoon, et al.
Published: (2026)
Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment
by: Senocak, Arda, et al.
Published: (2024)
by: Senocak, Arda, et al.
Published: (2024)
MonoWAD: Weather-Adaptive Diffusion Model for Robust Monocular 3D Object Detection
by: Oh, Youngmin, et al.
Published: (2024)
by: Oh, Youngmin, et al.
Published: (2024)
InterRVOS: Interaction-aware Referring Video Object Segmentation
by: Jin, Woojeong, et al.
Published: (2025)
by: Jin, Woojeong, et al.
Published: (2025)
SoundBrush: Sound as a Brush for Visual Scene Editing
by: Sung-Bin, Kim, et al.
Published: (2024)
by: Sung-Bin, Kim, et al.
Published: (2024)
Seeing Speech and Sound: Distinguishing and Locating Audios in Visual Scenes
by: Ryu, Hyeonggon, et al.
Published: (2025)
by: Ryu, Hyeonggon, et al.
Published: (2025)
Semi-Supervised Audio-Visual Video Action Recognition with Audio Source Localization Guided Mixup
by: Kang, Seokun, et al.
Published: (2025)
by: Kang, Seokun, et al.
Published: (2025)
TV-LiVE: Training-Free, Text-Guided Video Editing via Layer Informed Vitality Exploitation
by: Kim, Min-Jung, et al.
Published: (2025)
by: Kim, Min-Jung, et al.
Published: (2025)
Continuous Degradation Modeling via Latent Flow Matching for Real-World Super-Resolution
by: Kim, Hyeonjae, et al.
Published: (2026)
by: Kim, Hyeonjae, et al.
Published: (2026)
InsertAnywhere: Bridging 4D Scene Geometry and Diffusion Models for Realistic Video Object Insertion
by: Jin, Hoiyeong, et al.
Published: (2025)
by: Jin, Hoiyeong, et al.
Published: (2025)
Enhancing Weakly Supervised Video Grounding via Diverse Inference Strategies for Boundary and Prediction Selection
by: Kim, Sunoh, et al.
Published: (2025)
by: Kim, Sunoh, et al.
Published: (2025)
AVOID: The Adverse Visual Conditions Dataset with Obstacles for Driving Scene Understanding
by: Jeong, Jongoh, et al.
Published: (2025)
by: Jeong, Jongoh, et al.
Published: (2025)
RA-SSU: Towards Fine-Grained Audio-Visual Learning with Region-Aware Sound Source Understanding
by: Sun, Muyi, et al.
Published: (2026)
by: Sun, Muyi, et al.
Published: (2026)
EXOT: Exit-aware Object Tracker for Safe Robotic Manipulation of Moving Object
by: Kim, Hyunseo, et al.
Published: (2023)
by: Kim, Hyunseo, et al.
Published: (2023)
UM-Depth : Uncertainty Masked Self-Supervised Monocular Depth Estimation with Visual Odometry
by: Um, Tae-Wook, et al.
Published: (2025)
by: Um, Tae-Wook, et al.
Published: (2025)
Audio-Visual Camera Pose Estimation with Passive Scene Sounds and In-the-Wild Video
by: Adebi, Daniel, et al.
Published: (2025)
by: Adebi, Daniel, et al.
Published: (2025)
Breaking the Visual Shortcuts in Multimodal Knowledge-Based Visual Question Answering
by: Lee, Dosung, et al.
Published: (2025)
by: Lee, Dosung, et al.
Published: (2025)
AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models
by: Sung-Bin, Kim, et al.
Published: (2024)
by: Sung-Bin, Kim, et al.
Published: (2024)
Do We Need Perfect Data? Leveraging Noise for Domain Generalized Segmentation
by: Kim, Taeyeong, et al.
Published: (2025)
by: Kim, Taeyeong, et al.
Published: (2025)
STELLA: Continual Audio-Video Pre-training with Spatio-Temporal Localized Alignment
by: Lee, Jaewoo, et al.
Published: (2023)
by: Lee, Jaewoo, et al.
Published: (2023)
ObjectGS: Object-aware Scene Reconstruction and Scene Understanding via Gaussian Splatting
by: Zhu, Ruijie, et al.
Published: (2025)
by: Zhu, Ruijie, et al.
Published: (2025)
Learning from Oblivion: Predicting Knowledge Overflowed Weights via Retrodiction of Forgetting
by: Jang, Jinhyeok, et al.
Published: (2025)
by: Jang, Jinhyeok, et al.
Published: (2025)
ORIDa: Object-centric Real-world Image Composition Dataset
by: Kim, Jinwoo, et al.
Published: (2025)
by: Kim, Jinwoo, et al.
Published: (2025)
Locality-aware Cross-modal Correspondence Learning for Dense Audio-Visual Events Localization
by: Xing, Ling, et al.
Published: (2024)
by: Xing, Ling, et al.
Published: (2024)
Diffusion-Based sRGB Real Noise Generation via Prompt-Driven Noise Representation Learning
by: Ko, Jaekyun, et al.
Published: (2026)
by: Ko, Jaekyun, et al.
Published: (2026)
Locality-Aware Zero-Shot Human-Object Interaction Detection
by: Kim, Sanghyun, et al.
Published: (2025)
by: Kim, Sanghyun, et al.
Published: (2025)
Patch-level Sounding Object Tracking for Audio-Visual Question Answering
by: Li, Zhangbin, et al.
Published: (2024)
by: Li, Zhangbin, et al.
Published: (2024)
Crab: A Unified Audio-Visual Scene Understanding Model with Explicit Cooperation
by: Du, Henghui, et al.
Published: (2025)
by: Du, Henghui, et al.
Published: (2025)
Multispectral Pedestrian Detection with Sparsely Annotated Label
by: Lee, Chan, et al.
Published: (2025)
by: Lee, Chan, et al.
Published: (2025)
VLCounter: Text-aware Visual Representation for Zero-Shot Object Counting
by: Kang, Seunggu, et al.
Published: (2023)
by: Kang, Seunggu, et al.
Published: (2023)
Similar Items
-
Learning to Visually Localize Sound Sources from Mixtures without Prior Source Knowledge
by: Kim, Dongjin, et al.
Published: (2024) -
Watch Video, Catch Keyword: Context-aware Keyword Attention for Moment Retrieval and Highlight Detection
by: Um, Sung Jin, et al.
Published: (2025) -
Generate, Analyze, and Refine: Training-Free Sound Source Localization via MLLM Meta-Reasoning
by: Park, Subin, et al.
Published: (2026) -
See, Rank, and Filter: Important Word-Aware Clip Filtering via Scene Understanding for Moment Retrieval and Highlight Detection
by: Lee, YuEun, et al.
Published: (2025) -
From Adaptation to Generalization: Adaptive Visual Prompting for Medical Image Segmentation
by: Çetinkaya, Evren, et al.
Published: (2026)