Generate, Analyze, and Refine: Training-Free Sound Source Localization via MLLM Meta-Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Park, Subin, Kim, Jung Uk |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Object-aware Sound Source Localization via Audio-Visual Scene Understanding
by: Um, Sung Jin, et al.
Published: (2025)
by: Um, Sung Jin, et al.
Published: (2025)
Learning to Visually Localize Sound Sources from Mixtures without Prior Source Knowledge
by: Kim, Dongjin, et al.
Published: (2024)
by: Kim, Dongjin, et al.
Published: (2024)
See, Rank, and Filter: Important Word-Aware Clip Filtering via Scene Understanding for Moment Retrieval and Highlight Detection
by: Lee, YuEun, et al.
Published: (2025)
by: Lee, YuEun, et al.
Published: (2025)
Improving Sound Source Localization with Joint Slot Attention on Image and Audio
by: Kim, Inho, et al.
Published: (2025)
by: Kim, Inho, et al.
Published: (2025)
MonoSAOD: Monocular 3D Object Detection with Sparsely Annotated Label
by: Jung, Junyoung, et al.
Published: (2026)
by: Jung, Junyoung, et al.
Published: (2026)
Leveraging Textual Compositional Reasoning for Robust Change Captioning
by: Park, Kyu Ri, et al.
Published: (2025)
by: Park, Kyu Ri, et al.
Published: (2025)
Towards Model-Agnostic Dataset Condensation by Heterogeneous Models
by: Moon, Jun-Yeong, et al.
Published: (2024)
by: Moon, Jun-Yeong, et al.
Published: (2024)
Filter, Correlate, Compress: Training-Free Token Reduction for MLLM Acceleration
by: Han, Yuhang, et al.
Published: (2024)
by: Han, Yuhang, et al.
Published: (2024)
Training-Free Refinement of Flow Matching with Divergence-based Sampling
by: Cha, Yeonwoo, et al.
Published: (2026)
by: Cha, Yeonwoo, et al.
Published: (2026)
FreeVA: Offline MLLM as Training-Free Video Assistant
by: Wu, Wenhao
Published: (2024)
by: Wu, Wenhao
Published: (2024)
Multispectral Pedestrian Detection with Sparsely Annotated Label
by: Lee, Chan, et al.
Published: (2025)
by: Lee, Chan, et al.
Published: (2025)
Enhancing Sound Source Localization via False Negative Elimination
by: Song, Zengjie, et al.
Published: (2024)
by: Song, Zengjie, et al.
Published: (2024)
EraseLoRA: MLLM-Driven Foreground Exclusion and Background Subtype Aggregation for Dataset-Free Object Removal
by: Jo, Sanghyun, et al.
Published: (2025)
by: Jo, Sanghyun, et al.
Published: (2025)
VisionTrim: Unified Vision Token Compression for Training-Free MLLM Acceleration
by: Yu, Hanxun, et al.
Published: (2026)
by: Yu, Hanxun, et al.
Published: (2026)
Do We Need Perfect Data? Leveraging Noise for Domain Generalized Segmentation
by: Kim, Taeyeong, et al.
Published: (2025)
by: Kim, Taeyeong, et al.
Published: (2025)
A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images
by: Lee, Jaeseong, et al.
Published: (2025)
by: Lee, Jaeseong, et al.
Published: (2025)
ELITE: Efficient Gaussian Head Avatar from a Monocular Video via Learned Initialization and TEst-time Generative Adaptation
by: Youwang, Kim, et al.
Published: (2026)
by: Youwang, Kim, et al.
Published: (2026)
Interaction-Consistent Object Removal via MLLM-Based Reasoning
by: Huang, Ching-Kai, et al.
Published: (2026)
by: Huang, Ching-Kai, et al.
Published: (2026)
Group3D: MLLM-Driven Semantic Grouping for Open-Vocabulary 3D Object Detection
by: Kim, Youbin, et al.
Published: (2026)
by: Kim, Youbin, et al.
Published: (2026)
ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language Models
by: Wu, Mingrui, et al.
Published: (2024)
by: Wu, Mingrui, et al.
Published: (2024)
Empowering Source-Free Domain Adaptation via MLLM-Guided Reliability-Based Curriculum Learning
by: Chen, Dongjie, et al.
Published: (2024)
by: Chen, Dongjie, et al.
Published: (2024)
DTG-Restore: Training-Free Diffusion Refinement for Generative Video Super-Resolution
by: Yesiltepe, Hidir, et al.
Published: (2026)
by: Yesiltepe, Hidir, et al.
Published: (2026)
TV-LiVE: Training-Free, Text-Guided Video Editing via Layer Informed Vitality Exploitation
by: Kim, Min-Jung, et al.
Published: (2025)
by: Kim, Min-Jung, et al.
Published: (2025)
TAMEing Long Contexts in Personalization: Towards Training-Free and State-Aware MLLM Personalized Assistant
by: Hong, Rongpei, et al.
Published: (2025)
by: Hong, Rongpei, et al.
Published: (2025)
ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement
by: Huang, Runhui, et al.
Published: (2025)
by: Huang, Runhui, et al.
Published: (2025)
Learning from Oblivion: Predicting Knowledge Overflowed Weights via Retrodiction of Forgetting
by: Jang, Jinhyeok, et al.
Published: (2025)
by: Jang, Jinhyeok, et al.
Published: (2025)
CAD-MLLM: Unifying Multimodality-Conditioned CAD Generation With MLLM
by: Xu, Jingwei, et al.
Published: (2024)
by: Xu, Jingwei, et al.
Published: (2024)
Training-Free Personalization via Retrieval and Reasoning on Fingerprints
by: Das, Deepayan, et al.
Published: (2025)
by: Das, Deepayan, et al.
Published: (2025)
From Adaptation to Generalization: Adaptive Visual Prompting for Medical Image Segmentation
by: Çetinkaya, Evren, et al.
Published: (2026)
by: Çetinkaya, Evren, et al.
Published: (2026)
$h$-control: Training-Free Camera Control via Block-Conditional Gibbs Refinement
by: Wang, Yuzhu, et al.
Published: (2026)
by: Wang, Yuzhu, et al.
Published: (2026)
DiffuseSlide: Training-Free High Frame Rate Video Generation Diffusion
by: Hwang, Geunmin, et al.
Published: (2025)
by: Hwang, Geunmin, et al.
Published: (2025)
MonoWAD: Weather-Adaptive Diffusion Model for Robust Monocular 3D Object Detection
by: Oh, Youngmin, et al.
Published: (2024)
by: Oh, Youngmin, et al.
Published: (2024)
Survey on AI-Generated Media Detection: From Non-MLLM to MLLM
by: Zou, Yueying, et al.
Published: (2025)
by: Zou, Yueying, et al.
Published: (2025)
Fundus-R1: Training a Fundus-Reading MLLM with Knowledge-Aware Reasoning on Public Data
by: Deng, Yuchuan, et al.
Published: (2026)
by: Deng, Yuchuan, et al.
Published: (2026)
Learning Trimodal Relation for Audio-Visual Question Answering with Missing Modality
by: Park, Kyu Ri, et al.
Published: (2024)
by: Park, Kyu Ri, et al.
Published: (2024)
A Training-Free Style-aligned Image Generation with Scale-wise Autoregressive Model
by: Park, Jihun, et al.
Published: (2025)
by: Park, Jihun, et al.
Published: (2025)
Deep Forcing: Training-Free Long Video Generation with Deep Sink and Participative Compression
by: Yi, Jung, et al.
Published: (2025)
by: Yi, Jung, et al.
Published: (2025)
Reasoning Guided Embeddings: Leveraging MLLM Reasoning for Improved Multimodal Retrieval
by: Liu, Chunxu, et al.
Published: (2025)
by: Liu, Chunxu, et al.
Published: (2025)
Prompt Learning via Meta-Regularization
by: Park, Jinyoung, et al.
Published: (2024)
by: Park, Jinyoung, et al.
Published: (2024)
Boosting MLLM Reasoning with Text-Debiased Hint-GRPO
by: Huang, Qihan, et al.
Published: (2025)
by: Huang, Qihan, et al.
Published: (2025)
Similar Items
-
Object-aware Sound Source Localization via Audio-Visual Scene Understanding
by: Um, Sung Jin, et al.
Published: (2025) -
Learning to Visually Localize Sound Sources from Mixtures without Prior Source Knowledge
by: Kim, Dongjin, et al.
Published: (2024) -
See, Rank, and Filter: Important Word-Aware Clip Filtering via Scene Understanding for Moment Retrieval and Highlight Detection
by: Lee, YuEun, et al.
Published: (2025) -
Improving Sound Source Localization with Joint Slot Attention on Image and Audio
by: Kim, Inho, et al.
Published: (2025) -
MonoSAOD: Monocular 3D Object Detection with Sparsely Annotated Label
by: Jung, Junyoung, et al.
Published: (2026)