Learning Visual Affordance from Audio
Fuente:
arXiv
Saved in:
| Main Authors: | Lu, Lidong, Chen, Guo, Wei, Zhu, Liu, Yicheng, Lu, Tong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
by: Lu, Lidong, et al.
Published: (2025)
by: Lu, Lidong, et al.
Published: (2025)
Potential Field as Scene Affordance for Behavior Change-Based Visual Risk Object Identification
by: Pao, Pang-Yuan, et al.
Published: (2024)
by: Pao, Pang-Yuan, et al.
Published: (2024)
Complementary and Contrastive Learning for Audio-Visual Segmentation
by: Gong, Sitong, et al.
Published: (2025)
by: Gong, Sitong, et al.
Published: (2025)
Visual-Geometric Collaborative Guidance for Affordance Learning
by: Luo, Hongchen, et al.
Published: (2024)
by: Luo, Hongchen, et al.
Published: (2024)
CorrAdaptor: Adaptive Local Context Learning for Correspondence Pruning
by: Zhu, Wei, et al.
Published: (2024)
by: Zhu, Wei, et al.
Published: (2024)
Semantic Noise Reduction via Teacher-Guided Dual-Path Audio-Visual Representation Learning
by: Wang, Linge, et al.
Published: (2026)
by: Wang, Linge, et al.
Published: (2026)
R-AVST: Empowering Video-LLMs with Fine-Grained Spatio-Temporal Reasoning in Complex Audio-Visual Scenarios
by: Zhu, Lu, et al.
Published: (2025)
by: Zhu, Lu, et al.
Published: (2025)
The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio
by: Leng, Sicong, et al.
Published: (2024)
by: Leng, Sicong, et al.
Published: (2024)
Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision
by: He, Yuping, et al.
Published: (2025)
by: He, Yuping, et al.
Published: (2025)
GEAL: Generalizable 3D Affordance Learning with Cross-Modal Consistency
by: Lu, Dongyue, et al.
Published: (2024)
by: Lu, Dongyue, et al.
Published: (2024)
video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
by: Sun, Guangzhi, et al.
Published: (2024)
by: Sun, Guangzhi, et al.
Published: (2024)
AffordMatcher: Affordance Learning in 3D Scenes from Visual Signifiers
by: Vu, Nghia, et al.
Published: (2026)
by: Vu, Nghia, et al.
Published: (2026)
Frequency-Domain Decomposition and Recomposition for Robust Audio-Visual Segmentation
by: Shen, Yunzhe, et al.
Published: (2025)
by: Shen, Yunzhe, et al.
Published: (2025)
Informative Text-Image Alignment for Visual Affordance Learning with Foundation Models
by: Zhang, Qian, et al.
Published: (2025)
by: Zhang, Qian, et al.
Published: (2025)
EchoingPixels: Cross-Modal Adaptive Token Reduction for Efficient Audio-Visual LLMs
by: Gong, Chao, et al.
Published: (2025)
by: Gong, Chao, et al.
Published: (2025)
AffordanceSAM: Segment Anything Once More in Affordance Grounding
by: Jiang, Dengyang, et al.
Published: (2025)
by: Jiang, Dengyang, et al.
Published: (2025)
Towards Multimodal Lifelong Understanding: A Dataset and Agentic Baseline
by: Chen, Guo, et al.
Published: (2026)
by: Chen, Guo, et al.
Published: (2026)
InteractiveOmni: A Unified Omni-modal Model for Audio-Visual Multi-turn Dialogue
by: Tong, Wenwen, et al.
Published: (2025)
by: Tong, Wenwen, et al.
Published: (2025)
Epsilon: Exploring Comprehensive Visual-Semantic Projection for Multi-Label Zero-Shot Learning
by: Liu, Ziming, et al.
Published: (2024)
by: Liu, Ziming, et al.
Published: (2024)
CRAFT: A Neuro-Symbolic Framework for Visual Functional Affordance Grounding
by: Chen, Zhou, et al.
Published: (2025)
by: Chen, Zhou, et al.
Published: (2025)
Visual Affordance Prediction: Survey and Reproducibility
by: Apicella, Tommaso, et al.
Published: (2025)
by: Apicella, Tommaso, et al.
Published: (2025)
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing
by: Wang, Langyu, et al.
Published: (2025)
by: Wang, Langyu, et al.
Published: (2025)
AVS-Net: Audio-Visual Scale Net for Self-supervised Monocular Metric Depth Estimation
by: Liu, Xiaohu, et al.
Published: (2024)
by: Liu, Xiaohu, et al.
Published: (2024)
Affordance-First Decomposition for Continual Learning in Video-Language Understanding
by: Xu, Mengzhu, et al.
Published: (2025)
by: Xu, Mengzhu, et al.
Published: (2025)
AVS-Mamba: Exploring Temporal and Multi-modal Mamba for Audio-Visual Segmentation
by: Gong, Sitong, et al.
Published: (2025)
by: Gong, Sitong, et al.
Published: (2025)
Learning 2D Invariant Affordance Knowledge for 3D Affordance Grounding
by: Gao, Xianqiang, et al.
Published: (2024)
by: Gao, Xianqiang, et al.
Published: (2024)
Selective Noise Suppression and Discriminative Mutual Interaction for Robust Audio-Visual Segmentation
by: Peng, Kai, et al.
Published: (2026)
by: Peng, Kai, et al.
Published: (2026)
AffordanceLLM: Grounding Affordance from Vision Language Models
by: Qian, Shengyi, et al.
Published: (2024)
by: Qian, Shengyi, et al.
Published: (2024)
Object Affordance Recognition and Grounding via Multi-scale Cross-modal Representation Learning
by: Wan, Xinhang, et al.
Published: (2025)
by: Wan, Xinhang, et al.
Published: (2025)
Panoramic Affordance Prediction
by: Zhang, Zixin, et al.
Published: (2026)
by: Zhang, Zixin, et al.
Published: (2026)
Beyond Text: Frozen Large Language Models in Visual Signal Comprehension
by: Zhu, Lei, et al.
Published: (2024)
by: Zhu, Lei, et al.
Published: (2024)
Learning Visual Generative Priors without Text
by: Ma, Shuailei, et al.
Published: (2024)
by: Ma, Shuailei, et al.
Published: (2024)
FitControler: Toward Fit-Aware Virtual Try-On
by: Yang, Lu, et al.
Published: (2025)
by: Yang, Lu, et al.
Published: (2025)
SpongeBob: Sync-Aware Harmonious Audio-Visual Generative Editing
by: Liang, Sen, et al.
Published: (2026)
by: Liang, Sen, et al.
Published: (2026)
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing
by: Wang, Langyu, et al.
Published: (2024)
by: Wang, Langyu, et al.
Published: (2024)
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
by: Chen, Guo, et al.
Published: (2024)
by: Chen, Guo, et al.
Published: (2024)
Robust Audio-Visual Segmentation via Audio-Guided Visual Convergent Alignment
by: Liu, Chen, et al.
Published: (2025)
by: Liu, Chen, et al.
Published: (2025)
Affordance-R1: Reinforcement Learning for Generalizable Affordance Reasoning in Multimodal Large Language Model
by: Wang, Hanqing, et al.
Published: (2025)
by: Wang, Hanqing, et al.
Published: (2025)
Native Audio-Visual Alignment for Generation
by: Ji, Longbin, et al.
Published: (2026)
by: Ji, Longbin, et al.
Published: (2026)
Uncertainty Estimation in Instance Segmentation of Affordances via Bayesian Visual Transformers
by: Mur-Labadia, Lorenzo, et al.
Published: (2026)
by: Mur-Labadia, Lorenzo, et al.
Published: (2026)
Similar Items
-
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
by: Lu, Lidong, et al.
Published: (2025) -
Potential Field as Scene Affordance for Behavior Change-Based Visual Risk Object Identification
by: Pao, Pang-Yuan, et al.
Published: (2024) -
Complementary and Contrastive Learning for Audio-Visual Segmentation
by: Gong, Sitong, et al.
Published: (2025) -
Visual-Geometric Collaborative Guidance for Affordance Learning
by: Luo, Hongchen, et al.
Published: (2024) -
CorrAdaptor: Adaptive Local Context Learning for Correspondence Pruning
by: Zhu, Wei, et al.
Published: (2024)