Asymmetric Hierarchical Anchoring for Audio-Visual Joint Representation: Resolving Information Allocation Ambiguity for Robust Cross-Modal Generalization
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Bixing, Zhao, Yuhong, Ye, Zongli, Lian, Jiachen, Yue, Xiangyu, Anumanchipalli, Gopala |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Hierarchical Spoken Language Dysfluency Modeling
by: Lian, Jiachen, et al.
Published: (2024)
by: Lian, Jiachen, et al.
Published: (2024)
Unsupervised TTS Acoustic Modeling for TTS with Conditional Disentangled Sequential VAE
by: Lian, Jiachen, et al.
Published: (2022)
by: Lian, Jiachen, et al.
Published: (2022)
Rebuilding ROME : Resolving Model Collapse during Sequential Model Editing
by: Gupta, Akshat, et al.
Published: (2024)
by: Gupta, Akshat, et al.
Published: (2024)
Speech World Model: Causal State-Action Planning with Explicit Reasoning for Speech
by: Zhou, Xuanru, et al.
Published: (2025)
by: Zhou, Xuanru, et al.
Published: (2025)
Audio Texture Manipulation by Exemplar-Based Analogy
by: Cheng, Kan Jen, et al.
Published: (2025)
by: Cheng, Kan Jen, et al.
Published: (2025)
Self-Supervised Audio-Visual Soundscape Stylization
by: Li, Tingle, et al.
Published: (2024)
by: Li, Tingle, et al.
Published: (2024)
Teaching Machines to Speak Using Articulatory Control
by: Anand, Akshay, et al.
Published: (2025)
by: Anand, Akshay, et al.
Published: (2025)
TART: A Comprehensive Tool for Technique-Aware Audio-to-Tab Guitar Transcription
by: Gupta, Akshaj, et al.
Published: (2025)
by: Gupta, Akshaj, et al.
Published: (2025)
Sylber: Syllabic Embedding Representation of Speech from Raw Audio
by: Cho, Cheol Jun, et al.
Published: (2024)
by: Cho, Cheol Jun, et al.
Published: (2024)
A Unified Framework for Model Editing
by: Gupta, Akshat, et al.
Published: (2024)
by: Gupta, Akshat, et al.
Published: (2024)
Is Bigger Edit Batch Size Always Better? -- An Empirical Study on Model Editing with Llama-3
by: Yoon, Junsang, et al.
Published: (2024)
by: Yoon, Junsang, et al.
Published: (2024)
Geometric Interpretation of Layer Normalization and a Comparative Analysis with RMSNorm
by: Gupta, Akshat, et al.
Published: (2024)
by: Gupta, Akshat, et al.
Published: (2024)
Model Editing at Scale leads to Gradual and Catastrophic Forgetting
by: Gupta, Akshat, et al.
Published: (2024)
by: Gupta, Akshat, et al.
Published: (2024)
StyleStream: Real-Time Zero-Shot Voice Style Conversion
by: Liu, Yisi, et al.
Published: (2026)
by: Liu, Yisi, et al.
Published: (2026)
Self-Assessment Tests are Unreliable Measures of LLM Personality
by: Gupta, Akshat, et al.
Published: (2023)
by: Gupta, Akshat, et al.
Published: (2023)
HuPER: A Human-Inspired Framework for Phonetic Perception
by: Guo, Chenxu, et al.
Published: (2026)
by: Guo, Chenxu, et al.
Published: (2026)
AV-data2vec: Self-supervised Learning of Audio-Visual Speech Representations with Contextualized Target Representations
by: Lian, Jiachen, et al.
Published: (2023)
by: Lian, Jiachen, et al.
Published: (2023)
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues
by: Zhou, Dingkun, et al.
Published: (2025)
by: Zhou, Dingkun, et al.
Published: (2025)
RLBind: Adversarial-Invariant Cross-Modal Alignment for Unified Robust Embeddings
by: Lu, Yuhong
Published: (2025)
by: Lu, Yuhong
Published: (2025)
Time and Tokens: Benchmarking End-to-End Speech Dysfluency Detection
by: Zhou, Xuanru, et al.
Published: (2024)
by: Zhou, Xuanru, et al.
Published: (2024)
Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities
by: Lin, Guan-Ting, et al.
Published: (2025)
by: Lin, Guan-Ting, et al.
Published: (2025)
Sounding that Object: Interactive Object-Aware Image to Audio Generation
by: Li, Tingle, et al.
Published: (2025)
by: Li, Tingle, et al.
Published: (2025)
Heterogeneous Mean Field Games and Local Well-posedness
by: Qiao, Bixing
Published: (2025)
by: Qiao, Bixing
Published: (2025)
UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions
by: Zhang, Guozhen, et al.
Published: (2025)
by: Zhang, Guozhen, et al.
Published: (2025)
How Do LLMs Use Their Depth?
by: Gupta, Akshat, et al.
Published: (2025)
by: Gupta, Akshat, et al.
Published: (2025)
Efficient Knowledge Editing via Minimal Precomputation
by: Gupta, Akshat, et al.
Published: (2025)
by: Gupta, Akshat, et al.
Published: (2025)
Improving Joint Audio-Video Generation with Cross-Modal Context Learning
by: Ma, Bingqi, et al.
Published: (2026)
by: Ma, Bingqi, et al.
Published: (2026)
Cross-Modal Bottleneck Fusion For Noise Robust Audio-Visual Speech Recognition
by: Ok, Seaone, et al.
Published: (2026)
by: Ok, Seaone, et al.
Published: (2026)
PROMISE: Prompt-Attentive Hierarchical Contrastive Learning for Robust Cross-Modal Representation with Missing Modalities
by: Chen, Jiajun, et al.
Published: (2025)
by: Chen, Jiajun, et al.
Published: (2025)
Towards Robust Evaluation of Visual Activity Recognition: Resolving Verb Ambiguity with Sense Clustering
by: Yao, Louie Hong, et al.
Published: (2025)
by: Yao, Louie Hong, et al.
Published: (2025)
The Sound of Simulation: Learning Multimodal Sim-to-Real Robot Policies with Generative Audio
by: Wang, Renhao, et al.
Published: (2025)
by: Wang, Renhao, et al.
Published: (2025)
Audio-Visual Cross-Modal Compression for Generative Face Video Coding
by: Xu, Youmin, et al.
Published: (2025)
by: Xu, Youmin, et al.
Published: (2025)
Dysfluent WFST: A Framework for Zero-Shot Speech Dysfluency Transcription and Detection
by: Guo, Chenxu, et al.
Published: (2025)
by: Guo, Chenxu, et al.
Published: (2025)
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation
by: Tan, Weiting, et al.
Published: (2025)
by: Tan, Weiting, et al.
Published: (2025)
VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering
by: Meng, Yiran, et al.
Published: (2025)
by: Meng, Yiran, et al.
Published: (2025)
Seamless Dysfluent Speech Text Alignment for Disordered Speech Analysis
by: Ye, Zongli, et al.
Published: (2025)
by: Ye, Zongli, et al.
Published: (2025)
Aligning Audio-Visual Joint Representations with an Agentic Workflow
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation Learning
by: Qian, Chengxuan, et al.
Published: (2025)
by: Qian, Chengxuan, et al.
Published: (2025)
Anchoring the Eigengap: Cross-Modal Spectral Stabilization for Sample-Efficient Representation Learning
by: Dhinagar, Nikhil J., et al.
Published: (2026)
by: Dhinagar, Nikhil J., et al.
Published: (2026)
RT-VC: Real-Time Zero-Shot Voice Conversion with Speech Articulatory Coding
by: Liu, Yisi, et al.
Published: (2025)
by: Liu, Yisi, et al.
Published: (2025)
Similar Items
-
Towards Hierarchical Spoken Language Dysfluency Modeling
by: Lian, Jiachen, et al.
Published: (2024) -
Unsupervised TTS Acoustic Modeling for TTS with Conditional Disentangled Sequential VAE
by: Lian, Jiachen, et al.
Published: (2022) -
Rebuilding ROME : Resolving Model Collapse during Sequential Model Editing
by: Gupta, Akshat, et al.
Published: (2024) -
Speech World Model: Causal State-Action Planning with Explicit Reasoning for Speech
by: Zhou, Xuanru, et al.
Published: (2025) -
Audio Texture Manipulation by Exemplar-Based Analogy
by: Cheng, Kan Jen, et al.
Published: (2025)