MoLT: Mixture of Layer-Wise Tokens for Efficient Audio-Visual Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rho, Kyeongha, Lee, Hyeongkeun, Cho, Jae Won, Chung, Joon Son |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport
von: Rho, Kyeongha, et al.
Veröffentlicht: (2025)
von: Rho, Kyeongha, et al.
Veröffentlicht: (2025)
EquiAV: Leveraging Equivariance for Audio-Visual Contrastive Learning
von: Kim, Jongsuk, et al.
Veröffentlicht: (2024)
von: Kim, Jongsuk, et al.
Veröffentlicht: (2024)
Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment
von: Senocak, Arda, et al.
Veröffentlicht: (2024)
von: Senocak, Arda, et al.
Veröffentlicht: (2024)
Plug-and-Steer: Decoupling Separation and Selection in Audio-Visual Target Speaker Extraction
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
Cinematic Audio Source Separation Using Visual Cues
von: Zhang, Kang, et al.
Veröffentlicht: (2026)
von: Zhang, Kang, et al.
Veröffentlicht: (2026)
Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2026)
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2026)
DASH: Dynamic Audio-Driven Semantic Chunking for Efficient Omnimodal Token Compression
von: Li, Bingzhou, et al.
Veröffentlicht: (2026)
von: Li, Bingzhou, et al.
Veröffentlicht: (2026)
Schrodinger Audio-Visual Editor: Object-Level Audiovisual Removal
von: Xu, Weihan, et al.
Veröffentlicht: (2025)
von: Xu, Weihan, et al.
Veröffentlicht: (2025)
Cross-Modal Binary Attention: An Energy-Efficient Fusion Framework for Audio-Visual Learning
von: Saleh, Mohamed, et al.
Veröffentlicht: (2026)
von: Saleh, Mohamed, et al.
Veröffentlicht: (2026)
MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
Audio-Guided Visual Perception for Audio-Visual Navigation
von: Wang, Yi, et al.
Veröffentlicht: (2025)
von: Wang, Yi, et al.
Veröffentlicht: (2025)
AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers
von: Araujo, Edson, et al.
Veröffentlicht: (2026)
von: Araujo, Edson, et al.
Veröffentlicht: (2026)
Audio-Visual World Models: Towards Multisensory Imagination in Sight and Sound
von: Wang, Jiahua, et al.
Veröffentlicht: (2025)
von: Wang, Jiahua, et al.
Veröffentlicht: (2025)
OmniForcing: Unleashing Real-time Joint Audio-Visual Generation
von: Su, Yaofeng, et al.
Veröffentlicht: (2026)
von: Su, Yaofeng, et al.
Veröffentlicht: (2026)
SonoWorld: From One Image to a 3D Audio-Visual Scene
von: Jin, Derong, et al.
Veröffentlicht: (2026)
von: Jin, Derong, et al.
Veröffentlicht: (2026)
MultiSoundGen: Video-to-Audio Generation for Multi-Event Scenarios via SlowFast Contrastive Audio-Visual Pretraining and Direct Preference Optimization
von: Yang, Jianxuan, et al.
Veröffentlicht: (2025)
von: Yang, Jianxuan, et al.
Veröffentlicht: (2025)
READ-Net: Clarifying Emotional Ambiguity via Adaptive Feature Recalibration for Audio-Visual Depression Detection
von: Chen, Chenglizhao, et al.
Veröffentlicht: (2026)
von: Chen, Chenglizhao, et al.
Veröffentlicht: (2026)
AV-Edit: Multimodal Generative Sound Effect Editing via Audio-Visual Semantic Joint Control
von: Guo, Xinyue, et al.
Veröffentlicht: (2025)
von: Guo, Xinyue, et al.
Veröffentlicht: (2025)
SeeingSounds: Learning Audio-to-Visual Alignment via Text
von: Carnemolla, Simone, et al.
Veröffentlicht: (2025)
von: Carnemolla, Simone, et al.
Veröffentlicht: (2025)
Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
AudioStory: Generating Long-Form Narrative Audio with Large Language Models
von: Guo, Yuxin, et al.
Veröffentlicht: (2025)
von: Guo, Yuxin, et al.
Veröffentlicht: (2025)
Learning to Visually Localize Sound Sources from Mixtures without Prior Source Knowledge
von: Kim, Dongjin, et al.
Veröffentlicht: (2024)
von: Kim, Dongjin, et al.
Veröffentlicht: (2024)
MA-AVT: Modality Alignment for Parameter-Efficient Audio-Visual Transformers
von: Mahmud, Tanvir, et al.
Veröffentlicht: (2024)
von: Mahmud, Tanvir, et al.
Veröffentlicht: (2024)
AccKV: Towards Efficient Audio-Video LLMs Inference via Adaptive-Focusing and Cross-Calibration KV Cache Optimization
von: Jiang, Zhonghua, et al.
Veröffentlicht: (2025)
von: Jiang, Zhonghua, et al.
Veröffentlicht: (2025)
Hierarchical Semantic Correlation-Aware Masked Autoencoder for Unsupervised Audio-Visual Representation Learning
von: Zeng, Donghuo, et al.
Veröffentlicht: (2026)
von: Zeng, Donghuo, et al.
Veröffentlicht: (2026)
LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV
von: Liu, Tengfei, et al.
Veröffentlicht: (2026)
von: Liu, Tengfei, et al.
Veröffentlicht: (2026)
PAVAS: Physics-Aware Video-to-Audio Synthesis
von: Hyun-Bin, Oh, et al.
Veröffentlicht: (2025)
von: Hyun-Bin, Oh, et al.
Veröffentlicht: (2025)
Learning Self-Supervised Audio-Visual Representations for Sound Recommendations
von: Krishnamurthy, Sudha
Veröffentlicht: (2024)
von: Krishnamurthy, Sudha
Veröffentlicht: (2024)
DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information
von: Nakada, Shota, et al.
Veröffentlicht: (2024)
von: Nakada, Shota, et al.
Veröffentlicht: (2024)
Spiking Tucker Fusion Transformer for Audio-Visual Zero-Shot Learning
von: Li, Wenrui, et al.
Veröffentlicht: (2024)
von: Li, Wenrui, et al.
Veröffentlicht: (2024)
Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence
von: Liao, Junchao, et al.
Veröffentlicht: (2026)
von: Liao, Junchao, et al.
Veröffentlicht: (2026)
JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation
von: Liu, Kai, et al.
Veröffentlicht: (2026)
von: Liu, Kai, et al.
Veröffentlicht: (2026)
Omni2Sound: Towards Unified Video-Text-to-Audio Generation
von: Dai, Yusheng, et al.
Veröffentlicht: (2026)
von: Dai, Yusheng, et al.
Veröffentlicht: (2026)
3D Audio-Visual Segmentation
von: Sokolov, Artem, et al.
Veröffentlicht: (2024)
von: Sokolov, Artem, et al.
Veröffentlicht: (2024)
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text
von: Pian, Weiguo, et al.
Veröffentlicht: (2026)
von: Pian, Weiguo, et al.
Veröffentlicht: (2026)
UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation
von: Li, Hebeizi, et al.
Veröffentlicht: (2026)
von: Li, Hebeizi, et al.
Veröffentlicht: (2026)
Discrepancy-Aware Attention Network for Enhanced Audio-Visual Zero-Shot Learning
von: Yu, RunLin, et al.
Veröffentlicht: (2024)
von: Yu, RunLin, et al.
Veröffentlicht: (2024)
TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation
von: Liu, Xiangyu, et al.
Veröffentlicht: (2026)
von: Liu, Xiangyu, et al.
Veröffentlicht: (2026)
Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation
von: Chen, Yuheng, et al.
Veröffentlicht: (2026)
von: Chen, Yuheng, et al.
Veröffentlicht: (2026)
ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling
von: Yang, Jianxuan, et al.
Veröffentlicht: (2026)
von: Yang, Jianxuan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport
von: Rho, Kyeongha, et al.
Veröffentlicht: (2025) -
EquiAV: Leveraging Equivariance for Audio-Visual Contrastive Learning
von: Kim, Jongsuk, et al.
Veröffentlicht: (2024) -
Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment
von: Senocak, Arda, et al.
Veröffentlicht: (2024) -
Plug-and-Steer: Decoupling Separation and Selection in Audio-Visual Target Speaker Extraction
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026) -
Cinematic Audio Source Separation Using Visual Cues
von: Zhang, Kang, et al.
Veröffentlicht: (2026)