LLaVA-NeuMT: Selective Layer-Neuron Modulation for Efficient Multilingual Multimodal Translation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wei, Jingxuan, Jia, Caijun, Chen, Qi, Cai, Yujun, Sun, Linzhuang, Zhang, Xiangxiang, Wu, Gaowei, Yu, Bihui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs
von: Caffagni, Davide, et al.
Veröffentlicht: (2024)
von: Caffagni, Davide, et al.
Veröffentlicht: (2024)
Can Sound Replace Vision in LLaVA With Token Substitution?
von: Vosoughi, Ali, et al.
Veröffentlicht: (2025)
von: Vosoughi, Ali, et al.
Veröffentlicht: (2025)
ResearchPulse: Building Method-Experiment Chains through Multi-Document Scientific Inference
von: Chen, Qi, et al.
Veröffentlicht: (2025)
von: Chen, Qi, et al.
Veröffentlicht: (2025)
A Survey on Image-text Multimodal Models
von: Guo, Ruifeng, et al.
Veröffentlicht: (2023)
von: Guo, Ruifeng, et al.
Veröffentlicht: (2023)
LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs
von: Sun, Boyuan, et al.
Veröffentlicht: (2025)
von: Sun, Boyuan, et al.
Veröffentlicht: (2025)
LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning
von: Cocchi, Federico, et al.
Veröffentlicht: (2025)
von: Cocchi, Federico, et al.
Veröffentlicht: (2025)
Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning
von: Cheng, Zebang, et al.
Veröffentlicht: (2024)
von: Cheng, Zebang, et al.
Veröffentlicht: (2024)
Virbo: Multimodal Multilingual Avatar Video Generation in Digital Marketing
von: Zhang, Juan, et al.
Veröffentlicht: (2024)
von: Zhang, Juan, et al.
Veröffentlicht: (2024)
Concept Drift Guided LayerNorm Tuning for Efficient Multimodal Metaphor Identification
von: Qian, Wenhao, et al.
Veröffentlicht: (2025)
von: Qian, Wenhao, et al.
Veröffentlicht: (2025)
MIDI-LLaMA: An Instruction-Following Multimodal LLM for Symbolic Music Understanding
von: Yang, Meng, et al.
Veröffentlicht: (2026)
von: Yang, Meng, et al.
Veröffentlicht: (2026)
LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via a Hybrid Architecture
von: Wang, Xidong, et al.
Veröffentlicht: (2024)
von: Wang, Xidong, et al.
Veröffentlicht: (2024)
TraveLLaMA: A Multimodal Travel Assistant with Large-Scale Dataset and Structured Reasoning
von: Chu, Meng, et al.
Veröffentlicht: (2025)
von: Chu, Meng, et al.
Veröffentlicht: (2025)
DAT: Dual-Aware Adaptive Transmission for Efficient Multimodal LLM Inference in Edge-Cloud Systems
von: Guo, Qi, et al.
Veröffentlicht: (2026)
von: Guo, Qi, et al.
Veröffentlicht: (2026)
AxiomVision: Accuracy-Guaranteed Adaptive Visual Model Selection for Perspective-Aware Video Analytics
von: Dai, Xiangxiang, et al.
Veröffentlicht: (2024)
von: Dai, Xiangxiang, et al.
Veröffentlicht: (2024)
Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding
von: Wang, Shaoguang, et al.
Veröffentlicht: (2026)
von: Wang, Shaoguang, et al.
Veröffentlicht: (2026)
From Natural Alignment to Conditional Controllability in Multimodal Dialogue
von: Jin, Zeyu, et al.
Veröffentlicht: (2026)
von: Jin, Zeyu, et al.
Veröffentlicht: (2026)
MInD: Improving Multimodal Sentiment Analysis via Multimodal Information Disentanglement
von: Dai, Weichen, et al.
Veröffentlicht: (2024)
von: Dai, Weichen, et al.
Veröffentlicht: (2024)
Bridging Discrete and Continuous: A Multimodal Strategy for Complex Emotion Detection
von: Jia, Jiehui, et al.
Veröffentlicht: (2024)
von: Jia, Jiehui, et al.
Veröffentlicht: (2024)
DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian Culture
von: Maji, Arijit, et al.
Veröffentlicht: (2025)
von: Maji, Arijit, et al.
Veröffentlicht: (2025)
NeuSaver: Neural Adaptive Power Consumption Optimization for Mobile Video Streaming
von: Park, Kyoungjun, et al.
Veröffentlicht: (2021)
von: Park, Kyoungjun, et al.
Veröffentlicht: (2021)
Geoint-R1: Formalizing Multimodal Geometric Reasoning with Dynamic Auxiliary Constructions
von: Wei, Jingxuan, et al.
Veröffentlicht: (2025)
von: Wei, Jingxuan, et al.
Veröffentlicht: (2025)
Multimodal Interaction Modeling via Self-Supervised Multi-Task Learning for Review Helpfulness Prediction
von: Gong, HongLin, et al.
Veröffentlicht: (2024)
von: Gong, HongLin, et al.
Veröffentlicht: (2024)
Is One-Shot In-Context Learning Helpful for Data Selection in Task-Specific Fine-Tuning of Multimodal LLMs?
von: An, Xiao, et al.
Veröffentlicht: (2026)
von: An, Xiao, et al.
Veröffentlicht: (2026)
Enhancing Neural Adaptive Wireless Video Streaming via Lower-Layer Information Exposure and Online Tuning
von: Zhao, Lingzhi, et al.
Veröffentlicht: (2025)
von: Zhao, Lingzhi, et al.
Veröffentlicht: (2025)
MM-InstructEval: Zero-Shot Evaluation of (Multimodal) Large Language Models on Multimodal Reasoning Tasks
von: Yang, Xiaocui, et al.
Veröffentlicht: (2024)
von: Yang, Xiaocui, et al.
Veröffentlicht: (2024)
SACRED: A Faithful Annotated Multimedia Multimodal Multilingual Dataset for Classifying Connectedness Types in Online Spirituality
von: Guan, Qinghao, et al.
Veröffentlicht: (2026)
von: Guan, Qinghao, et al.
Veröffentlicht: (2026)
MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
SZTU-CMU at MER2024: Improving Emotion-LLaMA with Conv-Attention for Multimodal Emotion Recognition
von: Cheng, Zebang, et al.
Veröffentlicht: (2024)
von: Cheng, Zebang, et al.
Veröffentlicht: (2024)
PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning
von: Lyu, Yibo, et al.
Veröffentlicht: (2025)
von: Lyu, Yibo, et al.
Veröffentlicht: (2025)
GGAvatar: Reconstructing Garment-Separated 3D Gaussian Splatting Avatars from Monocular Video
von: Chen, Jingxuan
Veröffentlicht: (2024)
von: Chen, Jingxuan
Veröffentlicht: (2024)
Multimodal Unlearnable Examples: Protecting Data against Multimodal Contrastive Learning
von: Liu, Xinwei, et al.
Veröffentlicht: (2024)
von: Liu, Xinwei, et al.
Veröffentlicht: (2024)
KeyVideoLLM: Towards Large-scale Video Keyframe Selection
von: Liang, Hao, et al.
Veröffentlicht: (2024)
von: Liang, Hao, et al.
Veröffentlicht: (2024)
Differential Mental Disorder Detection with Psychology-Inspired Multimodal Stimuli
von: Zhou, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Zhou, Zhiyuan, et al.
Veröffentlicht: (2026)
Zero-Shot Relational Learning for Multimodal Knowledge Graphs
von: Cai, Rui, et al.
Veröffentlicht: (2024)
von: Cai, Rui, et al.
Veröffentlicht: (2024)
LLaQo: Towards a Query-Based Coach in Expressive Music Performance Assessment
von: Zhang, Huan, et al.
Veröffentlicht: (2024)
von: Zhang, Huan, et al.
Veröffentlicht: (2024)
Multimodal Classification and Out-of-distribution Detection for Multimodal Intent Understanding
von: Zhang, Hanlei, et al.
Veröffentlicht: (2024)
von: Zhang, Hanlei, et al.
Veröffentlicht: (2024)
Interdisciplinary Translations: Sensory Perception as a Universal Language
von: Kang, Xindi, et al.
Veröffentlicht: (2024)
von: Kang, Xindi, et al.
Veröffentlicht: (2024)
Stable Multimodal Graph Unlearning via Feature-Dimension Aware Quantile Selection
von: Zhou, Jingjing, et al.
Veröffentlicht: (2026)
von: Zhou, Jingjing, et al.
Veröffentlicht: (2026)
Dark Side of Modalities: Reinforced Multimodal Distillation for Multimodal Knowledge Graph Reasoning
von: Zhao, Yu, et al.
Veröffentlicht: (2025)
von: Zhao, Yu, et al.
Veröffentlicht: (2025)
Hyperbolic Multimodal Generative Representation Learning for Generalized Zero-Shot Multimodal Information Extraction
von: Zhou, Baohang, et al.
Veröffentlicht: (2026)
von: Zhou, Baohang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs
von: Caffagni, Davide, et al.
Veröffentlicht: (2024) -
Can Sound Replace Vision in LLaVA With Token Substitution?
von: Vosoughi, Ali, et al.
Veröffentlicht: (2025) -
ResearchPulse: Building Method-Experiment Chains through Multi-Document Scientific Inference
von: Chen, Qi, et al.
Veröffentlicht: (2025) -
A Survey on Image-text Multimodal Models
von: Guo, Ruifeng, et al.
Veröffentlicht: (2023) -
LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs
von: Sun, Boyuan, et al.
Veröffentlicht: (2025)