LLaVA-FA: Learning Fourier Approximation for Compressing Large Multimodal Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zheng, Pengcheng, Zhang, Chaoning, Mo, Jiarong, Li, GuoHui, Zhang, Jiaquan, Zhang, Jiahao, Cao, Sihan, Zheng, Sheng, Qin, Caiyan, Wang, Guoqing, Yang, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings
von: Yan, Dawei, et al.
Veröffentlicht: (2024)
von: Yan, Dawei, et al.
Veröffentlicht: (2024)
Language-Guided Token Compression with Reinforcement Learning in Large Vision-Language Models
von: Cao, Sihan, et al.
Veröffentlicht: (2026)
von: Cao, Sihan, et al.
Veröffentlicht: (2026)
MC-LLaVA: Multi-Concept Personalized Vision-Language Model
von: An, Ruichuan, et al.
Veröffentlicht: (2025)
von: An, Ruichuan, et al.
Veröffentlicht: (2025)
MC-LLaVA: Multi-Concept Personalized Vision-Language Model
von: An, Ruichuan, et al.
Veröffentlicht: (2024)
von: An, Ruichuan, et al.
Veröffentlicht: (2024)
Relaxing Anchor-Frame Dominance for Mitigating Hallucinations in Video Large Language Models
von: Liu, Zijian, et al.
Veröffentlicht: (2026)
von: Liu, Zijian, et al.
Veröffentlicht: (2026)
Topology-Aware Layer Pruning for Large Vision-Language Models
von: Zheng, Pengcheng, et al.
Veröffentlicht: (2026)
von: Zheng, Pengcheng, et al.
Veröffentlicht: (2026)
LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
von: Zhang, Shaolei, et al.
Veröffentlicht: (2025)
von: Zhang, Shaolei, et al.
Veröffentlicht: (2025)
Immunizing 3D Gaussian Generative Models Against Unauthorized Fine-Tuning via Attribute-Space Traps
von: Zhang, Jianwei, et al.
Veröffentlicht: (2026)
von: Zhang, Jianwei, et al.
Veröffentlicht: (2026)
RCP: Representation Consistency Pruner for Mitigating Distribution Shift in Large Vision-Language Models
von: Zhang, Jianwei, et al.
Veröffentlicht: (2026)
von: Zhang, Jianwei, et al.
Veröffentlicht: (2026)
Agri-LLaVA: Knowledge-Infused Large Multimodal Assistant on Agricultural Pests and Diseases
von: Wang, Liqiong, et al.
Veröffentlicht: (2024)
von: Wang, Liqiong, et al.
Veröffentlicht: (2024)
Rethinking Input Domains in Physics-Informed Neural Networks via Geometric Compactification Mappings
von: Huang, Zhenzhen, et al.
Veröffentlicht: (2026)
von: Huang, Zhenzhen, et al.
Veröffentlicht: (2026)
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
von: Shu, Fangxun, et al.
Veröffentlicht: (2024)
von: Shu, Fangxun, et al.
Veröffentlicht: (2024)
Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
von: Zhang, Yi-Fan, et al.
Veröffentlicht: (2024)
von: Zhang, Yi-Fan, et al.
Veröffentlicht: (2024)
Cosmos-LLaVA: Chatting with the Visual Cosmos-LLaVA: Görselle Sohbet Etmek
von: Zeer, Ahmed, et al.
Veröffentlicht: (2024)
von: Zeer, Ahmed, et al.
Veröffentlicht: (2024)
LLaVA-Read: Enhancing Reading Ability of Multimodal Language Models
von: Zhang, Ruiyi, et al.
Veröffentlicht: (2024)
von: Zhang, Ruiyi, et al.
Veröffentlicht: (2024)
LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
von: An, Xiang, et al.
Veröffentlicht: (2025)
von: An, Xiang, et al.
Veröffentlicht: (2025)
LLaVA-Critic: Learning to Evaluate Multimodal Models
von: Xiong, Tianyi, et al.
Veröffentlicht: (2024)
von: Xiong, Tianyi, et al.
Veröffentlicht: (2024)
From Similarity to Structure: Training-free LLM Context Compression with Hybrid Graph Priors
von: Zhou, Yitian, et al.
Veröffentlicht: (2026)
von: Zhou, Yitian, et al.
Veröffentlicht: (2026)
MLLM-LLaVA-FL: Multimodal Large Language Model Assisted Federated Learning
von: Zhang, Jianyi, et al.
Veröffentlicht: (2024)
von: Zhang, Jianyi, et al.
Veröffentlicht: (2024)
LLaVA-OneVision: Easy Visual Task Transfer
von: Li, Bo, et al.
Veröffentlicht: (2024)
von: Li, Bo, et al.
Veröffentlicht: (2024)
Fast SAM2 with Text-Driven Token Pruning
von: Mandal, Avilasha, et al.
Veröffentlicht: (2025)
von: Mandal, Avilasha, et al.
Veröffentlicht: (2025)
LLaVAC: Fine-tuning LLaVA as a Multimodal Sentiment Classifier
von: Chay-intr, T., et al.
Veröffentlicht: (2025)
von: Chay-intr, T., et al.
Veröffentlicht: (2025)
Amharic LLaMA and LLaVA: Multimodal LLMs for Low Resource Languages
von: Andersland, Michael
Veröffentlicht: (2024)
von: Andersland, Michael
Veröffentlicht: (2024)
Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models
von: Shi, Wenhao, et al.
Veröffentlicht: (2024)
von: Shi, Wenhao, et al.
Veröffentlicht: (2024)
Purrfessor: A Fine-tuned Multimodal LLaVA Diet Health Chatbot
von: Lu, Linqi, et al.
Veröffentlicht: (2024)
von: Lu, Linqi, et al.
Veröffentlicht: (2024)
WSI-LLaVA: A Multimodal Large Language Model for Whole Slide Image
von: Liang, Yuci, et al.
Veröffentlicht: (2024)
von: Liang, Yuci, et al.
Veröffentlicht: (2024)
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence
von: An, Xiang, et al.
Veröffentlicht: (2026)
von: An, Xiang, et al.
Veröffentlicht: (2026)
LLaVA-KD: A Framework of Distilling Multimodal Large Language Models
von: Cai, Yuxuan, et al.
Veröffentlicht: (2024)
von: Cai, Yuxuan, et al.
Veröffentlicht: (2024)
Geometric Neural Operators via Lie Group-Constrained Latent Dynamics
von: Zhang, Jiaquan, et al.
Veröffentlicht: (2026)
von: Zhang, Jiaquan, et al.
Veröffentlicht: (2026)
Power-LLaVA: Large Language and Vision Assistant for Power Transmission Line Inspection
von: Wang, Jiahao, et al.
Veröffentlicht: (2024)
von: Wang, Jiahao, et al.
Veröffentlicht: (2024)
Causal-LLaVA: Causal Disentanglement for Mitigating Hallucination in Multimodal Large Language Models
von: Hu, Xinmiao, et al.
Veröffentlicht: (2025)
von: Hu, Xinmiao, et al.
Veröffentlicht: (2025)
LLaVA-UHD v3: Progressive Visual Compression for Efficient Native-Resolution Encoding in MLLMs
von: Sun, Shichu, et al.
Veröffentlicht: (2025)
von: Sun, Shichu, et al.
Veröffentlicht: (2025)
Enhance Image-to-Image Generation with LLaVA-generated Prompts
von: Ding, Zhicheng, et al.
Veröffentlicht: (2024)
von: Ding, Zhicheng, et al.
Veröffentlicht: (2024)
Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs
von: Caffagni, Davide, et al.
Veröffentlicht: (2024)
von: Caffagni, Davide, et al.
Veröffentlicht: (2024)
LLaVA-Video: Video Instruction Tuning With Synthetic Data
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2024)
LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information
von: Wang, Ke, et al.
Veröffentlicht: (2024)
von: Wang, Ke, et al.
Veröffentlicht: (2024)
LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs
von: Sun, Boyuan, et al.
Veröffentlicht: (2025)
von: Sun, Boyuan, et al.
Veröffentlicht: (2025)
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
von: Li, Feng, et al.
Veröffentlicht: (2024)
von: Li, Feng, et al.
Veröffentlicht: (2024)
Continual LLaVA: Continual Instruction Tuning in Large Vision-Language Models
von: Cao, Meng, et al.
Veröffentlicht: (2024)
von: Cao, Meng, et al.
Veröffentlicht: (2024)
LLaVA-NeuMT: Selective Layer-Neuron Modulation for Efficient Multilingual Multimodal Translation
von: Wei, Jingxuan, et al.
Veröffentlicht: (2025)
von: Wei, Jingxuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings
von: Yan, Dawei, et al.
Veröffentlicht: (2024) -
Language-Guided Token Compression with Reinforcement Learning in Large Vision-Language Models
von: Cao, Sihan, et al.
Veröffentlicht: (2026) -
MC-LLaVA: Multi-Concept Personalized Vision-Language Model
von: An, Ruichuan, et al.
Veröffentlicht: (2025) -
MC-LLaVA: Multi-Concept Personalized Vision-Language Model
von: An, Ruichuan, et al.
Veröffentlicht: (2024) -
Relaxing Anchor-Frame Dominance for Mitigating Hallucinations in Video Large Language Models
von: Liu, Zijian, et al.
Veröffentlicht: (2026)