LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via a Hybrid Architecture
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Xidong, Song, Dingjie, Chen, Shunian, Chen, Junyin, Cai, Zhenyang, Zhang, Chen, Sun, Lichao, Wang, Benyou |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs
di: Song, Dingjie, et al.
Pubblicazione: (2024)
di: Song, Dingjie, et al.
Pubblicazione: (2024)
Both Text and Images Leaked! A Systematic Analysis of Data Contamination in Multimodal LLM
di: Song, Dingjie, et al.
Pubblicazione: (2024)
di: Song, Dingjie, et al.
Pubblicazione: (2024)
LLaVA-NeuMT: Selective Layer-Neuron Modulation for Efficient Multilingual Multimodal Translation
di: Wei, Jingxuan, et al.
Pubblicazione: (2025)
di: Wei, Jingxuan, et al.
Pubblicazione: (2025)
Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization
di: Lai, Zhengzhao, et al.
Pubblicazione: (2025)
di: Lai, Zhengzhao, et al.
Pubblicazione: (2025)
Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs
di: Caffagni, Davide, et al.
Pubblicazione: (2024)
di: Caffagni, Davide, et al.
Pubblicazione: (2024)
TraveLLaMA: A Multimodal Travel Assistant with Large-Scale Dataset and Structured Reasoning
di: Chu, Meng, et al.
Pubblicazione: (2025)
di: Chu, Meng, et al.
Pubblicazione: (2025)
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
di: Chen, Junying, et al.
Pubblicazione: (2025)
di: Chen, Junying, et al.
Pubblicazione: (2025)
MileBench: Benchmarking MLLMs in Long Context
di: Song, Dingjie, et al.
Pubblicazione: (2024)
di: Song, Dingjie, et al.
Pubblicazione: (2024)
LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs
di: Sun, Boyuan, et al.
Pubblicazione: (2025)
di: Sun, Boyuan, et al.
Pubblicazione: (2025)
ShizhenGPT: Towards Multimodal LLMs for Traditional Chinese Medicine
di: Chen, Junying, et al.
Pubblicazione: (2025)
di: Chen, Junying, et al.
Pubblicazione: (2025)
Med-Banana-50K: A Cross-modality Large-Scale Dataset for Text-guided Medical Image Editing
di: Chen, Zhihui, et al.
Pubblicazione: (2025)
di: Chen, Zhihui, et al.
Pubblicazione: (2025)
LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning
di: Cocchi, Federico, et al.
Pubblicazione: (2025)
di: Cocchi, Federico, et al.
Pubblicazione: (2025)
Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging
di: Cai, Zhenyang, et al.
Pubblicazione: (2024)
di: Cai, Zhenyang, et al.
Pubblicazione: (2024)
Hand1000: Generating Realistic Hands from Text with Only 1,000 Images
di: Zhang, Haozhuo, et al.
Pubblicazione: (2024)
di: Zhang, Haozhuo, et al.
Pubblicazione: (2024)
Spatiotemporal Graph Guided Multi-modal Network for Livestreaming Product Retrieval
di: Hu, Xiaowan, et al.
Pubblicazione: (2024)
di: Hu, Xiaowan, et al.
Pubblicazione: (2024)
AIM: Let Any Multi-modal Large Language Models Embrace Efficient In-Context Learning
di: Gao, Jun, et al.
Pubblicazione: (2024)
di: Gao, Jun, et al.
Pubblicazione: (2024)
Tile Classification Based Viewport Prediction with Multi-modal Fusion Transformer
di: Zhang, Zhihao, et al.
Pubblicazione: (2023)
di: Zhang, Zhihao, et al.
Pubblicazione: (2023)
MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos
di: Wang, Rongsheng, et al.
Pubblicazione: (2025)
di: Wang, Rongsheng, et al.
Pubblicazione: (2025)
Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed Inputs
di: Ding, Peng, et al.
Pubblicazione: (2024)
di: Ding, Peng, et al.
Pubblicazione: (2024)
HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale
di: Chen, Junying, et al.
Pubblicazione: (2024)
di: Chen, Junying, et al.
Pubblicazione: (2024)
Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe Retrieval
di: Wang, Qing, et al.
Pubblicazione: (2025)
di: Wang, Qing, et al.
Pubblicazione: (2025)
OmnixR: Evaluating Omni-modality Language Models on Reasoning across Modalities
di: Chen, Lichang, et al.
Pubblicazione: (2024)
di: Chen, Lichang, et al.
Pubblicazione: (2024)
See or Guess: Counterfactually Regularized Image Captioning
di: Cao, Qian, et al.
Pubblicazione: (2024)
di: Cao, Qian, et al.
Pubblicazione: (2024)
Overcome Modal Bias in Multi-modal Federated Learning via Balanced Modality Selection
di: Fan, Yunfeng, et al.
Pubblicazione: (2023)
di: Fan, Yunfeng, et al.
Pubblicazione: (2023)
IBMEA: Exploring Variational Information Bottleneck for Multi-modal Entity Alignment
di: Su, Taoyu, et al.
Pubblicazione: (2024)
di: Su, Taoyu, et al.
Pubblicazione: (2024)
HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer
di: Cai, Qi, et al.
Pubblicazione: (2025)
di: Cai, Qi, et al.
Pubblicazione: (2025)
Not All Attention is Needed: Parameter and Computation Efficient Transfer Learning for Multi-modal Large Language Models
di: Wu, Qiong, et al.
Pubblicazione: (2024)
di: Wu, Qiong, et al.
Pubblicazione: (2024)
Improving Multi-modal Large Language Model through Boosting Vision Capabilities
di: Sun, Yanpeng, et al.
Pubblicazione: (2024)
di: Sun, Yanpeng, et al.
Pubblicazione: (2024)
MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
di: Wang, Chengyao, et al.
Pubblicazione: (2025)
di: Wang, Chengyao, et al.
Pubblicazione: (2025)
G4G:A Generic Framework for High Fidelity Talking Face Generation with Fine-grained Intra-modal Alignment
di: Zhang, Juan, et al.
Pubblicazione: (2024)
di: Zhang, Juan, et al.
Pubblicazione: (2024)
HDCompression: Hybrid-Diffusion Image Compression for Ultra-Low Bitrates
di: Lu, Lei, et al.
Pubblicazione: (2025)
di: Lu, Lei, et al.
Pubblicazione: (2025)
Learning Generalizable and Efficient Image Watermarking via Hierarchical Two-Stage Optimization
di: Liu, Ke, et al.
Pubblicazione: (2025)
di: Liu, Ke, et al.
Pubblicazione: (2025)
GMFVAD: Using Grained Multi-modal Feature to Improve Video Anomaly Detection
di: Dai, Guangyu, et al.
Pubblicazione: (2025)
di: Dai, Guangyu, et al.
Pubblicazione: (2025)
OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video
di: Pu, Junfu, et al.
Pubblicazione: (2026)
di: Pu, Junfu, et al.
Pubblicazione: (2026)
LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model
di: Sun, Yirong, et al.
Pubblicazione: (2025)
di: Sun, Yirong, et al.
Pubblicazione: (2025)
End-to-End RGB-IR Joint Image Compression With Channel-wise Cross-modality Entropy Model
di: Wang, Haofeng, et al.
Pubblicazione: (2025)
di: Wang, Haofeng, et al.
Pubblicazione: (2025)
Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding
di: Wang, Shaoguang, et al.
Pubblicazione: (2026)
di: Wang, Shaoguang, et al.
Pubblicazione: (2026)
Towards Efficient Low-rate Image Compression with Frequency-aware Diffusion Prior Refinement
di: Xia, Yichong, et al.
Pubblicazione: (2026)
di: Xia, Yichong, et al.
Pubblicazione: (2026)
Mixture-of-Prompt-Experts for Multi-modal Semantic Understanding
di: Wu, Zichen, et al.
Pubblicazione: (2024)
di: Wu, Zichen, et al.
Pubblicazione: (2024)
CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception
di: Li, Liupeng, et al.
Pubblicazione: (2026)
di: Li, Liupeng, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs
di: Song, Dingjie, et al.
Pubblicazione: (2024) -
Both Text and Images Leaked! A Systematic Analysis of Data Contamination in Multimodal LLM
di: Song, Dingjie, et al.
Pubblicazione: (2024) -
LLaVA-NeuMT: Selective Layer-Neuron Modulation for Efficient Multilingual Multimodal Translation
di: Wei, Jingxuan, et al.
Pubblicazione: (2025) -
Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization
di: Lai, Zhengzhao, et al.
Pubblicazione: (2025) -
Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs
di: Caffagni, Davide, et al.
Pubblicazione: (2024)