VoCo-LLaMA: Towards Vision Compression with Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Xubing, Gan, Yukang, Huang, Xiaoke, Ge, Yixiao, Tang, Yansong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models
by: Ye, Xubing, et al.
Published: (2024)
by: Ye, Xubing, et al.
Published: (2024)
Adaptive-VoCo: Complexity-Aware Visual Token Compression for Vision-Language Models
by: Guo, Xiaoyang, et al.
Published: (2025)
by: Guo, Xiaoyang, et al.
Published: (2025)
VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks
by: Chu, Xiangxiang, et al.
Published: (2024)
by: Chu, Xiangxiang, et al.
Published: (2024)
Adapting LLaMA Decoder to Vision Transformer
by: Wang, Jiahao, et al.
Published: (2024)
by: Wang, Jiahao, et al.
Published: (2024)
LLaMA-Reg: Using LLaMA 2 for Unsupervised Medical Image Registration
by: Ma, Mingrui, et al.
Published: (2024)
by: Ma, Mingrui, et al.
Published: (2024)
Dia-LLaMA: Towards Large Language Model-driven CT Report Generation
by: Chen, Zhixuan, et al.
Published: (2024)
by: Chen, Zhixuan, et al.
Published: (2024)
LLaMA Pro: Progressive LLaMA with Block Expansion
by: Wu, Chengyue, et al.
Published: (2024)
by: Wu, Chengyue, et al.
Published: (2024)
EMO-LLaMA: Enhancing Facial Emotion Understanding with Instruction Tuning
by: Xing, Bohao, et al.
Published: (2024)
by: Xing, Bohao, et al.
Published: (2024)
EDVD-LLaMA: Explainable Deepfake Video Detection via Multimodal Large Language Model Reasoning
by: Sun, Haoran, et al.
Published: (2025)
by: Sun, Haoran, et al.
Published: (2025)
Vista-LLaMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens
by: Ma, Fan, et al.
Published: (2023)
by: Ma, Fan, et al.
Published: (2023)
Efficient LLaMA-3.2-Vision by Trimming Cross-attended Visual Features
by: Lee, Jewon, et al.
Published: (2025)
by: Lee, Jewon, et al.
Published: (2025)
LLaMA-XR: A Novel Framework for Radiology Report Generation using LLaMA and QLoRA Fine Tuning
by: Jahangir, Md. Zihad Bin, et al.
Published: (2025)
by: Jahangir, Md. Zihad Bin, et al.
Published: (2025)
Multimodal Medical Disease Classification with LLaMA II
by: Gapp, Christian, et al.
Published: (2024)
by: Gapp, Christian, et al.
Published: (2024)
What If We Recaption Billions of Web Images with LLaMA-3?
by: Li, Xianhang, et al.
Published: (2024)
by: Li, Xianhang, et al.
Published: (2024)
LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
by: Zhang, Renrui, et al.
Published: (2023)
by: Zhang, Renrui, et al.
Published: (2023)
TraveLLaMA: A Multimodal Travel Assistant with Large-Scale Dataset and Structured Reasoning
by: Chu, Meng, et al.
Published: (2025)
by: Chu, Meng, et al.
Published: (2025)
SEED-Story: Multimodal Long Story Generation with Large Language Model
by: Yang, Shuai, et al.
Published: (2024)
by: Yang, Shuai, et al.
Published: (2024)
LLaMA-Excitor: General Instruction Tuning via Indirect Feature Interaction
by: Zou, Bo, et al.
Published: (2024)
by: Zou, Bo, et al.
Published: (2024)
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
by: Lin, Bin, et al.
Published: (2024)
by: Lin, Bin, et al.
Published: (2024)
DiCoDe: Diffusion-Compressed Deep Tokens for Autoregressive Video Generation with Language Models
by: Li, Yizhuo, et al.
Published: (2024)
by: Li, Yizhuo, et al.
Published: (2024)
LLaVAShield: Safeguarding Multimodal Multi-Turn Dialogues in Vision-Language Models
by: Huang, Guolei, et al.
Published: (2025)
by: Huang, Guolei, et al.
Published: (2025)
OccLLaMA: An Occupancy-Language-Action Generative World Model for Autonomous Driving
by: Wei, Julong, et al.
Published: (2024)
by: Wei, Julong, et al.
Published: (2024)
CoLLaVO: Crayon Large Language and Vision mOdel
by: Lee, Byung-Kwan, et al.
Published: (2024)
by: Lee, Byung-Kwan, et al.
Published: (2024)
SZTU-CMU at MER2024: Improving Emotion-LLaMA with Conv-Attention for Multimodal Emotion Recognition
by: Cheng, Zebang, et al.
Published: (2024)
by: Cheng, Zebang, et al.
Published: (2024)
LLaMA-Mesh: Unifying 3D Mesh Generation with Language Models
by: Wang, Zhengyi, et al.
Published: (2024)
by: Wang, Zhengyi, et al.
Published: (2024)
ViT3D Alignment of LLaMA3: 3D Medical Image Report Generation
by: Li, Siyou, et al.
Published: (2024)
by: Li, Siyou, et al.
Published: (2024)
ST-LLM: Large Language Models Are Effective Temporal Learners
by: Liu, Ruyang, et al.
Published: (2024)
by: Liu, Ruyang, et al.
Published: (2024)
LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
by: Xu, Guowei, et al.
Published: (2024)
by: Xu, Guowei, et al.
Published: (2024)
ViVo: A Dataset for Volumetric Video Reconstruction and Compression
by: Azzarelli, Adrian, et al.
Published: (2025)
by: Azzarelli, Adrian, et al.
Published: (2025)
High-Accuracy ECG Image Interpretation using Parameter-Efficient LoRA Fine-Tuning with Multimodal LLaMA 3.2
by: M, Nandakishor, et al.
Published: (2025)
by: M, Nandakishor, et al.
Published: (2025)
MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens
by: Yeo, Jeong Hun, et al.
Published: (2025)
by: Yeo, Jeong Hun, et al.
Published: (2025)
MCA-LLaVA: Manhattan Causal Attention for Reducing Hallucination in Large Vision-Language Models
by: Zhao, Qiyan, et al.
Published: (2025)
by: Zhao, Qiyan, et al.
Published: (2025)
Q-VLM: Post-training Quantization for Large Vision-Language Models
by: Wang, Changyuan, et al.
Published: (2024)
by: Wang, Changyuan, et al.
Published: (2024)
Continual LLaVA: Continual Instruction Tuning in Large Vision-Language Models
by: Cao, Meng, et al.
Published: (2024)
by: Cao, Meng, et al.
Published: (2024)
Surgical-LLaVA: Toward Surgical Scenario Understanding via Large Language and Vision Models
by: Jin, Juseong, et al.
Published: (2024)
by: Jin, Juseong, et al.
Published: (2024)
Once for Both: Single Stage of Importance and Sparsity Search for Vision Transformer Compression
by: Ye, Hancheng, et al.
Published: (2024)
by: Ye, Hancheng, et al.
Published: (2024)
LLaVA-VSD: Large Language-and-Vision Assistant for Visual Spatial Description
by: Jin, Yizhang, et al.
Published: (2024)
by: Jin, Yizhang, et al.
Published: (2024)
LLaVA-KD: A Framework of Distilling Multimodal Large Language Models
by: Cai, Yuxuan, et al.
Published: (2024)
by: Cai, Yuxuan, et al.
Published: (2024)
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence
by: An, Xiang, et al.
Published: (2026)
by: An, Xiang, et al.
Published: (2026)
HiRes-LLaVA: Restoring Fragmentation Input in High-Resolution Large Vision-Language Models
by: Huang, Runhui, et al.
Published: (2024)
by: Huang, Runhui, et al.
Published: (2024)
Similar Items
-
ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models
by: Ye, Xubing, et al.
Published: (2024) -
Adaptive-VoCo: Complexity-Aware Visual Token Compression for Vision-Language Models
by: Guo, Xiaoyang, et al.
Published: (2025) -
VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks
by: Chu, Xiangxiang, et al.
Published: (2024) -
Adapting LLaMA Decoder to Vision Transformer
by: Wang, Jiahao, et al.
Published: (2024) -
LLaMA-Reg: Using LLaMA 2 for Unsupervised Medical Image Registration
by: Ma, Mingrui, et al.
Published: (2024)