Adapting LLaMA Decoder to Vision Transformer
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Jiahao, Shao, Wenqi, Chen, Mengzhao, Wu, Chengyue, Liu, Yong, Wu, Taiqiang, Zhang, Kaipeng, Zhang, Songyang, Chen, Kai, Luo, Ping |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LiT: Delving into a Simple Linear Diffusion Transformer for Image Generation
von: Wang, Jiahao, et al.
Veröffentlicht: (2025)
von: Wang, Jiahao, et al.
Veröffentlicht: (2025)
VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks
von: Chu, Xiangxiang, et al.
Veröffentlicht: (2024)
von: Chu, Xiangxiang, et al.
Veröffentlicht: (2024)
LLaMA-Reg: Using LLaMA 2 for Unsupervised Medical Image Registration
von: Ma, Mingrui, et al.
Veröffentlicht: (2024)
von: Ma, Mingrui, et al.
Veröffentlicht: (2024)
Dia-LLaMA: Towards Large Language Model-driven CT Report Generation
von: Chen, Zhixuan, et al.
Veröffentlicht: (2024)
von: Chen, Zhixuan, et al.
Veröffentlicht: (2024)
VoCo-LLaMA: Towards Vision Compression with Large Language Models
von: Ye, Xubing, et al.
Veröffentlicht: (2024)
von: Ye, Xubing, et al.
Veröffentlicht: (2024)
LLaMA Pro: Progressive LLaMA with Block Expansion
von: Wu, Chengyue, et al.
Veröffentlicht: (2024)
von: Wu, Chengyue, et al.
Veröffentlicht: (2024)
Efficient LLaMA-3.2-Vision by Trimming Cross-attended Visual Features
von: Lee, Jewon, et al.
Veröffentlicht: (2025)
von: Lee, Jewon, et al.
Veröffentlicht: (2025)
Enhance-A-Video: Better Generated Video for Free
von: Luo, Yang, et al.
Veröffentlicht: (2025)
von: Luo, Yang, et al.
Veröffentlicht: (2025)
EMO-LLaMA: Enhancing Facial Emotion Understanding with Instruction Tuning
von: Xing, Bohao, et al.
Veröffentlicht: (2024)
von: Xing, Bohao, et al.
Veröffentlicht: (2024)
LLaMA-XR: A Novel Framework for Radiology Report Generation using LLaMA and QLoRA Fine Tuning
von: Jahangir, Md. Zihad Bin, et al.
Veröffentlicht: (2025)
von: Jahangir, Md. Zihad Bin, et al.
Veröffentlicht: (2025)
Multimodal Medical Disease Classification with LLaMA II
von: Gapp, Christian, et al.
Veröffentlicht: (2024)
von: Gapp, Christian, et al.
Veröffentlicht: (2024)
What If We Recaption Billions of Web Images with LLaMA-3?
von: Li, Xianhang, et al.
Veröffentlicht: (2024)
von: Li, Xianhang, et al.
Veröffentlicht: (2024)
Dynamic Multimodal Evaluation with Flexible Complexity by Vision-Language Bootstrapping
von: Yang, Yue, et al.
Veröffentlicht: (2024)
von: Yang, Yue, et al.
Veröffentlicht: (2024)
ViT3D Alignment of LLaMA3: 3D Medical Image Report Generation
von: Li, Siyou, et al.
Veröffentlicht: (2024)
von: Li, Siyou, et al.
Veröffentlicht: (2024)
EDVD-LLaMA: Explainable Deepfake Video Detection via Multimodal Large Language Model Reasoning
von: Sun, Haoran, et al.
Veröffentlicht: (2025)
von: Sun, Haoran, et al.
Veröffentlicht: (2025)
Vista-LLaMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens
von: Ma, Fan, et al.
Veröffentlicht: (2023)
von: Ma, Fan, et al.
Veröffentlicht: (2023)
Efficient High-Resolution Visual Representation Learning with State Space Model for Human Pose Estimation
von: Zhang, Hao, et al.
Veröffentlicht: (2024)
von: Zhang, Hao, et al.
Veröffentlicht: (2024)
LLaMA-Excitor: General Instruction Tuning via Indirect Feature Interaction
von: Zou, Bo, et al.
Veröffentlicht: (2024)
von: Zou, Bo, et al.
Veröffentlicht: (2024)
LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
von: Zhang, Renrui, et al.
Veröffentlicht: (2023)
von: Zhang, Renrui, et al.
Veröffentlicht: (2023)
MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions
von: Zhang, Hao, et al.
Veröffentlicht: (2024)
von: Zhang, Hao, et al.
Veröffentlicht: (2024)
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
von: Cheng, Zesen, et al.
Veröffentlicht: (2024)
von: Cheng, Zesen, et al.
Veröffentlicht: (2024)
SearchLVLMs: A Plug-and-Play Framework for Augmenting Large Vision-Language Models by Searching Up-to-Date Internet Knowledge
von: Li, Chuanhao, et al.
Veröffentlicht: (2024)
von: Li, Chuanhao, et al.
Veröffentlicht: (2024)
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
von: Zhang, Boqiang, et al.
Veröffentlicht: (2025)
von: Zhang, Boqiang, et al.
Veröffentlicht: (2025)
ChartAssisstant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning
von: Meng, Fanqing, et al.
Veröffentlicht: (2024)
von: Meng, Fanqing, et al.
Veröffentlicht: (2024)
Open-Vocabulary Animal Keypoint Detection with Semantic-feature Matching
von: Zhang, Hao, et al.
Veröffentlicht: (2023)
von: Zhang, Hao, et al.
Veröffentlicht: (2023)
ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification
von: He, Yefei, et al.
Veröffentlicht: (2024)
von: He, Yefei, et al.
Veröffentlicht: (2024)
SZTU-CMU at MER2024: Improving Emotion-LLaMA with Conv-Attention for Multimodal Emotion Recognition
von: Cheng, Zebang, et al.
Veröffentlicht: (2024)
von: Cheng, Zebang, et al.
Veröffentlicht: (2024)
LINA: Linear Autoregressive Image Generative Models with Continuous Tokens
von: Wang, Jiahao, et al.
Veröffentlicht: (2026)
von: Wang, Jiahao, et al.
Veröffentlicht: (2026)
SAMRefiner: Taming Segment Anything Model for Universal Mask Refinement
von: Lin, Yuqi, et al.
Veröffentlicht: (2025)
von: Lin, Yuqi, et al.
Veröffentlicht: (2025)
High-Accuracy ECG Image Interpretation using Parameter-Efficient LoRA Fine-Tuning with Multimodal LLaMA 3.2
von: M, Nandakishor, et al.
Veröffentlicht: (2025)
von: M, Nandakishor, et al.
Veröffentlicht: (2025)
From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language Models
von: Li, Rongjie, et al.
Veröffentlicht: (2024)
von: Li, Rongjie, et al.
Veröffentlicht: (2024)
Diffree: Text-Guided Shape Free Object Inpainting with Diffusion Model
von: Zhao, Lirui, et al.
Veröffentlicht: (2024)
von: Zhao, Lirui, et al.
Veröffentlicht: (2024)
MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
von: Meng, Fanqing, et al.
Veröffentlicht: (2024)
von: Meng, Fanqing, et al.
Veröffentlicht: (2024)
GUIOdyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices
von: Lu, Quanfeng, et al.
Veröffentlicht: (2024)
von: Lu, Quanfeng, et al.
Veröffentlicht: (2024)
MC-LLaVA: Multi-Concept Personalized Vision-Language Model
von: An, Ruichuan, et al.
Veröffentlicht: (2025)
von: An, Ruichuan, et al.
Veröffentlicht: (2025)
PhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image Models
von: Meng, Fanqing, et al.
Veröffentlicht: (2024)
von: Meng, Fanqing, et al.
Veröffentlicht: (2024)
TP-Eval: Tap Multimodal LLMs' Potential in Evaluation by Customizing Prompts
von: Xie, Yuxuan, et al.
Veröffentlicht: (2024)
von: Xie, Yuxuan, et al.
Veröffentlicht: (2024)
FiT: Flexible Vision Transformer for Diffusion Model
von: Lu, Zeyu, et al.
Veröffentlicht: (2024)
von: Lu, Zeyu, et al.
Veröffentlicht: (2024)
MC-LLaVA: Multi-Concept Personalized Vision-Language Model
von: An, Ruichuan, et al.
Veröffentlicht: (2024)
von: An, Ruichuan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LiT: Delving into a Simple Linear Diffusion Transformer for Image Generation
von: Wang, Jiahao, et al.
Veröffentlicht: (2025) -
VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks
von: Chu, Xiangxiang, et al.
Veröffentlicht: (2024) -
LLaMA-Reg: Using LLaMA 2 for Unsupervised Medical Image Registration
von: Ma, Mingrui, et al.
Veröffentlicht: (2024) -
Dia-LLaMA: Towards Large Language Model-driven CT Report Generation
von: Chen, Zhixuan, et al.
Veröffentlicht: (2024) -
VoCo-LLaMA: Towards Vision Compression with Large Language Models
von: Ye, Xubing, et al.
Veröffentlicht: (2024)