Back to the Barn with LLAMAs: Evolving Pretrained LLM Backbones in Finetuning Vision Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Horawalavithana, Sameera, Phillips, Lauren, Stewart, Ian, Munikoti, Sai, Pazdernik, Karl |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SCITUNE: Aligning Large Language Models with Human-Curated Scientific Multimodal Instructions
by: Horawalavithana, Sameera, et al.
Published: (2023)
by: Horawalavithana, Sameera, et al.
Published: (2023)
Surprisingly Fragile: Assessing and Addressing Prompt Instability in Multimodal Foundation Models
by: Stewart, Ian, et al.
Published: (2024)
by: Stewart, Ian, et al.
Published: (2024)
MULTISEISMO: A Multimodal Seismic Dataset and Model for Cross-Modal Seismic Understanding
by: Munikoti, Sai, et al.
Published: (2026)
by: Munikoti, Sai, et al.
Published: (2026)
Generalist Multimodal AI: A Review of Architectures, Challenges and Opportunities
by: Munikoti, Sai, et al.
Published: (2024)
by: Munikoti, Sai, et al.
Published: (2024)
Reward Design for Physical Reasoning in Vision-Language Models
by: Lilienthal, Derek, et al.
Published: (2026)
by: Lilienthal, Derek, et al.
Published: (2026)
Audit, Alignment, and Optimization of LM-Powered Subroutines with Application to Public Comment Processing
by: Raab, Reilly, et al.
Published: (2025)
by: Raab, Reilly, et al.
Published: (2025)
Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone
by: Ye, Jiacheng, et al.
Published: (2025)
by: Ye, Jiacheng, et al.
Published: (2025)
Benchmarking LLMs for Environmental Review and Permitting
by: Meyur, Rounak, et al.
Published: (2024)
by: Meyur, Rounak, et al.
Published: (2024)
Directional Concentration Uncertainty: A representational approach to uncertainty quantification for generative models
by: Chattopadhyay, Souradeep, et al.
Published: (2026)
by: Chattopadhyay, Souradeep, et al.
Published: (2026)
WeQA: A Benchmark for Retrieval Augmented Generation in Wind Energy Domain
by: Meyur, Rounak, et al.
Published: (2024)
by: Meyur, Rounak, et al.
Published: (2024)
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
by: Mitra, Chancharik, et al.
Published: (2025)
by: Mitra, Chancharik, et al.
Published: (2025)
Malayalam Sign Language Identification using Finetuned YOLOv8 and Computer Vision Techniques
by: K., Abhinand, et al.
Published: (2024)
by: K., Abhinand, et al.
Published: (2024)
TTL: Test-time Textual Learning for OOD Detection with Pretrained Vision-Language Models
by: Ye, Jinlun, et al.
Published: (2026)
by: Ye, Jinlun, et al.
Published: (2026)
Pretraining Vision-Language Model for Difference Visual Question Answering in Longitudinal Chest X-rays
by: Cho, Yeongjae, et al.
Published: (2024)
by: Cho, Yeongjae, et al.
Published: (2024)
TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones
by: Yuan, Zhengqing, et al.
Published: (2023)
by: Yuan, Zhengqing, et al.
Published: (2023)
FrEVL: Leveraging Frozen Pretrained Embeddings for Efficient Vision-Language Understanding
by: Bourigault, Emmanuelle, et al.
Published: (2025)
by: Bourigault, Emmanuelle, et al.
Published: (2025)
PhenoLIP: Integrating Phenotype Ontology Knowledge into Medical Vision-Language Pretraining
by: Liang, Cheng, et al.
Published: (2026)
by: Liang, Cheng, et al.
Published: (2026)
Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary Segmentation
by: Barsellotti, Luca, et al.
Published: (2024)
by: Barsellotti, Luca, et al.
Published: (2024)
Finetuned Multimodal Language Models Are High-Quality Image-Text Data Filters
by: Wang, Weizhi, et al.
Published: (2024)
by: Wang, Weizhi, et al.
Published: (2024)
EvolveNav: Empowering LLM-Based Vision-Language Navigation via Self-Improving Embodied Reasoning
by: Lin, Bingqian, et al.
Published: (2025)
by: Lin, Bingqian, et al.
Published: (2025)
Modeling Caption Diversity in Contrastive Vision-Language Pretraining
by: Lavoie, Samuel, et al.
Published: (2024)
by: Lavoie, Samuel, et al.
Published: (2024)
Maya: An Instruction Finetuned Multilingual Multimodal Model
by: Alam, Nahid, et al.
Published: (2024)
by: Alam, Nahid, et al.
Published: (2024)
Mordal: Automated Pretrained Model Selection for Vision Language Models
by: He, Shiqi, et al.
Published: (2025)
by: He, Shiqi, et al.
Published: (2025)
ETPNav: Evolving Topological Planning for Vision-Language Navigation in Continuous Environments
by: An, Dong, et al.
Published: (2023)
by: An, Dong, et al.
Published: (2023)
Vision-and-Language Navigation Generative Pretrained Transformer
by: Hanlin, Wen
Published: (2024)
by: Hanlin, Wen
Published: (2024)
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos
by: Hannan, Tanveer, et al.
Published: (2024)
by: Hannan, Tanveer, et al.
Published: (2024)
Evaluation of Cultural Competence of Vision-Language Models
by: Yadav, Srishti, et al.
Published: (2025)
by: Yadav, Srishti, et al.
Published: (2025)
VLSM-Adapter: Finetuning Vision-Language Segmentation Efficiently with Lightweight Blocks
by: Dhakal, Manish, et al.
Published: (2024)
by: Dhakal, Manish, et al.
Published: (2024)
Renaissance: Investigating the Pretraining of Vision-Language Encoders
by: Fields, Clayton, et al.
Published: (2024)
by: Fields, Clayton, et al.
Published: (2024)
Do VLMs Have a Moral Backbone? A Study on the Fragile Morality of Vision-Language Models
by: Liu, Zhining, et al.
Published: (2026)
by: Liu, Zhining, et al.
Published: (2026)
VisPlay: Self-Evolving Vision-Language Models from Images
by: He, Yicheng, et al.
Published: (2025)
by: He, Yicheng, et al.
Published: (2025)
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining
by: Zhang, Wenqi, et al.
Published: (2025)
by: Zhang, Wenqi, et al.
Published: (2025)
Language-Pretraining-Induced Bias: A Strong Foundation for General Vision Tasks
by: Luo, Yaxin, et al.
Published: (2026)
by: Luo, Yaxin, et al.
Published: (2026)
How Does Vision-Language Adaptation Impact the Safety of Vision Language Models?
by: Lee, Seongyun, et al.
Published: (2024)
by: Lee, Seongyun, et al.
Published: (2024)
Learning to Read Where to Look: Disease-Aware Vision-Language Pretraining for 3D CT
by: Ging, Simon, et al.
Published: (2026)
by: Ging, Simon, et al.
Published: (2026)
Conflict Adaptation in Vision-Language Models
by: Hu, Xiaoyang
Published: (2025)
by: Hu, Xiaoyang
Published: (2025)
Vision Language Models are Confused Tourists
by: Irawan, Patrick Amadeus, et al.
Published: (2025)
by: Irawan, Patrick Amadeus, et al.
Published: (2025)
LLM-CXR: Instruction-Finetuned LLM for CXR Image Understanding and Generation
by: Lee, Suhyeon, et al.
Published: (2023)
by: Lee, Suhyeon, et al.
Published: (2023)
FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models
by: Cai, Hengxing, et al.
Published: (2025)
by: Cai, Hengxing, et al.
Published: (2025)
Finetune Like You Pretrain: Boosting Zero-shot Adversarial Robustness in Vision-language Models
by: Xing, Songlong, et al.
Published: (2026)
by: Xing, Songlong, et al.
Published: (2026)
Similar Items
-
SCITUNE: Aligning Large Language Models with Human-Curated Scientific Multimodal Instructions
by: Horawalavithana, Sameera, et al.
Published: (2023) -
Surprisingly Fragile: Assessing and Addressing Prompt Instability in Multimodal Foundation Models
by: Stewart, Ian, et al.
Published: (2024) -
MULTISEISMO: A Multimodal Seismic Dataset and Model for Cross-Modal Seismic Understanding
by: Munikoti, Sai, et al.
Published: (2026) -
Generalist Multimodal AI: A Review of Architectures, Challenges and Opportunities
by: Munikoti, Sai, et al.
Published: (2024) -
Reward Design for Physical Reasoning in Vision-Language Models
by: Lilienthal, Derek, et al.
Published: (2026)