Surprisingly Fragile: Assessing and Addressing Prompt Instability in Multimodal Foundation Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Stewart, Ian, Horawalavithana, Sameera, Kennedy, Brendan, Munikoti, Sai, Pazdernik, Karl |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SCITUNE: Aligning Large Language Models with Human-Curated Scientific Multimodal Instructions
por: Horawalavithana, Sameera, et al.
Publicado: (2023)
por: Horawalavithana, Sameera, et al.
Publicado: (2023)
Surprise Calibration for Better In-Context Learning
por: Tan, Zhihang, et al.
Publicado: (2025)
por: Tan, Zhihang, et al.
Publicado: (2025)
Back to the Barn with LLAMAs: Evolving Pretrained LLM Backbones in Finetuning Vision Language Models
por: Horawalavithana, Sameera, et al.
Publicado: (2026)
por: Horawalavithana, Sameera, et al.
Publicado: (2026)
Whose wife is it anyway? Assessing bias against same-gender relationships in machine translation
por: Stewart, Ian, et al.
Publicado: (2024)
por: Stewart, Ian, et al.
Publicado: (2024)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
por: Ashuach, Tomer, et al.
Publicado: (2025)
por: Ashuach, Tomer, et al.
Publicado: (2025)
Named Entity Recognition for Address Extraction in Speech-to-Text Transcriptions Using Synthetic Data
por: Lajčinová, Bibiána, et al.
Publicado: (2024)
por: Lajčinová, Bibiána, et al.
Publicado: (2024)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
por: Peters, Sydney, et al.
Publicado: (2025)
por: Peters, Sydney, et al.
Publicado: (2025)
Refining Packing and Shuffling Strategies for Enhanced Performance in Generative Language Models
por: Chen, Yanbing, et al.
Publicado: (2024)
por: Chen, Yanbing, et al.
Publicado: (2024)
Have Multimodal Large Language Models (MLLMs) Really Learned to Tell the Time on Analog Clocks?
por: Fu, Tairan, et al.
Publicado: (2025)
por: Fu, Tairan, et al.
Publicado: (2025)
Scaling In, Not Up? Testing Thick Citation Context Analysis with GPT-5 and Fragile Prompts
por: Simons, Arno
Publicado: (2026)
por: Simons, Arno
Publicado: (2026)
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
por: Collado-Montañez, Jaime, et al.
Publicado: (2025)
por: Collado-Montañez, Jaime, et al.
Publicado: (2025)
MULTISEISMO: A Multimodal Seismic Dataset and Model for Cross-Modal Seismic Understanding
por: Munikoti, Sai, et al.
Publicado: (2026)
por: Munikoti, Sai, et al.
Publicado: (2026)
Learning Translations via Matrix Completion
por: Wijaya, Derry, et al.
Publicado: (2024)
por: Wijaya, Derry, et al.
Publicado: (2024)
AMELI: Enhancing Multimodal Entity Linking with Fine-Grained Attributes
por: Yao, Barry Menglong, et al.
Publicado: (2023)
por: Yao, Barry Menglong, et al.
Publicado: (2023)
Assessing RAG and HyDE on 1B vs. 4B-Parameter Gemma LLMs for Personal Assistants Integretion
por: Sorstkins, Andrejs
Publicado: (2025)
por: Sorstkins, Andrejs
Publicado: (2025)
MCP: A Control-Theoretic Orchestration Framework for Synergistic Efficiency and Interpretability in Multimodal Large Language Models
por: Zhang, Luyan
Publicado: (2025)
por: Zhang, Luyan
Publicado: (2025)
EmoS: A High-Fidelity Multimodal Benchmark for Fine-grained Streaming Emotional Understanding
por: Guo, Pengze, et al.
Publicado: (2026)
por: Guo, Pengze, et al.
Publicado: (2026)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
por: Saji, Alan, et al.
Publicado: (2025)
por: Saji, Alan, et al.
Publicado: (2025)
Rethinking Addressing in Language Models via Contexualized Equivariant Positional Encoding
por: Zhu, Jiajun, et al.
Publicado: (2025)
por: Zhu, Jiajun, et al.
Publicado: (2025)
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment
por: Bian, Zhipeng, et al.
Publicado: (2026)
por: Bian, Zhipeng, et al.
Publicado: (2026)
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
por: Smădu, Răzvan-Alexandru, et al.
Publicado: (2025)
por: Smădu, Răzvan-Alexandru, et al.
Publicado: (2025)
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
por: Tong, Jingqi, et al.
Publicado: (2025)
por: Tong, Jingqi, et al.
Publicado: (2025)
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
por: Dai, Song, et al.
Publicado: (2025)
por: Dai, Song, et al.
Publicado: (2025)
Generalist Multimodal AI: A Review of Architectures, Challenges and Opportunities
por: Munikoti, Sai, et al.
Publicado: (2024)
por: Munikoti, Sai, et al.
Publicado: (2024)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
por: Oketunji, Abiodun Finbarrs
Publicado: (2023)
por: Oketunji, Abiodun Finbarrs
Publicado: (2023)
KinyaColBERT: A Lexically Grounded Retrieval Model for Low-Resource Retrieval-Augmented Generation
por: Nzeyimana, Antoine, et al.
Publicado: (2025)
por: Nzeyimana, Antoine, et al.
Publicado: (2025)
Are Generative Models Underconfident? Better Quality Estimation with Boosted Model Probability
por: Dinh, Tu Anh, et al.
Publicado: (2025)
por: Dinh, Tu Anh, et al.
Publicado: (2025)
Adapting Multilingual Models to Code-Mixed Tasks via Model Merging
por: Kodali, Prashant, et al.
Publicado: (2025)
por: Kodali, Prashant, et al.
Publicado: (2025)
Prompt-Time Symbolic Knowledge Capture with Large Language Models
por: Çöplü, Tolga, et al.
Publicado: (2024)
por: Çöplü, Tolga, et al.
Publicado: (2024)
Exploring the Maze of Multilingual Modeling
por: Nezhad, Sina Bagheri, et al.
Publicado: (2023)
por: Nezhad, Sina Bagheri, et al.
Publicado: (2023)
LinkNER: Linking Local Named Entity Recognition Models to Large Language Models using Uncertainty
por: Zhang, Zhen, et al.
Publicado: (2024)
por: Zhang, Zhen, et al.
Publicado: (2024)
Precise Length Control in Large Language Models
por: Butcher, Bradley, et al.
Publicado: (2024)
por: Butcher, Bradley, et al.
Publicado: (2024)
What Drives Performance in Multilingual Language Models?
por: Nezhad, Sina Bagheri, et al.
Publicado: (2024)
por: Nezhad, Sina Bagheri, et al.
Publicado: (2024)
Investigating the Impact of Text Summarization on Topic Modeling
por: Khandelwal, Trishia
Publicado: (2024)
por: Khandelwal, Trishia
Publicado: (2024)
Improving QA Model Performance with Cartographic Inoculation
por: Chen, Allen, et al.
Publicado: (2024)
por: Chen, Allen, et al.
Publicado: (2024)
Large Language Models for Biomedical Article Classification
por: Proboszcz, Jakub, et al.
Publicado: (2026)
por: Proboszcz, Jakub, et al.
Publicado: (2026)
Arabic Hate Speech Identification and Masking in Social Media using Deep Learning Models and Pre-trained Models Fine-tuning
por: Doghmash, Salam Thabet, et al.
Publicado: (2025)
por: Doghmash, Salam Thabet, et al.
Publicado: (2025)
Is Our Chatbot Telling Lies? Assessing Correctness of an LLM-based Dutch Support Chatbot
por: Lassche, Herman, et al.
Publicado: (2024)
por: Lassche, Herman, et al.
Publicado: (2024)
Strategy Adaptation in Large Language Model Werewolf Agents
por: Nakamori, Fuya, et al.
Publicado: (2025)
por: Nakamori, Fuya, et al.
Publicado: (2025)
Test-Time Scaling of Reasoning Models for Machine Translation
por: Li, Zihao, et al.
Publicado: (2025)
por: Li, Zihao, et al.
Publicado: (2025)
Ejemplares similares
-
SCITUNE: Aligning Large Language Models with Human-Curated Scientific Multimodal Instructions
por: Horawalavithana, Sameera, et al.
Publicado: (2023) -
Surprise Calibration for Better In-Context Learning
por: Tan, Zhihang, et al.
Publicado: (2025) -
Back to the Barn with LLAMAs: Evolving Pretrained LLM Backbones in Finetuning Vision Language Models
por: Horawalavithana, Sameera, et al.
Publicado: (2026) -
Whose wife is it anyway? Assessing bias against same-gender relationships in machine translation
por: Stewart, Ian, et al.
Publicado: (2024) -
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
por: Ashuach, Tomer, et al.
Publicado: (2025)