Extreme Miscalibration and the Illusion of Adversarial Robustness
Fuente:
arXiv
Salvato in:
| Autori principali: | Raina, Vyas, Tan, Samson, Cevher, Volkan, Rawal, Aditya, Zha, Sheng, Karypis, George |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Revisiting SMoE Language Models by Evaluating Inefficiencies with Task Specific Expert Pruning
di: Sarkar, Soumajyoti, et al.
Pubblicazione: (2024)
di: Sarkar, Soumajyoti, et al.
Pubblicazione: (2024)
Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment
di: Raina, Vyas, et al.
Pubblicazione: (2024)
di: Raina, Vyas, et al.
Pubblicazione: (2024)
Learning to Generate Answers with Citations via Factual Consistency Models
di: Aly, Rami, et al.
Pubblicazione: (2024)
di: Aly, Rami, et al.
Pubblicazione: (2024)
Controlling Whisper: Universal Acoustic Adversarial Attacks to Control Speech Foundation Models
di: Raina, Vyas, et al.
Pubblicazione: (2024)
di: Raina, Vyas, et al.
Pubblicazione: (2024)
Differentially Private Bias-Term Fine-tuning of Foundation Models
di: Bu, Zhiqi, et al.
Pubblicazione: (2022)
di: Bu, Zhiqi, et al.
Pubblicazione: (2022)
Pack of LLMs: Model Fusion at Test-Time via Perplexity Optimization
di: Mavromatis, Costas, et al.
Pubblicazione: (2024)
di: Mavromatis, Costas, et al.
Pubblicazione: (2024)
Certified Robustness Under Bounded Levenshtein Distance
di: Rocamora, Elias Abad, et al.
Pubblicazione: (2025)
di: Rocamora, Elias Abad, et al.
Pubblicazione: (2025)
SemPool: Simple, robust, and interpretable KG pooling for enhancing language models
di: Mavromatis, Costas, et al.
Pubblicazione: (2024)
di: Mavromatis, Costas, et al.
Pubblicazione: (2024)
Chameleon: A Flexible Data-mixing Framework for Language Model Pretraining and Finetuning
di: Xie, Wanyun, et al.
Pubblicazione: (2025)
di: Xie, Wanyun, et al.
Pubblicazione: (2025)
Extending Input Contexts of Language Models through Training on Segmented Sequences
di: Karypis, Petros, et al.
Pubblicazione: (2023)
di: Karypis, Petros, et al.
Pubblicazione: (2023)
Universal Acoustic Adversarial Attacks for Flexible Control of Speech-LLMs
di: Ma, Rao, et al.
Pubblicazione: (2025)
di: Ma, Rao, et al.
Pubblicazione: (2025)
Revisiting Character-level Adversarial Attacks for Language Models
di: Rocamora, Elias Abad, et al.
Pubblicazione: (2024)
di: Rocamora, Elias Abad, et al.
Pubblicazione: (2024)
Sequence-level Large Language Model Training with Contrastive Preference Optimization
di: Feng, Zhili, et al.
Pubblicazione: (2025)
di: Feng, Zhili, et al.
Pubblicazione: (2025)
ESLM: Risk-Averse Selective Language Modeling for Efficient Pretraining
di: Bal, Melis Ilayda, et al.
Pubblicazione: (2025)
di: Bal, Melis Ilayda, et al.
Pubblicazione: (2025)
DEM: Distribution Edited Model for Training with Mixed Data Distributions
di: Ram, Dhananjay, et al.
Pubblicazione: (2024)
di: Ram, Dhananjay, et al.
Pubblicazione: (2024)
Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation Models
di: Raina, Vyas, et al.
Pubblicazione: (2024)
di: Raina, Vyas, et al.
Pubblicazione: (2024)
GNN-RAG: Graph Neural Retrieval for Large Language Model Reasoning
di: Mavromatis, Costas, et al.
Pubblicazione: (2024)
di: Mavromatis, Costas, et al.
Pubblicazione: (2024)
GRASP: Deterministic argument ranking in interaction graphs
di: Misra, Diganta, et al.
Pubblicazione: (2026)
di: Misra, Diganta, et al.
Pubblicazione: (2026)
Hallucination, Monofacts, and Miscalibration: An Empirical Investigation
di: Miao, Miranda Muqing, et al.
Pubblicazione: (2025)
di: Miao, Miranda Muqing, et al.
Pubblicazione: (2025)
Not Every Token Needs Forgetting: Selective Unlearning to Limit Change in Utility in Large Language Model Unlearning
di: Wan, Yixin, et al.
Pubblicazione: (2025)
di: Wan, Yixin, et al.
Pubblicazione: (2025)
LLM Task Interference: An Initial Study on the Impact of Task-Switch in Conversational History
di: Gupta, Akash, et al.
Pubblicazione: (2024)
di: Gupta, Akash, et al.
Pubblicazione: (2024)
Leveraging the Context through Multi-Round Interactions for Jailbreaking Attacks
di: Cheng, Yixin, et al.
Pubblicazione: (2024)
di: Cheng, Yixin, et al.
Pubblicazione: (2024)
Fine-Tuning Language Models on Multiple Datasets for Citation Intention Classification
di: Shui, Zeren, et al.
Pubblicazione: (2024)
di: Shui, Zeren, et al.
Pubblicazione: (2024)
The Confidence Dichotomy: Analyzing and Mitigating Miscalibration in Tool-Use Agents
di: Xuan, Weihao, et al.
Pubblicazione: (2026)
di: Xuan, Weihao, et al.
Pubblicazione: (2026)
Incongruent Positivity: When Miscalibrated Positivity Undermines Online Supportive Conversations
di: Almajed, Leen, et al.
Pubblicazione: (2025)
di: Almajed, Leen, et al.
Pubblicazione: (2025)
Funny or Persuasive, but Not Both: Evaluating Fine-Grained Multi-Concept Control in LLMs
di: Labroo, Arya, et al.
Pubblicazione: (2026)
di: Labroo, Arya, et al.
Pubblicazione: (2026)
Large Language Models are Miscalibrated In-Context Learners
di: Li, Chengzu, et al.
Pubblicazione: (2023)
di: Li, Chengzu, et al.
Pubblicazione: (2023)
Selective Rotary Position Embedding
di: Movahedi, Sajad, et al.
Pubblicazione: (2025)
di: Movahedi, Sajad, et al.
Pubblicazione: (2025)
Mapping from Meaning: Addressing the Miscalibration of Prompt-Sensitive Language Models
di: Cox, Kyle, et al.
Pubblicazione: (2025)
di: Cox, Kyle, et al.
Pubblicazione: (2025)
AIGeN: An Adversarial Approach for Instruction Generation in VLN
di: Rawal, Niyati, et al.
Pubblicazione: (2024)
di: Rawal, Niyati, et al.
Pubblicazione: (2024)
A Course Correction in Steerability Evaluation: Revealing Miscalibration and Side Effects in LLMs
di: Chang, Trenton, et al.
Pubblicazione: (2025)
di: Chang, Trenton, et al.
Pubblicazione: (2025)
On Overcoming Miscalibrated Conversational Priors in LLM-based Chatbots
di: Herlihy, Christine, et al.
Pubblicazione: (2024)
di: Herlihy, Christine, et al.
Pubblicazione: (2024)
Single-pass Detection of Jailbreaking Input in Large Language Models
di: Candogan, Leyla Naz, et al.
Pubblicazione: (2025)
di: Candogan, Leyla Naz, et al.
Pubblicazione: (2025)
Unlearning as multi-task optimization: A normalized gradient difference approach with an adaptive learning rate
di: Bu, Zhiqi, et al.
Pubblicazione: (2024)
di: Bu, Zhiqi, et al.
Pubblicazione: (2024)
The Illusion-Illusion: Vision Language Models See Illusions Where There are None
di: Ullman, Tomer
Pubblicazione: (2024)
di: Ullman, Tomer
Pubblicazione: (2024)
Blending Is All You Need: Cheaper, Better Alternative to Trillion-Parameters LLM
di: Lu, Xiaoding, et al.
Pubblicazione: (2024)
di: Lu, Xiaoding, et al.
Pubblicazione: (2024)
Multi-Step Alignment as Markov Games: An Optimistic Online Gradient Descent Approach with Convergence Guarantees
di: Wu, Yongtao, et al.
Pubblicazione: (2025)
di: Wu, Yongtao, et al.
Pubblicazione: (2025)
Beyond instruction-conditioning, MoTE: Mixture of Task Experts for Multi-task Embedding Models
di: Romero, Miguel, et al.
Pubblicazione: (2025)
di: Romero, Miguel, et al.
Pubblicazione: (2025)
False Fixed Points: Kantian Feedback, Stable Miscalibration, and Representational Compression in LLMs
di: Okutomi, Akira
Pubblicazione: (2025)
di: Okutomi, Akira
Pubblicazione: (2025)
Question-Based Retrieval using Atomic Units for Enterprise RAG
di: Raina, Vatsal, et al.
Pubblicazione: (2024)
di: Raina, Vatsal, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Revisiting SMoE Language Models by Evaluating Inefficiencies with Task Specific Expert Pruning
di: Sarkar, Soumajyoti, et al.
Pubblicazione: (2024) -
Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment
di: Raina, Vyas, et al.
Pubblicazione: (2024) -
Learning to Generate Answers with Citations via Factual Consistency Models
di: Aly, Rami, et al.
Pubblicazione: (2024) -
Controlling Whisper: Universal Acoustic Adversarial Attacks to Control Speech Foundation Models
di: Raina, Vyas, et al.
Pubblicazione: (2024) -
Differentially Private Bias-Term Fine-tuning of Foundation Models
di: Bu, Zhiqi, et al.
Pubblicazione: (2022)