LLM Targeted Underperformance Disproportionately Impacts Vulnerable Users
Fuente:
arXiv
Salvato in:
| Autori principali: | Poole-Dayan, Elinor, Roy, Deb, Kabbara, Jad |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
An AI-Powered Framework for Analyzing Collective Idea Evolution in Deliberative Assemblies
di: Poole-Dayan, Elinor, et al.
Pubblicazione: (2025)
di: Poole-Dayan, Elinor, et al.
Pubblicazione: (2025)
On the Relationship between Truth and Political Bias in Language Models
di: Fulay, Suyash, et al.
Pubblicazione: (2024)
di: Fulay, Suyash, et al.
Pubblicazione: (2024)
Computational Analysis of Conversation Dynamics through Participant Responsivity
di: Hughes, Margaret, et al.
Pubblicazione: (2025)
di: Hughes, Margaret, et al.
Pubblicazione: (2025)
Confidence Under the Hood: An Investigation into the Confidence-Probability Alignment in Large Language Models
di: Kumar, Abhishek, et al.
Pubblicazione: (2024)
di: Kumar, Abhishek, et al.
Pubblicazione: (2024)
PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits
di: Jiang, Hang, et al.
Pubblicazione: (2023)
di: Jiang, Hang, et al.
Pubblicazione: (2023)
Bridging Context Gaps: Enhancing Comprehension in Long-Form Social Conversations Through Contextualized Excerpts
di: Mohanty, Shrestha, et al.
Pubblicazione: (2024)
di: Mohanty, Shrestha, et al.
Pubblicazione: (2024)
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
di: van der Weij, Teun, et al.
Pubblicazione: (2024)
di: van der Weij, Teun, et al.
Pubblicazione: (2024)
User-LLM: Efficient LLM Contextualization with User Embeddings
di: Ning, Lin, et al.
Pubblicazione: (2024)
di: Ning, Lin, et al.
Pubblicazione: (2024)
On the Limitations of Language Targeted Pruning: Investigating the Calibration Language Impact in Multilingual LLM Pruning
di: Kurz, Simon, et al.
Pubblicazione: (2024)
di: Kurz, Simon, et al.
Pubblicazione: (2024)
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
RLTHF: Targeted Human Feedback for LLM Alignment
di: Xu, Yifei, et al.
Pubblicazione: (2025)
di: Xu, Yifei, et al.
Pubblicazione: (2025)
LLM See, LLM Do: Guiding Data Generation to Target Non-Differentiable Objectives
di: Shimabucoro, Luísa, et al.
Pubblicazione: (2024)
di: Shimabucoro, Luísa, et al.
Pubblicazione: (2024)
LLMAP: LLM-Assisted Multi-Objective Route Planning with User Preferences
di: Yuan, Liangqi, et al.
Pubblicazione: (2025)
di: Yuan, Liangqi, et al.
Pubblicazione: (2025)
From Exact Hits to Close Enough: Semantic Caching for LLM Embeddings
di: Biton, Dvir David, et al.
Pubblicazione: (2026)
di: Biton, Dvir David, et al.
Pubblicazione: (2026)
The Impact of Language Mixing on Bilingual LLM Reasoning
di: Li, Yihao, et al.
Pubblicazione: (2025)
di: Li, Yihao, et al.
Pubblicazione: (2025)
Transparent Screening for LLM Inference and Training Impacts
di: Pachot, Arnault, et al.
Pubblicazione: (2026)
di: Pachot, Arnault, et al.
Pubblicazione: (2026)
Building, Reusing, and Generalizing Abstract Representations from Concrete Sequences
di: Wu, Shuchen, et al.
Pubblicazione: (2024)
di: Wu, Shuchen, et al.
Pubblicazione: (2024)
Factual Inconsistency in Data-to-Text Generation Scales Exponentially with LLM Size: A Statistical Validation
di: Mahapatra, Joy, et al.
Pubblicazione: (2025)
di: Mahapatra, Joy, et al.
Pubblicazione: (2025)
GanitLLM: Difficulty-Aware Bengali Mathematical Reasoning through Curriculum-GRPO
di: Dipta, Shubhashis Roy, et al.
Pubblicazione: (2026)
di: Dipta, Shubhashis Roy, et al.
Pubblicazione: (2026)
Systematically Analyzing Prompt Injection Vulnerabilities in Diverse LLM Architectures
di: Benjamin, Victoria, et al.
Pubblicazione: (2024)
di: Benjamin, Victoria, et al.
Pubblicazione: (2024)
Are My Optimized Prompts Compromised? Exploring Vulnerabilities of LLM-based Optimizers
di: Zhao, Andrew, et al.
Pubblicazione: (2025)
di: Zhao, Andrew, et al.
Pubblicazione: (2025)
When Greedy Wins: Emergent Exploitation Bias in Meta-Bandit LLM Training
di: Chen, Sanxing, et al.
Pubblicazione: (2025)
di: Chen, Sanxing, et al.
Pubblicazione: (2025)
JuStRank: Benchmarking LLM Judges for System Ranking
di: Gera, Ariel, et al.
Pubblicazione: (2024)
di: Gera, Ariel, et al.
Pubblicazione: (2024)
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
di: Yehudai, Asaf, et al.
Pubblicazione: (2025)
di: Yehudai, Asaf, et al.
Pubblicazione: (2025)
Survey on Evaluation of LLM-based Agents
di: Yehudai, Asaf, et al.
Pubblicazione: (2025)
di: Yehudai, Asaf, et al.
Pubblicazione: (2025)
HindSight: Evaluating LLM-Generated Research Ideas via Future Impact
di: Jiang, Bo
Pubblicazione: (2026)
di: Jiang, Bo
Pubblicazione: (2026)
Breach in the Shield: Unveiling the Vulnerabilities of Large Language Models
di: Dai, Runpeng, et al.
Pubblicazione: (2025)
di: Dai, Runpeng, et al.
Pubblicazione: (2025)
UserBench: An Interactive Gym Environment for User-Centric Agents
di: Qian, Cheng, et al.
Pubblicazione: (2025)
di: Qian, Cheng, et al.
Pubblicazione: (2025)
In Vino Veritas and Vulnerabilities: Examining LLM Safety via Drunk Language Inducement
di: Shetty, Anudeex, et al.
Pubblicazione: (2026)
di: Shetty, Anudeex, et al.
Pubblicazione: (2026)
Stay Tuned: An Empirical Study of the Impact of Hyperparameters on LLM Tuning in Real-World Applications
di: Halfon, Alon, et al.
Pubblicazione: (2024)
di: Halfon, Alon, et al.
Pubblicazione: (2024)
Accidental Vulnerability: Factors in Fine-Tuning that Shift Model Safeguards
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
To Think or Not to Think: Exploring the Unthinking Vulnerability in Large Reasoning Models
di: Zhu, Zihao, et al.
Pubblicazione: (2025)
di: Zhu, Zihao, et al.
Pubblicazione: (2025)
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing
di: Tran, Thien Q., et al.
Pubblicazione: (2025)
di: Tran, Thien Q., et al.
Pubblicazione: (2025)
Controllable User Simulation
di: Tennenholtz, Guy, et al.
Pubblicazione: (2026)
di: Tennenholtz, Guy, et al.
Pubblicazione: (2026)
UserSumBench: A Benchmark Framework for Evaluating User Summarization Approaches
di: Wang, Chao, et al.
Pubblicazione: (2024)
di: Wang, Chao, et al.
Pubblicazione: (2024)
UserRL: Training Interactive User-Centric Agent via Reinforcement Learning
di: Qian, Cheng, et al.
Pubblicazione: (2025)
di: Qian, Cheng, et al.
Pubblicazione: (2025)
Who Gets Which Message? Auditing Demographic Bias in LLM-Generated Targeted Text
di: Islam, Tunazzina
Pubblicazione: (2026)
di: Islam, Tunazzina
Pubblicazione: (2026)
Advancing Jailbreak Strategies: A Hybrid Approach to Exploiting LLM Vulnerabilities and Bypassing Modern Defenses
di: Ahmed, Mohamed, et al.
Pubblicazione: (2025)
di: Ahmed, Mohamed, et al.
Pubblicazione: (2025)
BiasJailbreak:Analyzing Ethical Biases and Jailbreak Vulnerabilities in Large Language Models
di: Lee, Isack, et al.
Pubblicazione: (2024)
di: Lee, Isack, et al.
Pubblicazione: (2024)
Estimating LLM Consistency: A User Baseline vs Surrogate Metrics
di: Wu, Xiaoyuan, et al.
Pubblicazione: (2025)
di: Wu, Xiaoyuan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
An AI-Powered Framework for Analyzing Collective Idea Evolution in Deliberative Assemblies
di: Poole-Dayan, Elinor, et al.
Pubblicazione: (2025) -
On the Relationship between Truth and Political Bias in Language Models
di: Fulay, Suyash, et al.
Pubblicazione: (2024) -
Computational Analysis of Conversation Dynamics through Participant Responsivity
di: Hughes, Margaret, et al.
Pubblicazione: (2025) -
Confidence Under the Hood: An Investigation into the Confidence-Probability Alignment in Large Language Models
di: Kumar, Abhishek, et al.
Pubblicazione: (2024) -
PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits
di: Jiang, Hang, et al.
Pubblicazione: (2023)