LLM Targeted Underperformance Disproportionately Impacts Vulnerable Users
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Poole-Dayan, Elinor, Roy, Deb, Kabbara, Jad |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
An AI-Powered Framework for Analyzing Collective Idea Evolution in Deliberative Assemblies
par: Poole-Dayan, Elinor, et autres
Publié: (2025)
par: Poole-Dayan, Elinor, et autres
Publié: (2025)
On the Relationship between Truth and Political Bias in Language Models
par: Fulay, Suyash, et autres
Publié: (2024)
par: Fulay, Suyash, et autres
Publié: (2024)
Computational Analysis of Conversation Dynamics through Participant Responsivity
par: Hughes, Margaret, et autres
Publié: (2025)
par: Hughes, Margaret, et autres
Publié: (2025)
Confidence Under the Hood: An Investigation into the Confidence-Probability Alignment in Large Language Models
par: Kumar, Abhishek, et autres
Publié: (2024)
par: Kumar, Abhishek, et autres
Publié: (2024)
PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits
par: Jiang, Hang, et autres
Publié: (2023)
par: Jiang, Hang, et autres
Publié: (2023)
Bridging Context Gaps: Enhancing Comprehension in Long-Form Social Conversations Through Contextualized Excerpts
par: Mohanty, Shrestha, et autres
Publié: (2024)
par: Mohanty, Shrestha, et autres
Publié: (2024)
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
par: van der Weij, Teun, et autres
Publié: (2024)
par: van der Weij, Teun, et autres
Publié: (2024)
User-LLM: Efficient LLM Contextualization with User Embeddings
par: Ning, Lin, et autres
Publié: (2024)
par: Ning, Lin, et autres
Publié: (2024)
On the Limitations of Language Targeted Pruning: Investigating the Calibration Language Impact in Multilingual LLM Pruning
par: Kurz, Simon, et autres
Publié: (2024)
par: Kurz, Simon, et autres
Publié: (2024)
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
par: Pandey, Punya Syon, et autres
Publié: (2025)
par: Pandey, Punya Syon, et autres
Publié: (2025)
RLTHF: Targeted Human Feedback for LLM Alignment
par: Xu, Yifei, et autres
Publié: (2025)
par: Xu, Yifei, et autres
Publié: (2025)
LLM See, LLM Do: Guiding Data Generation to Target Non-Differentiable Objectives
par: Shimabucoro, Luísa, et autres
Publié: (2024)
par: Shimabucoro, Luísa, et autres
Publié: (2024)
LLMAP: LLM-Assisted Multi-Objective Route Planning with User Preferences
par: Yuan, Liangqi, et autres
Publié: (2025)
par: Yuan, Liangqi, et autres
Publié: (2025)
From Exact Hits to Close Enough: Semantic Caching for LLM Embeddings
par: Biton, Dvir David, et autres
Publié: (2026)
par: Biton, Dvir David, et autres
Publié: (2026)
The Impact of Language Mixing on Bilingual LLM Reasoning
par: Li, Yihao, et autres
Publié: (2025)
par: Li, Yihao, et autres
Publié: (2025)
Transparent Screening for LLM Inference and Training Impacts
par: Pachot, Arnault, et autres
Publié: (2026)
par: Pachot, Arnault, et autres
Publié: (2026)
Building, Reusing, and Generalizing Abstract Representations from Concrete Sequences
par: Wu, Shuchen, et autres
Publié: (2024)
par: Wu, Shuchen, et autres
Publié: (2024)
Factual Inconsistency in Data-to-Text Generation Scales Exponentially with LLM Size: A Statistical Validation
par: Mahapatra, Joy, et autres
Publié: (2025)
par: Mahapatra, Joy, et autres
Publié: (2025)
GanitLLM: Difficulty-Aware Bengali Mathematical Reasoning through Curriculum-GRPO
par: Dipta, Shubhashis Roy, et autres
Publié: (2026)
par: Dipta, Shubhashis Roy, et autres
Publié: (2026)
Systematically Analyzing Prompt Injection Vulnerabilities in Diverse LLM Architectures
par: Benjamin, Victoria, et autres
Publié: (2024)
par: Benjamin, Victoria, et autres
Publié: (2024)
Are My Optimized Prompts Compromised? Exploring Vulnerabilities of LLM-based Optimizers
par: Zhao, Andrew, et autres
Publié: (2025)
par: Zhao, Andrew, et autres
Publié: (2025)
When Greedy Wins: Emergent Exploitation Bias in Meta-Bandit LLM Training
par: Chen, Sanxing, et autres
Publié: (2025)
par: Chen, Sanxing, et autres
Publié: (2025)
JuStRank: Benchmarking LLM Judges for System Ranking
par: Gera, Ariel, et autres
Publié: (2024)
par: Gera, Ariel, et autres
Publié: (2024)
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
par: Yehudai, Asaf, et autres
Publié: (2025)
par: Yehudai, Asaf, et autres
Publié: (2025)
Survey on Evaluation of LLM-based Agents
par: Yehudai, Asaf, et autres
Publié: (2025)
par: Yehudai, Asaf, et autres
Publié: (2025)
HindSight: Evaluating LLM-Generated Research Ideas via Future Impact
par: Jiang, Bo
Publié: (2026)
par: Jiang, Bo
Publié: (2026)
Breach in the Shield: Unveiling the Vulnerabilities of Large Language Models
par: Dai, Runpeng, et autres
Publié: (2025)
par: Dai, Runpeng, et autres
Publié: (2025)
UserBench: An Interactive Gym Environment for User-Centric Agents
par: Qian, Cheng, et autres
Publié: (2025)
par: Qian, Cheng, et autres
Publié: (2025)
In Vino Veritas and Vulnerabilities: Examining LLM Safety via Drunk Language Inducement
par: Shetty, Anudeex, et autres
Publié: (2026)
par: Shetty, Anudeex, et autres
Publié: (2026)
Stay Tuned: An Empirical Study of the Impact of Hyperparameters on LLM Tuning in Real-World Applications
par: Halfon, Alon, et autres
Publié: (2024)
par: Halfon, Alon, et autres
Publié: (2024)
Accidental Vulnerability: Factors in Fine-Tuning that Shift Model Safeguards
par: Pandey, Punya Syon, et autres
Publié: (2025)
par: Pandey, Punya Syon, et autres
Publié: (2025)
To Think or Not to Think: Exploring the Unthinking Vulnerability in Large Reasoning Models
par: Zhu, Zihao, et autres
Publié: (2025)
par: Zhu, Zihao, et autres
Publié: (2025)
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing
par: Tran, Thien Q., et autres
Publié: (2025)
par: Tran, Thien Q., et autres
Publié: (2025)
Controllable User Simulation
par: Tennenholtz, Guy, et autres
Publié: (2026)
par: Tennenholtz, Guy, et autres
Publié: (2026)
UserSumBench: A Benchmark Framework for Evaluating User Summarization Approaches
par: Wang, Chao, et autres
Publié: (2024)
par: Wang, Chao, et autres
Publié: (2024)
UserRL: Training Interactive User-Centric Agent via Reinforcement Learning
par: Qian, Cheng, et autres
Publié: (2025)
par: Qian, Cheng, et autres
Publié: (2025)
Who Gets Which Message? Auditing Demographic Bias in LLM-Generated Targeted Text
par: Islam, Tunazzina
Publié: (2026)
par: Islam, Tunazzina
Publié: (2026)
Advancing Jailbreak Strategies: A Hybrid Approach to Exploiting LLM Vulnerabilities and Bypassing Modern Defenses
par: Ahmed, Mohamed, et autres
Publié: (2025)
par: Ahmed, Mohamed, et autres
Publié: (2025)
BiasJailbreak:Analyzing Ethical Biases and Jailbreak Vulnerabilities in Large Language Models
par: Lee, Isack, et autres
Publié: (2024)
par: Lee, Isack, et autres
Publié: (2024)
Estimating LLM Consistency: A User Baseline vs Surrogate Metrics
par: Wu, Xiaoyuan, et autres
Publié: (2025)
par: Wu, Xiaoyuan, et autres
Publié: (2025)
Documents similaires
-
An AI-Powered Framework for Analyzing Collective Idea Evolution in Deliberative Assemblies
par: Poole-Dayan, Elinor, et autres
Publié: (2025) -
On the Relationship between Truth and Political Bias in Language Models
par: Fulay, Suyash, et autres
Publié: (2024) -
Computational Analysis of Conversation Dynamics through Participant Responsivity
par: Hughes, Margaret, et autres
Publié: (2025) -
Confidence Under the Hood: An Investigation into the Confidence-Probability Alignment in Large Language Models
par: Kumar, Abhishek, et autres
Publié: (2024) -
PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits
par: Jiang, Hang, et autres
Publié: (2023)