Say It Another Way: Auditing LLMs with a User-Grounded Automated Paraphrasing Framework
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Chataigner, Cléa, Ma, Rebecca, Ganesh, Prakhar, Chen, Yuhao, Taïk, Afaf, Creager, Elliot, Farnadi, Golnoosh |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Multilingual Hallucination Gaps in Large Language Models
par: Chataigner, Cléa, et autres
Publié: (2024)
par: Chataigner, Cléa, et autres
Publié: (2024)
Systemizing Multiplicity: The Curious Case of Arbitrariness in Machine Learning
par: Ganesh, Prakhar, et autres
Publié: (2025)
par: Ganesh, Prakhar, et autres
Publié: (2025)
Towards More Realistic Extraction Attacks: An Adversarial Perspective
par: More, Yash, et autres
Publié: (2024)
par: More, Yash, et autres
Publié: (2024)
Fairness in Federated Learning: Fairness for Whom?
par: Taik, Afaf, et autres
Publié: (2025)
par: Taik, Afaf, et autres
Publié: (2025)
Promoting Fair Vaccination Strategies Through Influence Maximization: A Case Study on COVID-19 Spread
par: Neophytou, Nicola, et autres
Publié: (2024)
par: Neophytou, Nicola, et autres
Publié: (2024)
Differentially Private Clustered Federated Learning
par: Malekmohammadi, Saber, et autres
Publié: (2024)
par: Malekmohammadi, Saber, et autres
Publié: (2024)
Enhancing Privacy in the Early Detection of Sexual Predators Through Federated Learning and Differential Privacy
par: Chehbouni, Khaoula, et autres
Publié: (2025)
par: Chehbouni, Khaoula, et autres
Publié: (2025)
Rethinking Hallucinations: Correctness, Consistency, and Prompt Multiplicity
par: Ganesh, Prakhar, et autres
Publié: (2026)
par: Ganesh, Prakhar, et autres
Publié: (2026)
Data as a Lever: A Neighbouring Datasets Perspective on Predictive Multiplicity
par: Ganesh, Prakhar, et autres
Publié: (2025)
par: Ganesh, Prakhar, et autres
Publié: (2025)
From Representational Harms to Quality-of-Service Harms: A Case Study on Llama 2 Safety Safeguards
par: Chehbouni, Khaoula, et autres
Publié: (2024)
par: Chehbouni, Khaoula, et autres
Publié: (2024)
Fairness Incentives in Response to Unfair Dynamic Pricing
par: Thibodeau, Jesse, et autres
Publié: (2024)
par: Thibodeau, Jesse, et autres
Publié: (2024)
Balancing Profit and Fairness in Risk-Based Pricing Markets
par: Thibodeau, Jesse, et autres
Publié: (2025)
par: Thibodeau, Jesse, et autres
Publié: (2025)
Different Horses for Different Courses: Comparing Bias Mitigation Algorithms in ML
par: Ganesh, Prakhar, et autres
Publié: (2024)
par: Ganesh, Prakhar, et autres
Publié: (2024)
LoRA Provides Differential Privacy by Design via Random Sketching
par: Malekmohammadi, Saber, et autres
Publié: (2024)
par: Malekmohammadi, Saber, et autres
Publié: (2024)
The Cost of Arbitrariness for Individuals: Examining the Legal and Technical Challenges of Model Multiplicity
par: Ganesh, Prakhar, et autres
Publié: (2024)
par: Ganesh, Prakhar, et autres
Publié: (2024)
Beyond the Safety Bundle: Auditing the Helpful and Harmless Dataset
par: Chehbouni, Khaoula, et autres
Publié: (2024)
par: Chehbouni, Khaoula, et autres
Publié: (2024)
Crossing Boundaries: Leveraging Semantic Divergences to Explore Cultural Novelty in Cooking Recipes
par: Carichon, Florian, et autres
Publié: (2025)
par: Carichon, Florian, et autres
Publié: (2025)
Understanding Intrinsic Socioeconomic Biases in Large Language Models
par: Arzaghi, Mina, et autres
Publié: (2024)
par: Arzaghi, Mina, et autres
Publié: (2024)
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges
par: Chehbouni, Khaoula, et autres
Publié: (2025)
par: Chehbouni, Khaoula, et autres
Publié: (2025)
Multilingual Amnesia: On the Transferability of Unlearning in Multilingual LLMs
par: Farashah, Alireza Dehghanpour, et autres
Publié: (2026)
par: Farashah, Alireza Dehghanpour, et autres
Publié: (2026)
Mechanics of Bias and Reasoning: Interpreting the Impact of Chain-of-Thought Prompting on Gender Bias in LLMs
par: Pearman, Edie, et autres
Publié: (2026)
par: Pearman, Edie, et autres
Publié: (2026)
Reviving Your MNEME: Predicting The Side Effects of LLM Unlearning and Fine-Tuning via Sparse Model Diffing
par: Kassem, Aly M., et autres
Publié: (2025)
par: Kassem, Aly M., et autres
Publié: (2025)
Trust No Bot: Discovering Personal Disclosures in Human-LLM Conversations in the Wild
par: Mireshghallah, Niloofar, et autres
Publié: (2024)
par: Mireshghallah, Niloofar, et autres
Publié: (2024)
From Fragile to Certified: Wasserstein Audits of Group Fairness Under Distribution Shift
par: Ehyaei, Ahmad-Reza, et autres
Publié: (2025)
par: Ehyaei, Ahmad-Reza, et autres
Publié: (2025)
Intrinsic Meets Extrinsic Fairness: Assessing the Downstream Impact of Bias Mitigation in Large Language Models
par: Arzaghi', 'Mina, et autres
Publié: (2025)
par: Arzaghi', 'Mina, et autres
Publié: (2025)
Being Another Way
par: Klinger, Dustin D.
Publié: (2024)
par: Klinger, Dustin D.
Publié: (2024)
IDEAFix: Evaluation Framework for Creative Defixation Prompting in LLMs
par: Carichon, F., et autres
Publié: (2026)
par: Carichon, F., et autres
Publié: (2026)
You Didn't Have to Say It like That: Subliminal Learning from Faithful Paraphrases
par: Gisler, Isaia, et autres
Publié: (2026)
par: Gisler, Isaia, et autres
Publié: (2026)
Hallucination Detox: Sensitivity Dropout (SenD) for Large Language Model Training
par: Mohammadzadeh, Shahrad, et autres
Publié: (2024)
par: Mohammadzadeh, Shahrad, et autres
Publié: (2024)
Show, Don't Tell: Uncovering Implicit Character Portrayal using LLMs
par: Jaipersaud, Brandon, et autres
Publié: (2024)
par: Jaipersaud, Brandon, et autres
Publié: (2024)
Large Language Models Often Say One Thing and Do Another
par: Xu, Ruoxi, et autres
Publié: (2025)
par: Xu, Ruoxi, et autres
Publié: (2025)
Exploring Robustness of LLMs to Paraphrasing Based on Sociodemographic Factors
par: Arora, Pulkit, et autres
Publié: (2025)
par: Arora, Pulkit, et autres
Publié: (2025)
Assessing LLMs for Zero-shot Abstractive Summarization Through the Lens of Relevance Paraphrasing
par: Askari, Hadi, et autres
Publié: (2024)
par: Askari, Hadi, et autres
Publié: (2024)
Offscript: Automated Auditing of Instruction Adherence in LLMs
par: Clark, Nicholas, et autres
Publié: (2025)
par: Clark, Nicholas, et autres
Publié: (2025)
Mitigating Paraphrase Attacks on Machine-Text Detectors via Paraphrase Inversion
par: Soto, Rafael Rivera, et autres
Publié: (2024)
par: Soto, Rafael Rivera, et autres
Publié: (2024)
Auditing Support Strategies in LLMs through Grounded Multi-Turn Social Simulation
par: Star, Michelle, et autres
Publié: (2026)
par: Star, Michelle, et autres
Publié: (2026)
Paraphrase-Aligned Machine Translation
par: Chang, Ke-Ching, et autres
Publié: (2024)
par: Chang, Ke-Ching, et autres
Publié: (2024)
Advancing Cultural Inclusivity: Optimizing Embedding Spaces for Balanced Music Recommendations
par: Moradi, Armin, et autres
Publié: (2024)
par: Moradi, Armin, et autres
Publié: (2024)
Position: Cracking the Code of Cascading Disparity Towards Marginalized Communities
par: Farnadi, Golnoosh, et autres
Publié: (2024)
par: Farnadi, Golnoosh, et autres
Publié: (2024)
LLMs Show Surface-Form Brittleness Under Paraphrase Stress Tests
par: Carranza, Juan Miguel Navarro
Publié: (2025)
par: Carranza, Juan Miguel Navarro
Publié: (2025)
Documents similaires
-
Multilingual Hallucination Gaps in Large Language Models
par: Chataigner, Cléa, et autres
Publié: (2024) -
Systemizing Multiplicity: The Curious Case of Arbitrariness in Machine Learning
par: Ganesh, Prakhar, et autres
Publié: (2025) -
Towards More Realistic Extraction Attacks: An Adversarial Perspective
par: More, Yash, et autres
Publié: (2024) -
Fairness in Federated Learning: Fairness for Whom?
par: Taik, Afaf, et autres
Publié: (2025) -
Promoting Fair Vaccination Strategies Through Influence Maximization: A Case Study on COVID-19 Spread
par: Neophytou, Nicola, et autres
Publié: (2024)