Benchmarking and Improving LLM Robustness for Personalized Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Okite, Chimaobi, Deng, Naihao, Bodipati, Kiran, Hou, Huaidian, Chai, Joyce, Mihalcea, Rada |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking Table Instruction Tuning
by: Deng, Naihao, et al.
Published: (2025)
by: Deng, Naihao, et al.
Published: (2025)
LUCid: Redefining Relevance For Lifelong Personalization
by: Okite, Chimaobi, et al.
Published: (2026)
by: Okite, Chimaobi, et al.
Published: (2026)
Are Human Interactions Replicable by Generative Agents? A Case Study on Pronoun Usage in Hierarchical Interactions
by: Deng, Naihao, et al.
Published: (2025)
by: Deng, Naihao, et al.
Published: (2025)
SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models
by: Torres-Fonseca, Josue, et al.
Published: (2026)
by: Torres-Fonseca, Josue, et al.
Published: (2026)
Chumor 2.0: Towards Benchmarking Chinese Humor Understanding
by: He, Ruiqi, et al.
Published: (2024)
by: He, Ruiqi, et al.
Published: (2024)
$R^3$: "This is My SQL, Are You With Me?" A Consensus-Based Multi-Agent System for Text-to-SQL Tasks
by: Xia, Hanchen, et al.
Published: (2024)
by: Xia, Hanchen, et al.
Published: (2024)
Cross-cultural Inspiration Detection and Analysis in Real and LLM-generated Social Media Data
by: Ignat, Oana, et al.
Published: (2024)
by: Ignat, Oana, et al.
Published: (2024)
One Word at a Time: Incremental Completion Decomposition Breaks LLM Safety
by: Arif, Samee, et al.
Published: (2026)
by: Arif, Samee, et al.
Published: (2026)
Mind the (Belief) Gap: Group Identity in the World of LLMs
by: Borah, Angana, et al.
Published: (2025)
by: Borah, Angana, et al.
Published: (2025)
MAiDE-up: Multilingual Deception Detection of GPT-generated Hotel Reviews
by: Ignat, Oana, et al.
Published: (2024)
by: Ignat, Oana, et al.
Published: (2024)
Tables as Texts or Images: Evaluating the Table Reasoning Ability of LLMs and MLLMs
by: Deng, Naihao, et al.
Published: (2024)
by: Deng, Naihao, et al.
Published: (2024)
Table as Thought: Exploring Structured Thoughts in LLM Reasoning
by: Sun, Zhenjie, et al.
Published: (2025)
by: Sun, Zhenjie, et al.
Published: (2025)
DOTRAG: Retrieval-Time Reasoning Along Paths
by: Moore, Larnell, et al.
Published: (2026)
by: Moore, Larnell, et al.
Published: (2026)
Towards Algorithmic Fidelity: Mental Health Representation across Demographics in Synthetic vs. Human-generated Data
by: Mori, Shinka, et al.
Published: (2024)
by: Mori, Shinka, et al.
Published: (2024)
Belief-Sim: Towards Belief-Driven Simulation of Demographic Misinformation Susceptibility
by: Borah, Angana, et al.
Published: (2026)
by: Borah, Angana, et al.
Published: (2026)
When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas
by: Backmann, Steffen, et al.
Published: (2025)
by: Backmann, Steffen, et al.
Published: (2025)
VERVE: Template-based ReflectiVE Rewriting for MotiVational IntErviewing
by: Min, Do June, et al.
Published: (2023)
by: Min, Do June, et al.
Published: (2023)
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
by: Pandey, Punya Syon, et al.
Published: (2025)
by: Pandey, Punya Syon, et al.
Published: (2025)
CliniDial: A Naturally Occurring Multimodal Dialogue Dataset for Team Reflection in Action During Clinical Operation
by: Deng, Naihao, et al.
Published: (2025)
by: Deng, Naihao, et al.
Published: (2025)
The Power of Many: Multi-Agent Multimodal Models for Cultural Image Captioning
by: Bai, Longju, et al.
Published: (2024)
by: Bai, Longju, et al.
Published: (2024)
Annotations on a Budget: Leveraging Geo-Data Similarity to Balance Model Performance and Annotation Cost
by: Ignat, Oana, et al.
Published: (2024)
by: Ignat, Oana, et al.
Published: (2024)
The Generation Gap: Exploring Age Bias in the Value Systems of Large Language Models
by: Liu, Siyang, et al.
Published: (2024)
by: Liu, Siyang, et al.
Published: (2024)
A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
by: Lee, Andrew, et al.
Published: (2024)
by: Lee, Andrew, et al.
Published: (2024)
Implicit Personalization in Language Models: A Systematic Study
by: Jin, Zhijing, et al.
Published: (2024)
by: Jin, Zhijing, et al.
Published: (2024)
Deception Detection from Linguistic and Physiological Data Streams Using Bimodal Convolutional Neural Networks
by: Li, Panfeng, et al.
Published: (2023)
by: Li, Panfeng, et al.
Published: (2023)
Towards Implicit Bias Detection and Mitigation in Multi-Agent LLM Interactions
by: Borah, Angana, et al.
Published: (2024)
by: Borah, Angana, et al.
Published: (2024)
Uplifting Lower-Income Data: Strategies for Socioeconomic Perspective Shifts in Large Multi-modal Models
by: Nwatu, Joan, et al.
Published: (2024)
by: Nwatu, Joan, et al.
Published: (2024)
Babysit A Language Model From Scratch: Interactive Language Learning by Trials and Demonstrations
by: Ma, Ziqiao, et al.
Published: (2024)
by: Ma, Ziqiao, et al.
Published: (2024)
PTCBENCH: Benchmarking Contextual Stability of Personality Traits in LLM Systems
by: Yu, Jiongchi, et al.
Published: (2026)
by: Yu, Jiongchi, et al.
Published: (2026)
Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization
by: Majumder, Navonil, et al.
Published: (2024)
by: Majumder, Navonil, et al.
Published: (2024)
Voices of Her: Analyzing Gender Differences in the AI Publication World
by: Ding, Yiwen, et al.
Published: (2023)
by: Ding, Yiwen, et al.
Published: (2023)
Culture Affordance Atlas: Reconciling Object Diversity Through Functional Mapping
by: Nwatu, Joan, et al.
Published: (2025)
by: Nwatu, Joan, et al.
Published: (2025)
Bilateral Personalized Dialogue Generation with Contrastive Learning
by: Li, Bin, et al.
Published: (2021)
by: Li, Bin, et al.
Published: (2021)
Guided Profile Generation Improves Personalization with LLMs
by: Zhang, Jiarui
Published: (2024)
by: Zhang, Jiarui
Published: (2024)
Towards a Benchmark for Large Language Models for Business Process Management Tasks
by: Busch, Kiran, et al.
Published: (2024)
by: Busch, Kiran, et al.
Published: (2024)
AlpsBench: An LLM Personalization Benchmark for Real-Dialogue Memorization and Preference Alignment
by: Xiao, Jianfei, et al.
Published: (2026)
by: Xiao, Jianfei, et al.
Published: (2026)
Has It All Been Solved? Open NLP Research Questions Not Solved by Large Language Models
by: Ignat, Oana, et al.
Published: (2023)
by: Ignat, Oana, et al.
Published: (2023)
Towards A Holistic Landscape of Situated Theory of Mind in Large Language Models
by: Ma, Ziqiao, et al.
Published: (2023)
by: Ma, Ziqiao, et al.
Published: (2023)
Eeyore: Realistic Depression Simulation via Supervised and Preference Optimization
by: Liu, Siyang, et al.
Published: (2025)
by: Liu, Siyang, et al.
Published: (2025)
SALAD: Improving Robustness and Generalization through Contrastive Learning with Structure-Aware and LLM-Driven Augmented Data
by: Bae, Suyoung, et al.
Published: (2025)
by: Bae, Suyoung, et al.
Published: (2025)
Similar Items
-
Rethinking Table Instruction Tuning
by: Deng, Naihao, et al.
Published: (2025) -
LUCid: Redefining Relevance For Lifelong Personalization
by: Okite, Chimaobi, et al.
Published: (2026) -
Are Human Interactions Replicable by Generative Agents? A Case Study on Pronoun Usage in Hierarchical Interactions
by: Deng, Naihao, et al.
Published: (2025) -
SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models
by: Torres-Fonseca, Josue, et al.
Published: (2026) -
Chumor 2.0: Towards Benchmarking Chinese Humor Understanding
by: He, Ruiqi, et al.
Published: (2024)