Flipping the Dialogue: Training and Evaluating User Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Naous, Tarek, Laban, Philippe, Xu, Wei, Neville, Jennifer |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On The Origin of Cultural Biases in Language Models: From Pre-training Data to Linguistic Phenomena
von: Naous, Tarek, et al.
Veröffentlicht: (2025)
von: Naous, Tarek, et al.
Veröffentlicht: (2025)
LLMs Corrupt Your Documents When You Delegate
von: Laban, Philippe, et al.
Veröffentlicht: (2026)
von: Laban, Philippe, et al.
Veröffentlicht: (2026)
What are Foundation Models Cooking in the Post-Soviet World?
von: Lavrouk, Anton, et al.
Veröffentlicht: (2025)
von: Lavrouk, Anton, et al.
Veröffentlicht: (2025)
Having Beer after Prayer? Measuring Cultural Bias in Large Language Models
von: Naous, Tarek, et al.
Veröffentlicht: (2023)
von: Naous, Tarek, et al.
Veröffentlicht: (2023)
LLMs Get Lost In Multi-Turn Conversation
von: Laban, Philippe, et al.
Veröffentlicht: (2025)
von: Laban, Philippe, et al.
Veröffentlicht: (2025)
ReadMe++: Benchmarking Multilingual Language Models for Multi-Domain Readability Assessment
von: Naous, Tarek, et al.
Veröffentlicht: (2023)
von: Naous, Tarek, et al.
Veröffentlicht: (2023)
Reducing Privacy Risks in Online Self-Disclosures with Language Models
von: Dou, Yao, et al.
Veröffentlicht: (2023)
von: Dou, Yao, et al.
Veröffentlicht: (2023)
Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment
von: Laban, Philippe, et al.
Veröffentlicht: (2023)
von: Laban, Philippe, et al.
Veröffentlicht: (2023)
SPASM: Stable Persona-driven Agent Simulation for Multi-turn Dialogue Generation
von: Luo, Han, et al.
Veröffentlicht: (2026)
von: Luo, Han, et al.
Veröffentlicht: (2026)
CARE: Multilingual Human Preference Learning for Cultural Awareness
von: Guo, Geyang, et al.
Veröffentlicht: (2025)
von: Guo, Geyang, et al.
Veröffentlicht: (2025)
Stanceosaurus 2.0: Classifying Stance Towards Russian and Spanish Misinformation
von: Lavrouk, Anton, et al.
Veröffentlicht: (2024)
von: Lavrouk, Anton, et al.
Veröffentlicht: (2024)
AI-Slop to AI-Polish? Aligning Language Models through Edit-Based Writing Rewards and Test-time Computation
von: Chakrabarty, Tuhin, et al.
Veröffentlicht: (2025)
von: Chakrabarty, Tuhin, et al.
Veröffentlicht: (2025)
Large Language Models as User-Agents for Evaluating Task-Oriented-Dialogue Systems
von: Kazi, Taaha, et al.
Veröffentlicht: (2024)
von: Kazi, Taaha, et al.
Veröffentlicht: (2024)
Beyond Single-User Dialogue: Assessing Multi-User Dialogue State Tracking Capabilities of Large Language Models
von: Song, Sangmin, et al.
Veröffentlicht: (2025)
von: Song, Sangmin, et al.
Veröffentlicht: (2025)
SummExecEdit: A Factual Consistency Benchmark in Summarization with Executable Edits
von: Thorat, Onkar, et al.
Veröffentlicht: (2024)
von: Thorat, Onkar, et al.
Veröffentlicht: (2024)
Art or Artifice? Large Language Models and the False Promise of Creativity
von: Chakrabarty, Tuhin, et al.
Veröffentlicht: (2023)
von: Chakrabarty, Tuhin, et al.
Veröffentlicht: (2023)
Measuring and Mitigating the Distributional Gap Between Real and Simulated User Behaviors
von: Mehri, Shuhaib, et al.
Veröffentlicht: (2026)
von: Mehri, Shuhaib, et al.
Veröffentlicht: (2026)
MiniCheck: Efficient Fact-Checking of LLMs on Grounding Documents
von: Tang, Liyan, et al.
Veröffentlicht: (2024)
von: Tang, Liyan, et al.
Veröffentlicht: (2024)
Can AI writing be salvaged? Mitigating Idiosyncrasies and Improving Human-AI Alignment in the Writing Process through Edits
von: Chakrabarty, Tuhin, et al.
Veröffentlicht: (2024)
von: Chakrabarty, Tuhin, et al.
Veröffentlicht: (2024)
Do RAG Systems Cover What Matters? Evaluating and Optimizing Responses with Sub-Question Coverage
von: Xie, Kaige, et al.
Veröffentlicht: (2024)
von: Xie, Kaige, et al.
Veröffentlicht: (2024)
EvalAgent: Discovering Implicit Evaluation Criteria from the Web
von: Wadhwa, Manya, et al.
Veröffentlicht: (2025)
von: Wadhwa, Manya, et al.
Veröffentlicht: (2025)
User-Specific Dialogue Generation with User Profile-Aware Pre-Training Model and Parameter-Efficient Fine-Tuning
von: Otsuka, Atsushi, et al.
Veröffentlicht: (2024)
von: Otsuka, Atsushi, et al.
Veröffentlicht: (2024)
Simulating User Diversity in Task-Oriented Dialogue Systems using Large Language Models
von: Ahmad, Adnan, et al.
Veröffentlicht: (2025)
von: Ahmad, Adnan, et al.
Veröffentlicht: (2025)
RAD-Bench: Evaluating Large Language Models Capabilities in Retrieval Augmented Dialogues
von: Kuo, Tzu-Lin, et al.
Veröffentlicht: (2024)
von: Kuo, Tzu-Lin, et al.
Veröffentlicht: (2024)
User-Aware Active Knowledge Acquisition for Emotional Support Dialogue
von: Xu, Mufan, et al.
Veröffentlicht: (2026)
von: Xu, Mufan, et al.
Veröffentlicht: (2026)
Revealing User Familiarity Bias in Task-Oriented Dialogue via Interactive Evaluation
von: Kim, Takyoung, et al.
Veröffentlicht: (2023)
von: Kim, Takyoung, et al.
Veröffentlicht: (2023)
Beyond Vision: How Large Language Models Interpret Facial Expressions from Valence-Arousal Values
von: Mehra, Vaibhav, et al.
Veröffentlicht: (2025)
von: Mehra, Vaibhav, et al.
Veröffentlicht: (2025)
Commonsense Generation and Evaluation for Dialogue Systems using Large Language Models
von: Estecha-Garitagoitia, Marcos, et al.
Veröffentlicht: (2025)
von: Estecha-Garitagoitia, Marcos, et al.
Veröffentlicht: (2025)
Controlling Language Difficulty in Dialogues with Linguistic Features
von: Xu, Shuyao, et al.
Veröffentlicht: (2025)
von: Xu, Shuyao, et al.
Veröffentlicht: (2025)
An Analysis of User Behaviors for Objectively Evaluating Spoken Dialogue Systems
von: Inoue, Koji, et al.
Veröffentlicht: (2024)
von: Inoue, Koji, et al.
Veröffentlicht: (2024)
Summary of a Haystack: A Challenge to Long-Context LLMs and RAG Systems
von: Laban, Philippe, et al.
Veröffentlicht: (2024)
von: Laban, Philippe, et al.
Veröffentlicht: (2024)
Evaluating Large Language Models in Analysing Classroom Dialogue
von: Long, Yun, et al.
Veröffentlicht: (2024)
von: Long, Yun, et al.
Veröffentlicht: (2024)
LLMExplainer: Large Language Model based Bayesian Inference for Graph Explanation Generation
von: Zhang, Jiaxing, et al.
Veröffentlicht: (2024)
von: Zhang, Jiaxing, et al.
Veröffentlicht: (2024)
How Stylistic Similarity Shapes Preferences in Dialogue Dataset with User and Third Party Evaluations
von: Numaya, Ikumi, et al.
Veröffentlicht: (2025)
von: Numaya, Ikumi, et al.
Veröffentlicht: (2025)
Aligning Dialogue Agents with Global Feedback via Large Language Model Multimodal Reward Decomposition
von: Lee, Dong Won, et al.
Veröffentlicht: (2025)
von: Lee, Dong Won, et al.
Veröffentlicht: (2025)
A Comprehensive Analysis of the Effectiveness of Large Language Models as Automatic Dialogue Evaluators
von: Zhang, Chen, et al.
Veröffentlicht: (2023)
von: Zhang, Chen, et al.
Veröffentlicht: (2023)
Large Language Model based Situational Dialogues for Second Language Learning
von: Xu, Shuyao, et al.
Veröffentlicht: (2024)
von: Xu, Shuyao, et al.
Veröffentlicht: (2024)
Rethinking the Evaluation of Dialogue Systems: Effects of User Feedback on Crowdworkers and LLMs
von: Siro, Clemencia, et al.
Veröffentlicht: (2024)
von: Siro, Clemencia, et al.
Veröffentlicht: (2024)
An Interpretable and Crosslingual Method for Evaluating Second-Language Dialogues
von: Gao, Rena, et al.
Veröffentlicht: (2024)
von: Gao, Rena, et al.
Veröffentlicht: (2024)
Simulating Before Planning: Constructing Intrinsic User World Model for User-Tailored Dialogue Policy Planning
von: He, Tao, et al.
Veröffentlicht: (2025)
von: He, Tao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
On The Origin of Cultural Biases in Language Models: From Pre-training Data to Linguistic Phenomena
von: Naous, Tarek, et al.
Veröffentlicht: (2025) -
LLMs Corrupt Your Documents When You Delegate
von: Laban, Philippe, et al.
Veröffentlicht: (2026) -
What are Foundation Models Cooking in the Post-Soviet World?
von: Lavrouk, Anton, et al.
Veröffentlicht: (2025) -
Having Beer after Prayer? Measuring Cultural Bias in Large Language Models
von: Naous, Tarek, et al.
Veröffentlicht: (2023) -
LLMs Get Lost In Multi-Turn Conversation
von: Laban, Philippe, et al.
Veröffentlicht: (2025)