User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Saha, Sougata, Choudhury, Monojit |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reading between the Lines: Can LLMs Identify Cross-Cultural Communication Gaps?
by: Saha, Sougata, et al.
Published: (2025)
by: Saha, Sougata, et al.
Published: (2025)
Meta-Cultural Competence: Climbing the Right Hill of Cultural Awareness
by: Saha, Sougata, et al.
Published: (2025)
by: Saha, Sougata, et al.
Published: (2025)
To Generate or Discriminate? Methodological Considerations for Measuring Cultural Alignment in LLMs
by: Pandey, Saurabh Kumar, et al.
Published: (2026)
by: Pandey, Saurabh Kumar, et al.
Published: (2026)
Consolidating Strategies for Countering Hate Speech Using Persuasive Dialogues
by: Saha, Sougata, et al.
Published: (2024)
by: Saha, Sougata, et al.
Published: (2024)
Exploring Adapter Design Tradeoffs for Low Resource Music Generation
by: Mehta, Atharva, et al.
Published: (2025)
by: Mehta, Atharva, et al.
Published: (2025)
Ethical Reasoning and Moral Value Alignment of LLMs Depend on the Language we Prompt them in
by: Agarwal, Utkarsh, et al.
Published: (2024)
by: Agarwal, Utkarsh, et al.
Published: (2024)
Do Moral Judgment and Reasoning Capability of LLMs Change with Language? A Study using the Multilingual Defining Issues Test
by: Khandelwal, Aditi, et al.
Published: (2024)
by: Khandelwal, Aditi, et al.
Published: (2024)
Sacred or Synthetic? Evaluating LLM Reliability and Abstention for Religious Questions
by: Atif, Farah, et al.
Published: (2025)
by: Atif, Farah, et al.
Published: (2025)
Litmus (Re)Agent: A Benchmark and Agentic System for Predictive Evaluation of Multilingual Models
by: Mittal, Avni, et al.
Published: (2026)
by: Mittal, Avni, et al.
Published: (2026)
Evaluating Large Language Models for Health-related Queries with Presuppositions
by: Kaur, Navreet, et al.
Published: (2023)
by: Kaur, Navreet, et al.
Published: (2023)
Women, Infamous, and Exotic Beings: A Comparative Study of Honorific Usages in Wikipedia and LLMs for Bengali and Hindi
by: Mukherjee, Sourabrata, et al.
Published: (2025)
by: Mukherjee, Sourabrata, et al.
Published: (2025)
Missing Melodies: AI Music Generation and its "Nearly" Complete Omission of the Global South
by: Mehta, Atharva, et al.
Published: (2024)
by: Mehta, Atharva, et al.
Published: (2024)
Tuning Language Models for Robust Prediction of Diverse User Behaviors
by: Meng, Fanjin, et al.
Published: (2025)
by: Meng, Fanjin, et al.
Published: (2025)
LLMs Reading the Rhythms of Daily Life: Aligned Understanding for Behavior Prediction and Generation
by: Meng, Fanjin, et al.
Published: (2026)
by: Meng, Fanjin, et al.
Published: (2026)
Low-Cost Generation and Evaluation of Dictionary Example Sentences
by: Cai, Bill, et al.
Published: (2024)
by: Cai, Bill, et al.
Published: (2024)
Music for All: Representational Bias and Cross-Cultural Adaptability of Music Generation Models
by: Mehta, Atharva, et al.
Published: (2025)
by: Mehta, Atharva, et al.
Published: (2025)
From Human Judgements to Predictive Models: Unravelling Acceptability in Code-Mixed Sentences
by: Kodali, Prashant, et al.
Published: (2024)
by: Kodali, Prashant, et al.
Published: (2024)
"They are uncultured": Unveiling Covert Harms and Social Threats in LLM Generated Conversations
by: Dammu, Preetam Prabhu Srikar, et al.
Published: (2024)
by: Dammu, Preetam Prabhu Srikar, et al.
Published: (2024)
Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents
by: Chopra, Harshita, et al.
Published: (2026)
by: Chopra, Harshita, et al.
Published: (2026)
Towards Measuring and Modeling "Culture" in LLMs: A Survey
by: Adilazuarda, Muhammad Farid, et al.
Published: (2024)
by: Adilazuarda, Muhammad Farid, et al.
Published: (2024)
Evaluating LLMs for Visualization Generation and Understanding
by: Khan, Saadiq Rauf, et al.
Published: (2025)
by: Khan, Saadiq Rauf, et al.
Published: (2025)
Learning to Retrieve User History and Generate User Profiles for Personalized Persuasiveness Prediction
by: Park, Sejun, et al.
Published: (2026)
by: Park, Sejun, et al.
Published: (2026)
Towards Simulating Social Media Users with LLMs: Evaluating the Operational Validity of Conditioned Comment Prediction
by: Schwager, Nils, et al.
Published: (2026)
by: Schwager, Nils, et al.
Published: (2026)
Improved Generalized Planning with LLMs through Strategy Refinement and Reflection
by: Stein, Katharina, et al.
Published: (2025)
by: Stein, Katharina, et al.
Published: (2025)
Enabling Scalable Evaluation of Bias Patterns in Medical LLMs
by: Fayyaz, Hamed, et al.
Published: (2024)
by: Fayyaz, Hamed, et al.
Published: (2024)
CHORUS: Zero-shot Hierarchical Retrieval and Orchestration for Generating Linear Programming Code
by: Ahmed, Tasnim, et al.
Published: (2025)
by: Ahmed, Tasnim, et al.
Published: (2025)
The Two Sides of the Coin: Hallucination Generation and Detection with LLMs as Evaluators for LLMs
by: Bui, Anh Thu Maria, et al.
Published: (2024)
by: Bui, Anh Thu Maria, et al.
Published: (2024)
Dialogue Benchmark Generation from Knowledge Graphs with Cost-Effective Retrieval-Augmented LLMs
by: Omar, Reham, et al.
Published: (2025)
by: Omar, Reham, et al.
Published: (2025)
Automating Legal Interpretation with LLMs: Retrieval, Generation, and Evaluation
by: Luo, Kangcheng, et al.
Published: (2025)
by: Luo, Kangcheng, et al.
Published: (2025)
LLMs for Generating and Evaluating Counterfactuals: A Comprehensive Study
by: Nguyen, Van Bach, et al.
Published: (2024)
by: Nguyen, Van Bach, et al.
Published: (2024)
Towards Robust Evaluation of Unlearning in LLMs via Data Transformations
by: Joshi, Abhinav, et al.
Published: (2024)
by: Joshi, Abhinav, et al.
Published: (2024)
Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
by: Pres, Itamar, et al.
Published: (2024)
by: Pres, Itamar, et al.
Published: (2024)
Enhancing NLP Robustness and Generalization through LLM-Generated Contrast Sets: A Scalable Framework for Systematic Evaluation and Adversarial Training
by: Lin, Hender
Published: (2025)
by: Lin, Hender
Published: (2025)
Generating Leakage-Free Benchmarks for Robust RAG Evaluation
by: Liu, Jiayi, et al.
Published: (2026)
by: Liu, Jiayi, et al.
Published: (2026)
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
by: Lunardi, Riccardo, et al.
Published: (2025)
by: Lunardi, Riccardo, et al.
Published: (2025)
PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs
by: Yadav, Ankit, et al.
Published: (2024)
by: Yadav, Ankit, et al.
Published: (2024)
Toward Robust Multilingual Adaptation of LLMs for Low-Resource Languages
by: Li, Haolin, et al.
Published: (2025)
by: Li, Haolin, et al.
Published: (2025)
Generative Data Augmentation using LLMs improves Distributional Robustness in Question Answering
by: Chowdhury, Arijit Ghosh, et al.
Published: (2023)
by: Chowdhury, Arijit Ghosh, et al.
Published: (2023)
Robust Training for Conversational Question Answering Models with Reinforced Reformulation Generation
by: Kaiser, Magdalena, et al.
Published: (2023)
by: Kaiser, Magdalena, et al.
Published: (2023)
Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate
by: Chern, Steffi, et al.
Published: (2024)
by: Chern, Steffi, et al.
Published: (2024)
Similar Items
-
Reading between the Lines: Can LLMs Identify Cross-Cultural Communication Gaps?
by: Saha, Sougata, et al.
Published: (2025) -
Meta-Cultural Competence: Climbing the Right Hill of Cultural Awareness
by: Saha, Sougata, et al.
Published: (2025) -
To Generate or Discriminate? Methodological Considerations for Measuring Cultural Alignment in LLMs
by: Pandey, Saurabh Kumar, et al.
Published: (2026) -
Consolidating Strategies for Countering Hate Speech Using Persuasive Dialogues
by: Saha, Sougata, et al.
Published: (2024) -
Exploring Adapter Design Tradeoffs for Low Resource Music Generation
by: Mehta, Atharva, et al.
Published: (2025)