Eval4Sim: An Evaluation Framework for Persona Simulation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bao, Eliseo, Perez, Anxo, Wang, Xi, Parapar, Javier |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Explainable Depression Symptom Detection in Social Media
von: Bao, Eliseo, et al.
Veröffentlicht: (2023)
von: Bao, Eliseo, et al.
Veröffentlicht: (2023)
ReDSM5: A Reddit Dataset for DSM-5 Depression Detection
von: Bao, Eliseo, et al.
Veröffentlicht: (2025)
von: Bao, Eliseo, et al.
Veröffentlicht: (2025)
Learning Evidence of Depression Symptoms via Prompt Induction
von: Bao, Eliseo, et al.
Veröffentlicht: (2026)
von: Bao, Eliseo, et al.
Veröffentlicht: (2026)
TalkDep: Clinically Grounded LLM Personas for Conversation-Centric Depression Screening
von: Wang, Xi, et al.
Veröffentlicht: (2025)
von: Wang, Xi, et al.
Veröffentlicht: (2025)
Decoding Hate: Exploring Language Models' Reactions to Hate Speech
von: Piot, Paloma, et al.
Veröffentlicht: (2024)
von: Piot, Paloma, et al.
Veröffentlicht: (2024)
Towards Efficient and Explainable Hate Speech Detection via Model Distillation
von: Piot, Paloma, et al.
Veröffentlicht: (2024)
von: Piot, Paloma, et al.
Veröffentlicht: (2024)
PersonaEval: Are LLM Evaluators Human Enough to Judge Role-Play?
von: Zhou, Lingfeng, et al.
Veröffentlicht: (2025)
von: Zhou, Lingfeng, et al.
Veröffentlicht: (2025)
WATCHED: A Web AI Agent Tool for Combating Hate Speech by Expanding Data
von: Piot, Paloma, et al.
Veröffentlicht: (2025)
von: Piot, Paloma, et al.
Veröffentlicht: (2025)
PatientSim: A Persona-Driven Simulator for Realistic Doctor-Patient Interactions
von: Kyung, Daeun, et al.
Veröffentlicht: (2025)
von: Kyung, Daeun, et al.
Veröffentlicht: (2025)
BaZi-Based Character Simulation Benchmark: Evaluating AI on Temporal and Persona Reasoning
von: Zheng, Siyuan, et al.
Veröffentlicht: (2025)
von: Zheng, Siyuan, et al.
Veröffentlicht: (2025)
Who Is the Story About? Protagonist Entity Recognition in News
von: Gabín, Jorge, et al.
Veröffentlicht: (2025)
von: Gabín, Jorge, et al.
Veröffentlicht: (2025)
Personalisation or Prejudice? Addressing Geographic Bias in Hate Speech Detection using Debias Tuning in Large Language Models
von: Piot, Paloma, et al.
Veröffentlicht: (2025)
von: Piot, Paloma, et al.
Veröffentlicht: (2025)
Can LLMs Evaluate What They Cannot Annotate? Revisiting LLM Reliability in Hate Speech Detection
von: Piot, Paloma, et al.
Veröffentlicht: (2025)
von: Piot, Paloma, et al.
Veröffentlicht: (2025)
PartisanLens: A Multilingual Dataset of Hyperpartisan and Conspiratorial Immigration Narratives in European Media
von: Maggini, Michele Joshua, et al.
Veröffentlicht: (2026)
von: Maggini, Michele Joshua, et al.
Veröffentlicht: (2026)
Bridging Gaps in Hate Speech Detection: Meta-Collections and Benchmarks for Low-Resource Iberian Languages
von: Piot, Paloma, et al.
Veröffentlicht: (2025)
von: Piot, Paloma, et al.
Veröffentlicht: (2025)
MetaHate: A Dataset for Unifying Efforts on Hate Speech Detection
von: Piot, Paloma, et al.
Veröffentlicht: (2024)
von: Piot, Paloma, et al.
Veröffentlicht: (2024)
Lost in the Evidence? Reproducing Document Position and Context Size Effects in RAG
von: Gabín, Jorge, et al.
Veröffentlicht: (2026)
von: Gabín, Jorge, et al.
Veröffentlicht: (2026)
CoSER: A Comprehensive Literary Dataset and Framework for Training and Evaluating LLM Role-Playing and Persona Simulation
von: Wang, Xintao, et al.
Veröffentlicht: (2025)
von: Wang, Xintao, et al.
Veröffentlicht: (2025)
EvalCards: A Framework for Standardized Evaluation Reporting
von: Dhar, Ruchira, et al.
Veröffentlicht: (2025)
von: Dhar, Ruchira, et al.
Veröffentlicht: (2025)
FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation
von: He, Zheqi, et al.
Veröffentlicht: (2025)
von: He, Zheqi, et al.
Veröffentlicht: (2025)
The Need for a Socially-Grounded Persona Framework for User Simulation
von: Venkit, Pranav Narayanan, et al.
Veröffentlicht: (2026)
von: Venkit, Pranav Narayanan, et al.
Veröffentlicht: (2026)
DeepResearchEval: An Automated Framework for Deep Research Task Construction and Agentic Evaluation
von: Wang, Yibo, et al.
Veröffentlicht: (2026)
von: Wang, Yibo, et al.
Veröffentlicht: (2026)
EvalSense: A Framework for Domain-Specific LLM (Meta-)Evaluation
von: Dejl, Adam, et al.
Veröffentlicht: (2026)
von: Dejl, Adam, et al.
Veröffentlicht: (2026)
LifeSim: Long-Horizon User Life Simulator for Personalized Assistant Evaluation
von: Duan, Feiyu, et al.
Veröffentlicht: (2026)
von: Duan, Feiyu, et al.
Veröffentlicht: (2026)
ReviewEval: An Evaluation Framework for AI-Generated Reviews
von: Garg, Madhav Krishan, et al.
Veröffentlicht: (2025)
von: Garg, Madhav Krishan, et al.
Veröffentlicht: (2025)
Quantifying the Persona Effect in LLM Simulations
von: Hu, Tiancheng, et al.
Veröffentlicht: (2024)
von: Hu, Tiancheng, et al.
Veröffentlicht: (2024)
InterviewSim: A Scalable Framework for Interview-Grounded Personality Simulation
von: Li, Yu, et al.
Veröffentlicht: (2026)
von: Li, Yu, et al.
Veröffentlicht: (2026)
Eka-Eval: An Evaluation Framework for Low-Resource Multilingual Large Language Models
von: Sinha, Samridhi Raj, et al.
Veröffentlicht: (2025)
von: Sinha, Samridhi Raj, et al.
Veröffentlicht: (2025)
Tau-Eval: A Unified Evaluation Framework for Useful and Private Text Anonymization
von: Loiseau, Gabriel, et al.
Veröffentlicht: (2025)
von: Loiseau, Gabriel, et al.
Veröffentlicht: (2025)
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs
von: Wang, Ganghua, et al.
Veröffentlicht: (2025)
von: Wang, Ganghua, et al.
Veröffentlicht: (2025)
Evaluating Cultural Adaptability of a Large Language Model via Simulation of Synthetic Personas
von: Kwok, Louis, et al.
Veröffentlicht: (2024)
von: Kwok, Louis, et al.
Veröffentlicht: (2024)
AlphaEval: Evaluating Agents in Production
von: Lu, Pengrui, et al.
Veröffentlicht: (2026)
von: Lu, Pengrui, et al.
Veröffentlicht: (2026)
BatchEval: Towards Human-like Text Evaluation
von: Yuan, Peiwen, et al.
Veröffentlicht: (2023)
von: Yuan, Peiwen, et al.
Veröffentlicht: (2023)
PICon: A Multi-Turn Interrogation Framework for Evaluating Persona Agent Consistency
von: Kim, Minseo, et al.
Veröffentlicht: (2026)
von: Kim, Minseo, et al.
Veröffentlicht: (2026)
SceneJailEval: A Scenario-Adaptive Multi-Dimensional Framework for Jailbreak Evaluation
von: Jiang, Lai, et al.
Veröffentlicht: (2025)
von: Jiang, Lai, et al.
Veröffentlicht: (2025)
AraHalluEval: A Fine-grained Hallucination Evaluation Framework for Arabic LLMs
von: Alansari, Aisha, et al.
Veröffentlicht: (2025)
von: Alansari, Aisha, et al.
Veröffentlicht: (2025)
FreeEval: A Modular Framework for Trustworthy and Efficient Evaluation of Large Language Models
von: Yu, Zhuohao, et al.
Veröffentlicht: (2024)
von: Yu, Zhuohao, et al.
Veröffentlicht: (2024)
GenSim: Generating Robotic Simulation Tasks via Large Language Models
von: Wang, Lirui, et al.
Veröffentlicht: (2023)
von: Wang, Lirui, et al.
Veröffentlicht: (2023)
SocialSim: Towards Socialized Simulation of Emotional Support Conversation
von: Chen, Zhuang, et al.
Veröffentlicht: (2025)
von: Chen, Zhuang, et al.
Veröffentlicht: (2025)
HintEval: A Comprehensive Framework for Hint Generation and Evaluation for Questions
von: Mozafari, Jamshid, et al.
Veröffentlicht: (2025)
von: Mozafari, Jamshid, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Explainable Depression Symptom Detection in Social Media
von: Bao, Eliseo, et al.
Veröffentlicht: (2023) -
ReDSM5: A Reddit Dataset for DSM-5 Depression Detection
von: Bao, Eliseo, et al.
Veröffentlicht: (2025) -
Learning Evidence of Depression Symptoms via Prompt Induction
von: Bao, Eliseo, et al.
Veröffentlicht: (2026) -
TalkDep: Clinically Grounded LLM Personas for Conversation-Centric Depression Screening
von: Wang, Xi, et al.
Veröffentlicht: (2025) -
Decoding Hate: Exploring Language Models' Reactions to Hate Speech
von: Piot, Paloma, et al.
Veröffentlicht: (2024)