BaZi-Based Character Simulation Benchmark: Evaluating AI on Temporal and Persona Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zheng, Siyuan, Liu, Pai, Chen, Xi, Dong, Jizheng, Jia, Sihan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Eval4Sim: An Evaluation Framework for Persona Simulation
von: Bao, Eliseo, et al.
Veröffentlicht: (2026)
von: Bao, Eliseo, et al.
Veröffentlicht: (2026)
PersonaTwin: A Multi-Tier Prompt Conditioning Framework for Generating and Evaluating Personalized Digital Twins
von: Chen, Sihan, et al.
Veröffentlicht: (2025)
von: Chen, Sihan, et al.
Veröffentlicht: (2025)
BaziQA-Benchmark: Evaluating Symbolic and Temporally Compositional Reasoning in Large Language Models
von: Chen, Jiangxi, et al.
Veröffentlicht: (2026)
von: Chen, Jiangxi, et al.
Veröffentlicht: (2026)
Evaluating Computational Representations of Character: An Austen Character Similarity Benchmark
von: Yang, Funing, et al.
Veröffentlicht: (2024)
von: Yang, Funing, et al.
Veröffentlicht: (2024)
Crafting Customisable Characters with LLMs: A Persona-Driven Role-Playing Agent Framework
von: Yang, Bohao, et al.
Veröffentlicht: (2024)
von: Yang, Bohao, et al.
Veröffentlicht: (2024)
AI Hospital: Benchmarking Large Language Models in a Multi-agent Medical Interaction Simulator
von: Fan, Zhihao, et al.
Veröffentlicht: (2024)
von: Fan, Zhihao, et al.
Veröffentlicht: (2024)
PersonaLens: A Benchmark for Personalization Evaluation in Conversational AI Assistants
von: Zhao, Zheng, et al.
Veröffentlicht: (2025)
von: Zhao, Zheng, et al.
Veröffentlicht: (2025)
Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI
von: Maiya, Sharan, et al.
Veröffentlicht: (2025)
von: Maiya, Sharan, et al.
Veröffentlicht: (2025)
How Far are LLMs from Being Our Digital Twins? A Benchmark for Persona-Based Behavior Chain Simulation
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
OpenCharacter: Training Customizable Role-Playing LLMs with Large-Scale Synthetic Personas
von: Wang, Xiaoyang, et al.
Veröffentlicht: (2025)
von: Wang, Xiaoyang, et al.
Veröffentlicht: (2025)
Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
von: Fatemi, Bahare, et al.
Veröffentlicht: (2024)
von: Fatemi, Bahare, et al.
Veröffentlicht: (2024)
Persona Vectors: Monitoring and Controlling Character Traits in Language Models
von: Chen, Runjin, et al.
Veröffentlicht: (2025)
von: Chen, Runjin, et al.
Veröffentlicht: (2025)
AstroMind: A High-Fidelity Benchmark for Spacecraft Behavior Reasoning Based on Large Language Models
von: Liu, Hao, et al.
Veröffentlicht: (2026)
von: Liu, Hao, et al.
Veröffentlicht: (2026)
CharacterBench: Benchmarking Character Customization of Large Language Models
von: Zhou, Jinfeng, et al.
Veröffentlicht: (2024)
von: Zhou, Jinfeng, et al.
Veröffentlicht: (2024)
Beyond Static Benchmarks: Synthesizing Harmful Content via Persona-based Simulation for Robust Evaluation
von: Lee, Huije, et al.
Veröffentlicht: (2026)
von: Lee, Huije, et al.
Veröffentlicht: (2026)
ZiGong 1.0: A Large Language Model for Financial Credit
von: Lei, Yu, et al.
Veröffentlicht: (2025)
von: Lei, Yu, et al.
Veröffentlicht: (2025)
LaRe: Latent Refocusing for Multimodal Reasoning
von: Ma, Jizheng, et al.
Veröffentlicht: (2025)
von: Ma, Jizheng, et al.
Veröffentlicht: (2025)
Can Deception Detection Go Deeper? Dataset, Evaluation, and Benchmark for Deception Reasoning
von: Chen, Kang, et al.
Veröffentlicht: (2024)
von: Chen, Kang, et al.
Veröffentlicht: (2024)
MR-GSM8K: A Meta-Reasoning Benchmark for Large Language Model Evaluation
von: Zeng, Zhongshen, et al.
Veröffentlicht: (2023)
von: Zeng, Zhongshen, et al.
Veröffentlicht: (2023)
Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
von: Shin, Jisu, et al.
Veröffentlicht: (2025)
von: Shin, Jisu, et al.
Veröffentlicht: (2025)
PersonaMath: Boosting Mathematical Reasoning via Persona-Driven Data Augmentation
von: Luo, Jing, et al.
Veröffentlicht: (2024)
von: Luo, Jing, et al.
Veröffentlicht: (2024)
Beyond Perplexity: Character Distribution Signatures and the MDTA Benchmark for AI Text Detection
von: Narayanasamy, Priyadarshan, et al.
Veröffentlicht: (2026)
von: Narayanasamy, Priyadarshan, et al.
Veröffentlicht: (2026)
TRAVELER: A Benchmark for Evaluating Temporal Reasoning across Vague, Implicit and Explicit References
von: Kenneweg, Svenja, et al.
Veröffentlicht: (2025)
von: Kenneweg, Svenja, et al.
Veröffentlicht: (2025)
LTLBench: Towards Benchmarks for Evaluating Temporal Reasoning in Large Language Models
von: Tang, Weizhi, et al.
Veröffentlicht: (2024)
von: Tang, Weizhi, et al.
Veröffentlicht: (2024)
A Multi-Task Role-Playing Agent Capable of Imitating Character Linguistic Styles
von: Chen, Siyuan, et al.
Veröffentlicht: (2024)
von: Chen, Siyuan, et al.
Veröffentlicht: (2024)
CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation
von: Tu, Quan, et al.
Veröffentlicht: (2024)
von: Tu, Quan, et al.
Veröffentlicht: (2024)
CharacterGPT: A Persona Reconstruction Framework for Role-Playing Agents
von: Park, Jeiyoon, et al.
Veröffentlicht: (2024)
von: Park, Jeiyoon, et al.
Veröffentlicht: (2024)
TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models
von: Chu, Zheng, et al.
Veröffentlicht: (2023)
von: Chu, Zheng, et al.
Veröffentlicht: (2023)
Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations
von: Juneja, Prerna, et al.
Veröffentlicht: (2026)
von: Juneja, Prerna, et al.
Veröffentlicht: (2026)
A Persona-Based Evaluation Framework for Pluralistic Alignment in Generative AI
von: Karagoz, Atahan
Veröffentlicht: (2026)
von: Karagoz, Atahan
Veröffentlicht: (2026)
Benchmarking Temporal Reasoning and Alignment Across Chinese Dynasties
von: Wang, Zhenglin, et al.
Veröffentlicht: (2025)
von: Wang, Zhenglin, et al.
Veröffentlicht: (2025)
Quantifying the Persona Effect in LLM Simulations
von: Hu, Tiancheng, et al.
Veröffentlicht: (2024)
von: Hu, Tiancheng, et al.
Veröffentlicht: (2024)
TRAM: Benchmarking Temporal Reasoning for Large Language Models
von: Wang, Yuqing, et al.
Veröffentlicht: (2023)
von: Wang, Yuqing, et al.
Veröffentlicht: (2023)
Benchmarking and Confidence Evaluation of LALMs For Temporal Reasoning
von: Bhattacharya, Debarpan, et al.
Veröffentlicht: (2025)
von: Bhattacharya, Debarpan, et al.
Veröffentlicht: (2025)
Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing Styles
von: Truong, Kimberly Le, et al.
Veröffentlicht: (2025)
von: Truong, Kimberly Le, et al.
Veröffentlicht: (2025)
Evaluating Cultural Adaptability of a Large Language Model via Simulation of Synthetic Personas
von: Kwok, Louis, et al.
Veröffentlicht: (2024)
von: Kwok, Louis, et al.
Veröffentlicht: (2024)
Persona-Based Conversational AI: State of the Art and Challenges
von: Liu, Junfeng, et al.
Veröffentlicht: (2022)
von: Liu, Junfeng, et al.
Veröffentlicht: (2022)
How Far Are LLMs from Believable AI? A Benchmark for Evaluating the Believability of Human Behavior Simulation
von: Xiao, Yang, et al.
Veröffentlicht: (2023)
von: Xiao, Yang, et al.
Veröffentlicht: (2023)
Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History
von: Kim, Serin, et al.
Veröffentlicht: (2026)
von: Kim, Serin, et al.
Veröffentlicht: (2026)
SimulBench: Evaluating Language Models with Creative Simulation Tasks
von: Jia, Qi, et al.
Veröffentlicht: (2024)
von: Jia, Qi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Eval4Sim: An Evaluation Framework for Persona Simulation
von: Bao, Eliseo, et al.
Veröffentlicht: (2026) -
PersonaTwin: A Multi-Tier Prompt Conditioning Framework for Generating and Evaluating Personalized Digital Twins
von: Chen, Sihan, et al.
Veröffentlicht: (2025) -
BaziQA-Benchmark: Evaluating Symbolic and Temporally Compositional Reasoning in Large Language Models
von: Chen, Jiangxi, et al.
Veröffentlicht: (2026) -
Evaluating Computational Representations of Character: An Austen Character Similarity Benchmark
von: Yang, Funing, et al.
Veröffentlicht: (2024) -
Crafting Customisable Characters with LLMs: A Persona-Driven Role-Playing Agent Framework
von: Yang, Bohao, et al.
Veröffentlicht: (2024)