PICon: A Multi-Turn Interrogation Framework for Evaluating Persona Agent Consistency
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Minseo, Im, Sujeong, Choi, Junseong, Lee, Junhee, Shim, Chaeeun, Hong, Hwajung, Choi, Edward |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Error-Free EHRs: Reasoning-Intensive Consistency Verification Between Clinical Notes and Structured Tables in Electronic Health Records
by: Kwon, Yeonsu, et al.
Published: (2026)
by: Kwon, Yeonsu, et al.
Published: (2026)
R2-KG: General-Purpose Dual-Agent Framework for Reliable Reasoning on Knowledge Graphs
by: Jo, Sumin, et al.
Published: (2025)
by: Jo, Sumin, et al.
Published: (2025)
MapCoder-Lite: Distilling Multi-Agent Coding into a Single Small LLM
by: Lee, Woongkyu, et al.
Published: (2025)
by: Lee, Woongkyu, et al.
Published: (2025)
Integrated Framework for LLM Evaluation with Answer Generation
by: Lee, Sujeong, et al.
Published: (2025)
by: Lee, Sujeong, et al.
Published: (2025)
General-Purpose Retrieval-Enhanced Medical Prediction Model Using Near-Infinite History
by: Kim, Junu, et al.
Published: (2023)
by: Kim, Junu, et al.
Published: (2023)
MoCoRP: Modeling Consistent Relations between Persona and Response for Persona-based Dialogue
by: Lee, Kyungro, et al.
Published: (2025)
by: Lee, Kyungro, et al.
Published: (2025)
ProPerSim: Developing Proactive and Personalized AI Assistants through User-Assistant Simulation
by: Kim, Jiho, et al.
Published: (2025)
by: Kim, Jiho, et al.
Published: (2025)
Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning
by: Abdulhai, Marwa, et al.
Published: (2025)
by: Abdulhai, Marwa, et al.
Published: (2025)
Feed-O-Meter: Investigating AI-Generated Mentee Personas as Interactive Agents for Scaffolding Design Feedback Practice
by: Lim, Hyunseung, et al.
Published: (2025)
by: Lim, Hyunseung, et al.
Published: (2025)
HuDEx: Integrating Hallucination Detection and Explainability for Enhancing the Reliability of LLM responses
by: Lee, Sujeong, et al.
Published: (2025)
by: Lee, Sujeong, et al.
Published: (2025)
SpeakerSleuth: Can Large Audio-Language Models Judge Speaker Consistency across Multi-turn Dialogues?
by: Lee, Jonggeun, et al.
Published: (2026)
by: Lee, Jonggeun, et al.
Published: (2026)
Mind the Blind Spots: A Focus-Level Evaluation Framework for LLM Reviews
by: Shin, Hyungyu, et al.
Published: (2025)
by: Shin, Hyungyu, et al.
Published: (2025)
Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations
by: Juneja, Prerna, et al.
Published: (2026)
by: Juneja, Prerna, et al.
Published: (2026)
Belief in Authority: Impact of Authority in Multi-Agent Evaluation Framework
by: Choi, Junhyuk, et al.
Published: (2026)
by: Choi, Junhyuk, et al.
Published: (2026)
Evaluating Temporal Consistency in Multi-Turn Language Models
by: Atri, Yash Kumar, et al.
Published: (2026)
by: Atri, Yash Kumar, et al.
Published: (2026)
ECG-Reasoning-Benchmark: A Benchmark for Evaluating Clinical Reasoning Capabilities in ECG Interpretation
by: Oh, Jungwoo, et al.
Published: (2026)
by: Oh, Jungwoo, et al.
Published: (2026)
Re-Ex: Revising after Explanation Reduces the Factual Errors in LLM Responses
by: Kim, Juyeon, et al.
Published: (2024)
by: Kim, Juyeon, et al.
Published: (2024)
Can Separators Improve Chain-of-Thought Prompting?
by: Park, Yoonjeong, et al.
Published: (2024)
by: Park, Yoonjeong, et al.
Published: (2024)
Understanding Human-Multi-Agent Team Formation for Creative Work
by: Lim, Hyunseung, et al.
Published: (2026)
by: Lim, Hyunseung, et al.
Published: (2026)
LegalSearchLM: Rethinking Legal Case Retrieval as Legal Elements Generation
by: Kim, Chaeeun, et al.
Published: (2025)
by: Kim, Chaeeun, et al.
Published: (2025)
Interrogating LLM design under a fair learning doctrine
by: Wei, Johnny Tian-Zheng, et al.
Published: (2025)
by: Wei, Johnny Tian-Zheng, et al.
Published: (2025)
H-AdminSim: A Multi-Agent Simulator for Realistic Hospital Administrative Workflows with FHIR Integration
by: Lee, Jun-Min, et al.
Published: (2026)
by: Lee, Jun-Min, et al.
Published: (2026)
DiTTO-LLM: Framework for Discovering Topic-based Technology Opportunities via Large Language Model
by: Kim, Wonyoung, et al.
Published: (2025)
by: Kim, Wonyoung, et al.
Published: (2025)
Trans-EnV: A Framework for Evaluating the Linguistic Robustness of LLMs Against English Varieties
by: Lee, Jiyoung, et al.
Published: (2025)
by: Lee, Jiyoung, et al.
Published: (2025)
LabTOP: A Unified Model for Lab Test Outcome Prediction on Electronic Health Records
by: Im, Sujeong, et al.
Published: (2025)
by: Im, Sujeong, et al.
Published: (2025)
Examining Identity Drift in Conversations of LLM Agents
by: Choi, Junhyuk, et al.
Published: (2024)
by: Choi, Junhyuk, et al.
Published: (2024)
Remember Your Trace: Memory-Guided Long-Horizon Agentic Framework for Consistent and Hierarchical Repository-Level Code Documentation
by: Bae, Suyoung, et al.
Published: (2026)
by: Bae, Suyoung, et al.
Published: (2026)
Linq-Embed-Mistral Technical Report
by: Choi, Chanyeol, et al.
Published: (2024)
by: Choi, Chanyeol, et al.
Published: (2024)
Mitigating Length Bias in RLHF through a Causal Lens
by: Kim, Hyeonji, et al.
Published: (2025)
by: Kim, Hyeonji, et al.
Published: (2025)
Tool-MAD: A Multi-Agent Debate Framework for Fact Verification with Diverse Tool Augmentation and Adaptive Retrieval
by: Jeong, Seyeon, et al.
Published: (2026)
by: Jeong, Seyeon, et al.
Published: (2026)
Evaluating the Consistency of LLM Evaluators
by: Lee, Noah, et al.
Published: (2024)
by: Lee, Noah, et al.
Published: (2024)
PatientSim: A Persona-Driven Simulator for Realistic Doctor-Patient Interactions
by: Kyung, Daeun, et al.
Published: (2025)
by: Kyung, Daeun, et al.
Published: (2025)
A Large-Scale Real-World Evaluation of LLM-Based Virtual Teaching Assistant
by: Kweon, Sunjun, et al.
Published: (2025)
by: Kweon, Sunjun, et al.
Published: (2025)
Stark: Social Long-Term Multi-Modal Conversation with Persona Commonsense Knowledge
by: Lee, Young-Jun, et al.
Published: (2024)
by: Lee, Young-Jun, et al.
Published: (2024)
Overthinking Loops in Agents: A Structural Risk via MCP Tools
by: Lee, Yohan, et al.
Published: (2026)
by: Lee, Yohan, et al.
Published: (2026)
Cross-lingual Transfer for Automatic Question Generation by Learning Interrogative Structures in Target Languages
by: Hwang, Seonjeong, et al.
Published: (2024)
by: Hwang, Seonjeong, et al.
Published: (2024)
Fact-Consistency Evaluation of Text-to-SQL Generation for Business Intelligence Using Exaone 3.5
by: Choi, Jeho
Published: (2025)
by: Choi, Jeho
Published: (2025)
Pedagogical Alignment for Vision-Language-Action Models: A Comprehensive Framework for Data, Architecture, and Evaluation in Education
by: Lee, Unggi, et al.
Published: (2026)
by: Lee, Unggi, et al.
Published: (2026)
SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents
by: Seo, Gyuhyeon, et al.
Published: (2025)
by: Seo, Gyuhyeon, et al.
Published: (2025)
InfiniPot: Infinite Context Processing on Memory-Constrained LLMs
by: Kim, Minsoo, et al.
Published: (2024)
by: Kim, Minsoo, et al.
Published: (2024)
Similar Items
-
Towards Error-Free EHRs: Reasoning-Intensive Consistency Verification Between Clinical Notes and Structured Tables in Electronic Health Records
by: Kwon, Yeonsu, et al.
Published: (2026) -
R2-KG: General-Purpose Dual-Agent Framework for Reliable Reasoning on Knowledge Graphs
by: Jo, Sumin, et al.
Published: (2025) -
MapCoder-Lite: Distilling Multi-Agent Coding into a Single Small LLM
by: Lee, Woongkyu, et al.
Published: (2025) -
Integrated Framework for LLM Evaluation with Answer Generation
by: Lee, Sujeong, et al.
Published: (2025) -
General-Purpose Retrieval-Enhanced Medical Prediction Model Using Near-Infinite History
by: Kim, Junu, et al.
Published: (2023)