PTCBENCH: Benchmarking Contextual Stability of Personality Traits in LLM Systems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yu, Jiongchi, Ma, Yuhan, Zhang, Xiaoyu, Wang, Junjie, Hu, Qiang, Shen, Chao, Xie, Xiaofei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AutoEmpirical: LLM-Based Automated Research for Empirical Software Fault Analysis
von: Yu, Jiongchi, et al.
Veröffentlicht: (2025)
von: Yu, Jiongchi, et al.
Veröffentlicht: (2025)
The Foundation Cracks: A Comprehensive Study on Bugs and Testing Practices in LLM Libraries
von: Jiang, Weipeng, et al.
Veröffentlicht: (2025)
von: Jiang, Weipeng, et al.
Veröffentlicht: (2025)
Chimera: Harnessing Multi-Agent LLMs for Automatic Insider Threat Simulation
von: Yu, Jiongchi, et al.
Veröffentlicht: (2025)
von: Yu, Jiongchi, et al.
Veröffentlicht: (2025)
DynaTrust: Defending Multi-Agent Systems Against Sleeper Agents via Dynamic Trust Graphs
von: Li, Yu, et al.
Veröffentlicht: (2026)
von: Li, Yu, et al.
Veröffentlicht: (2026)
Exposing Product Bias in LLM Investment Recommendation
von: Zhi, Yuhan, et al.
Veröffentlicht: (2025)
von: Zhi, Yuhan, et al.
Veröffentlicht: (2025)
Variation is the Key: A Variation-Based Framework for LLM-Generated Text Detection
von: Li, Xuecong, et al.
Veröffentlicht: (2026)
von: Li, Xuecong, et al.
Veröffentlicht: (2026)
HYDRA: Model Factorization Framework for Black-Box LLM Personalization
von: Zhuang, Yuchen, et al.
Veröffentlicht: (2024)
von: Zhuang, Yuchen, et al.
Veröffentlicht: (2024)
Exploring the Personality Traits of LLMs through Latent Features Steering
von: Yang, Shu, et al.
Veröffentlicht: (2024)
von: Yang, Shu, et al.
Veröffentlicht: (2024)
PsychAdapter: Adapting LLM Transformers to Reflect Traits, Personality and Mental Health
von: Vu, Huy, et al.
Veröffentlicht: (2024)
von: Vu, Huy, et al.
Veröffentlicht: (2024)
CTC-Assisted LLM-Based Contextual ASR
von: Yang, Guanrou, et al.
Veröffentlicht: (2024)
von: Yang, Guanrou, et al.
Veröffentlicht: (2024)
AlpsBench: An LLM Personalization Benchmark for Real-Dialogue Memorization and Preference Alignment
von: Xiao, Jianfei, et al.
Veröffentlicht: (2026)
von: Xiao, Jianfei, et al.
Veröffentlicht: (2026)
Hybrid OCR-LLM Framework for Enterprise-Scale Document Information Extraction Under Copy-heavy Task
von: Wang, Zilong, et al.
Veröffentlicht: (2025)
von: Wang, Zilong, et al.
Veröffentlicht: (2025)
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations
von: Li, Li, et al.
Veröffentlicht: (2025)
von: Li, Li, et al.
Veröffentlicht: (2025)
Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History
von: Kim, Serin, et al.
Veröffentlicht: (2026)
von: Kim, Serin, et al.
Veröffentlicht: (2026)
Selection-Based Vulnerabilities: Clean-Label Backdoor Attacks in Active Learning
von: Zhi, Yuhan, et al.
Veröffentlicht: (2025)
von: Zhi, Yuhan, et al.
Veröffentlicht: (2025)
Benchmarking and Improving LLM Robustness for Personalized Generation
von: Okite, Chimaobi, et al.
Veröffentlicht: (2025)
von: Okite, Chimaobi, et al.
Veröffentlicht: (2025)
Personality as Relational Infrastructure: User Perceptions of Personality-Trait-Infused LLM Messaging
von: Hofer, Dominik P., et al.
Veröffentlicht: (2026)
von: Hofer, Dominik P., et al.
Veröffentlicht: (2026)
Predicting the Big Five Personality Traits in Chinese Counselling Dialogues Using Large Language Models
von: Yan, Yang, et al.
Veröffentlicht: (2024)
von: Yan, Yang, et al.
Veröffentlicht: (2024)
Contextualized Privacy Defense for LLM Agents
von: Wen, Yule, et al.
Veröffentlicht: (2026)
von: Wen, Yule, et al.
Veröffentlicht: (2026)
Eliciting Personality Traits in Large Language Models
von: Hilliard, Airlie, et al.
Veröffentlicht: (2024)
von: Hilliard, Airlie, et al.
Veröffentlicht: (2024)
Personalized Turn-Level User Conversation Satisfaction Benchmark
von: Wang, Zhefan, et al.
Veröffentlicht: (2026)
von: Wang, Zhefan, et al.
Veröffentlicht: (2026)
A Comparative Study of Large Language Models and Human Personality Traits
von: Jiaqi, Wang, et al.
Veröffentlicht: (2025)
von: Jiaqi, Wang, et al.
Veröffentlicht: (2025)
PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits
von: Jiang, Hang, et al.
Veröffentlicht: (2023)
von: Jiang, Hang, et al.
Veröffentlicht: (2023)
Persona Dynamics: Unveiling the Impact of Personality Traits on Agents in Text-Based Games
von: Lim, Seungwon, et al.
Veröffentlicht: (2025)
von: Lim, Seungwon, et al.
Veröffentlicht: (2025)
PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
von: Wang, Yidong, et al.
Veröffentlicht: (2023)
von: Wang, Yidong, et al.
Veröffentlicht: (2023)
Advancing and Benchmarking Personalized Tool Invocation for LLMs
von: Huang, Xu, et al.
Veröffentlicht: (2025)
von: Huang, Xu, et al.
Veröffentlicht: (2025)
MPCI-Bench: A Benchmark for Multimodal Pairwise Contextual Integrity Evaluation of Language Model Agents
von: Wang, Shouju, et al.
Veröffentlicht: (2026)
von: Wang, Shouju, et al.
Veröffentlicht: (2026)
Identifying and Manipulating Personality Traits in LLMs Through Activation Engineering
von: Allbert, Rumi, et al.
Veröffentlicht: (2024)
von: Allbert, Rumi, et al.
Veröffentlicht: (2024)
ES-MemEval: Benchmarking Conversational Agents on Personalized Long-Term Emotional Support
von: Chen, Tiantian, et al.
Veröffentlicht: (2026)
von: Chen, Tiantian, et al.
Veröffentlicht: (2026)
BioKGBench: A Knowledge Graph Checking Benchmark of AI Agent for Biomedical Science
von: Lin, Xinna, et al.
Veröffentlicht: (2024)
von: Lin, Xinna, et al.
Veröffentlicht: (2024)
User-LLM: Efficient LLM Contextualization with User Embeddings
von: Ning, Lin, et al.
Veröffentlicht: (2024)
von: Ning, Lin, et al.
Veröffentlicht: (2024)
CTSM: Combining Trait and State Emotions for Empathetic Response Model
von: Yufeng, Wang, et al.
Veröffentlicht: (2024)
von: Yufeng, Wang, et al.
Veröffentlicht: (2024)
Leveraging Implicit Sentiments: Enhancing Reliability and Validity in Psychological Trait Evaluation of LLMs
von: Ma, Huanhuan, et al.
Veröffentlicht: (2025)
von: Ma, Huanhuan, et al.
Veröffentlicht: (2025)
OP-Bench: Benchmarking Over-Personalization for Memory-Augmented Personalized Conversational Agents
von: Hu, Yulin, et al.
Veröffentlicht: (2026)
von: Hu, Yulin, et al.
Veröffentlicht: (2026)
SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation
von: Yu, Wenyi, et al.
Veröffentlicht: (2025)
von: Yu, Wenyi, et al.
Veröffentlicht: (2025)
ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning
von: Potamitis, Nearchos, et al.
Veröffentlicht: (2025)
von: Potamitis, Nearchos, et al.
Veröffentlicht: (2025)
JudgeAgent: Beyond Static Benchmarks for Knowledge-Driven and Dynamic LLM Evaluation
von: Shi, Zhichao, et al.
Veröffentlicht: (2025)
von: Shi, Zhichao, et al.
Veröffentlicht: (2025)
Inclusivity in Large Language Models: Personality Traits and Gender Bias in Scientific Abstracts
von: Pervez, Naseela, et al.
Veröffentlicht: (2024)
von: Pervez, Naseela, et al.
Veröffentlicht: (2024)
SRFUND: A Multi-Granularity Hierarchical Structure Reconstruction Benchmark in Form Understanding
von: Ma, Jiefeng, et al.
Veröffentlicht: (2024)
von: Ma, Jiefeng, et al.
Veröffentlicht: (2024)
StableMask: Refining Causal Masking in Decoder-only Transformer
von: Yin, Qingyu, et al.
Veröffentlicht: (2024)
von: Yin, Qingyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
AutoEmpirical: LLM-Based Automated Research for Empirical Software Fault Analysis
von: Yu, Jiongchi, et al.
Veröffentlicht: (2025) -
The Foundation Cracks: A Comprehensive Study on Bugs and Testing Practices in LLM Libraries
von: Jiang, Weipeng, et al.
Veröffentlicht: (2025) -
Chimera: Harnessing Multi-Agent LLMs for Automatic Insider Threat Simulation
von: Yu, Jiongchi, et al.
Veröffentlicht: (2025) -
DynaTrust: Defending Multi-Agent Systems Against Sleeper Agents via Dynamic Trust Graphs
von: Li, Yu, et al.
Veröffentlicht: (2026) -
Exposing Product Bias in LLM Investment Recommendation
von: Zhi, Yuhan, et al.
Veröffentlicht: (2025)