Self-Assessment Tests are Unreliable Measures of LLM Personality
Fuente:
arXiv
Saved in:
| Main Authors: | Gupta, Akshat, Song, Xiaoyang, Anumanchipalli, Gopala |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rebuilding ROME : Resolving Model Collapse during Sequential Model Editing
by: Gupta, Akshat, et al.
Published: (2024)
by: Gupta, Akshat, et al.
Published: (2024)
A Unified Framework for Model Editing
by: Gupta, Akshat, et al.
Published: (2024)
by: Gupta, Akshat, et al.
Published: (2024)
Is Bigger Edit Batch Size Always Better? -- An Empirical Study on Model Editing with Llama-3
by: Yoon, Junsang, et al.
Published: (2024)
by: Yoon, Junsang, et al.
Published: (2024)
Geometric Interpretation of Layer Normalization and a Comparative Analysis with RMSNorm
by: Gupta, Akshat, et al.
Published: (2024)
by: Gupta, Akshat, et al.
Published: (2024)
Model Editing at Scale leads to Gradual and Catastrophic Forgetting
by: Gupta, Akshat, et al.
Published: (2024)
by: Gupta, Akshat, et al.
Published: (2024)
How Do LLMs Use Their Depth?
by: Gupta, Akshat, et al.
Published: (2025)
by: Gupta, Akshat, et al.
Published: (2025)
Efficient Knowledge Editing via Minimal Precomputation
by: Gupta, Akshat, et al.
Published: (2025)
by: Gupta, Akshat, et al.
Published: (2025)
Identifying Multiple Personalities in Large Language Models with External Evaluation
by: Song, Xiaoyang, et al.
Published: (2024)
by: Song, Xiaoyang, et al.
Published: (2024)
The Oracle Has Spoken: A Multi-Aspect Evaluation of Dialogue in Pythia
by: Chen, Zixun, et al.
Published: (2025)
by: Chen, Zixun, et al.
Published: (2025)
Evolutionary Strategies lead to Catastrophic Forgetting in LLMs
by: Abdi, Immanuel, et al.
Published: (2026)
by: Abdi, Immanuel, et al.
Published: (2026)
Lifelong Knowledge Editing requires Better Regularization
by: Gupta, Akshat, et al.
Published: (2025)
by: Gupta, Akshat, et al.
Published: (2025)
Norm Growth and Stability Challenges in Localized Sequential Knowledge Editing
by: Gupta, Akshat, et al.
Published: (2025)
by: Gupta, Akshat, et al.
Published: (2025)
PokerBench: Training Large Language Models to become Professional Poker Players
by: Zhuang, Richard, et al.
Published: (2025)
by: Zhuang, Richard, et al.
Published: (2025)
FineZip : Pushing the Limits of Large Language Models for Practical Lossless Text Compression
by: Mittu, Fazal, et al.
Published: (2024)
by: Mittu, Fazal, et al.
Published: (2024)
Unsupervised TTS Acoustic Modeling for TTS with Conditional Disentangled Sequential VAE
by: Lian, Jiachen, et al.
Published: (2022)
by: Lian, Jiachen, et al.
Published: (2022)
Coding Speech through Vocal Tract Kinematics
by: Cho, Cheol Jun, et al.
Published: (2024)
by: Cho, Cheol Jun, et al.
Published: (2024)
Enabling Conversational Behavior Reasoning Capabilities in Full-Duplex Speech
by: Pan, Shuchang, et al.
Published: (2025)
by: Pan, Shuchang, et al.
Published: (2025)
It Takes Two: A Dual Stage Approach for Terminology-Aware Translation
by: Jaswal, Akshat Singh
Published: (2025)
by: Jaswal, Akshat Singh
Published: (2025)
From Test-Taking to Test-Making: Examining LLM Authoring of Commonsense Assessment Items
by: Roemmele, Melissa, et al.
Published: (2024)
by: Roemmele, Melissa, et al.
Published: (2024)
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
by: Mayne, Harry, et al.
Published: (2025)
by: Mayne, Harry, et al.
Published: (2025)
Towards Hierarchical Spoken Language Dysfluency Modeling
by: Lian, Jiachen, et al.
Published: (2024)
by: Lian, Jiachen, et al.
Published: (2024)
Persistent Instability in LLM's Personality Measurements: Effects of Scale, Reasoning, and Conversation History
by: Tosato, Tommaso, et al.
Published: (2025)
by: Tosato, Tommaso, et al.
Published: (2025)
Beyond Self-Reports: Multi-Observer Agents for Personality Assessment in Large Language Models
by: Huang, Yin Jou, et al.
Published: (2025)
by: Huang, Yin Jou, et al.
Published: (2025)
"All that Glitters": Approaches to Evaluations with Unreliable Model and Human Annotations
by: Hardy, Michael
Published: (2024)
by: Hardy, Michael
Published: (2024)
OckBench: Measuring the Efficiency of LLM Reasoning
by: Du, Zheng, et al.
Published: (2025)
by: Du, Zheng, et al.
Published: (2025)
Beyond Single-shot Writing: Deep Research Agents are Unreliable at Multi-turn Report Revision
by: Chen, Bingsen, et al.
Published: (2026)
by: Chen, Bingsen, et al.
Published: (2026)
Personalized LLM Decoding via Contrasting Personal Preference
by: Bu, Hyungjune, et al.
Published: (2025)
by: Bu, Hyungjune, et al.
Published: (2025)
Relying on the Unreliable: The Impact of Language Models' Reluctance to Express Uncertainty
by: Zhou, Kaitlyn, et al.
Published: (2024)
by: Zhou, Kaitlyn, et al.
Published: (2024)
Self-Improving LLM Agents at Test-Time
by: Acikgoz, Emre Can, et al.
Published: (2025)
by: Acikgoz, Emre Can, et al.
Published: (2025)
Generative LLM Powered Conversational AI Application for Personalized Risk Assessment: A Case Study in COVID-19
by: Roshani, Mohammad Amin, et al.
Published: (2024)
by: Roshani, Mohammad Amin, et al.
Published: (2024)
Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
by: Wei, Tianxin, et al.
Published: (2025)
by: Wei, Tianxin, et al.
Published: (2025)
Krutrim LLM: Multilingual Foundational Model for over a Billion People
by: Kallappa, Aditya, et al.
Published: (2025)
by: Kallappa, Aditya, et al.
Published: (2025)
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
by: Spiliopoulou, Evangelia, et al.
Published: (2025)
by: Spiliopoulou, Evangelia, et al.
Published: (2025)
StyleStream: Real-Time Zero-Shot Voice Style Conversion
by: Liu, Yisi, et al.
Published: (2026)
by: Liu, Yisi, et al.
Published: (2026)
Tool Preferences in Agentic LLMs are Unreliable
by: Faghih, Kazem, et al.
Published: (2025)
by: Faghih, Kazem, et al.
Published: (2025)
From Personal to Collective: On the Role of Local and Global Memory in LLM Personalization
by: Wang, Zehong, et al.
Published: (2025)
by: Wang, Zehong, et al.
Published: (2025)
Silicon Bureaucracy and AI Test-Oriented Education: Contamination Sensitivity and Score Confidence in LLM Benchmarks
by: Song, Yiliang, et al.
Published: (2026)
by: Song, Yiliang, et al.
Published: (2026)
Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination
by: Wu, Mingqi, et al.
Published: (2025)
by: Wu, Mingqi, et al.
Published: (2025)
Adaptive Self-Supervised Learning Strategies for Dynamic On-Device LLM Personalization
by: Mendoza, Rafael, et al.
Published: (2024)
by: Mendoza, Rafael, et al.
Published: (2024)
Self-signals Driven Multi-LLM Debate for Efficient and Accurate Reasoning
by: Chen, Xuhang, et al.
Published: (2025)
by: Chen, Xuhang, et al.
Published: (2025)
Similar Items
-
Rebuilding ROME : Resolving Model Collapse during Sequential Model Editing
by: Gupta, Akshat, et al.
Published: (2024) -
A Unified Framework for Model Editing
by: Gupta, Akshat, et al.
Published: (2024) -
Is Bigger Edit Batch Size Always Better? -- An Empirical Study on Model Editing with Llama-3
by: Yoon, Junsang, et al.
Published: (2024) -
Geometric Interpretation of Layer Normalization and a Comparative Analysis with RMSNorm
by: Gupta, Akshat, et al.
Published: (2024) -
Model Editing at Scale leads to Gradual and Catastrophic Forgetting
by: Gupta, Akshat, et al.
Published: (2024)