Evaluation Drift in LLM Personality Induction: Are We Moving the Goalpost?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rajput, Prateek, Song, Yewei, Olatunji, Iyiola E., Klein, Jacques, Bissyandé, Tegawendé F. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Measuring LLM Code Generation Stability via Structural Entropy
von: Song, Yewei, et al.
Veröffentlicht: (2025)
von: Song, Yewei, et al.
Veröffentlicht: (2025)
Dynamic Stability of LLM-Generated Code
von: Rajput, Prateek, et al.
Veröffentlicht: (2025)
von: Rajput, Prateek, et al.
Veröffentlicht: (2025)
Correctness isnt Efficiency: Runtime Memory Divergence in LLM-Generated Code
von: Rajput, Prateek, et al.
Veröffentlicht: (2026)
von: Rajput, Prateek, et al.
Veröffentlicht: (2026)
Revisiting Code Similarity Evaluation with Abstract Syntax Tree Edit Distance
von: Song, Yewei, et al.
Veröffentlicht: (2024)
von: Song, Yewei, et al.
Veröffentlicht: (2024)
How Secure is Secure Code Generation? Adversarial Prompts Put LLM Defenses to the Test
von: Tessa, Melissa, et al.
Veröffentlicht: (2026)
von: Tessa, Melissa, et al.
Veröffentlicht: (2026)
Predictable Confabulations: Factual Recall by LLMs Scales with Model Size and Topic Frequency
von: Smith, Matthew L., et al.
Veröffentlicht: (2026)
von: Smith, Matthew L., et al.
Veröffentlicht: (2026)
Empirical Evaluation of PDF Parsing and Chunking for Financial Question Answering with RAG
von: Bachyr, Omar El, et al.
Veröffentlicht: (2026)
von: Bachyr, Omar El, et al.
Veröffentlicht: (2026)
CallNavi, A Challenge and Empirical Study on LLM Function Calling and Routing
von: Song, Yewei, et al.
Veröffentlicht: (2025)
von: Song, Yewei, et al.
Veröffentlicht: (2025)
Why Low-Resource NLP Needs More Than Cross-Lingual Transfer: Lessons Learned from Luxembourgish
von: Philippy, Fred, et al.
Veröffentlicht: (2026)
von: Philippy, Fred, et al.
Veröffentlicht: (2026)
LuxEmbedder: A Cross-Lingual Approach to Enhanced Luxembourgish Sentence Embeddings
von: Philippy, Fred, et al.
Veröffentlicht: (2024)
von: Philippy, Fred, et al.
Veröffentlicht: (2024)
Reinforcement Learning-Guided Chain-of-Draft for Token-Efficient Code Generation
von: Tang, Xunzhu, et al.
Veröffentlicht: (2025)
von: Tang, Xunzhu, et al.
Veröffentlicht: (2025)
Is Small Language Model the Silver Bullet to Low-Resource Languages Machine Translation?
von: Song, Yewei, et al.
Veröffentlicht: (2025)
von: Song, Yewei, et al.
Veröffentlicht: (2025)
Memorization or Interpolation ? Detecting LLM Memorization through Input Perturbation Analysis
von: Djiré, Albérick Euraste, et al.
Veröffentlicht: (2025)
von: Djiré, Albérick Euraste, et al.
Veröffentlicht: (2025)
Beyond Real Faces: Synthetic Datasets Can Achieve Reliable Recognition Performance without Privacy Compromise
von: Borsukiewicz, Paweł, et al.
Veröffentlicht: (2025)
von: Borsukiewicz, Paweł, et al.
Veröffentlicht: (2025)
Do Large Language Models Grasp The Grammar? Evidence from Grammar-Book-Guided Probing in Luxembourgish
von: Li, Lujun, et al.
Veröffentlicht: (2025)
von: Li, Lujun, et al.
Veröffentlicht: (2025)
LuxInstruct: A Cross-Lingual Instruction Tuning Dataset For Luxembourgish
von: Philippy, Fred, et al.
Veröffentlicht: (2025)
von: Philippy, Fred, et al.
Veröffentlicht: (2025)
Enhancing Small Language Models for Cross-Lingual Generalized Zero-Shot Classification with Soft Prompt Tuning
von: Philippy, Fred, et al.
Veröffentlicht: (2025)
von: Philippy, Fred, et al.
Veröffentlicht: (2025)
LaFiCMIL: Rethinking Large File Classification from the Perspective of Correlated Multiple Instance Learning
von: Sun, Tiezhu, et al.
Veröffentlicht: (2023)
von: Sun, Tiezhu, et al.
Veröffentlicht: (2023)
Soft Prompt Tuning for Cross-Lingual Transfer: When Less is More
von: Philippy, Fred, et al.
Veröffentlicht: (2024)
von: Philippy, Fred, et al.
Veröffentlicht: (2024)
Explainable AI for Analyzing Person-Specific Patterns in Facial Recognition Tasks
von: Borsukiewicz, Paweł Jakub, et al.
Veröffentlicht: (2025)
von: Borsukiewicz, Paweł Jakub, et al.
Veröffentlicht: (2025)
Boosting Open-Source LLMs for Program Repair via Reasoning Transfer and LLM-Guided Reinforcement Learning
von: Tang, Xunzhu, et al.
Veröffentlicht: (2025)
von: Tang, Xunzhu, et al.
Veröffentlicht: (2025)
Using Contrastive Learning to Improve Two-Way Reasoning in Large Language Models: The Obfuscation Task as a Case Study
von: Nikiema, Serge Lionel, et al.
Veröffentlicht: (2025)
von: Nikiema, Serge Lionel, et al.
Veröffentlicht: (2025)
Characterizing Build Compromises Through Vulnerability Disclosure Analysis
von: Diao, Maimouna Tamah, et al.
Veröffentlicht: (2025)
von: Diao, Maimouna Tamah, et al.
Veröffentlicht: (2025)
AndroLibZoo: A Reliable Dataset of Libraries Based on Software Dependency Analysis
von: Samhi, Jordan, et al.
Veröffentlicht: (2023)
von: Samhi, Jordan, et al.
Veröffentlicht: (2023)
Evaluating Large Language Models in detecting Secrets in Android Apps
von: Alecci, Marco, et al.
Veröffentlicht: (2025)
von: Alecci, Marco, et al.
Veröffentlicht: (2025)
How Small Transformation Expose the Weakness of Semantic Similarity Measures
von: Nikiema, Serge Lionel, et al.
Veröffentlicht: (2025)
von: Nikiema, Serge Lionel, et al.
Veröffentlicht: (2025)
Adversarial Camouflage
von: Borsukiewicz, Paweł, et al.
Veröffentlicht: (2026)
von: Borsukiewicz, Paweł, et al.
Veröffentlicht: (2026)
MalLoc: Toward Fine-grained Android Malicious Payload Localization via LLMs
von: Sun, Tiezhu, et al.
Veröffentlicht: (2025)
von: Sun, Tiezhu, et al.
Veröffentlicht: (2025)
(In)Security of Mobile Apps in Developing Countries: A Systematic Literature Review
von: Diallo, Alioune, et al.
Veröffentlicht: (2024)
von: Diallo, Alioune, et al.
Veröffentlicht: (2024)
From Rookie to Expert: Manipulating LLMs for Automated Vulnerability Exploitation in Enterprise Software
von: Diouf, Moustapha Awwalou, et al.
Veröffentlicht: (2025)
von: Diouf, Moustapha Awwalou, et al.
Veröffentlicht: (2025)
The Necessity of Setting Temperature in LLM-as-a-Judge
von: Li, Lujun, et al.
Veröffentlicht: (2026)
von: Li, Lujun, et al.
Veröffentlicht: (2026)
Software Security in Software-Defined Networking: A Systematic Literature Review
von: Diouf, Moustapha Awwalou, et al.
Veröffentlicht: (2025)
von: Diouf, Moustapha Awwalou, et al.
Veröffentlicht: (2025)
CodeAgent: Autonomous Communicative Agents for Code Review
von: Tang, Xunzhu, et al.
Veröffentlicht: (2024)
von: Tang, Xunzhu, et al.
Veröffentlicht: (2024)
Enhancing Text-to-SQL Translation for Financial System Design
von: Song, Yewei, et al.
Veröffentlicht: (2023)
von: Song, Yewei, et al.
Veröffentlicht: (2023)
Where Are We? Evaluating LLM Performance on African Languages
von: Adebara, Ife, et al.
Veröffentlicht: (2025)
von: Adebara, Ife, et al.
Veröffentlicht: (2025)
Evaluating Performance Drift from Model Switching in Multi-Turn LLM Systems
von: Khraishi, Raad, et al.
Veröffentlicht: (2026)
von: Khraishi, Raad, et al.
Veröffentlicht: (2026)
Neuron-based Personality Trait Induction in Large Language Models
von: Deng, Jia, et al.
Veröffentlicht: (2024)
von: Deng, Jia, et al.
Veröffentlicht: (2024)
How Value Induction Reshapes LLM Behaviour
von: Arora, Arnav, et al.
Veröffentlicht: (2026)
von: Arora, Arnav, et al.
Veröffentlicht: (2026)
The Struggles of LLMs in Cross-lingual Code Clone Detection
von: Moumoula, Micheline Bénédicte, et al.
Veröffentlicht: (2024)
von: Moumoula, Micheline Bénédicte, et al.
Veröffentlicht: (2024)
Path Drift in Large Reasoning Models:How First-Person Commitments Override Safety
von: Huang, Yuyi, et al.
Veröffentlicht: (2025)
von: Huang, Yuyi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Measuring LLM Code Generation Stability via Structural Entropy
von: Song, Yewei, et al.
Veröffentlicht: (2025) -
Dynamic Stability of LLM-Generated Code
von: Rajput, Prateek, et al.
Veröffentlicht: (2025) -
Correctness isnt Efficiency: Runtime Memory Divergence in LLM-Generated Code
von: Rajput, Prateek, et al.
Veröffentlicht: (2026) -
Revisiting Code Similarity Evaluation with Abstract Syntax Tree Edit Distance
von: Song, Yewei, et al.
Veröffentlicht: (2024) -
How Secure is Secure Code Generation? Adversarial Prompts Put LLM Defenses to the Test
von: Tessa, Melissa, et al.
Veröffentlicht: (2026)