A Looming Replication Crisis in Evaluating Behavior in Language Models? Evidence and Solutions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Vaugrante, Laurène, Niepert, Mathias, Hagendorff, Thilo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Compromising Honesty and Harmlessness in Language Models via Deception Attacks
von: Vaugrante, Laurène, et al.
Veröffentlicht: (2025)
von: Vaugrante, Laurène, et al.
Veröffentlicht: (2025)
Emergently Misaligned Language Models Show Behavioral Self-Awareness That Shifts With Subsequent Realignment
von: Vaugrante, Laurène, et al.
Veröffentlicht: (2026)
von: Vaugrante, Laurène, et al.
Veröffentlicht: (2026)
Deception Abilities Emerged in Large Language Models
von: Hagendorff, Thilo
Veröffentlicht: (2023)
von: Hagendorff, Thilo
Veröffentlicht: (2023)
Large Reasoning Models Are Autonomous Jailbreak Agents
von: Hagendorff, Thilo, et al.
Veröffentlicht: (2025)
von: Hagendorff, Thilo, et al.
Veröffentlicht: (2025)
Mapping the Ethics of Generative AI: A Comprehensive Scoping Review
von: Hagendorff, Thilo
Veröffentlicht: (2024)
von: Hagendorff, Thilo
Veröffentlicht: (2024)
"Dark Triad" Model Organisms of Misalignment: Narrow Fine-Tuning Mirrors Human Antisocial Behavior
von: Lulla, Roshni, et al.
Veröffentlicht: (2026)
von: Lulla, Roshni, et al.
Veröffentlicht: (2026)
On the Inevitability of Left-Leaning Political Bias in Aligned Language Models
von: Hagendorff, Thilo
Veröffentlicht: (2025)
von: Hagendorff, Thilo
Veröffentlicht: (2025)
Evaluation Awareness in Language Models Has Limited Effect on Behaviour
von: Knecht, Amelie, et al.
Veröffentlicht: (2026)
von: Knecht, Amelie, et al.
Veröffentlicht: (2026)
Beyond Chains of Thought: Benchmarking Latent-Space Reasoning Abilities in Large Language Models
von: Hagendorff, Thilo, et al.
Veröffentlicht: (2025)
von: Hagendorff, Thilo, et al.
Veröffentlicht: (2025)
Fairness Hacking: The Malicious Practice of Shrouding Unfairness in Algorithms
von: Meding, Kristof, et al.
Veröffentlicht: (2023)
von: Meding, Kristof, et al.
Veröffentlicht: (2023)
Shadow-Loom: Causal Reasoning over Graphical World Models of Narratives
von: Wilmot, David
Veröffentlicht: (2026)
von: Wilmot, David
Veröffentlicht: (2026)
ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences
von: Nguyen, Bang, et al.
Veröffentlicht: (2026)
von: Nguyen, Bang, et al.
Veröffentlicht: (2026)
Machine Psychology
von: Hagendorff, Thilo, et al.
Veröffentlicht: (2023)
von: Hagendorff, Thilo, et al.
Veröffentlicht: (2023)
PRIDE -- Parameter-Efficient Reduction of Identity Discrimination for Equality in LLMs
von: Menke, Maluna, et al.
Veröffentlicht: (2025)
von: Menke, Maluna, et al.
Veröffentlicht: (2025)
Speciesism in AI: Evaluating Discrimination Against Animals in Large Language Models
von: Jotautaitė, Monika, et al.
Veröffentlicht: (2025)
von: Jotautaitė, Monika, et al.
Veröffentlicht: (2025)
Let the Results Speak: A Replication-First Paradigm for LLM Behavioral Benchmarking
von: Yuming, et al.
Veröffentlicht: (2026)
von: Yuming, et al.
Veröffentlicht: (2026)
Replicating Human Social Perception in Generative AI: Evaluating the Valence-Dominance Model
von: Gurkan, Necdet, et al.
Veröffentlicht: (2025)
von: Gurkan, Necdet, et al.
Veröffentlicht: (2025)
Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
Evaluating Large Language Models in Crisis Detection: A Real-World Benchmark from Psychological Support Hotlines
von: Deng, Guifeng, et al.
Veröffentlicht: (2025)
von: Deng, Guifeng, et al.
Veröffentlicht: (2025)
Style over Substance: Distilled Language Models Reason Via Stylistic Replication
von: Lippmann, Philip, et al.
Veröffentlicht: (2025)
von: Lippmann, Philip, et al.
Veröffentlicht: (2025)
PaperBench: Evaluating AI's Ability to Replicate AI Research
von: Starace, Giulio, et al.
Veröffentlicht: (2025)
von: Starace, Giulio, et al.
Veröffentlicht: (2025)
Behavioral Bias of Vision-Language Models: A Behavioral Finance View
von: Xiao, Yuhang, et al.
Veröffentlicht: (2024)
von: Xiao, Yuhang, et al.
Veröffentlicht: (2024)
SMART: Scalable Mesh-free Aerodynamic Simulations from Raw Geometries using a Transformer-based Surrogate Model
von: Hagnberger, Jan, et al.
Veröffentlicht: (2026)
von: Hagnberger, Jan, et al.
Veröffentlicht: (2026)
Mechanistic Behavior Editing of Language Models
von: Singh, Joykirat, et al.
Veröffentlicht: (2024)
von: Singh, Joykirat, et al.
Veröffentlicht: (2024)
Behavioral Fingerprinting of Large Language Models
von: Pei, Zehua, et al.
Veröffentlicht: (2025)
von: Pei, Zehua, et al.
Veröffentlicht: (2025)
CBT-Bench: Evaluating Large Language Models on Assisting Cognitive Behavior Therapy
von: Zhang, Mian, et al.
Veröffentlicht: (2024)
von: Zhang, Mian, et al.
Veröffentlicht: (2024)
Preference-Based Gradient Estimation for ML-Guided Approximate Combinatorial Optimization
von: Mielke, Arman, et al.
Veröffentlicht: (2025)
von: Mielke, Arman, et al.
Veröffentlicht: (2025)
Behavior-Aware Item Modeling via Dynamic Procedural Solution Representations for Knowledge Tracing
von: Seo, Jun, et al.
Veröffentlicht: (2026)
von: Seo, Jun, et al.
Veröffentlicht: (2026)
Dimension-Level Intent Fidelity Evaluation for Large Language Models: Evidence from Structured Prompt Ablation
von: Peng, GAng
Veröffentlicht: (2026)
von: Peng, GAng
Veröffentlicht: (2026)
Refusal Behavior in Large Language Models: A Nonlinear Perspective
von: Hildebrandt, Fabian, et al.
Veröffentlicht: (2025)
von: Hildebrandt, Fabian, et al.
Veröffentlicht: (2025)
Zero-Shot Classification of Crisis Tweets Using Instruction-Finetuned Large Language Models
von: McDaniel, Emma, et al.
Veröffentlicht: (2024)
von: McDaniel, Emma, et al.
Veröffentlicht: (2024)
Evaluating Language Models' Evaluations of Games
von: Collins, Katherine M., et al.
Veröffentlicht: (2025)
von: Collins, Katherine M., et al.
Veröffentlicht: (2025)
ChatGPT Alternative Solutions: Large Language Models Survey
von: Alipour, Hanieh, et al.
Veröffentlicht: (2024)
von: Alipour, Hanieh, et al.
Veröffentlicht: (2024)
The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior
von: Wang, Angelina, et al.
Veröffentlicht: (2025)
von: Wang, Angelina, et al.
Veröffentlicht: (2025)
A Dynamic Fusion Model for Consistent Crisis Response
von: Song, Xiaoying, et al.
Veröffentlicht: (2025)
von: Song, Xiaoying, et al.
Veröffentlicht: (2025)
A Lightweight Multi Aspect Controlled Text Generation Solution For Large Language Models
von: Zhang, Chenyang, et al.
Veröffentlicht: (2024)
von: Zhang, Chenyang, et al.
Veröffentlicht: (2024)
The System Hallucination Scale (SHS): A Minimal yet Effective Human-Centered Instrument for Evaluating Hallucination-Related Behavior in Large Language Models
von: Müller, Heimo, et al.
Veröffentlicht: (2026)
von: Müller, Heimo, et al.
Veröffentlicht: (2026)
OLMES: A Standard for Language Model Evaluations
von: Gu, Yuling, et al.
Veröffentlicht: (2024)
von: Gu, Yuling, et al.
Veröffentlicht: (2024)
A Survey on Evaluation of Large Language Models
von: Chang, Yupeng, et al.
Veröffentlicht: (2023)
von: Chang, Yupeng, et al.
Veröffentlicht: (2023)
Behavior Trees Enable Structured Programming of Language Model Agents
von: Kelley, Richard
Veröffentlicht: (2024)
von: Kelley, Richard
Veröffentlicht: (2024)
Ähnliche Einträge
-
Compromising Honesty and Harmlessness in Language Models via Deception Attacks
von: Vaugrante, Laurène, et al.
Veröffentlicht: (2025) -
Emergently Misaligned Language Models Show Behavioral Self-Awareness That Shifts With Subsequent Realignment
von: Vaugrante, Laurène, et al.
Veröffentlicht: (2026) -
Deception Abilities Emerged in Large Language Models
von: Hagendorff, Thilo
Veröffentlicht: (2023) -
Large Reasoning Models Are Autonomous Jailbreak Agents
von: Hagendorff, Thilo, et al.
Veröffentlicht: (2025) -
Mapping the Ethics of Generative AI: A Comprehensive Scoping Review
von: Hagendorff, Thilo
Veröffentlicht: (2024)