Does Self-Evaluation Enable Wireheading in Language Models?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Africa, David Demitri, Ting, Hans Ethan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness
von: Ivanov, Igor, et al.
Veröffentlicht: (2026)
von: Ivanov, Igor, et al.
Veröffentlicht: (2026)
Steering Awareness: Detecting Activation Steering from Within
von: Rivera, Joshua Fonseca, et al.
Veröffentlicht: (2025)
von: Rivera, Joshua Fonseca, et al.
Veröffentlicht: (2025)
Investigating ReLoRA: Effects on the Learning Dynamics of Small Language Models
von: Weiss, Yuval, et al.
Veröffentlicht: (2025)
von: Weiss, Yuval, et al.
Veröffentlicht: (2025)
Learning Dynamics of Meta-Learning in Small Model Pretraining
von: Africa, David Demitri, et al.
Veröffentlicht: (2025)
von: Africa, David Demitri, et al.
Veröffentlicht: (2025)
Consistency Training while Mitigating Obfuscation via Rate Matching
von: Imran, Sohaib, et al.
Veröffentlicht: (2026)
von: Imran, Sohaib, et al.
Veröffentlicht: (2026)
Pico: A Modular Framework for Hypothesis-Driven Small Language Model Research
von: Martinez, Richard Diehl, et al.
Veröffentlicht: (2025)
von: Martinez, Richard Diehl, et al.
Veröffentlicht: (2025)
Meta-Pretraining for Zero-Shot Cross-Lingual Named Entity Recognition in Low-Resource Philippine Languages
von: Africa, David Demitri, et al.
Veröffentlicht: (2025)
von: Africa, David Demitri, et al.
Veröffentlicht: (2025)
Learning Modular Exponentiation with Transformers
von: Africa, David Demitri, et al.
Veröffentlicht: (2025)
von: Africa, David Demitri, et al.
Veröffentlicht: (2025)
Batayan: A Filipino NLP benchmark for evaluating Large Language Models
von: Montalan, Jann Railey, et al.
Veröffentlicht: (2025)
von: Montalan, Jann Railey, et al.
Veröffentlicht: (2025)
No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes
von: Cencerrado, Iván Vicente Moreno, et al.
Veröffentlicht: (2025)
von: Cencerrado, Iván Vicente Moreno, et al.
Veröffentlicht: (2025)
Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-time
von: Tan, Daniel, et al.
Veröffentlicht: (2025)
von: Tan, Daniel, et al.
Veröffentlicht: (2025)
Identifying a Circuit for Verb Conjugation in GPT-2
von: Africa, David Demitri
Veröffentlicht: (2025)
von: Africa, David Demitri
Veröffentlicht: (2025)
Intelligence Analysis of Language Models
von: Galanti, Liane, et al.
Veröffentlicht: (2024)
von: Galanti, Liane, et al.
Veröffentlicht: (2024)
Language Models can Self-Improve at State-Value Estimation for Better Search
von: Mendes, Ethan, et al.
Veröffentlicht: (2025)
von: Mendes, Ethan, et al.
Veröffentlicht: (2025)
Anticipatory Evaluation of Language Models
von: Park, Jungsoo, et al.
Veröffentlicht: (2025)
von: Park, Jungsoo, et al.
Veröffentlicht: (2025)
Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate
von: Chern, Steffi, et al.
Veröffentlicht: (2024)
von: Chern, Steffi, et al.
Veröffentlicht: (2024)
Generalization or Memorization? Brittleness Testing for Chess-Trained Language Models
von: Tang, Ethan
Veröffentlicht: (2026)
von: Tang, Ethan
Veröffentlicht: (2026)
Does Thought Require Sensory Grounding? From Pure Thinkers to Large Language Models
von: Chalmers, David J.
Veröffentlicht: (2024)
von: Chalmers, David J.
Veröffentlicht: (2024)
The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity?
von: Hägele, Alexander, et al.
Veröffentlicht: (2026)
von: Hägele, Alexander, et al.
Veröffentlicht: (2026)
Self-Destructive Language Model
von: Wang, Yuhui, et al.
Veröffentlicht: (2025)
von: Wang, Yuhui, et al.
Veröffentlicht: (2025)
Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks
von: Li, Yuangang, et al.
Veröffentlicht: (2026)
von: Li, Yuangang, et al.
Veröffentlicht: (2026)
Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
von: Rocamonde, Juan, et al.
Veröffentlicht: (2023)
von: Rocamonde, Juan, et al.
Veröffentlicht: (2023)
Evaluating and Enhancing Large Language Models for Novelty Assessment in Scholarly Publications
von: Lin, Ethan, et al.
Veröffentlicht: (2024)
von: Lin, Ethan, et al.
Veröffentlicht: (2024)
Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment
von: Tice, Cameron, et al.
Veröffentlicht: (2026)
von: Tice, Cameron, et al.
Veröffentlicht: (2026)
How Many Parameters Does it Take to Change a Light Bulb? Evaluating Performance in Self-Play of Conversational Games as a Function of Model Characteristics
von: Bhavsar, Nidhir, et al.
Veröffentlicht: (2024)
von: Bhavsar, Nidhir, et al.
Veröffentlicht: (2024)
Does In-IDE Calibration of Large Language Models work at Scale?
von: Koohestani, Roham, et al.
Veröffentlicht: (2025)
von: Koohestani, Roham, et al.
Veröffentlicht: (2025)
Collapse of Self-trained Language Models
von: Herel, David, et al.
Veröffentlicht: (2024)
von: Herel, David, et al.
Veröffentlicht: (2024)
OCCULT: Evaluating Large Language Models for Offensive Cyber Operation Capabilities
von: Kouremetis, Michael, et al.
Veröffentlicht: (2025)
von: Kouremetis, Michael, et al.
Veröffentlicht: (2025)
Are Audio-Language Models Listening? Audio-Specialist Heads for Adaptive Audio Steering
von: Glazer, Neta, et al.
Veröffentlicht: (2026)
von: Glazer, Neta, et al.
Veröffentlicht: (2026)
AgenticEval: Toward Agentic and Self-Evolving Safety Evaluation of Large Language Models
von: Wang, Yixu, et al.
Veröffentlicht: (2025)
von: Wang, Yixu, et al.
Veröffentlicht: (2025)
What Does the AI Doctor Value? Auditing Pluralism in the Clinical Ethics of Language Models
von: Chandak, Payal, et al.
Veröffentlicht: (2026)
von: Chandak, Payal, et al.
Veröffentlicht: (2026)
Extensive Self-Contrast Enables Feedback-Free Language Model Alignment
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
Enterprise Large Language Model Evaluation Benchmark
von: Wang, Liya, et al.
Veröffentlicht: (2025)
von: Wang, Liya, et al.
Veröffentlicht: (2025)
CPSDBench: A Large Language Model Evaluation Benchmark and Baseline for Chinese Public Security Domain
von: Tong, Xin, et al.
Veröffentlicht: (2024)
von: Tong, Xin, et al.
Veröffentlicht: (2024)
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models
von: Pombal, José, et al.
Veröffentlicht: (2026)
von: Pombal, José, et al.
Veröffentlicht: (2026)
Efficient Self-Evaluation for Diffusion Language Models via Sequence Regeneration
von: Zhong, Linhao, et al.
Veröffentlicht: (2026)
von: Zhong, Linhao, et al.
Veröffentlicht: (2026)
Equifinality in Mixture of Experts: Routing Topology Does Not Determine Language Modeling Quality
von: Ternovtsii, Ivan, et al.
Veröffentlicht: (2026)
von: Ternovtsii, Ivan, et al.
Veröffentlicht: (2026)
A Survey on Self-Evolution of Large Language Models
von: Tao, Zhengwei, et al.
Veröffentlicht: (2024)
von: Tao, Zhengwei, et al.
Veröffentlicht: (2024)
Does Inference Scaling Improve Reasoning Faithfulness? A Multi-Model Analysis of Self-Consistency Tradeoffs
von: Mehta, Deep
Veröffentlicht: (2026)
von: Mehta, Deep
Veröffentlicht: (2026)
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness
von: Ivanov, Igor, et al.
Veröffentlicht: (2026) -
Steering Awareness: Detecting Activation Steering from Within
von: Rivera, Joshua Fonseca, et al.
Veröffentlicht: (2025) -
Investigating ReLoRA: Effects on the Learning Dynamics of Small Language Models
von: Weiss, Yuval, et al.
Veröffentlicht: (2025) -
Learning Dynamics of Meta-Learning in Small Model Pretraining
von: Africa, David Demitri, et al.
Veröffentlicht: (2025) -
Consistency Training while Mitigating Obfuscation via Rate Matching
von: Imran, Sohaib, et al.
Veröffentlicht: (2026)