The Elicitation Game: Evaluating Capability Elicitation Techniques
Fuente:
arXiv
Saved in:
| Main Authors: | Hofstätter, Felix, van der Weij, Teun, Teoh, Jayden, Djoneva, Rada, Bartsch, Henning, Ward, Francis Rhys |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
by: van der Weij, Teun, et al.
Published: (2024)
by: van der Weij, Teun, et al.
Published: (2024)
Extending Activation Steering to Broad Skills and Multiple Behaviours
by: van der Weij, Teun, et al.
Published: (2024)
by: van der Weij, Teun, et al.
Published: (2024)
Eliciting Fine-Tuned Transformer Capabilities via Inference-Time Techniques
by: Sharma, Asankhaya
Published: (2025)
by: Sharma, Asankhaya
Published: (2025)
Reasoning Under Pressure: How do Training Incentives Influence Chain-of-Thought Monitorability?
by: MacDermott, Matt, et al.
Published: (2025)
by: MacDermott, Matt, et al.
Published: (2025)
Towards a Theory of AI Personhood
by: Ward, Francis Rhys
Published: (2025)
by: Ward, Francis Rhys
Published: (2025)
Eliciting Numerical Predictive Distributions of LLMs Without Autoregression
by: Piskorz, Julianna, et al.
Published: (2026)
by: Piskorz, Julianna, et al.
Published: (2026)
Equitable Evaluation via Elicitation
by: Du, Elbert, et al.
Published: (2026)
by: Du, Elbert, et al.
Published: (2026)
Causal Preference Elicitation
by: Bonilla, Edwin V., et al.
Published: (2026)
by: Bonilla, Edwin V., et al.
Published: (2026)
Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs
by: Naderi, Nariman, et al.
Published: (2025)
by: Naderi, Nariman, et al.
Published: (2025)
On Distinguishing Capability Elicitation from Capability Creation in Post-Training: A Free-Energy Perspective
by: Li, Yuhao, et al.
Published: (2026)
by: Li, Yuhao, et al.
Published: (2026)
Preference Elicitation for Offline Reinforcement Learning
by: Pace, Alizée, et al.
Published: (2024)
by: Pace, Alizée, et al.
Published: (2024)
Personalized Algorithmic Recourse with Preference Elicitation
by: De Toni, Giovanni, et al.
Published: (2022)
by: De Toni, Giovanni, et al.
Published: (2022)
Diversity of Thought Elicits Stronger Reasoning Capabilities in Multi-Agent Debate Frameworks
by: Hegazy, Mahmood
Published: (2024)
by: Hegazy, Mahmood
Published: (2024)
CTRL-ALT-DECEIT: Sabotage Evaluations for Automated AI R&D
by: Ward, Francis Rhys, et al.
Published: (2025)
by: Ward, Francis Rhys, et al.
Published: (2025)
ElicitationGPT: Text Elicitation Mechanisms via Language Models
by: Wu, Yifan, et al.
Published: (2024)
by: Wu, Yifan, et al.
Published: (2024)
Eliciting Harmful Capabilities by Fine-Tuning On Safeguarded Outputs
by: Kaunismaa, Jackson, et al.
Published: (2026)
by: Kaunismaa, Jackson, et al.
Published: (2026)
Uncertainty in Action: Confidence Elicitation in Embodied Agents
by: Yu, Tianjiao, et al.
Published: (2025)
by: Yu, Tianjiao, et al.
Published: (2025)
Incentivizing Truthful Language Models via Peer Elicitation Games
by: Chen, Baiting, et al.
Published: (2025)
by: Chen, Baiting, et al.
Published: (2025)
Lookahead Unmasking Elicits Accurate Decoding in Diffusion Language Models
by: Lee, Sanghyun, et al.
Published: (2025)
by: Lee, Sanghyun, et al.
Published: (2025)
Can LLMs Assist Expert Elicitation for Probabilistic Causal Modeling?
by: Shaposhnyk, Olha, et al.
Published: (2025)
by: Shaposhnyk, Olha, et al.
Published: (2025)
Embedding by Elicitation: Dynamic Representations for Bayesian Optimization of System Prompts
by: Lin, Zhiyuan Jerry, et al.
Published: (2026)
by: Lin, Zhiyuan Jerry, et al.
Published: (2026)
Three Concrete Challenges and Two Hopes for the Safety of Unsupervised Elicitation
by: Canavan, Callum, et al.
Published: (2026)
by: Canavan, Callum, et al.
Published: (2026)
Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models
by: Tice, Cameron, et al.
Published: (2024)
by: Tice, Cameron, et al.
Published: (2024)
Eliciting Language Model Behaviors with Investigator Agents
by: Li, Xiang Lisa, et al.
Published: (2025)
by: Li, Xiang Lisa, et al.
Published: (2025)
Online Rubrics Elicitation from Pairwise Comparisons
by: Rezaei, MohammadHossein, et al.
Published: (2025)
by: Rezaei, MohammadHossein, et al.
Published: (2025)
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
by: Kutasov, Jonathan, et al.
Published: (2025)
by: Kutasov, Jonathan, et al.
Published: (2025)
Synthetic Error Injection Fails to Elicit Self-Correction In Language Models
by: Wu, David X., et al.
Published: (2025)
by: Wu, David X., et al.
Published: (2025)
Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
by: Zhang, Shenao, et al.
Published: (2024)
by: Zhang, Shenao, et al.
Published: (2024)
Synthesize, Partition, then Adapt: Eliciting Diverse Samples from Foundation Models
by: Wen, Yeming, et al.
Published: (2024)
by: Wen, Yeming, et al.
Published: (2024)
Password-Activated Shutdown Protocols for Misaligned Frontier Agents
by: Williams, Kai, et al.
Published: (2025)
by: Williams, Kai, et al.
Published: (2025)
Verbal Process Supervision Elicits Better Coding Agents
by: Chen, Hao-Yuan, et al.
Published: (2025)
by: Chen, Hao-Yuan, et al.
Published: (2025)
Adaptive Elicitation of Latent Information Using Natural Language
by: Wang, Jimmy, et al.
Published: (2025)
by: Wang, Jimmy, et al.
Published: (2025)
Weak-to-Strong Elicitation via Mismatched Wrong Drafts
by: Deng, Wei
Published: (2026)
by: Deng, Wei
Published: (2026)
On Eliciting Syntax from Language Models via Hashing
by: Wang, Yiran, et al.
Published: (2024)
by: Wang, Yiran, et al.
Published: (2024)
Eliciting Latent Knowledge from Quirky Language Models
by: Mallen, Alex, et al.
Published: (2023)
by: Mallen, Alex, et al.
Published: (2023)
Reasoning Elicitation in Language Models via Counterfactual Feedback
by: Hüyük, Alihan, et al.
Published: (2024)
by: Hüyük, Alihan, et al.
Published: (2024)
Technical Report: Evaluating Goal Drift in Language Model Agents
by: Arike, Rauno, et al.
Published: (2025)
by: Arike, Rauno, et al.
Published: (2025)
A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning
by: Liu, Licheng, et al.
Published: (2025)
by: Liu, Licheng, et al.
Published: (2025)
Causality Elicitation from Large Language Models
by: Kameyama, Takashi, et al.
Published: (2026)
by: Kameyama, Takashi, et al.
Published: (2026)
Preference Elicitation for Multi-objective Combinatorial Optimization with Active Learning and Maximum Likelihood Estimation
by: Defresne, Marianne, et al.
Published: (2025)
by: Defresne, Marianne, et al.
Published: (2025)
Similar Items
-
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
by: van der Weij, Teun, et al.
Published: (2024) -
Extending Activation Steering to Broad Skills and Multiple Behaviours
by: van der Weij, Teun, et al.
Published: (2024) -
Eliciting Fine-Tuned Transformer Capabilities via Inference-Time Techniques
by: Sharma, Asankhaya
Published: (2025) -
Reasoning Under Pressure: How do Training Incentives Influence Chain-of-Thought Monitorability?
by: MacDermott, Matt, et al.
Published: (2025) -
Towards a Theory of AI Personhood
by: Ward, Francis Rhys
Published: (2025)