LLMs Encode Their Failures: Predicting Success from Pre-Generation Activations
Fuente:
arXiv
Saved in:
| Main Authors: | Lugoloobi, William, Foster, Thomas, Bankes, William, Russell, Chris |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLMs Encode How Difficult Problems Are
by: Lugoloobi, William, et al.
Published: (2025)
by: Lugoloobi, William, et al.
Published: (2025)
Known By Their Actions: Fingerprinting LLM Browser Agents via UI Traces
by: Lugoloobi, William, et al.
Published: (2026)
by: Lugoloobi, William, et al.
Published: (2026)
Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
by: Karvonen, Adam, et al.
Published: (2025)
by: Karvonen, Adam, et al.
Published: (2025)
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
by: Mayne, Harry, et al.
Published: (2025)
by: Mayne, Harry, et al.
Published: (2025)
Replay Failures as Successes: Sample-Efficient Reinforcement Learning for Instruction Following
by: Zhang, Kongcheng, et al.
Published: (2025)
by: Zhang, Kongcheng, et al.
Published: (2025)
Do LLMs Encode Functional Importance of Reasoning Tokens?
by: Singh, Janvijay, et al.
Published: (2026)
by: Singh, Janvijay, et al.
Published: (2026)
A Surprising Failure? Multimodal LLMs and the NLVR Challenge
by: Wu, Anne, et al.
Published: (2024)
by: Wu, Anne, et al.
Published: (2024)
Extracting Unlearned Information from LLMs with Activation Steering
by: Seyitoğlu, Atakan, et al.
Published: (2024)
by: Seyitoğlu, Atakan, et al.
Published: (2024)
Memory Injections: Correcting Multi-Hop Reasoning Failures during Inference in Transformer-Based Language Models
by: Sakarvadia, Mansi, et al.
Published: (2023)
by: Sakarvadia, Mansi, et al.
Published: (2023)
Between Circuits and Chomsky: Pre-pretraining on Formal Languages Imparts Linguistic Biases
by: Hu, Michael Y., et al.
Published: (2025)
by: Hu, Michael Y., et al.
Published: (2025)
Self-Evaluating LLMs for Multi-Step Tasks: Stepwise Confidence Estimation for Failure Detection
by: Mavi, Vaibhav, et al.
Published: (2025)
by: Mavi, Vaibhav, et al.
Published: (2025)
AutoScale: Scale-Aware Data Mixing for Pre-Training LLMs
by: Kang, Feiyang, et al.
Published: (2024)
by: Kang, Feiyang, et al.
Published: (2024)
Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge
by: Shen, Xuan, et al.
Published: (2023)
by: Shen, Xuan, et al.
Published: (2023)
AdapterSwap: Continuous Training of LLMs with Data Removal and Access-Control Guarantees
by: Fleshman, William, et al.
Published: (2024)
by: Fleshman, William, et al.
Published: (2024)
CAST: Compositional Analysis via Spectral Tracking for Understanding Transformer Layer Functions
by: Fu, Zihao, et al.
Published: (2025)
by: Fu, Zihao, et al.
Published: (2025)
LLMem: Estimating GPU Memory Usage for Fine-Tuning Pre-Trained LLMs
by: Kim, Taeho, et al.
Published: (2024)
by: Kim, Taeho, et al.
Published: (2024)
From Associations to Activations: Comparing Behavioral and Hidden-State Semantic Geometry in LLMs
by: Schiekiera, Louis, et al.
Published: (2026)
by: Schiekiera, Louis, et al.
Published: (2026)
LoRA-Augmented Generation (LAG) for Knowledge-Intensive Language Tasks
by: Fleshman, William, et al.
Published: (2025)
by: Fleshman, William, et al.
Published: (2025)
Do Music Generation Models Encode Music Theory?
by: Wei, Megan, et al.
Published: (2024)
by: Wei, Megan, et al.
Published: (2024)
Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation
by: Manvi, Rohin, et al.
Published: (2024)
by: Manvi, Rohin, et al.
Published: (2024)
Scaling Laws for Predicting Downstream Performance in LLMs
by: Chen, Yangyi, et al.
Published: (2024)
by: Chen, Yangyi, et al.
Published: (2024)
RLPF: Reinforcement Learning from Prediction Feedback for User Summarization with LLMs
by: Wu, Jiaxing, et al.
Published: (2024)
by: Wu, Jiaxing, et al.
Published: (2024)
Translating Expert Intuition into Quantifiable Features: Encode Investigator Domain Knowledge via LLM for Enhanced Predictive Analytics
by: Jing, Phoebe, et al.
Published: (2024)
by: Jing, Phoebe, et al.
Published: (2024)
Group Representational Position Encoding
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
Prompt-Counterfactual Explanations for Generative AI System Behavior
by: Goethals, Sofie, et al.
Published: (2026)
by: Goethals, Sofie, et al.
Published: (2026)
Learning a Generative Meta-Model of LLM Activations
by: Luo, Grace, et al.
Published: (2026)
by: Luo, Grace, et al.
Published: (2026)
Observational Scaling Laws and the Predictability of Language Model Performance
by: Ruan, Yangjun, et al.
Published: (2024)
by: Ruan, Yangjun, et al.
Published: (2024)
On Subjective Uncertainty Quantification and Calibration in Natural Language Generation
by: Wang, Ziyu, et al.
Published: (2024)
by: Wang, Ziyu, et al.
Published: (2024)
RoLoRA: Fine-tuning Rotated Outlier-free LLMs for Effective Weight-Activation Quantization
by: Huang, Xijie, et al.
Published: (2024)
by: Huang, Xijie, et al.
Published: (2024)
LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs?
by: Ziomek, Juliusz, et al.
Published: (2026)
by: Ziomek, Juliusz, et al.
Published: (2026)
AdaExplore: Failure-Driven Adaptation and Diversity-Preserving Search for Efficient Kernel Generation
by: Du, Weihua, et al.
Published: (2026)
by: Du, Weihua, et al.
Published: (2026)
Asynchronous and Segmented Bidirectional Encoding for NMT
by: Yang, Jingpu, et al.
Published: (2024)
by: Yang, Jingpu, et al.
Published: (2024)
PoPreRo: A New Dataset for Popularity Prediction of Romanian Reddit Posts
by: Rogoz, Ana-Cristina, et al.
Published: (2024)
by: Rogoz, Ana-Cristina, et al.
Published: (2024)
Beyond Confidence: Rethinking Self-Assessments for Performance Prediction in LLMs
by: Bhattacharyya, Sree, et al.
Published: (2026)
by: Bhattacharyya, Sree, et al.
Published: (2026)
InversionView: A General-Purpose Method for Reading Information from Neural Activations
by: Huang, Xinting, et al.
Published: (2024)
by: Huang, Xinting, et al.
Published: (2024)
How Data Inter-connectivity Shapes LLMs Unlearning: A Structural Unlearning Perspective
by: Qiu, Xinchi, et al.
Published: (2024)
by: Qiu, Xinchi, et al.
Published: (2024)
Learning to Reason at the Frontier of Learnability
by: Foster, Thomas, et al.
Published: (2025)
by: Foster, Thomas, et al.
Published: (2025)
Learning from Failures in Multi-Attempt Reinforcement Learning
by: Chung, Stephen, et al.
Published: (2025)
by: Chung, Stephen, et al.
Published: (2025)
Sparse Autoencoder Decomposition of Clinical Sequence Model Representations: Feature Complexity, Task Specialisation, and Mortality Prediction
by: Sainsbury, Chris, et al.
Published: (2026)
by: Sainsbury, Chris, et al.
Published: (2026)
Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers
by: Barron, Joshua, et al.
Published: (2025)
by: Barron, Joshua, et al.
Published: (2025)
Similar Items
-
LLMs Encode How Difficult Problems Are
by: Lugoloobi, William, et al.
Published: (2025) -
Known By Their Actions: Fingerprinting LLM Browser Agents via UI Traces
by: Lugoloobi, William, et al.
Published: (2026) -
Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
by: Karvonen, Adam, et al.
Published: (2025) -
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
by: Mayne, Harry, et al.
Published: (2025) -
Replay Failures as Successes: Sample-Efficient Reinforcement Learning for Instruction Following
by: Zhang, Kongcheng, et al.
Published: (2025)