LLMs Encode Their Failures: Predicting Success from Pre-Generation Activations
Fuente:
arXiv
Guardado en:
| Autores principales: | Lugoloobi, William, Foster, Thomas, Bankes, William, Russell, Chris |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LLMs Encode How Difficult Problems Are
por: Lugoloobi, William, et al.
Publicado: (2025)
por: Lugoloobi, William, et al.
Publicado: (2025)
Known By Their Actions: Fingerprinting LLM Browser Agents via UI Traces
por: Lugoloobi, William, et al.
Publicado: (2026)
por: Lugoloobi, William, et al.
Publicado: (2026)
Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
por: Karvonen, Adam, et al.
Publicado: (2025)
por: Karvonen, Adam, et al.
Publicado: (2025)
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
por: Mayne, Harry, et al.
Publicado: (2025)
por: Mayne, Harry, et al.
Publicado: (2025)
Replay Failures as Successes: Sample-Efficient Reinforcement Learning for Instruction Following
por: Zhang, Kongcheng, et al.
Publicado: (2025)
por: Zhang, Kongcheng, et al.
Publicado: (2025)
Do LLMs Encode Functional Importance of Reasoning Tokens?
por: Singh, Janvijay, et al.
Publicado: (2026)
por: Singh, Janvijay, et al.
Publicado: (2026)
A Surprising Failure? Multimodal LLMs and the NLVR Challenge
por: Wu, Anne, et al.
Publicado: (2024)
por: Wu, Anne, et al.
Publicado: (2024)
Extracting Unlearned Information from LLMs with Activation Steering
por: Seyitoğlu, Atakan, et al.
Publicado: (2024)
por: Seyitoğlu, Atakan, et al.
Publicado: (2024)
Memory Injections: Correcting Multi-Hop Reasoning Failures during Inference in Transformer-Based Language Models
por: Sakarvadia, Mansi, et al.
Publicado: (2023)
por: Sakarvadia, Mansi, et al.
Publicado: (2023)
Between Circuits and Chomsky: Pre-pretraining on Formal Languages Imparts Linguistic Biases
por: Hu, Michael Y., et al.
Publicado: (2025)
por: Hu, Michael Y., et al.
Publicado: (2025)
Self-Evaluating LLMs for Multi-Step Tasks: Stepwise Confidence Estimation for Failure Detection
por: Mavi, Vaibhav, et al.
Publicado: (2025)
por: Mavi, Vaibhav, et al.
Publicado: (2025)
AutoScale: Scale-Aware Data Mixing for Pre-Training LLMs
por: Kang, Feiyang, et al.
Publicado: (2024)
por: Kang, Feiyang, et al.
Publicado: (2024)
Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge
por: Shen, Xuan, et al.
Publicado: (2023)
por: Shen, Xuan, et al.
Publicado: (2023)
AdapterSwap: Continuous Training of LLMs with Data Removal and Access-Control Guarantees
por: Fleshman, William, et al.
Publicado: (2024)
por: Fleshman, William, et al.
Publicado: (2024)
CAST: Compositional Analysis via Spectral Tracking for Understanding Transformer Layer Functions
por: Fu, Zihao, et al.
Publicado: (2025)
por: Fu, Zihao, et al.
Publicado: (2025)
LLMem: Estimating GPU Memory Usage for Fine-Tuning Pre-Trained LLMs
por: Kim, Taeho, et al.
Publicado: (2024)
por: Kim, Taeho, et al.
Publicado: (2024)
From Associations to Activations: Comparing Behavioral and Hidden-State Semantic Geometry in LLMs
por: Schiekiera, Louis, et al.
Publicado: (2026)
por: Schiekiera, Louis, et al.
Publicado: (2026)
LoRA-Augmented Generation (LAG) for Knowledge-Intensive Language Tasks
por: Fleshman, William, et al.
Publicado: (2025)
por: Fleshman, William, et al.
Publicado: (2025)
Do Music Generation Models Encode Music Theory?
por: Wei, Megan, et al.
Publicado: (2024)
por: Wei, Megan, et al.
Publicado: (2024)
Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation
por: Manvi, Rohin, et al.
Publicado: (2024)
por: Manvi, Rohin, et al.
Publicado: (2024)
Scaling Laws for Predicting Downstream Performance in LLMs
por: Chen, Yangyi, et al.
Publicado: (2024)
por: Chen, Yangyi, et al.
Publicado: (2024)
RLPF: Reinforcement Learning from Prediction Feedback for User Summarization with LLMs
por: Wu, Jiaxing, et al.
Publicado: (2024)
por: Wu, Jiaxing, et al.
Publicado: (2024)
Translating Expert Intuition into Quantifiable Features: Encode Investigator Domain Knowledge via LLM for Enhanced Predictive Analytics
por: Jing, Phoebe, et al.
Publicado: (2024)
por: Jing, Phoebe, et al.
Publicado: (2024)
Group Representational Position Encoding
por: Zhang, Yifan, et al.
Publicado: (2025)
por: Zhang, Yifan, et al.
Publicado: (2025)
Prompt-Counterfactual Explanations for Generative AI System Behavior
por: Goethals, Sofie, et al.
Publicado: (2026)
por: Goethals, Sofie, et al.
Publicado: (2026)
Learning a Generative Meta-Model of LLM Activations
por: Luo, Grace, et al.
Publicado: (2026)
por: Luo, Grace, et al.
Publicado: (2026)
Observational Scaling Laws and the Predictability of Language Model Performance
por: Ruan, Yangjun, et al.
Publicado: (2024)
por: Ruan, Yangjun, et al.
Publicado: (2024)
On Subjective Uncertainty Quantification and Calibration in Natural Language Generation
por: Wang, Ziyu, et al.
Publicado: (2024)
por: Wang, Ziyu, et al.
Publicado: (2024)
RoLoRA: Fine-tuning Rotated Outlier-free LLMs for Effective Weight-Activation Quantization
por: Huang, Xijie, et al.
Publicado: (2024)
por: Huang, Xijie, et al.
Publicado: (2024)
LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs?
por: Ziomek, Juliusz, et al.
Publicado: (2026)
por: Ziomek, Juliusz, et al.
Publicado: (2026)
AdaExplore: Failure-Driven Adaptation and Diversity-Preserving Search for Efficient Kernel Generation
por: Du, Weihua, et al.
Publicado: (2026)
por: Du, Weihua, et al.
Publicado: (2026)
Asynchronous and Segmented Bidirectional Encoding for NMT
por: Yang, Jingpu, et al.
Publicado: (2024)
por: Yang, Jingpu, et al.
Publicado: (2024)
PoPreRo: A New Dataset for Popularity Prediction of Romanian Reddit Posts
por: Rogoz, Ana-Cristina, et al.
Publicado: (2024)
por: Rogoz, Ana-Cristina, et al.
Publicado: (2024)
Beyond Confidence: Rethinking Self-Assessments for Performance Prediction in LLMs
por: Bhattacharyya, Sree, et al.
Publicado: (2026)
por: Bhattacharyya, Sree, et al.
Publicado: (2026)
InversionView: A General-Purpose Method for Reading Information from Neural Activations
por: Huang, Xinting, et al.
Publicado: (2024)
por: Huang, Xinting, et al.
Publicado: (2024)
How Data Inter-connectivity Shapes LLMs Unlearning: A Structural Unlearning Perspective
por: Qiu, Xinchi, et al.
Publicado: (2024)
por: Qiu, Xinchi, et al.
Publicado: (2024)
Learning to Reason at the Frontier of Learnability
por: Foster, Thomas, et al.
Publicado: (2025)
por: Foster, Thomas, et al.
Publicado: (2025)
Learning from Failures in Multi-Attempt Reinforcement Learning
por: Chung, Stephen, et al.
Publicado: (2025)
por: Chung, Stephen, et al.
Publicado: (2025)
Sparse Autoencoder Decomposition of Clinical Sequence Model Representations: Feature Complexity, Task Specialisation, and Mortality Prediction
por: Sainsbury, Chris, et al.
Publicado: (2026)
por: Sainsbury, Chris, et al.
Publicado: (2026)
Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers
por: Barron, Joshua, et al.
Publicado: (2025)
por: Barron, Joshua, et al.
Publicado: (2025)
Ejemplares similares
-
LLMs Encode How Difficult Problems Are
por: Lugoloobi, William, et al.
Publicado: (2025) -
Known By Their Actions: Fingerprinting LLM Browser Agents via UI Traces
por: Lugoloobi, William, et al.
Publicado: (2026) -
Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
por: Karvonen, Adam, et al.
Publicado: (2025) -
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
por: Mayne, Harry, et al.
Publicado: (2025) -
Replay Failures as Successes: Sample-Efficient Reinforcement Learning for Instruction Following
por: Zhang, Kongcheng, et al.
Publicado: (2025)