Saved in:
| Main Authors: | Vergara-Browne, Tomás, Patil, Darshan, Titov, Ivan, Reddy, Siva, Pimentel, Tiago, Mosbach, Marius |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2602.15829 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Forecasting Downstream Performance of LLMs With Proxy Metrics
by: Patel, Arkil, et al.
Published: (2026)
by: Patel, Arkil, et al.
Published: (2026)
Build the web for agents, not agents for the web
by: Lù, Xing Han, et al.
Published: (2025)
by: Lù, Xing Han, et al.
Published: (2025)
Value Drifts: Tracing Value Alignment During LLM Post-Training
by: Bhatia, Mehar, et al.
Published: (2025)
by: Bhatia, Mehar, et al.
Published: (2025)
Do Generalisation Results Generalise?
by: Boglioni, Matteo, et al.
Published: (2025)
by: Boglioni, Matteo, et al.
Published: (2025)
Revisiting the Superficial Alignment Hypothesis
by: Raghavendra, Mohit, et al.
Published: (2024)
by: Raghavendra, Mohit, et al.
Published: (2024)
Superficial Safety Alignment Hypothesis
by: Li, Jianwei, et al.
Published: (2024)
by: Li, Jianwei, et al.
Published: (2024)
Not All Data Are Unlearned Equally
by: Krishnan, Aravind, et al.
Published: (2025)
by: Krishnan, Aravind, et al.
Published: (2025)
Eigenpruning: an Interpretability-Inspired PEFT Method
by: Vergara-Browne, Tomás, et al.
Published: (2024)
by: Vergara-Browne, Tomás, et al.
Published: (2024)
The Illusion of Superposition? A Principled Analysis of Latent Thinking in Language Models
by: Rizvi-Martel, Michael, et al.
Published: (2026)
by: Rizvi-Martel, Michael, et al.
Published: (2026)
From Insights to Actions: The Impact of Interpretability and Analysis Research on NLP
by: Mosbach, Marius, et al.
Published: (2024)
by: Mosbach, Marius, et al.
Published: (2024)
What's New in My Data? Novelty Exploration via Contrastive Generation
by: Isonuma, Masaru, et al.
Published: (2024)
by: Isonuma, Masaru, et al.
Published: (2024)
The Cancellation Hypothesis in Critic-Free RL: From Outcome Rewards to Token Credits
by: Cheng, Tianhao, et al.
Published: (2026)
by: Cheng, Tianhao, et al.
Published: (2026)
Structured Distillation of Web Agent Capabilities Enables Generalization
by: Lù, Xing Han, et al.
Published: (2026)
by: Lù, Xing Han, et al.
Published: (2026)
Test-Time Alignment via Hypothesis Reweighting
by: Lee, Yoonho, et al.
Published: (2024)
by: Lee, Yoonho, et al.
Published: (2024)
Prompt-responsive Object Retrieval with Memory-augmented Student-Teacher Learning
by: Mosbach, Malte, et al.
Published: (2025)
by: Mosbach, Malte, et al.
Published: (2025)
Grasp Anything: Combining Teacher-Augmented Policy Gradient Learning with Instance Segmentation to Grasp Arbitrary Objects
by: Mosbach, Malte, et al.
Published: (2024)
by: Mosbach, Malte, et al.
Published: (2024)
Intelligent Switching for Reset-Free RL
by: Patil, Darshan, et al.
Published: (2024)
by: Patil, Darshan, et al.
Published: (2024)
Knowledge Offloading: Decomposing LLMs into Sparse Backbones and Memory Modules
by: Galliamov, Karim, et al.
Published: (2026)
by: Galliamov, Karim, et al.
Published: (2026)
Enhancing RLHF with Human Gaze Modeling
by: Galliamov, Karim, et al.
Published: (2025)
by: Galliamov, Karim, et al.
Published: (2025)
CoPeP: Benchmarking Continual Pretraining for Protein Language Models
by: Patil, Darshan, et al.
Published: (2026)
by: Patil, Darshan, et al.
Published: (2026)
Robust Reward Alignment via Hypothesis Space Batch Cutting
by: Xie, Zhixian, et al.
Published: (2025)
by: Xie, Zhixian, et al.
Published: (2025)
BRIDGE: Predicting Human Task Completion Time From Model Performance
by: Liu, Fengyuan, et al.
Published: (2026)
by: Liu, Fengyuan, et al.
Published: (2026)
Faithfulness Measurable Masked Language Models
by: Madsen, Andreas, et al.
Published: (2023)
by: Madsen, Andreas, et al.
Published: (2023)
Autoencoding Conditional Neural Processes for Representation Learning
by: Prokhorov, Victor, et al.
Published: (2023)
by: Prokhorov, Victor, et al.
Published: (2023)
Mitigating Copy Bias in In-Context Learning through Neuron Pruning
by: Ali, Ameen, et al.
Published: (2024)
by: Ali, Ameen, et al.
Published: (2024)
Operationalising Rawlsian Ethics for Fairness in Norm-Learning Agents
by: Woodgate, Jessica, et al.
Published: (2024)
by: Woodgate, Jessica, et al.
Published: (2024)
Non-parametric Hypothesis Tests for Distributional Group Symmetry
by: Chiu, Kenny, et al.
Published: (2023)
by: Chiu, Kenny, et al.
Published: (2023)
Convergence and Divergence of Language Models under Different Random Seeds
by: Fehlauer, Finlay, et al.
Published: (2025)
by: Fehlauer, Finlay, et al.
Published: (2025)
Understanding the Influence of Synthetic Data for Text Embedders
by: Springer, Jacob Mitchell, et al.
Published: (2025)
by: Springer, Jacob Mitchell, et al.
Published: (2025)
Safety Alignment Can Be Not Superficial With Explicit Safety Signals
by: Li, Jianwei, et al.
Published: (2025)
by: Li, Jianwei, et al.
Published: (2025)
Some Theoretical Results on Layerwise Effective Dimension Oscillations in Finite Width ReLU Networks
by: Makwana, Darshan
Published: (2025)
by: Makwana, Darshan
Published: (2025)
Are self-explanations from Large Language Models faithful?
by: Madsen, Andreas, et al.
Published: (2024)
by: Madsen, Andreas, et al.
Published: (2024)
Shared Doubt: Zero-shot Cross-Lingual Confidence Estimation for Language Models
by: Kyriakou, Athina, et al.
Published: (2026)
by: Kyriakou, Athina, et al.
Published: (2026)
Joint Localization and Activation Editing for Low-Resource Fine-Tuning
by: Lai, Wen, et al.
Published: (2025)
by: Lai, Wen, et al.
Published: (2025)
What explains the success of cross-modal fine-tuning with ORCA?
by: García-de-Herreros, Paloma, et al.
Published: (2024)
by: García-de-Herreros, Paloma, et al.
Published: (2024)
Weasel: Out-of-Domain Generalization for Web Agents via Importance-Diversity Data Selection
by: Zadeh, Fatemeh Pesaran, et al.
Published: (2026)
by: Zadeh, Fatemeh Pesaran, et al.
Published: (2026)
Tokenisation via Convex Relaxations
by: Tempus, Jan, et al.
Published: (2026)
by: Tempus, Jan, et al.
Published: (2026)
The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?
by: Sutter, Denis, et al.
Published: (2025)
by: Sutter, Denis, et al.
Published: (2025)
Cache & Distil: Optimising API Calls to Large Language Models
by: Ramírez, Guillem, et al.
Published: (2023)
by: Ramírez, Guillem, et al.
Published: (2023)
Detecting and Pruning Prominent but Detrimental Neurons in Large Language Models
by: Ali, Ameen, et al.
Published: (2025)
by: Ali, Ameen, et al.
Published: (2025)
Similar Items
-
Forecasting Downstream Performance of LLMs With Proxy Metrics
by: Patel, Arkil, et al.
Published: (2026) -
Build the web for agents, not agents for the web
by: Lù, Xing Han, et al.
Published: (2025) -
Value Drifts: Tracing Value Alignment During LLM Post-Training
by: Bhatia, Mehar, et al.
Published: (2025) -
Do Generalisation Results Generalise?
by: Boglioni, Matteo, et al.
Published: (2025) -
Revisiting the Superficial Alignment Hypothesis
by: Raghavendra, Mohit, et al.
Published: (2024)