Operationalising the Superficial Alignment Hypothesis via Task Complexity
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Vergara-Browne, Tomás, Patil, Darshan, Titov, Ivan, Reddy, Siva, Pimentel, Tiago, Mosbach, Marius |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Forecasting Downstream Performance of LLMs With Proxy Metrics
von: Patel, Arkil, et al.
Veröffentlicht: (2026)
von: Patel, Arkil, et al.
Veröffentlicht: (2026)
Build the web for agents, not agents for the web
von: Lù, Xing Han, et al.
Veröffentlicht: (2025)
von: Lù, Xing Han, et al.
Veröffentlicht: (2025)
Value Drifts: Tracing Value Alignment During LLM Post-Training
von: Bhatia, Mehar, et al.
Veröffentlicht: (2025)
von: Bhatia, Mehar, et al.
Veröffentlicht: (2025)
Revisiting the Superficial Alignment Hypothesis
von: Raghavendra, Mohit, et al.
Veröffentlicht: (2024)
von: Raghavendra, Mohit, et al.
Veröffentlicht: (2024)
Do Generalisation Results Generalise?
von: Boglioni, Matteo, et al.
Veröffentlicht: (2025)
von: Boglioni, Matteo, et al.
Veröffentlicht: (2025)
Superficial Safety Alignment Hypothesis
von: Li, Jianwei, et al.
Veröffentlicht: (2024)
von: Li, Jianwei, et al.
Veröffentlicht: (2024)
Not All Data Are Unlearned Equally
von: Krishnan, Aravind, et al.
Veröffentlicht: (2025)
von: Krishnan, Aravind, et al.
Veröffentlicht: (2025)
The Illusion of Superposition? A Principled Analysis of Latent Thinking in Language Models
von: Rizvi-Martel, Michael, et al.
Veröffentlicht: (2026)
von: Rizvi-Martel, Michael, et al.
Veröffentlicht: (2026)
Eigenpruning: an Interpretability-Inspired PEFT Method
von: Vergara-Browne, Tomás, et al.
Veröffentlicht: (2024)
von: Vergara-Browne, Tomás, et al.
Veröffentlicht: (2024)
What's New in My Data? Novelty Exploration via Contrastive Generation
von: Isonuma, Masaru, et al.
Veröffentlicht: (2024)
von: Isonuma, Masaru, et al.
Veröffentlicht: (2024)
Test-Time Alignment via Hypothesis Reweighting
von: Lee, Yoonho, et al.
Veröffentlicht: (2024)
von: Lee, Yoonho, et al.
Veröffentlicht: (2024)
The Cancellation Hypothesis in Critic-Free RL: From Outcome Rewards to Token Credits
von: Cheng, Tianhao, et al.
Veröffentlicht: (2026)
von: Cheng, Tianhao, et al.
Veröffentlicht: (2026)
From Insights to Actions: The Impact of Interpretability and Analysis Research on NLP
von: Mosbach, Marius, et al.
Veröffentlicht: (2024)
von: Mosbach, Marius, et al.
Veröffentlicht: (2024)
Structured Distillation of Web Agent Capabilities Enables Generalization
von: Lù, Xing Han, et al.
Veröffentlicht: (2026)
von: Lù, Xing Han, et al.
Veröffentlicht: (2026)
Robust Reward Alignment via Hypothesis Space Batch Cutting
von: Xie, Zhixian, et al.
Veröffentlicht: (2025)
von: Xie, Zhixian, et al.
Veröffentlicht: (2025)
Prompt-responsive Object Retrieval with Memory-augmented Student-Teacher Learning
von: Mosbach, Malte, et al.
Veröffentlicht: (2025)
von: Mosbach, Malte, et al.
Veröffentlicht: (2025)
Grasp Anything: Combining Teacher-Augmented Policy Gradient Learning with Instance Segmentation to Grasp Arbitrary Objects
von: Mosbach, Malte, et al.
Veröffentlicht: (2024)
von: Mosbach, Malte, et al.
Veröffentlicht: (2024)
Knowledge Offloading: Decomposing LLMs into Sparse Backbones and Memory Modules
von: Galliamov, Karim, et al.
Veröffentlicht: (2026)
von: Galliamov, Karim, et al.
Veröffentlicht: (2026)
Enhancing RLHF with Human Gaze Modeling
von: Galliamov, Karim, et al.
Veröffentlicht: (2025)
von: Galliamov, Karim, et al.
Veröffentlicht: (2025)
Intelligent Switching for Reset-Free RL
von: Patil, Darshan, et al.
Veröffentlicht: (2024)
von: Patil, Darshan, et al.
Veröffentlicht: (2024)
CoPeP: Benchmarking Continual Pretraining for Protein Language Models
von: Patil, Darshan, et al.
Veröffentlicht: (2026)
von: Patil, Darshan, et al.
Veröffentlicht: (2026)
Operationalising Rawlsian Ethics for Fairness in Norm-Learning Agents
von: Woodgate, Jessica, et al.
Veröffentlicht: (2024)
von: Woodgate, Jessica, et al.
Veröffentlicht: (2024)
Non-parametric Hypothesis Tests for Distributional Group Symmetry
von: Chiu, Kenny, et al.
Veröffentlicht: (2023)
von: Chiu, Kenny, et al.
Veröffentlicht: (2023)
Faithfulness Measurable Masked Language Models
von: Madsen, Andreas, et al.
Veröffentlicht: (2023)
von: Madsen, Andreas, et al.
Veröffentlicht: (2023)
Some Theoretical Results on Layerwise Effective Dimension Oscillations in Finite Width ReLU Networks
von: Makwana, Darshan
Veröffentlicht: (2025)
von: Makwana, Darshan
Veröffentlicht: (2025)
BRIDGE: Predicting Human Task Completion Time From Model Performance
von: Liu, Fengyuan, et al.
Veröffentlicht: (2026)
von: Liu, Fengyuan, et al.
Veröffentlicht: (2026)
Autoencoding Conditional Neural Processes for Representation Learning
von: Prokhorov, Victor, et al.
Veröffentlicht: (2023)
von: Prokhorov, Victor, et al.
Veröffentlicht: (2023)
Mitigating Copy Bias in In-Context Learning through Neuron Pruning
von: Ali, Ameen, et al.
Veröffentlicht: (2024)
von: Ali, Ameen, et al.
Veröffentlicht: (2024)
Safety Alignment Can Be Not Superficial With Explicit Safety Signals
von: Li, Jianwei, et al.
Veröffentlicht: (2025)
von: Li, Jianwei, et al.
Veröffentlicht: (2025)
Weasel: Out-of-Domain Generalization for Web Agents via Importance-Diversity Data Selection
von: Zadeh, Fatemeh Pesaran, et al.
Veröffentlicht: (2026)
von: Zadeh, Fatemeh Pesaran, et al.
Veröffentlicht: (2026)
Are self-explanations from Large Language Models faithful?
von: Madsen, Andreas, et al.
Veröffentlicht: (2024)
von: Madsen, Andreas, et al.
Veröffentlicht: (2024)
Shared Doubt: Zero-shot Cross-Lingual Confidence Estimation for Language Models
von: Kyriakou, Athina, et al.
Veröffentlicht: (2026)
von: Kyriakou, Athina, et al.
Veröffentlicht: (2026)
Joint Localization and Activation Editing for Low-Resource Fine-Tuning
von: Lai, Wen, et al.
Veröffentlicht: (2025)
von: Lai, Wen, et al.
Veröffentlicht: (2025)
Towards true discovery of the differential equations
von: Hvatov, Alexander, et al.
Veröffentlicht: (2023)
von: Hvatov, Alexander, et al.
Veröffentlicht: (2023)
Convergence and Divergence of Language Models under Different Random Seeds
von: Fehlauer, Finlay, et al.
Veröffentlicht: (2025)
von: Fehlauer, Finlay, et al.
Veröffentlicht: (2025)
The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?
von: Sutter, Denis, et al.
Veröffentlicht: (2025)
von: Sutter, Denis, et al.
Veröffentlicht: (2025)
Tokenisation via Convex Relaxations
von: Tempus, Jan, et al.
Veröffentlicht: (2026)
von: Tempus, Jan, et al.
Veröffentlicht: (2026)
The Master Key Hypothesis: Unlocking Cross-Model Capability Transfer via Linear Subspace Alignment
von: Balasubramanian, Rishab, et al.
Veröffentlicht: (2026)
von: Balasubramanian, Rishab, et al.
Veröffentlicht: (2026)
Cache & Distil: Optimising API Calls to Large Language Models
von: Ramírez, Guillem, et al.
Veröffentlicht: (2023)
von: Ramírez, Guillem, et al.
Veröffentlicht: (2023)
Detecting and Pruning Prominent but Detrimental Neurons in Large Language Models
von: Ali, Ameen, et al.
Veröffentlicht: (2025)
von: Ali, Ameen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Forecasting Downstream Performance of LLMs With Proxy Metrics
von: Patel, Arkil, et al.
Veröffentlicht: (2026) -
Build the web for agents, not agents for the web
von: Lù, Xing Han, et al.
Veröffentlicht: (2025) -
Value Drifts: Tracing Value Alignment During LLM Post-Training
von: Bhatia, Mehar, et al.
Veröffentlicht: (2025) -
Revisiting the Superficial Alignment Hypothesis
von: Raghavendra, Mohit, et al.
Veröffentlicht: (2024) -
Do Generalisation Results Generalise?
von: Boglioni, Matteo, et al.
Veröffentlicht: (2025)