Beyond the Rosetta Stone: Unification Forces in Generalization Dynamics
Fuente:
arXiv
Saved in:
| Main Authors: | Blum, Carter, Filippova, Katja, Yuan, Ann, Ghandeharioun, Asma, Zimmert, Julian, Zhang, Fred, Hoffmann, Jessica, Linzen, Tal, Wattenberg, Martin, Dixon, Lucas, Geva, Mor |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Think Before You Lie: How Reasoning Leads to Honesty
by: Yuan, Ann, et al.
Published: (2026)
by: Yuan, Ann, et al.
Published: (2026)
Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models
by: Ghandeharioun, Asma, et al.
Published: (2024)
by: Ghandeharioun, Asma, et al.
Published: (2024)
Language Models Struggle to Use Representations Learned In-Context
by: Lepori, Michael A., et al.
Published: (2026)
by: Lepori, Michael A., et al.
Published: (2026)
Who's asking? User personas and the mechanics of latent misalignment
by: Ghandeharioun, Asma, et al.
Published: (2024)
by: Ghandeharioun, Asma, et al.
Published: (2024)
When Can Transformers Count to n?
by: Yehudai, Gilad, et al.
Published: (2024)
by: Yehudai, Gilad, et al.
Published: (2024)
Interpretability Illusions in the Generalization of Simplified Models
by: Friedman, Dan, et al.
Published: (2023)
by: Friedman, Dan, et al.
Published: (2023)
Beyond Translation: The Rosetta Stone and the Remaking of Ptolemaic Egyptian Identity
by: Revista, Zen, et al.
Published: (2025)
by: Revista, Zen, et al.
Published: (2025)
A Rosetta Stone for AI Benchmarks
by: Ho, Anson, et al.
Published: (2025)
by: Ho, Anson, et al.
Published: (2025)
Rosetta Stone of Neural Mass Models
by: Castaldo, Francesca, et al.
Published: (2025)
by: Castaldo, Francesca, et al.
Published: (2025)
Do Language Models' Words Refer?
by: Mandelkern, Matthew, et al.
Published: (2023)
by: Mandelkern, Matthew, et al.
Published: (2023)
SPAWNing Structural Priming Predictions from a Cognitively Motivated Parser
by: Prasad, Grusha, et al.
Published: (2024)
by: Prasad, Grusha, et al.
Published: (2024)
A Rosetta Stone for Wilson Line Defects
by: Julius, Julius, et al.
Published: (2025)
by: Julius, Julius, et al.
Published: (2025)
Estimating Knowledge in Large Language Models Without Generating a Single Token
by: Gottesman, Daniela, et al.
Published: (2024)
by: Gottesman, Daniela, et al.
Published: (2024)
Inferring Functionality of Attention Heads from their Parameters
by: Elhelo, Amit, et al.
Published: (2024)
by: Elhelo, Amit, et al.
Published: (2024)
Part‐time subsidies and maternal reemployment: Evidence from a difference‐in‐differences analysis
by: Franziska Zimmert, et al.
Published: (2024)
by: Franziska Zimmert, et al.
Published: (2024)
Stroke Lesions as a Rosetta Stone for Language Model Interpretability
by: Fridriksson, Julius, et al.
Published: (2026)
by: Fridriksson, Julius, et al.
Published: (2026)
Optimal cross-learning for contextual bandits with unknown context distributions
by: Schneider, Jon, et al.
Published: (2024)
by: Schneider, Jon, et al.
Published: (2024)
Constructing Interpretable Features from Compositional Neuron Groups
by: Shafran, Or, et al.
Published: (2025)
by: Shafran, Or, et al.
Published: (2025)
Language Models Encode Numbers Using Digit Representations in Base 10
by: Levy, Amit Arnold, et al.
Published: (2024)
by: Levy, Amit Arnold, et al.
Published: (2024)
Towards Unifying Interpretability and Control: Evaluation via Intervention
by: Bhalla, Usha, et al.
Published: (2024)
by: Bhalla, Usha, et al.
Published: (2024)
Racing Thoughts: Explaining Contextualization Errors in Large Language Models
by: Lepori, Michael A., et al.
Published: (2024)
by: Lepori, Michael A., et al.
Published: (2024)
Syntax as a Rosetta Stone: Universal Dependencies for In-Context Coptic Translation
by: Purushothama, Abhishek, et al.
Published: (2026)
by: Purushothama, Abhishek, et al.
Published: (2026)
Evaluating In-Context Translation with Synchronous Context-Free Grammar Transduction
by: Petty, Jackson, et al.
Published: (2026)
by: Petty, Jackson, et al.
Published: (2026)
Entailment Semantics Can Be Extracted from an Ideal Language Model
by: Merrill, William, et al.
Published: (2022)
by: Merrill, William, et al.
Published: (2022)
Deconstructing sentence disambiguation by joint latent modeling of reading paradigms: LLM surprisal is not enough
by: Paape, Dario, et al.
Published: (2026)
by: Paape, Dario, et al.
Published: (2026)
Manipulating language models' training data to study syntactic constraint learning: the case of English passivization
by: Leong, Cara Su-Yi, et al.
Published: (2024)
by: Leong, Cara Su-Yi, et al.
Published: (2024)
Can LLMs Introspect? A Reality Check
by: Singh, Shashwat, et al.
Published: (2026)
by: Singh, Shashwat, et al.
Published: (2026)
Why are language models less surprised than humans? Testing the Parse Multiplicity Mismatch Hypothesis
by: Timkey, William, et al.
Published: (2026)
by: Timkey, William, et al.
Published: (2026)
Analysis of respiratory events in obstructive sleep apnea syndrome: Inter-relations and association to simple nocturnal features
by: H. Ghandeharioun
Published: (2016)
by: H. Ghandeharioun
Published: (2016)
Incentive-compatible Bandits: Importance Weighting No More
by: Zimmert, Julian, et al.
Published: (2024)
by: Zimmert, Julian, et al.
Published: (2024)
A Rosetta Stone Hypothesis for Neurophenomenology: Mathematical Predictions from Predictive Processing
by: Da Costa, Lancelot, et al.
Published: (2024)
by: Da Costa, Lancelot, et al.
Published: (2024)
Data-driven Circuit Discovery for Interpretability of Language Models
by: Rai, Daking, et al.
Published: (2026)
by: Rai, Daking, et al.
Published: (2026)
Eliciting Textual Descriptions from Representations of Continuous Prompts
by: Ramati, Dana, et al.
Published: (2024)
by: Ramati, Dana, et al.
Published: (2024)
Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers
by: Yona, Gal, et al.
Published: (2024)
by: Yona, Gal, et al.
Published: (2024)
Preventing Rogue Agents Improves Multi-Agent Collaboration
by: Barbi, Ohav, et al.
Published: (2025)
by: Barbi, Ohav, et al.
Published: (2025)
Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?
by: Yona, Gal, et al.
Published: (2024)
by: Yona, Gal, et al.
Published: (2024)
Hallucinations Undermine Trust; Metacognition is a Way Forward
by: Yona, Gal, et al.
Published: (2026)
by: Yona, Gal, et al.
Published: (2026)
Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts
by: Ahrac, Sagi, et al.
Published: (2026)
by: Ahrac, Sagi, et al.
Published: (2026)
How Does Code Pretraining Affect Language Model Task Performance?
by: Petty, Jackson, et al.
Published: (2024)
by: Petty, Jackson, et al.
Published: (2024)
JADES -- The Rosetta Stone of JWST-discovered AGN: deciphering the intriguing nature of early AGN
by: Juodžbalis, Ignas, et al.
Published: (2024)
by: Juodžbalis, Ignas, et al.
Published: (2024)
Similar Items
-
Think Before You Lie: How Reasoning Leads to Honesty
by: Yuan, Ann, et al.
Published: (2026) -
Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models
by: Ghandeharioun, Asma, et al.
Published: (2024) -
Language Models Struggle to Use Representations Learned In-Context
by: Lepori, Michael A., et al.
Published: (2026) -
Who's asking? User personas and the mechanics of latent misalignment
by: Ghandeharioun, Asma, et al.
Published: (2024) -
When Can Transformers Count to n?
by: Yehudai, Gilad, et al.
Published: (2024)