Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention
Fuente:
arXiv
Guardado en:
| Autores principales: | Huang, Jing, Wurgaft, Daniel, Bansal, Rachit, Ruis, Laura, Saphra, Naomi, Alvarez-Melis, David, Lampinen, Andrew Kyle, Potts, Christopher, Lubana, Ekdeep Singh |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
por: Pres, Itamar, et al.
Publicado: (2024)
por: Pres, Itamar, et al.
Publicado: (2024)
Sometimes I am a Tree: Data Drives Unstable Hierarchical Generalization
por: Qin, Tian, et al.
Publicado: (2024)
por: Qin, Tian, et al.
Publicado: (2024)
In-Context Learning Strategies Emerge Rationally
por: Wurgaft, Daniel, et al.
Publicado: (2025)
por: Wurgaft, Daniel, et al.
Publicado: (2025)
How Do LLMs Persuade? Linear Probes Can Uncover Persuasion Dynamics in Multi-Turn Conversations
por: Jaipersaud, Brandon, et al.
Publicado: (2025)
por: Jaipersaud, Brandon, et al.
Publicado: (2025)
Abrupt Learning in Transformers: A Case Study on Matrix Completion
por: Gopalani, Pulkit, et al.
Publicado: (2024)
por: Gopalani, Pulkit, et al.
Publicado: (2024)
Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic Task
por: Okawa, Maya, et al.
Publicado: (2023)
por: Okawa, Maya, et al.
Publicado: (2023)
Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering
por: Bigelow, Eric, et al.
Publicado: (2025)
por: Bigelow, Eric, et al.
Publicado: (2025)
Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks
por: Ramesh, Rahul, et al.
Publicado: (2023)
por: Ramesh, Rahul, et al.
Publicado: (2023)
Random Scaling of Emergent Capabilities
por: Zhao, Rosie, et al.
Publicado: (2025)
por: Zhao, Rosie, et al.
Publicado: (2025)
Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space
por: Bigelow, Eric, et al.
Publicado: (2026)
por: Bigelow, Eric, et al.
Publicado: (2026)
Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics
por: Zur, Amir, et al.
Publicado: (2025)
por: Zur, Amir, et al.
Publicado: (2025)
Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders
por: Bohacek, Matyas, et al.
Publicado: (2025)
por: Bohacek, Matyas, et al.
Publicado: (2025)
Analyzing (In)Abilities of SAEs via Formal Languages
por: Menon, Abhinav, et al.
Publicado: (2024)
por: Menon, Abhinav, et al.
Publicado: (2024)
Can Interpretation Predict Behavior on Unseen Data?
por: Li, Victoria R., et al.
Publicado: (2025)
por: Li, Victoria R., et al.
Publicado: (2025)
Emergence of Hidden Capabilities: Exploring Learning Dynamics in Concept Space
por: Park, Core Francisco, et al.
Publicado: (2024)
por: Park, Core Francisco, et al.
Publicado: (2024)
The Shape of Beliefs: Geometry, Dynamics, and Interventions along Representation Manifolds of Language Models' Posteriors
por: Sarfati, Raphaël, et al.
Publicado: (2026)
por: Sarfati, Raphaël, et al.
Publicado: (2026)
A Percolation Model of Emergence: Analyzing Transformers Trained on a Formal Language
por: Lubana, Ekdeep Singh, et al.
Publicado: (2024)
por: Lubana, Ekdeep Singh, et al.
Publicado: (2024)
Competition Dynamics Shape Algorithmic Phases of In-Context Learning
por: Park, Core Francisco, et al.
Publicado: (2024)
por: Park, Core Francisco, et al.
Publicado: (2024)
Mechanistic?
por: Saphra, Naomi, et al.
Publicado: (2024)
por: Saphra, Naomi, et al.
Publicado: (2024)
From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
por: Mueller, Aaron, et al.
Publicado: (2025)
por: Mueller, Aaron, et al.
Publicado: (2025)
Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept Geometry
por: Hindupur, Sai Sumedh R., et al.
Publicado: (2025)
por: Hindupur, Sai Sumedh R., et al.
Publicado: (2025)
From Flat to Hierarchical: Extracting Sparse Representations with Matching Pursuit
por: Costa, Valérie, et al.
Publicado: (2025)
por: Costa, Valérie, et al.
Publicado: (2025)
Evaluating Sparse Autoencoders: From Shallow Design to Matching Pursuit
por: Costa, Valérie, et al.
Publicado: (2025)
por: Costa, Valérie, et al.
Publicado: (2025)
Features as Rewards: Scalable Supervision for Open-Ended Tasks via Interpretability
por: Prasad, Aaditya Vikram, et al.
Publicado: (2026)
por: Prasad, Aaditya Vikram, et al.
Publicado: (2026)
Adam or Gauss-Newton? A Comparative Study In Terms of Basis Alignment and SGD Noise
por: Liu, Bingbin, et al.
Publicado: (2025)
por: Liu, Bingbin, et al.
Publicado: (2025)
Representation Shattering in Transformers: A Synthetic Study with Knowledge Editing
por: Nishi, Kento, et al.
Publicado: (2024)
por: Nishi, Kento, et al.
Publicado: (2024)
The Impact of Off-Policy Training Data on Probe Generalisation
por: Kirch, Nathalie, et al.
Publicado: (2025)
por: Kirch, Nathalie, et al.
Publicado: (2025)
Hidden Breakthroughs in Language Model Training
por: Kangaslahti, Sara, et al.
Publicado: (2025)
por: Kangaslahti, Sara, et al.
Publicado: (2025)
Do Sparse Autoencoders Capture Concept Manifolds?
por: Bhalla, Usha, et al.
Publicado: (2026)
por: Bhalla, Usha, et al.
Publicado: (2026)
Beneath the Surface: Investigating LLMs' Capabilities for Communicating with Subtext
por: Ahuja, Kabir, et al.
Publicado: (2026)
por: Ahuja, Kabir, et al.
Publicado: (2026)
Arithmetic in the Wild: Llama uses Base-10 Addition to Reason About Cyclic Concepts
por: Feucht, Sheridan, et al.
Publicado: (2026)
por: Feucht, Sheridan, et al.
Publicado: (2026)
In-Context Learning Dynamics with Random Binary Sequences
por: Bigelow, Eric J., et al.
Publicado: (2023)
por: Bigelow, Eric J., et al.
Publicado: (2023)
Swing-by Dynamics in Concept Learning and Compositional Generalization
por: Yang, Yongyi, et al.
Publicado: (2024)
por: Yang, Yongyi, et al.
Publicado: (2024)
ChatGPT Doesn't Trust Chargers Fans: Guardrail Sensitivity in Context
por: Li, Victoria R., et al.
Publicado: (2024)
por: Li, Victoria R., et al.
Publicado: (2024)
Los fondos de inversión social y la reestructuración económica de América latina
por: José. Wurgaft
Publicado: (1992)
por: José. Wurgaft
Publicado: (1992)
Les fonds d'investissement social et la restructuration économique en Amérique latine
por: José. Wurgaft
Publicado: (1992)
por: José. Wurgaft
Publicado: (1992)
Social investment funds and economic restructuring in Latin America
por: José. Wurgaft
Publicado: (1992)
por: José. Wurgaft
Publicado: (1992)
ICLR: In-Context Learning of Representations
por: Park, Core Francisco, et al.
Publicado: (2024)
por: Park, Core Francisco, et al.
Publicado: (2024)
A Taxonomy of Transcendence
por: Abreu, Natalie, et al.
Publicado: (2025)
por: Abreu, Natalie, et al.
Publicado: (2025)
First Tragedy, then Parse: History Repeats Itself in the New Era of Large Language Models
por: Saphra, Naomi, et al.
Publicado: (2023)
por: Saphra, Naomi, et al.
Publicado: (2023)
Ejemplares similares
-
Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
por: Pres, Itamar, et al.
Publicado: (2024) -
Sometimes I am a Tree: Data Drives Unstable Hierarchical Generalization
por: Qin, Tian, et al.
Publicado: (2024) -
In-Context Learning Strategies Emerge Rationally
por: Wurgaft, Daniel, et al.
Publicado: (2025) -
How Do LLMs Persuade? Linear Probes Can Uncover Persuasion Dynamics in Multi-Turn Conversations
por: Jaipersaud, Brandon, et al.
Publicado: (2025) -
Abrupt Learning in Transformers: A Case Study on Matrix Completion
por: Gopalani, Pulkit, et al.
Publicado: (2024)