Towards Spectroscopy: Susceptibility Clusters in Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Gordon, Andrew, Baker, Garrett, Wang, George, Snell, William, van Wingerden, Stan, Murfet, Daniel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Embryology of a Language Model
por: Wang, George, et al.
Publicado: (2025)
por: Wang, George, et al.
Publicado: (2025)
Structural Inference: Interpreting Small Language Models with Susceptibilities
por: Baker, Garrett, et al.
Publicado: (2025)
por: Baker, Garrett, et al.
Publicado: (2025)
Differentiation and Specialization of Attention Heads via the Refined Local Learning Coefficient
por: Wang, George, et al.
Publicado: (2024)
por: Wang, George, et al.
Publicado: (2024)
Compressibility Measures Complexity: Minimum Description Length Meets Singular Learning Theory
por: Urdshals, Einar, et al.
Publicado: (2025)
por: Urdshals, Einar, et al.
Publicado: (2025)
Patterning: The Dual of Interpretability
por: Wang, George, et al.
Publicado: (2026)
por: Wang, George, et al.
Publicado: (2026)
Susceptibilities and Patterning: A Primer on Linear Response in Bayesian Learning
por: Elliott, Chris, et al.
Publicado: (2026)
por: Elliott, Chris, et al.
Publicado: (2026)
Interpreting Reinforcement Learning Agents with Susceptibilities
por: Elliott, Chris, et al.
Publicado: (2026)
por: Elliott, Chris, et al.
Publicado: (2026)
Modes of Sequence Models and Learning Coefficients
por: Chen, Zhongtian, et al.
Publicado: (2025)
por: Chen, Zhongtian, et al.
Publicado: (2025)
Linear Response Estimators for Singular Statistical Models
por: Elliott, Chris, et al.
Publicado: (2026)
por: Elliott, Chris, et al.
Publicado: (2026)
In-Context Clustering with Large Language Models
por: Wang, Ying, et al.
Publicado: (2025)
por: Wang, Ying, et al.
Publicado: (2025)
Programs as Singularities
por: Murfet, Daniel, et al.
Publicado: (2025)
por: Murfet, Daniel, et al.
Publicado: (2025)
Dynamics of Transient Structure in In-Context Linear Regression Transformers
por: Carroll, Liam, et al.
Publicado: (2025)
por: Carroll, Liam, et al.
Publicado: (2025)
Loss Landscape Degeneracy and Stagewise Development in Transformers
por: Hoogland, Jesse, et al.
Publicado: (2024)
por: Hoogland, Jesse, et al.
Publicado: (2024)
The Local Learning Coefficient: A Singularity-Aware Complexity Measure
por: Lau, Edmund, et al.
Publicado: (2023)
por: Lau, Edmund, et al.
Publicado: (2023)
Stagewise Reinforcement Learning and the Geometry of the Regret Landscape
por: Elliott, Chris, et al.
Publicado: (2026)
por: Elliott, Chris, et al.
Publicado: (2026)
You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation
por: Lehalleur, Simon Pepin, et al.
Publicado: (2025)
por: Lehalleur, Simon Pepin, et al.
Publicado: (2025)
Meta-Learning at Scale for Large Language Models via Low-Rank Amortized Bayesian Meta-Learning
por: Zhang, Liyi, et al.
Publicado: (2025)
por: Zhang, Liyi, et al.
Publicado: (2025)
Large Language Models Are Zero-Shot Time Series Forecasters
por: Gruver, Nate, et al.
Publicado: (2023)
por: Gruver, Nate, et al.
Publicado: (2023)
Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
por: Snell, Charlie, et al.
Publicado: (2024)
por: Snell, Charlie, et al.
Publicado: (2024)
Deep Learning is Not So Mysterious or Different
por: Wilson, Andrew Gordon
Publicado: (2025)
por: Wilson, Andrew Gordon
Publicado: (2025)
Conformal Prediction as Bayesian Quadrature
por: Snell, Jake C., et al.
Publicado: (2025)
por: Snell, Jake C., et al.
Publicado: (2025)
Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay
por: Marek, Martin, et al.
Publicado: (2026)
por: Marek, Martin, et al.
Publicado: (2026)
SimVPv2: Towards Simple yet Powerful Spatiotemporal Predictive Learning
por: Tan, Cheng, et al.
Publicado: (2022)
por: Tan, Cheng, et al.
Publicado: (2022)
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM
por: Fan, Zehao, et al.
Publicado: (2025)
por: Fan, Zehao, et al.
Publicado: (2025)
Unlocking Tokens as Data Points for Generalization Bounds on Larger Language Models
por: Lotfi, Sanae, et al.
Publicado: (2024)
por: Lotfi, Sanae, et al.
Publicado: (2024)
ADNAC: Audio Denoiser using Neural Audio Codec
por: Jimon, Daniel, et al.
Publicado: (2025)
por: Jimon, Daniel, et al.
Publicado: (2025)
Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful
por: Marek, Martin, et al.
Publicado: (2025)
por: Marek, Martin, et al.
Publicado: (2025)
Non-Vacuous Generalization Bounds for Large Language Models
por: Lotfi, Sanae, et al.
Publicado: (2023)
por: Lotfi, Sanae, et al.
Publicado: (2023)
Beyond the Academic Monoculture: A Unified Framework and Industrial Perspective for Attributed Graph Clustering
por: Liu, Yunhui, et al.
Publicado: (2026)
por: Liu, Yunhui, et al.
Publicado: (2026)
Q-Learning with Clustered-SMART (cSMART) Data: Examining Moderators in the Construction of Clustered Adaptive Interventions
por: Song, Yao, et al.
Publicado: (2025)
por: Song, Yao, et al.
Publicado: (2025)
Prompt Risk Control: A Rigorous Framework for Responsible Deployment of Large Language Models
por: Zollo, Thomas P., et al.
Publicado: (2023)
por: Zollo, Thomas P., et al.
Publicado: (2023)
Direct Profit Estimation Using Uplift Modeling under Clustered Network Interference
por: Akker, Bram van den
Publicado: (2025)
por: Akker, Bram van den
Publicado: (2025)
On the Reproducibility of "FairCLIP: Harnessing Fairness in Vision-Language Learning''
por: Bakker, Hua Chang, et al.
Publicado: (2025)
por: Bakker, Hua Chang, et al.
Publicado: (2025)
A Nonparametric Discrete Hawkes Model with a Collapsed Gaussian-Process Prior
por: Brisley, Trinnhallen, et al.
Publicado: (2025)
por: Brisley, Trinnhallen, et al.
Publicado: (2025)
Towards Multimodal Graph Large Language Model
por: Wang, Xin, et al.
Publicado: (2025)
por: Wang, Xin, et al.
Publicado: (2025)
Predicting Emergent Capabilities by Finetuning
por: Snell, Charlie, et al.
Publicado: (2024)
por: Snell, Charlie, et al.
Publicado: (2024)
Machine Learning for Raman Spectroscopy-based Cyber-Marine Fish Biochemical Composition Analysis
por: Zhou, Yun, et al.
Publicado: (2024)
por: Zhou, Yun, et al.
Publicado: (2024)
LaTable: Towards Large Tabular Models
por: van Breugel, Boris, et al.
Publicado: (2024)
por: van Breugel, Boris, et al.
Publicado: (2024)
Coding historical causes of death data with Large Language Models
por: Pedersen, Bjørn, et al.
Publicado: (2024)
por: Pedersen, Bjørn, et al.
Publicado: (2024)
Mechanistic Exploration of Backdoored Large Language Model Attention Patterns
por: Baker, Mohammed Abu, et al.
Publicado: (2025)
por: Baker, Mohammed Abu, et al.
Publicado: (2025)
Ejemplares similares
-
Embryology of a Language Model
por: Wang, George, et al.
Publicado: (2025) -
Structural Inference: Interpreting Small Language Models with Susceptibilities
por: Baker, Garrett, et al.
Publicado: (2025) -
Differentiation and Specialization of Attention Heads via the Refined Local Learning Coefficient
por: Wang, George, et al.
Publicado: (2024) -
Compressibility Measures Complexity: Minimum Description Length Meets Singular Learning Theory
por: Urdshals, Einar, et al.
Publicado: (2025) -
Patterning: The Dual of Interpretability
por: Wang, George, et al.
Publicado: (2026)