Circuit Stability Characterizes Language Model Generalization
Fuente:
arXiv
Salvato in:
| Autore principale: | Sun, Alan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Algorithmic Phase Transitions in Language Models: A Mechanistic Case Study of Arithmetic
di: Sun, Alan, et al.
Pubblicazione: (2024)
di: Sun, Alan, et al.
Pubblicazione: (2024)
Characterizing Learning Curves During Language Model Pre-Training: Learning, Forgetting, and Stability
di: Chang, Tyler A., et al.
Pubblicazione: (2023)
di: Chang, Tyler A., et al.
Pubblicazione: (2023)
How Much Do Circuits Tell Us? Measuring the Consistency and Specificity of Language Model Circuits
di: Li, Michael, et al.
Pubblicazione: (2026)
di: Li, Michael, et al.
Pubblicazione: (2026)
On Characterizations for Language Generation: Interplay of Hallucinations, Breadth, and Stability
di: Kalavasis, Alkis, et al.
Pubblicazione: (2024)
di: Kalavasis, Alkis, et al.
Pubblicazione: (2024)
Understanding Language Model Circuits through Knowledge Editing
di: Ge, Huaizhi, et al.
Pubblicazione: (2024)
di: Ge, Huaizhi, et al.
Pubblicazione: (2024)
Architecture, Not Scale: Circuit Localization in Large Language Models
di: Venkatesh, Sohan
Pubblicazione: (2026)
di: Venkatesh, Sohan
Pubblicazione: (2026)
Characterizing Truthfulness in Large Language Model Generations with Local Intrinsic Dimension
di: Yin, Fan, et al.
Pubblicazione: (2024)
di: Yin, Fan, et al.
Pubblicazione: (2024)
Characterizing Memorization in Diffusion Language Models: Generalized Extraction and Sampling Effects
di: Luo, Xiaoyu, et al.
Pubblicazione: (2026)
di: Luo, Xiaoyu, et al.
Pubblicazione: (2026)
Mechanistic Circuit-Based Knowledge Editing in Large Language Models
di: Zhao, Tianyi, et al.
Pubblicazione: (2026)
di: Zhao, Tianyi, et al.
Pubblicazione: (2026)
CircuitSynth: Reliable Synthetic Data Generation
di: Cheng, Zehua, et al.
Pubblicazione: (2026)
di: Cheng, Zehua, et al.
Pubblicazione: (2026)
Electronic Circuit Principles of Large Language Models
di: Chen, Qiguang, et al.
Pubblicazione: (2025)
di: Chen, Qiguang, et al.
Pubblicazione: (2025)
CktFormalizer: Autoformalization of Natural Language into Circuit Representations
di: Xiong, Jing, et al.
Pubblicazione: (2026)
di: Xiong, Jing, et al.
Pubblicazione: (2026)
Triggers Hijack Language Circuits: A Mechanistic Analysis of Backdoor Behaviors in Large Language Models
di: Lasnier, Théo, et al.
Pubblicazione: (2026)
di: Lasnier, Théo, et al.
Pubblicazione: (2026)
Applying Large Language Models to Characterize Public Narratives
di: Poole-Dayan, Elinor, et al.
Pubblicazione: (2025)
di: Poole-Dayan, Elinor, et al.
Pubblicazione: (2025)
Characterizing the Expressivity of Fixed-Precision Transformer Language Models
di: Li, Jiaoda, et al.
Pubblicazione: (2025)
di: Li, Jiaoda, et al.
Pubblicazione: (2025)
Characterizing the Role of Similarity in the Property Inferences of Language Models
di: Rodriguez, Juan Diego, et al.
Pubblicazione: (2024)
di: Rodriguez, Juan Diego, et al.
Pubblicazione: (2024)
Circuit Component Reuse Across Tasks in Transformer Language Models
di: Merullo, Jack, et al.
Pubblicazione: (2023)
di: Merullo, Jack, et al.
Pubblicazione: (2023)
Characterizing Selective Refusal Bias in Large Language Models
di: Khorramrouz, Adel, et al.
Pubblicazione: (2025)
di: Khorramrouz, Adel, et al.
Pubblicazione: (2025)
Representational and Behavioral Stability of Truth in Large Language Models
di: Dies, Samantha, et al.
Pubblicazione: (2025)
di: Dies, Samantha, et al.
Pubblicazione: (2025)
Lemma Dilemma: On Lemma Generation Without Domain- or Language-Specific Training Data
di: Toporkov, Olia, et al.
Pubblicazione: (2025)
di: Toporkov, Olia, et al.
Pubblicazione: (2025)
FakeGPT: Fake News Generation, Explanation and Detection of Large Language Models
di: Huang, Yue, et al.
Pubblicazione: (2023)
di: Huang, Yue, et al.
Pubblicazione: (2023)
Boosting Disfluency Detection with Large Language Model as Disfluency Generator
di: Cheng, Zhenrong, et al.
Pubblicazione: (2024)
di: Cheng, Zhenrong, et al.
Pubblicazione: (2024)
Detecting and Characterizing Planning in Language Models
di: Nainani, Jatin, et al.
Pubblicazione: (2025)
di: Nainani, Jatin, et al.
Pubblicazione: (2025)
Evaluating the Generation Capabilities of Large Chinese Language Models
di: Zeng, Hui, et al.
Pubblicazione: (2023)
di: Zeng, Hui, et al.
Pubblicazione: (2023)
NEO-BENCH: Evaluating Robustness of Large Language Models with Neologisms
di: Zheng, Jonathan, et al.
Pubblicazione: (2024)
di: Zheng, Jonathan, et al.
Pubblicazione: (2024)
Language Drift in Multilingual Retrieval-Augmented Generation: Characterization and Decoding-Time Mitigation
di: Li, Bo, et al.
Pubblicazione: (2025)
di: Li, Bo, et al.
Pubblicazione: (2025)
Discursive Circuits: How Do Language Models Understand Discourse Relations?
di: Miao, Yisong, et al.
Pubblicazione: (2025)
di: Miao, Yisong, et al.
Pubblicazione: (2025)
AlphaMaze: Enhancing Large Language Models' Spatial Intelligence via GRPO
di: Dao, Alan, et al.
Pubblicazione: (2025)
di: Dao, Alan, et al.
Pubblicazione: (2025)
Circuit Compositions: Exploring Modular Structures in Transformer-Based Language Models
di: Mondorf, Philipp, et al.
Pubblicazione: (2024)
di: Mondorf, Philipp, et al.
Pubblicazione: (2024)
Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models
di: O'Neill, Charles, et al.
Pubblicazione: (2024)
di: O'Neill, Charles, et al.
Pubblicazione: (2024)
The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models
di: Lu, Christina, et al.
Pubblicazione: (2026)
di: Lu, Christina, et al.
Pubblicazione: (2026)
Prompt Stability Scoring for Text Annotation with Large Language Models
di: Barrie, Christopher, et al.
Pubblicazione: (2024)
di: Barrie, Christopher, et al.
Pubblicazione: (2024)
Tending Towards Stability: Convergence Challenges in Small Language Models
di: Martinez, Richard Diehl, et al.
Pubblicazione: (2024)
di: Martinez, Richard Diehl, et al.
Pubblicazione: (2024)
BabyHGRN: Exploring RNNs for Sample-Efficient Training of Language Models
di: Haller, Patrick, et al.
Pubblicazione: (2024)
di: Haller, Patrick, et al.
Pubblicazione: (2024)
Effective Large Language Model Adaptation for Improved Grounding and Citation Generation
di: Ye, Xi, et al.
Pubblicazione: (2023)
di: Ye, Xi, et al.
Pubblicazione: (2023)
Language Model Circuits Are Sparse in the Neuron Basis
di: Arora, Aryaman, et al.
Pubblicazione: (2026)
di: Arora, Aryaman, et al.
Pubblicazione: (2026)
Dissecting the Ledger: Locating and Suppressing "Liar Circuits" in Financial Large Language Models
di: Mirajkar, Soham
Pubblicazione: (2025)
di: Mirajkar, Soham
Pubblicazione: (2025)
Unsupervised Distractor Generation via Large Language Model Distilling and Counterfactual Contrastive Decoding
di: Qu, Fanyi, et al.
Pubblicazione: (2024)
di: Qu, Fanyi, et al.
Pubblicazione: (2024)
Algorithmic Stability in Infinite Dimensions: Characterizing Unconditional Convergence in Banach Spaces
di: Spyra, Przemysław
Pubblicazione: (2026)
di: Spyra, Przemysław
Pubblicazione: (2026)
BEAR: A Unified Framework for Evaluating Relational Knowledge in Causal and Masked Language Models
di: Wiland, Jacek, et al.
Pubblicazione: (2024)
di: Wiland, Jacek, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Algorithmic Phase Transitions in Language Models: A Mechanistic Case Study of Arithmetic
di: Sun, Alan, et al.
Pubblicazione: (2024) -
Characterizing Learning Curves During Language Model Pre-Training: Learning, Forgetting, and Stability
di: Chang, Tyler A., et al.
Pubblicazione: (2023) -
How Much Do Circuits Tell Us? Measuring the Consistency and Specificity of Language Model Circuits
di: Li, Michael, et al.
Pubblicazione: (2026) -
On Characterizations for Language Generation: Interplay of Hallucinations, Breadth, and Stability
di: Kalavasis, Alkis, et al.
Pubblicazione: (2024) -
Understanding Language Model Circuits through Knowledge Editing
di: Ge, Huaizhi, et al.
Pubblicazione: (2024)