Is Grokking Worthwhile? Functional Analysis and Transferability of Generalization Circuits in Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | He, Kaiyu, Mian, Zhang, Wu, Peilin, Du, Xinya, Chen, Zhiyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GEAR: A General Evaluation Framework for Abductive Reasoning
von: He, Kaiyu, et al.
Veröffentlicht: (2025)
von: He, Kaiyu, et al.
Veröffentlicht: (2025)
HiPRAG: Hierarchical Process Rewards for Efficient Agentic Retrieval Augmented Generation
von: Wu, Peilin, et al.
Veröffentlicht: (2025)
von: Wu, Peilin, et al.
Veröffentlicht: (2025)
Search Wisely: Mitigating Sub-optimal Agentic Searches By Reducing Uncertainty
von: Wu, Peilin, et al.
Veröffentlicht: (2025)
von: Wu, Peilin, et al.
Veröffentlicht: (2025)
IDEA: Enhancing the Rule Learning Ability of Large Language Model Agent through Induction, Deduction, and Abduction
von: He, Kaiyu, et al.
Veröffentlicht: (2024)
von: He, Kaiyu, et al.
Veröffentlicht: (2024)
From Reasoning to Learning: A Survey on Hypothesis Discovery and Rule Learning with Large Language Models
von: He, Kaiyu, et al.
Veröffentlicht: (2025)
von: He, Kaiyu, et al.
Veröffentlicht: (2025)
Do Retrieval-Augmented Language Models Adapt to Varying User Needs?
von: Wu, Peilin, et al.
Veröffentlicht: (2025)
von: Wu, Peilin, et al.
Veröffentlicht: (2025)
AMARIS: A Memory-Augmented Rubric Improvement System for Rubric-Based Reinforcement Learning
von: Wu, Peilin, et al.
Veröffentlicht: (2026)
von: Wu, Peilin, et al.
Veröffentlicht: (2026)
PRD: Peer Rank and Discussion Improve Large Language Model based Evaluations
von: Li, Ruosen, et al.
Veröffentlicht: (2023)
von: Li, Ruosen, et al.
Veröffentlicht: (2023)
LDC: Learning to Generate Research Idea with Dynamic Control
von: Li, Ruochen, et al.
Veröffentlicht: (2024)
von: Li, Ruochen, et al.
Veröffentlicht: (2024)
Long Story Generation via Knowledge Graph and Literary Theory
von: Shi, Ge, et al.
Veröffentlicht: (2025)
von: Shi, Ge, et al.
Veröffentlicht: (2025)
Logical Reasoning over Natural Language as Knowledge Representation: A Survey
von: Yang, Zonglin, et al.
Veröffentlicht: (2023)
von: Yang, Zonglin, et al.
Veröffentlicht: (2023)
Extrapolation by Association: Length Generalization Transfer in Transformers
von: Cai, Ziyang, et al.
Veröffentlicht: (2025)
von: Cai, Ziyang, et al.
Veröffentlicht: (2025)
MLR-Copilot: Autonomous Machine Learning Research based on Large Language Models Agents
von: Li, Ruochen, et al.
Veröffentlicht: (2024)
von: Li, Ruochen, et al.
Veröffentlicht: (2024)
Large Language Models for Automated Open-domain Scientific Hypotheses Discovery
von: Yang, Zonglin, et al.
Veröffentlicht: (2023)
von: Yang, Zonglin, et al.
Veröffentlicht: (2023)
FlowLM: Few-Step Language Modeling via Diffusion-to-Flow Adaptation
von: Zhang, Runzhe, et al.
Veröffentlicht: (2026)
von: Zhang, Runzhe, et al.
Veröffentlicht: (2026)
FASTopic: Pretrained Transformer is a Fast, Adaptive, Stable, and Transferable Topic Model
von: Wu, Xiaobao, et al.
Veröffentlicht: (2024)
von: Wu, Xiaobao, et al.
Veröffentlicht: (2024)
Critical Data Size of Language Models from a Grokking Perspective
von: Zhu, Xuekai, et al.
Veröffentlicht: (2024)
von: Zhu, Xuekai, et al.
Veröffentlicht: (2024)
RLAC: Reinforcement Learning with Adversarial Critic for Free-Form Generation Tasks
von: Wu, Mian, et al.
Veröffentlicht: (2025)
von: Wu, Mian, et al.
Veröffentlicht: (2025)
Kwai-STaR: Transform LLMs into State-Transition Reasoners
von: Lu, Xingyu, et al.
Veröffentlicht: (2024)
von: Lu, Xingyu, et al.
Veröffentlicht: (2024)
LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research
von: Yan, Shuo, et al.
Veröffentlicht: (2025)
von: Yan, Shuo, et al.
Veröffentlicht: (2025)
How do Transformer Embeddings Represent Compositions? A Functional Analysis
von: Nagar, Aishik, et al.
Veröffentlicht: (2025)
von: Nagar, Aishik, et al.
Veröffentlicht: (2025)
Towards Reliable Detection of LLM-Generated Texts: A Comprehensive Evaluation Framework with CUDRT
von: Tao, Zhen, et al.
Veröffentlicht: (2024)
von: Tao, Zhen, et al.
Veröffentlicht: (2024)
DSBench: How Far Are Data Science Agents from Becoming Data Science Experts?
von: Jing, Liqiang, et al.
Veröffentlicht: (2024)
von: Jing, Liqiang, et al.
Veröffentlicht: (2024)
CircuitSynth: Reliable Synthetic Data Generation
von: Cheng, Zehua, et al.
Veröffentlicht: (2026)
von: Cheng, Zehua, et al.
Veröffentlicht: (2026)
CBT-Bench: Evaluating Large Language Models on Assisting Cognitive Behavior Therapy
von: Zhang, Mian, et al.
Veröffentlicht: (2024)
von: Zhang, Mian, et al.
Veröffentlicht: (2024)
Language Models as Inductive Reasoners
von: Yang, Zonglin, et al.
Veröffentlicht: (2022)
von: Yang, Zonglin, et al.
Veröffentlicht: (2022)
Zero-Shot Continuous Prompt Transfer: Generalizing Task Semantics Across Language Models
von: Wu, Zijun, et al.
Veröffentlicht: (2023)
von: Wu, Zijun, et al.
Veröffentlicht: (2023)
Transformer Circuit Faithfulness Metrics are not Robust
von: Miller, Joseph, et al.
Veröffentlicht: (2024)
von: Miller, Joseph, et al.
Veröffentlicht: (2024)
Measuring Diversity in Synthetic Datasets
von: Zhu, Yuchang, et al.
Veröffentlicht: (2025)
von: Zhu, Yuchang, et al.
Veröffentlicht: (2025)
Towards Understanding Fine-Tuning Mechanisms of LLMs via Circuit Analysis
von: Wang, Xu, et al.
Veröffentlicht: (2025)
von: Wang, Xu, et al.
Veröffentlicht: (2025)
CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning
von: He, Zheqi, et al.
Veröffentlicht: (2024)
von: He, Zheqi, et al.
Veröffentlicht: (2024)
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
von: Huan, Maggie, et al.
Veröffentlicht: (2025)
von: Huan, Maggie, et al.
Veröffentlicht: (2025)
FCKT: Fine-Grained Cross-Task Knowledge Transfer with Semantic Contrastive Learning for Targeted Sentiment Analysis
von: Chen, Wei, et al.
Veröffentlicht: (2025)
von: Chen, Wei, et al.
Veröffentlicht: (2025)
SynGhost: Invisible and Universal Task-agnostic Backdoor Attack via Syntactic Transfer
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2024)
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2024)
Circuit Complexity Bounds for RoPE-based Transformer Architecture
von: Chen, Bo, et al.
Veröffentlicht: (2024)
von: Chen, Bo, et al.
Veröffentlicht: (2024)
Do LLMs "Feel"? Emotion Circuits Discovery and Control
von: Wang, Chenxi, et al.
Veröffentlicht: (2025)
von: Wang, Chenxi, et al.
Veröffentlicht: (2025)
A Comprehensive Framework for Semantic Similarity Analysis of Human and AI-Generated Text Using Transformer Architectures and Ensemble Techniques
von: Gao, Lifu, et al.
Veröffentlicht: (2025)
von: Gao, Lifu, et al.
Veröffentlicht: (2025)
Generative Explore-Exploit: Training-free Optimization of Generative Recommender Systems using LLM Optimizers
von: Senel, Lütfi Kerem, et al.
Veröffentlicht: (2024)
von: Senel, Lütfi Kerem, et al.
Veröffentlicht: (2024)
Beyond Transfer Accuracy: Faithful Circuits for Controlled Low-Resource Adaptation
von: Nur'aini, Khumaisa, et al.
Veröffentlicht: (2026)
von: Nur'aini, Khumaisa, et al.
Veröffentlicht: (2026)
LogQuant: Log-Distributed 2-Bit Quantization of KV Cache with Superior Accuracy Preservation
von: Chen, Han, et al.
Veröffentlicht: (2025)
von: Chen, Han, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
GEAR: A General Evaluation Framework for Abductive Reasoning
von: He, Kaiyu, et al.
Veröffentlicht: (2025) -
HiPRAG: Hierarchical Process Rewards for Efficient Agentic Retrieval Augmented Generation
von: Wu, Peilin, et al.
Veröffentlicht: (2025) -
Search Wisely: Mitigating Sub-optimal Agentic Searches By Reducing Uncertainty
von: Wu, Peilin, et al.
Veröffentlicht: (2025) -
IDEA: Enhancing the Rule Learning Ability of Large Language Model Agent through Induction, Deduction, and Abduction
von: He, Kaiyu, et al.
Veröffentlicht: (2024) -
From Reasoning to Learning: A Survey on Hypothesis Discovery and Rule Learning with Large Language Models
von: He, Kaiyu, et al.
Veröffentlicht: (2025)