Complexity Control Facilitates Reasoning-Based Compositional Generalization in Transformers
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Zhongwang, Lin, Pengxiao, Wang, Zhiwei, Zhang, Yaoyu, Xu, Zhi-Qin John |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Reasoning Bias of Next Token Prediction Training
di: Lin, Pengxiao, et al.
Pubblicazione: (2025)
di: Lin, Pengxiao, et al.
Pubblicazione: (2025)
Initialization is Critical to Whether Transformers Fit Composite Functions by Reasoning or Memorizing
di: Zhang, Zhongwang, et al.
Pubblicazione: (2024)
di: Zhang, Zhongwang, et al.
Pubblicazione: (2024)
An Analysis for Reasoning Bias of Language Models with Small Initialization
di: Yao, Junjie, et al.
Pubblicazione: (2025)
di: Yao, Junjie, et al.
Pubblicazione: (2025)
Scalable Complexity Control Facilitates Reasoning Ability of LLMs
di: Hang, Liangkai, et al.
Pubblicazione: (2025)
di: Hang, Liangkai, et al.
Pubblicazione: (2025)
Understanding the Language Model to Solve the Symbolic Multi-Step Reasoning Problem from the Perspective of Buffer Mechanism
di: Wang, Zhiwei, et al.
Pubblicazione: (2024)
di: Wang, Zhiwei, et al.
Pubblicazione: (2024)
Anchor function: a type of benchmark functions for studying language models
di: Zhang, Zhongwang, et al.
Pubblicazione: (2024)
di: Zhang, Zhongwang, et al.
Pubblicazione: (2024)
Local Linear Recovery Guarantee of Deep Neural Networks at Overparameterization
di: Zhang, Yaoyu, et al.
Pubblicazione: (2024)
di: Zhang, Yaoyu, et al.
Pubblicazione: (2024)
Achilles' Heel of Mamba: Essential difficulties of the Mamba architecture demonstrated by synthetic data
di: Chen, Tianyi, et al.
Pubblicazione: (2025)
di: Chen, Tianyi, et al.
Pubblicazione: (2025)
Identity Bridge: Enabling Implicit Reasoning via Shared Latent Memory
di: Lin, Pengxiao, et al.
Pubblicazione: (2025)
di: Lin, Pengxiao, et al.
Pubblicazione: (2025)
Loss Jump During Loss Switch in Solving PDEs with Neural Networks
di: Wang, Zhiwei, et al.
Pubblicazione: (2024)
di: Wang, Zhiwei, et al.
Pubblicazione: (2024)
A rationale from frequency perspective for grokking in training neural network
di: Zhou, Zhangchen, et al.
Pubblicazione: (2024)
di: Zhou, Zhangchen, et al.
Pubblicazione: (2024)
ComplexFormer: Disruptively Advancing Transformer Inference Ability via Head-Specific Complex Vector Attention
di: Shao, Jintian, et al.
Pubblicazione: (2025)
di: Shao, Jintian, et al.
Pubblicazione: (2025)
Loss Spike in Training Neural Networks
di: Li, Xiaolong, et al.
Pubblicazione: (2023)
di: Li, Xiaolong, et al.
Pubblicazione: (2023)
Embedding Principle in Depth for the Loss Landscape Analysis of Deep Neural Networks
di: Bai, Zhiwei, et al.
Pubblicazione: (2022)
di: Bai, Zhiwei, et al.
Pubblicazione: (2022)
Feature Resemblance: Towards a Theoretical Understanding of Analogical Reasoning in Transformers
di: Xu, Ruichen, et al.
Pubblicazione: (2026)
di: Xu, Ruichen, et al.
Pubblicazione: (2026)
Generative Evaluation of Complex Reasoning in Large Language Models
di: Lin, Haowei, et al.
Pubblicazione: (2025)
di: Lin, Haowei, et al.
Pubblicazione: (2025)
Concept Algebra for (Score-Based) Text-Controlled Generative Models
di: Wang, Zihao, et al.
Pubblicazione: (2023)
di: Wang, Zihao, et al.
Pubblicazione: (2023)
Compositional Reasoning with Transformers, RNNs, and Chain of Thought
di: Yehudai, Gilad, et al.
Pubblicazione: (2025)
di: Yehudai, Gilad, et al.
Pubblicazione: (2025)
Self-Infilling Code Generation
di: Zheng, Lin, et al.
Pubblicazione: (2023)
di: Zheng, Lin, et al.
Pubblicazione: (2023)
Generalization or Hallucination? Understanding Out-of-Context Reasoning in Transformers
di: Huang, Yixiao, et al.
Pubblicazione: (2025)
di: Huang, Yixiao, et al.
Pubblicazione: (2025)
XFinBench: Benchmarking LLMs in Complex Financial Problem Solving and Reasoning
di: Zhang, Zhihan, et al.
Pubblicazione: (2025)
di: Zhang, Zhihan, et al.
Pubblicazione: (2025)
Focus and Dilution: The Multi-stage Learning Process of Attention
di: Chen, Zheng-An, et al.
Pubblicazione: (2026)
di: Chen, Zheng-An, et al.
Pubblicazione: (2026)
A Complexity-Based Theory of Compositionality
di: Elmoznino, Eric, et al.
Pubblicazione: (2024)
di: Elmoznino, Eric, et al.
Pubblicazione: (2024)
Beyond Autoregression: Discrete Diffusion for Complex Reasoning and Planning
di: Ye, Jiacheng, et al.
Pubblicazione: (2024)
di: Ye, Jiacheng, et al.
Pubblicazione: (2024)
Complex Logical Instruction Generation
di: Zhang, Mian, et al.
Pubblicazione: (2025)
di: Zhang, Mian, et al.
Pubblicazione: (2025)
Circuit Compositions: Exploring Modular Structures in Transformer-Based Language Models
di: Mondorf, Philipp, et al.
Pubblicazione: (2024)
di: Mondorf, Philipp, et al.
Pubblicazione: (2024)
Reinforcing General Reasoning without Verifiers
di: Zhou, Xiangxin, et al.
Pubblicazione: (2025)
di: Zhou, Xiangxin, et al.
Pubblicazione: (2025)
Light-IF: Endowing LLMs with Generalizable Reasoning via Preview and Self-Checking for Complex Instruction Following
di: Wang, Chenyang, et al.
Pubblicazione: (2025)
di: Wang, Chenyang, et al.
Pubblicazione: (2025)
Retrieval-Reasoning Large Language Model-based Synthetic Clinical Trial Generation
di: Xu, Zerui, et al.
Pubblicazione: (2024)
di: Xu, Zerui, et al.
Pubblicazione: (2024)
Weights to Code: Extracting Interpretable Algorithms from the Discrete Transformer
di: Zhang, Yifan, et al.
Pubblicazione: (2026)
di: Zhang, Yifan, et al.
Pubblicazione: (2026)
TRACE Back from the Future: A Probabilistic Reasoning Approach to Controllable Language Generation
di: Weng, Gwen Yidou, et al.
Pubblicazione: (2025)
di: Weng, Gwen Yidou, et al.
Pubblicazione: (2025)
Reasoning in Transformers -- Mitigating Spurious Correlations and Reasoning Shortcuts
di: Enström, Daniel, et al.
Pubblicazione: (2024)
di: Enström, Daniel, et al.
Pubblicazione: (2024)
Gained in Translation: Privileged Pairwise Judges Enhance Multilingual Reasoning
di: Sutawika, Lintang, et al.
Pubblicazione: (2026)
di: Sutawika, Lintang, et al.
Pubblicazione: (2026)
Interpreting and Controlling LLM Reasoning through Integrated Policy Gradient
di: Li, Changming, et al.
Pubblicazione: (2026)
di: Li, Changming, et al.
Pubblicazione: (2026)
On the Query Complexity of Verifier-Assisted Language Generation
di: Botta, Edoardo, et al.
Pubblicazione: (2025)
di: Botta, Edoardo, et al.
Pubblicazione: (2025)
Benchmarking and Understanding Compositional Relational Reasoning of LLMs
di: Ni, Ruikang, et al.
Pubblicazione: (2024)
di: Ni, Ruikang, et al.
Pubblicazione: (2024)
An overview of condensation phenomenon in deep learning
di: Xu, Zhi-Qin John, et al.
Pubblicazione: (2025)
di: Xu, Zhi-Qin John, et al.
Pubblicazione: (2025)
Overview frequency principle/spectral bias in deep learning
di: Xu, Zhi-Qin John, et al.
Pubblicazione: (2022)
di: Xu, Zhi-Qin John, et al.
Pubblicazione: (2022)
Laying the Foundation First? Investigating the Generalization from Atomic Skills to Complex Reasoning Tasks
di: Huang, Yuncheng, et al.
Pubblicazione: (2024)
di: Huang, Yuncheng, et al.
Pubblicazione: (2024)
Transformers meet Neural Algorithmic Reasoners
di: Bounsi, Wilfried, et al.
Pubblicazione: (2024)
di: Bounsi, Wilfried, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Reasoning Bias of Next Token Prediction Training
di: Lin, Pengxiao, et al.
Pubblicazione: (2025) -
Initialization is Critical to Whether Transformers Fit Composite Functions by Reasoning or Memorizing
di: Zhang, Zhongwang, et al.
Pubblicazione: (2024) -
An Analysis for Reasoning Bias of Language Models with Small Initialization
di: Yao, Junjie, et al.
Pubblicazione: (2025) -
Scalable Complexity Control Facilitates Reasoning Ability of LLMs
di: Hang, Liangkai, et al.
Pubblicazione: (2025) -
Understanding the Language Model to Solve the Symbolic Multi-Step Reasoning Problem from the Perspective of Buffer Mechanism
di: Wang, Zhiwei, et al.
Pubblicazione: (2024)