Understanding Dynamic Compute Allocation in Recurrent Transformers
Fuente:
arXiv
Salvato in:
| Autori principali: | Moosa, Ibraheem Muhammad, Lohit, Suhas, Wang, Ye, Chatterjee, Moitreya, Yin, Wenpeng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Programmatic Video Prediction Using Large Language Models
di: Tang, Hao, et al.
Pubblicazione: (2025)
di: Tang, Hao, et al.
Pubblicazione: (2025)
COMPACT: Common-token Optimized Model Pruning Across Channels and Tokens
di: Kwek, Eugene, et al.
Pubblicazione: (2025)
di: Kwek, Eugene, et al.
Pubblicazione: (2025)
A Systematic Study of Retrieval Pipeline Design for Retrieval-Augmented Medical Question Answering
di: Sultana, Nusrat, et al.
Pubblicazione: (2026)
di: Sultana, Nusrat, et al.
Pubblicazione: (2026)
Evaluating Large Vision-and-Language Models on Children's Mathematical Olympiads
di: Cherian, Anoop, et al.
Pubblicazione: (2024)
di: Cherian, Anoop, et al.
Pubblicazione: (2024)
MT-Ranker: Reference-free machine translation evaluation by inter-system ranking
di: Moosa, Ibraheem Muhammad, et al.
Pubblicazione: (2024)
di: Moosa, Ibraheem Muhammad, et al.
Pubblicazione: (2024)
Spectral Logit Sculpting: Adaptive Low-Rank Logit Transformation for Controlled Text Generation
di: Li, Jin, et al.
Pubblicazione: (2025)
di: Li, Jin, et al.
Pubblicazione: (2025)
Do Efficient Transformers Really Save Computation?
di: Yang, Kai, et al.
Pubblicazione: (2024)
di: Yang, Kai, et al.
Pubblicazione: (2024)
Direct-Inverse Prompting: Analyzing LLMs' Discriminative Capacity in Self-Improving Generation
di: Ahn, Jihyun Janice, et al.
Pubblicazione: (2024)
di: Ahn, Jihyun Janice, et al.
Pubblicazione: (2024)
Latent Chain-of-Thought? Decoding the Depth-Recurrent Transformer
di: Lu, Wenquan, et al.
Pubblicazione: (2025)
di: Lu, Wenquan, et al.
Pubblicazione: (2025)
Thinking Deeper, Not Longer: Depth-Recurrent Transformers for Compositional Generalization
di: Chen, Hung-Hsuan
Pubblicazione: (2026)
di: Chen, Hung-Hsuan
Pubblicazione: (2026)
Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers
di: Kohli, Harsh, et al.
Pubblicazione: (2026)
di: Kohli, Harsh, et al.
Pubblicazione: (2026)
Contrastive Instruction Tuning
di: Yan, Tianyi Lorena, et al.
Pubblicazione: (2024)
di: Yan, Tianyi Lorena, et al.
Pubblicazione: (2024)
Linear Chain Transformation: Expanding Optimization Dynamics for Fine-Tuning Large Language Models
di: Wang, Yulong, et al.
Pubblicazione: (2024)
di: Wang, Yulong, et al.
Pubblicazione: (2024)
CorrSynth -- A Correlated Sampling Method for Diverse Dataset Generation from LLMs
di: Kowshik, Suhas S, et al.
Pubblicazione: (2024)
di: Kowshik, Suhas S, et al.
Pubblicazione: (2024)
Budgeted LoRA: Distillation as Structured Compute Allocation for Efficient Inference
di: Sabry, Mohammed, et al.
Pubblicazione: (2026)
di: Sabry, Mohammed, et al.
Pubblicazione: (2026)
Learning How Hard to Think: Input-Adaptive Allocation of LM Computation
di: Damani, Mehul, et al.
Pubblicazione: (2024)
di: Damani, Mehul, et al.
Pubblicazione: (2024)
RecurrentGemma: Moving Past Transformers for Efficient Open Language Models
di: Botev, Aleksandar, et al.
Pubblicazione: (2024)
di: Botev, Aleksandar, et al.
Pubblicazione: (2024)
TableDART: Dynamic Adaptive Multi-Modal Routing for Table Understanding
di: Xing, Xiaobo, et al.
Pubblicazione: (2025)
di: Xing, Xiaobo, et al.
Pubblicazione: (2025)
Are Transformers Able to Reason by Connecting Separated Knowledge in Training Data?
di: Yin, Yutong, et al.
Pubblicazione: (2025)
di: Yin, Yutong, et al.
Pubblicazione: (2025)
Reframing Data Value for Large Language Models Through the Lens of Plausibility
di: Rammal, Mohamad Rida, et al.
Pubblicazione: (2024)
di: Rammal, Mohamad Rida, et al.
Pubblicazione: (2024)
Understanding Transformers via N-gram Statistics
di: Nguyen, Timothy
Pubblicazione: (2024)
di: Nguyen, Timothy
Pubblicazione: (2024)
Adaptive Computation Pruning for the Forgetting Transformer
di: Lin, Zhixuan, et al.
Pubblicazione: (2025)
di: Lin, Zhixuan, et al.
Pubblicazione: (2025)
When More is Less: Understanding Chain-of-Thought Length in LLMs
di: Wu, Yuyang, et al.
Pubblicazione: (2025)
di: Wu, Yuyang, et al.
Pubblicazione: (2025)
Associative Recurrent Memory Transformer
di: Rodkin, Ivan, et al.
Pubblicazione: (2024)
di: Rodkin, Ivan, et al.
Pubblicazione: (2024)
Compute-Constrained Data Selection
di: Yin, Junjie Oscar, et al.
Pubblicazione: (2024)
di: Yin, Junjie Oscar, et al.
Pubblicazione: (2024)
Understanding Chain-of-Thought in LLMs through Information Theory
di: Ton, Jean-Francois, et al.
Pubblicazione: (2024)
di: Ton, Jean-Francois, et al.
Pubblicazione: (2024)
Hidden Heroes and Gradient Bloats: Layer-Wise Redundancy Inverts Attribution in Transformers
di: Ye, Donald
Pubblicazione: (2026)
di: Ye, Donald
Pubblicazione: (2026)
Understanding Reasoning in Chain-of-Thought from the Hopfieldian View
di: Hu, Lijie, et al.
Pubblicazione: (2024)
di: Hu, Lijie, et al.
Pubblicazione: (2024)
AssemblyBench: Physics-Aware Assembly of Complex Industrial Objects
di: Li, Danrui, et al.
Pubblicazione: (2026)
di: Li, Danrui, et al.
Pubblicazione: (2026)
Recurrent Knowledge Identification and Fusion for Language Model Continual Learning
di: Feng, Yujie, et al.
Pubblicazione: (2025)
di: Feng, Yujie, et al.
Pubblicazione: (2025)
Dynamic Universal Approximation Theory: The Basic Theory for Transformer-based Large Language Models
di: Wang, Wei, et al.
Pubblicazione: (2024)
di: Wang, Wei, et al.
Pubblicazione: (2024)
CAST: Compositional Analysis via Spectral Tracking for Understanding Transformer Layer Functions
di: Fu, Zihao, et al.
Pubblicazione: (2025)
di: Fu, Zihao, et al.
Pubblicazione: (2025)
Optimal Decay Spectra for Linear Recurrences
di: Cao, Yang
Pubblicazione: (2026)
di: Cao, Yang
Pubblicazione: (2026)
JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention
di: Tian, Yuandong, et al.
Pubblicazione: (2023)
di: Tian, Yuandong, et al.
Pubblicazione: (2023)
Variation in Verification: Understanding Verification Dynamics in Large Language Models
di: Zhou, Yefan, et al.
Pubblicazione: (2025)
di: Zhou, Yefan, et al.
Pubblicazione: (2025)
Understanding In-context Learning of Addition via Activation Subspaces
di: Hu, Xinyan, et al.
Pubblicazione: (2025)
di: Hu, Xinyan, et al.
Pubblicazione: (2025)
Attention Needs to Focus: A Unified Perspective on Attention Allocation
di: Fu, Zichuan, et al.
Pubblicazione: (2026)
di: Fu, Zichuan, et al.
Pubblicazione: (2026)
SafePred: A Predictive Guardrail for Computer-Using Agents via World Models
di: Chen, Yurun, et al.
Pubblicazione: (2026)
di: Chen, Yurun, et al.
Pubblicazione: (2026)
Unifying Learning Dynamics and Generalization in Transformers Scaling Law
di: Yang, Chiwun
Pubblicazione: (2025)
di: Yang, Chiwun
Pubblicazione: (2025)
MGM: Global Understanding of Audience Overlap Graphs for Predicting the Factuality and the Bias of News Media
di: Manzoor, Muhammad Arslan, et al.
Pubblicazione: (2024)
di: Manzoor, Muhammad Arslan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Programmatic Video Prediction Using Large Language Models
di: Tang, Hao, et al.
Pubblicazione: (2025) -
COMPACT: Common-token Optimized Model Pruning Across Channels and Tokens
di: Kwek, Eugene, et al.
Pubblicazione: (2025) -
A Systematic Study of Retrieval Pipeline Design for Retrieval-Augmented Medical Question Answering
di: Sultana, Nusrat, et al.
Pubblicazione: (2026) -
Evaluating Large Vision-and-Language Models on Children's Mathematical Olympiads
di: Cherian, Anoop, et al.
Pubblicazione: (2024) -
MT-Ranker: Reference-free machine translation evaluation by inter-system ranking
di: Moosa, Ibraheem Muhammad, et al.
Pubblicazione: (2024)