Adaptive Loops and Memory in Transformers: Think Harder or Know More?
Fuente:
arXiv
Salvato in:
| Autori principali: | Frey, Markus, Shomali, Behzad, Bashir, Ali Hamza, Berghaus, David, Koehler, Joachim, Ali, Mehdi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Dual-Path Architecture for Scaling Compute and Capacity in LLMs
di: Frey, Markus, et al.
Pubblicazione: (2026)
di: Frey, Markus, et al.
Pubblicazione: (2026)
Is continuous CoT better suited for multi-lingual reasoning?
di: Bashir, Ali Hamza, et al.
Pubblicazione: (2026)
di: Bashir, Ali Hamza, et al.
Pubblicazione: (2026)
Domain-Adaptation through Synthetic Data: Fine-Tuning Large Language Models for German Law
di: Bashir, Ali Hamza, et al.
Pubblicazione: (2026)
di: Bashir, Ali Hamza, et al.
Pubblicazione: (2026)
Code-Switched Language Identification is Harder Than You Think
di: Burchell, Laurie, et al.
Pubblicazione: (2024)
di: Burchell, Laurie, et al.
Pubblicazione: (2024)
Think Multilingual, Not Harder: A Data-Efficient Framework for Teaching Reasoning Models to Code-Switch
di: Lin, Eleanor M., et al.
Pubblicazione: (2026)
di: Lin, Eleanor M., et al.
Pubblicazione: (2026)
Think Less, Know More: State-Aware Reasoning Compression with Knowledge Guidance for Efficient Reasoning
di: Sui, Yi, et al.
Pubblicazione: (2026)
di: Sui, Yi, et al.
Pubblicazione: (2026)
Once Upon a Time: Interactive Learning for Storytelling with Small Language Models
di: Martins, Jonas Mayer, et al.
Pubblicazione: (2025)
di: Martins, Jonas Mayer, et al.
Pubblicazione: (2025)
Harder Tasks Need More Experts: Dynamic Routing in MoE Models
di: Huang, Quzhe, et al.
Pubblicazione: (2024)
di: Huang, Quzhe, et al.
Pubblicazione: (2024)
The Harder The Better: Maintaining Supervised Fine-tuning Generalization with Less but Harder Data
di: Shang, Zhaoyang, et al.
Pubblicazione: (2025)
di: Shang, Zhaoyang, et al.
Pubblicazione: (2025)
Think Beyond Size: Adaptive Prompting for More Effective Reasoning
di: R, Kamesh
Pubblicazione: (2024)
di: R, Kamesh
Pubblicazione: (2024)
The Zero-Step Thinking: An Empirical Study of Mode Selection as Harder Early Exit in Reasoning Models
di: Tan, Yuqiao, et al.
Pubblicazione: (2025)
di: Tan, Yuqiao, et al.
Pubblicazione: (2025)
Inner Thinking Transformer: Leveraging Dynamic Depth Scaling to Foster Adaptive Internal Thinking
di: Chen, Yilong, et al.
Pubblicazione: (2025)
di: Chen, Yilong, et al.
Pubblicazione: (2025)
Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models
di: Vendrell, Victor Conchello, et al.
Pubblicazione: (2026)
di: Vendrell, Victor Conchello, et al.
Pubblicazione: (2026)
Do We Know What LLMs Don't Know? A Study of Consistency in Knowledge Probing
di: Zhao, Raoyuan, et al.
Pubblicazione: (2025)
di: Zhao, Raoyuan, et al.
Pubblicazione: (2025)
Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers
di: Kohli, Harsh, et al.
Pubblicazione: (2026)
di: Kohli, Harsh, et al.
Pubblicazione: (2026)
Read More, Think More: Revisiting Observation Reduction for Web Agents
di: Enomoto, Masafumi, et al.
Pubblicazione: (2026)
di: Enomoto, Masafumi, et al.
Pubblicazione: (2026)
LLMs Know More About Numbers than They Can Say
di: Yuchi, Fengting, et al.
Pubblicazione: (2026)
di: Yuchi, Fengting, et al.
Pubblicazione: (2026)
What Large Language Models Know and What People Think They Know
di: Steyvers, Mark, et al.
Pubblicazione: (2024)
di: Steyvers, Mark, et al.
Pubblicazione: (2024)
Using Perspectival Words Is Harder Than Vocabulary Words for Humans and Even More So for Multimodal Language Models
di: Dong, Dota Tianai, et al.
Pubblicazione: (2025)
di: Dong, Dota Tianai, et al.
Pubblicazione: (2025)
Can Thinking Models Think to Detect Hateful Memes?
di: Kmainasi, Mohamed Bayan, et al.
Pubblicazione: (2026)
di: Kmainasi, Mohamed Bayan, et al.
Pubblicazione: (2026)
S2D: Sorted Speculative Decoding For More Efficient Deployment of Nested Large Language Models
di: Kavehzadeh, Parsa, et al.
Pubblicazione: (2024)
di: Kavehzadeh, Parsa, et al.
Pubblicazione: (2024)
To Know is to Construct: Schema-Constrained Generation for Agent Memory
di: Zheng, Lei, et al.
Pubblicazione: (2026)
di: Zheng, Lei, et al.
Pubblicazione: (2026)
MirrorStories: Reflecting Diversity through Personalized Narrative Generation with Large Language Models
di: Yunusov, Sarfaroz, et al.
Pubblicazione: (2024)
di: Yunusov, Sarfaroz, et al.
Pubblicazione: (2024)
Reasoning Gets Harder for LLMs Inside A Dialogue
di: Kartáč, Ivan, et al.
Pubblicazione: (2026)
di: Kartáč, Ivan, et al.
Pubblicazione: (2026)
Know More, Know Clearer: A Meta-Cognitive Framework for Knowledge Augmentation in Large Language Models
di: Chen, Hao, et al.
Pubblicazione: (2026)
di: Chen, Hao, et al.
Pubblicazione: (2026)
Think More, Hallucinate Less: Mitigating Hallucinations via Dual Process of Fast and Slow Thinking
di: Cheng, Xiaoxue, et al.
Pubblicazione: (2025)
di: Cheng, Xiaoxue, et al.
Pubblicazione: (2025)
Think Before You Act: Decision Transformers with Working Memory
di: Kang, Jikun, et al.
Pubblicazione: (2023)
di: Kang, Jikun, et al.
Pubblicazione: (2023)
Thinking Out Loud: Do Reasoning Models Know When They're Right?
di: Zeng, Qingcheng, et al.
Pubblicazione: (2025)
di: Zeng, Qingcheng, et al.
Pubblicazione: (2025)
Overtrained Language Models Are Harder to Fine-Tune
di: Springer, Jacob Mitchell, et al.
Pubblicazione: (2025)
di: Springer, Jacob Mitchell, et al.
Pubblicazione: (2025)
Memory Dial: A Training Framework for Controllable Memorization in Language Models
di: Zhang, Xiangbo, et al.
Pubblicazione: (2026)
di: Zhang, Xiangbo, et al.
Pubblicazione: (2026)
EchoAtt: Attend, Copy, then Adjust for More Efficient Large Language Models
di: Rajabzadeh, Hossein, et al.
Pubblicazione: (2024)
di: Rajabzadeh, Hossein, et al.
Pubblicazione: (2024)
STALE: Can LLM Agents Know When Their Memories Are No Longer Valid?
di: Chao, Hanxiang, et al.
Pubblicazione: (2026)
di: Chao, Hanxiang, et al.
Pubblicazione: (2026)
Towards More Standardized AI Evaluation: From Models to Agents
di: Filali, Ali El, et al.
Pubblicazione: (2026)
di: Filali, Ali El, et al.
Pubblicazione: (2026)
Think Right, Not More: Test-Time Scaling for Numerical Claim Verification
di: Chungkham, Primakov, et al.
Pubblicazione: (2025)
di: Chungkham, Primakov, et al.
Pubblicazione: (2025)
Do LLMs Know They Are Being Tested? Evaluation Awareness and Incentive-Sensitive Failures in GPT-OSS-20B
di: Ahmed, Nisar, et al.
Pubblicazione: (2025)
di: Ahmed, Nisar, et al.
Pubblicazione: (2025)
UrduBench: An Urdu Reasoning Benchmark using Contextually Ensembled Translations with Human-in-the-Loop
di: Shafique, Muhammad Ali, et al.
Pubblicazione: (2026)
di: Shafique, Muhammad Ali, et al.
Pubblicazione: (2026)
GenKnowSub: Improving Modularity and Reusability of LLMs through General Knowledge Subtraction
di: Bagherifard, Mohammadtaha, et al.
Pubblicazione: (2025)
di: Bagherifard, Mohammadtaha, et al.
Pubblicazione: (2025)
"Knowing When You Don't Know": A Multilingual Relevance Assessment Dataset for Robust Retrieval-Augmented Generation
di: Thakur, Nandan, et al.
Pubblicazione: (2023)
di: Thakur, Nandan, et al.
Pubblicazione: (2023)
Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing
di: Berghaus, David, et al.
Pubblicazione: (2025)
di: Berghaus, David, et al.
Pubblicazione: (2025)
Can Language Models Be Tricked by Language Illusions? Easier with Syntax, Harder with Semantics
di: Zhang, Yuhan, et al.
Pubblicazione: (2023)
di: Zhang, Yuhan, et al.
Pubblicazione: (2023)
Documenti analoghi
-
A Dual-Path Architecture for Scaling Compute and Capacity in LLMs
di: Frey, Markus, et al.
Pubblicazione: (2026) -
Is continuous CoT better suited for multi-lingual reasoning?
di: Bashir, Ali Hamza, et al.
Pubblicazione: (2026) -
Domain-Adaptation through Synthetic Data: Fine-Tuning Large Language Models for German Law
di: Bashir, Ali Hamza, et al.
Pubblicazione: (2026) -
Code-Switched Language Identification is Harder Than You Think
di: Burchell, Laurie, et al.
Pubblicazione: (2024) -
Think Multilingual, Not Harder: A Data-Efficient Framework for Teaching Reasoning Models to Code-Switch
di: Lin, Eleanor M., et al.
Pubblicazione: (2026)