Working Memory Constraints Scaffold Learning in Transformers under Data Scarcity
Fuente:
arXiv
Saved in:
| Main Authors: | Madhyastha, Pranava, Adamcova, Dagmar |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Cognitively Grounded Bayesian Framework for Misinformation Susceptibility
by: Madhyastha, Pranava
Published: (2026)
by: Madhyastha, Pranava
Published: (2026)
$\texttt{SEM-CTRL}$: Semantically Controlled Decoding
by: Albinhassan, Mohammad, et al.
Published: (2025)
by: Albinhassan, Mohammad, et al.
Published: (2025)
Learning and Enforcing Context-Sensitive Control for LLMs
by: Albinhassan, Mohammad, et al.
Published: (2026)
by: Albinhassan, Mohammad, et al.
Published: (2026)
An Empirical Study of Conformal Prediction in LLM with ASP Scaffolds for Robust Reasoning
by: Kaur, Navdeep, et al.
Published: (2025)
by: Kaur, Navdeep, et al.
Published: (2025)
Think Before You Act: Decision Transformers with Working Memory
by: Kang, Jikun, et al.
Published: (2023)
by: Kang, Jikun, et al.
Published: (2023)
HiTZ at VarDial 2025 NorSID: Overcoming Data Scarcity with Language Transfer and Automatic Data Annotation
by: Bengoetxea, Jaione, et al.
Published: (2024)
by: Bengoetxea, Jaione, et al.
Published: (2024)
InvestAlign: Overcoming Data Scarcity in Aligning Large Language Models with Investor Decision-Making Processes under Herd Behavior
by: Wang, Huisheng, et al.
Published: (2025)
by: Wang, Huisheng, et al.
Published: (2025)
An Evolved Universal Transformer Memory
by: Cetin, Edoardo, et al.
Published: (2024)
by: Cetin, Edoardo, et al.
Published: (2024)
Mixture of Chapters: Scaling Learnt Memory in Transformers
by: Tibrewal, Tasmay Pankaj, et al.
Published: (2026)
by: Tibrewal, Tasmay Pankaj, et al.
Published: (2026)
Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models
by: Vendrell, Victor Conchello, et al.
Published: (2026)
by: Vendrell, Victor Conchello, et al.
Published: (2026)
MetaSynth: Meta-Prompting-Driven Agentic Scaffolds for Diverse Synthetic Data Generation
by: Riaz, Haris, et al.
Published: (2025)
by: Riaz, Haris, et al.
Published: (2025)
MetaState: Persistent Working Memory Enhances Reasoning in Discrete Diffusion Language Models
by: Xia, Kejing, et al.
Published: (2026)
by: Xia, Kejing, et al.
Published: (2026)
POET-X: Memory-efficient LLM Training by Scaling Orthogonal Transformation
by: Qiu, Zeju, et al.
Published: (2026)
by: Qiu, Zeju, et al.
Published: (2026)
HGMEM: Hypergraph-based Working Memory to Improve Multi-step RAG for Long-Context Complex Relational Modeling
by: Zhou, Chulun, et al.
Published: (2025)
by: Zhou, Chulun, et al.
Published: (2025)
Building Safe and Deployable Clinical Natural Language Processing under Temporal Leakage Constraints
by: Cho, Ha Na, et al.
Published: (2026)
by: Cho, Ha Na, et al.
Published: (2026)
Venn Diagram Prompting : Accelerating Comprehension with Scaffolding Effect
by: Mahendru, Sakshi, et al.
Published: (2024)
by: Mahendru, Sakshi, et al.
Published: (2024)
Emergence of Episodic Memory in Transformers: Characterizing Changes in Temporal Structure of Attention Scores During Training
by: Mistry, Deven Mahesh, et al.
Published: (2025)
by: Mistry, Deven Mahesh, et al.
Published: (2025)
Language Diffusion Models are Associative Memories Capable of Retrieving Unseen Data
by: Pham, Bao, et al.
Published: (2026)
by: Pham, Bao, et al.
Published: (2026)
Memory Injections: Correcting Multi-Hop Reasoning Failures during Inference in Transformer-Based Language Models
by: Sakarvadia, Mansi, et al.
Published: (2023)
by: Sakarvadia, Mansi, et al.
Published: (2023)
Transformers Struggle to Learn to Search
by: Saparov, Abulhair, et al.
Published: (2024)
by: Saparov, Abulhair, et al.
Published: (2024)
Mini-batch Coresets for Memory-efficient Language Model Training on Data Mixtures
by: Nguyen, Dang, et al.
Published: (2024)
by: Nguyen, Dang, et al.
Published: (2024)
Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning
by: Zhang, Xichen, et al.
Published: (2025)
by: Zhang, Xichen, et al.
Published: (2025)
From Instructions to Constraints: Language Model Alignment with Automatic Constraint Verification
by: Wang, Fei, et al.
Published: (2024)
by: Wang, Fei, et al.
Published: (2024)
Language Model Memory and Memory Models for Language
by: Badger, Benjamin L.
Published: (2026)
by: Badger, Benjamin L.
Published: (2026)
Learning Query-Aware Budget-Tier Routing for Runtime Agent Memory
by: Zhang, Haozhen, et al.
Published: (2026)
by: Zhang, Haozhen, et al.
Published: (2026)
MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents
by: Zhang, Haozhen, et al.
Published: (2026)
by: Zhang, Haozhen, et al.
Published: (2026)
Are Transformers Able to Reason by Connecting Separated Knowledge in Training Data?
by: Yin, Yutong, et al.
Published: (2025)
by: Yin, Yutong, et al.
Published: (2025)
Classification of Hope in Textual Data using Transformer-Based Models
by: Ijezue, Chukwuebuka Fortunate, et al.
Published: (2025)
by: Ijezue, Chukwuebuka Fortunate, et al.
Published: (2025)
Theoretical Constraints on the Expressive Power of $\mathsf{RoPE}$-based Tensor Attention Transformers
by: Li, Xiaoyu, et al.
Published: (2024)
by: Li, Xiaoyu, et al.
Published: (2024)
Time and Memory Trade-off of KV-Cache Compression in Tensor Transformer Decoding
by: Chen, Yifang, et al.
Published: (2025)
by: Chen, Yifang, et al.
Published: (2025)
Learning a Decision Tree Algorithm with Transformers
by: Zhuang, Yufan, et al.
Published: (2024)
by: Zhuang, Yufan, et al.
Published: (2024)
FOREVER: Forgetting Curve-Inspired Memory Replay for Language Model Continual Learning
by: Feng, Yujie, et al.
Published: (2026)
by: Feng, Yujie, et al.
Published: (2026)
SibylSense: Adaptive Rubric Learning via Memory Tuning and Adversarial Probing
by: Xu, Yifei, et al.
Published: (2026)
by: Xu, Yifei, et al.
Published: (2026)
Self-Supervised Transformers as Iterative Solution Improvers for Constraint Satisfaction
by: Xu, Yudong W., et al.
Published: (2025)
by: Xu, Yudong W., et al.
Published: (2025)
TourPlanner: A Competitive Consensus Framework with Constraint-Gated Reinforcement Learning for Travel Planning
by: Wang, Yinuo, et al.
Published: (2026)
by: Wang, Yinuo, et al.
Published: (2026)
JTCSE: Joint Tensor-Modulus Constraints and Cross-Attention for Unsupervised Contrastive Learning of Sentence Embeddings
by: Zong, Tianyu, et al.
Published: (2025)
by: Zong, Tianyu, et al.
Published: (2025)
Learning to Reason under Off-Policy Guidance
by: Yan, Jianhao, et al.
Published: (2025)
by: Yan, Jianhao, et al.
Published: (2025)
Associative Recurrent Memory Transformer
by: Rodkin, Ivan, et al.
Published: (2024)
by: Rodkin, Ivan, et al.
Published: (2024)
MolMem: Memory-Augmented Agentic Reinforcement Learning for Sample-Efficient Molecular Optimization
by: Wang, Ziqing, et al.
Published: (2026)
by: Wang, Ziqing, et al.
Published: (2026)
Continual Knowledge Updating in LLM Systems: Learning Through Multi-Timescale Memory Dynamics
by: Pattichis, Andreas, et al.
Published: (2026)
by: Pattichis, Andreas, et al.
Published: (2026)
Similar Items
-
A Cognitively Grounded Bayesian Framework for Misinformation Susceptibility
by: Madhyastha, Pranava
Published: (2026) -
$\texttt{SEM-CTRL}$: Semantically Controlled Decoding
by: Albinhassan, Mohammad, et al.
Published: (2025) -
Learning and Enforcing Context-Sensitive Control for LLMs
by: Albinhassan, Mohammad, et al.
Published: (2026) -
An Empirical Study of Conformal Prediction in LLM with ASP Scaffolds for Robust Reasoning
by: Kaur, Navdeep, et al.
Published: (2025) -
Think Before You Act: Decision Transformers with Working Memory
by: Kang, Jikun, et al.
Published: (2023)