Investigating Recurrent Transformers with Dynamic Halt
Fuente:
arXiv
Saved in:
| Main Authors: | Chowdhury, Jishnu Ray, Caragea, Cornelia |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Design Space Between Transformers and Recursive Neural Nets
by: Chowdhury, Jishnu Ray, et al.
Published: (2024)
by: Chowdhury, Jishnu Ray, et al.
Published: (2024)
Learning Sequence Attractors in Recurrent Networks with Hidden Neurons
by: Lu, Yao, et al.
Published: (2024)
by: Lu, Yao, et al.
Published: (2024)
Mamba-PTQ: Outlier Channels in Recurrent Large Language Models
by: Pierro, Alessandro, et al.
Published: (2024)
by: Pierro, Alessandro, et al.
Published: (2024)
Fuzzy Recurrent Stochastic Configuration Networks for Industrial Data Analytics
by: Wang, Dianhui, et al.
Published: (2024)
by: Wang, Dianhui, et al.
Published: (2024)
Scalable Learning in Structured Recurrent Spiking Neural Networks without Backpropagation
by: Tang, Bo, et al.
Published: (2026)
by: Tang, Bo, et al.
Published: (2026)
A Scalable Hybrid Training Approach for Recurrent Spiking Neural Networks
by: Baronig, Maximilian, et al.
Published: (2025)
by: Baronig, Maximilian, et al.
Published: (2025)
Sample-based Dynamic Hierarchical Transformer with Layer and Head Flexibility via Contextual Bandit
by: Meng, Fanfei, et al.
Published: (2023)
by: Meng, Fanfei, et al.
Published: (2023)
Recurrent Neural Networks Learn to Store and Generate Sequences using Non-Linear Representations
by: Csordás, Róbert, et al.
Published: (2024)
by: Csordás, Róbert, et al.
Published: (2024)
Structure of Artificial Neural Networks -- Empirical Investigations
by: Stier, Julian
Published: (2024)
by: Stier, Julian
Published: (2024)
Attending to Graph Transformers
by: Müller, Luis, et al.
Published: (2023)
by: Müller, Luis, et al.
Published: (2023)
Structure Development in List-Sorting Transformers
by: Urdshals, Einar, et al.
Published: (2025)
by: Urdshals, Einar, et al.
Published: (2025)
Empirical Investigation into Configuring Echo State Networks for Representative Benchmark Problem Domains
by: Weborg, Brooke R., et al.
Published: (2025)
by: Weborg, Brooke R., et al.
Published: (2025)
MoEUT: Mixture-of-Experts Universal Transformers
by: Csordás, Róbert, et al.
Published: (2024)
by: Csordás, Róbert, et al.
Published: (2024)
Understanding Transformer Optimization via Gradient Heterogeneity
by: Tomihari, Akiyoshi, et al.
Published: (2025)
by: Tomihari, Akiyoshi, et al.
Published: (2025)
Spiking Point Transformer for Point Cloud Classification
by: Wu, Peixi, et al.
Published: (2025)
by: Wu, Peixi, et al.
Published: (2025)
SGHormer: An Energy-Saving Graph Transformer Driven by Spikes
by: Zhang, Huizhe, et al.
Published: (2024)
by: Zhang, Huizhe, et al.
Published: (2024)
General-Purpose In-Context Learning by Meta-Learning Transformers
by: Kirsch, Louis, et al.
Published: (2022)
by: Kirsch, Louis, et al.
Published: (2022)
QSViT: A Methodology for Quantizing Spiking Vision Transformers
by: Putra, Rachmad Vidya Wicaksana, et al.
Published: (2025)
by: Putra, Rachmad Vidya Wicaksana, et al.
Published: (2025)
Decision SpikeFormer: Spike-Driven Transformer for Decision Making
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
The Dragon Hatchling: The Missing Link between the Transformer and Models of the Brain
by: Kosowski, Adrian, et al.
Published: (2025)
by: Kosowski, Adrian, et al.
Published: (2025)
Weight Decay Regimes in Grokking Transformers: Cheap Online Diagnostics
by: Verma, Lucky
Published: (2026)
by: Verma, Lucky
Published: (2026)
Why "classic" Transformers are shallow and how to make them go deep
by: Yu, Yueyao, et al.
Published: (2023)
by: Yu, Yueyao, et al.
Published: (2023)
Dynamic Reinforcement Learning for Actors
by: Shibata, Katsunari
Published: (2025)
by: Shibata, Katsunari
Published: (2025)
Sparsity Moves Computation: How FFN Architecture Reshapes Attention in Small Transformers
by: Smithline, Gabriel, et al.
Published: (2026)
by: Smithline, Gabriel, et al.
Published: (2026)
Dynamic Spiking Framework for Graph Neural Networks
by: Yin, Nan, et al.
Published: (2023)
by: Yin, Nan, et al.
Published: (2023)
Enabling Robust In-Context Memory and Rapid Task Adaptation in Transformers with Hebbian and Gradient-Based Plasticity
by: Chaudhary, Siddharth
Published: (2025)
by: Chaudhary, Siddharth
Published: (2025)
Dynamic Design of Machine Learning Pipelines via Metalearning
by: Alcobaça, Edesio, et al.
Published: (2025)
by: Alcobaça, Edesio, et al.
Published: (2025)
Decoding Listeners Identity: Person Identification from EEG Signals Using a Lightweight Spiking Transformer
by: Lin, Zheyuan, et al.
Published: (2025)
by: Lin, Zheyuan, et al.
Published: (2025)
Predicting Deterioration in Mild Cognitive Impairment with Survival Transformers, Extreme Gradient Boosting and Cox Proportional Hazard Modelling
by: Musto, Henry, et al.
Published: (2024)
by: Musto, Henry, et al.
Published: (2024)
Provably Optimal Memory Capacity for Modern Hopfield Models: Transformer-Compatible Dense Associative Memories as Spherical Codes
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
Synaptic Activation and Dual Liquid Dynamics for Interpretable Bio-Inspired Models
by: Farsang, Mónika, et al.
Published: (2026)
by: Farsang, Mónika, et al.
Published: (2026)
CosineGate: Semantic Dynamic Routing via Cosine Incompatibility in Residual Networks
by: Thota, Yogeswar Reddy
Published: (2025)
by: Thota, Yogeswar Reddy
Published: (2025)
Agent-GWO: Collaborative Agents for Dynamic Prompt Optimization in Large Language Models
by: Wang, Xudong, et al.
Published: (2026)
by: Wang, Xudong, et al.
Published: (2026)
Multiscale Astrocyte Network Calcium Dynamics for Biologically Plausible Intelligence in Anomaly Detection
by: Iskar, Berk, et al.
Published: (2025)
by: Iskar, Berk, et al.
Published: (2025)
Recursive Dynamics in Fast-Weights Homeostatic Reentry Networks: Toward Reflective Intelligence
by: Chae, B. G.
Published: (2025)
by: Chae, B. G.
Published: (2025)
GOAL: Graph-based Objective-Aligned Diffusion Solvers for Dynamic Multi-Objective Optimization
by: Li, Xingyu
Published: (2026)
by: Li, Xingyu
Published: (2026)
Izhikevich-Inspired Temporal Dynamics for Enhancing Privacy, Efficiency, and Transferability in Spiking Neural Networks
by: Moshruba, Ayana, et al.
Published: (2025)
by: Moshruba, Ayana, et al.
Published: (2025)
SiGNN: A Spike-induced Graph Neural Network for Dynamic Graph Representation Learning
by: Chen, Dong, et al.
Published: (2024)
by: Chen, Dong, et al.
Published: (2024)
Unveiling the Potential of Spiking Dynamics in Graph Representation Learning through Spatial-Temporal Normalization and Coding Strategies
by: Xu, Mingkun, et al.
Published: (2024)
by: Xu, Mingkun, et al.
Published: (2024)
Rethinking LLM-Driven Heuristic Design: Generating Efficient and Specialized Solvers via Dynamics-Aware Optimization
by: Wang, Rongzheng, et al.
Published: (2026)
by: Wang, Rongzheng, et al.
Published: (2026)
Similar Items
-
On the Design Space Between Transformers and Recursive Neural Nets
by: Chowdhury, Jishnu Ray, et al.
Published: (2024) -
Learning Sequence Attractors in Recurrent Networks with Hidden Neurons
by: Lu, Yao, et al.
Published: (2024) -
Mamba-PTQ: Outlier Channels in Recurrent Large Language Models
by: Pierro, Alessandro, et al.
Published: (2024) -
Fuzzy Recurrent Stochastic Configuration Networks for Industrial Data Analytics
by: Wang, Dianhui, et al.
Published: (2024) -
Scalable Learning in Structured Recurrent Spiking Neural Networks without Backpropagation
by: Tang, Bo, et al.
Published: (2026)