Interpreting Affine Recurrence Learning in GPT-style Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Bhargav, Samarth, Gu, Alexander |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Transfer Learning of Surrogate Models: Integrating Domain Warping and Affine Transformations
by: Pan, Shuaiqun, et al.
Published: (2025)
by: Pan, Shuaiqun, et al.
Published: (2025)
Likelihood-based Sensor Calibration using Affine Transformation
by: Machhamer, Rüdiger, et al.
Published: (2023)
by: Machhamer, Rüdiger, et al.
Published: (2023)
nGPT: Normalized Transformer with Representation Learning on the Hypersphere
by: Loshchilov, Ilya, et al.
Published: (2024)
by: Loshchilov, Ilya, et al.
Published: (2024)
Recurrent Action Transformer with Memory
by: Cherepanov, Egor, et al.
Published: (2023)
by: Cherepanov, Egor, et al.
Published: (2023)
DeepFilter: A Transformer-style Framework for Accurate and Efficient Process Monitoring
by: Wang, Hao, et al.
Published: (2025)
by: Wang, Hao, et al.
Published: (2025)
Transfer Learning of Surrogate Models via Domain Affine Transformation Across Synthetic and Real-World Benchmarks
by: Pan, Shuaiqun, et al.
Published: (2025)
by: Pan, Shuaiqun, et al.
Published: (2025)
Feature-Guided SAE Steering for Refusal-Rate Control using Contrasting Prompts
by: Bhargav, Samaksh, et al.
Published: (2025)
by: Bhargav, Samaksh, et al.
Published: (2025)
READ: Recurrent Adaptation of Large Transformers
by: Nguyen, John, et al.
Published: (2023)
by: Nguyen, John, et al.
Published: (2023)
Protein Secondary Structure Prediction Using 3D Graphs and Relation-Aware Message Passing Transformers
by: Varshney, Disha, et al.
Published: (2025)
by: Varshney, Disha, et al.
Published: (2025)
From Alexnet to Transformers: Measuring the Non-linearity of Deep Neural Networks with Affine Optimal Transport
by: Bouniot, Quentin, et al.
Published: (2023)
by: Bouniot, Quentin, et al.
Published: (2023)
StyleBench: Evaluating thinking styles in Large Language Models
by: Guo, Junyu, et al.
Published: (2025)
by: Guo, Junyu, et al.
Published: (2025)
Recurrent Reinforcement Learning with Memoroids
by: Morad, Steven, et al.
Published: (2024)
by: Morad, Steven, et al.
Published: (2024)
RingFormer: Rethinking Recurrent Transformer with Adaptive Level Signals
by: Heo, Jaemu, et al.
Published: (2025)
by: Heo, Jaemu, et al.
Published: (2025)
GraphGPT: Generative Pre-trained Graph Eulerian Transformer
by: Zhao, Qifang, et al.
Published: (2023)
by: Zhao, Qifang, et al.
Published: (2023)
Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism
by: Bick, Aviv, et al.
Published: (2025)
by: Bick, Aviv, et al.
Published: (2025)
GPU Memory Requirement Prediction for Deep Learning Task Based on Bidirectional Gated Recurrent Unit Optimization Transformer
by: Wang, Chao, et al.
Published: (2025)
by: Wang, Chao, et al.
Published: (2025)
Mechanistic Interpretability of GPT-like Models on Summarization Tasks
by: Mishra, Anurag
Published: (2025)
by: Mishra, Anurag
Published: (2025)
TrajGPT-R: Generating Urban Mobility Trajectory with Reinforcement Learning-Enhanced Generative Pre-trained Transformer
by: Wang, Jiawei, et al.
Published: (2026)
by: Wang, Jiawei, et al.
Published: (2026)
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning
by: Chen, Yutong, et al.
Published: (2025)
by: Chen, Yutong, et al.
Published: (2025)
Llamba: Scaling Distilled Recurrent Models for Efficient Language Processing
by: Bick, Aviv, et al.
Published: (2025)
by: Bick, Aviv, et al.
Published: (2025)
Learning Unbiased Cluster Descriptors for Interpretable Imbalanced Concept Drift Detection
by: Zhang, Yiqun, et al.
Published: (2026)
by: Zhang, Yiqun, et al.
Published: (2026)
Positive Unlabeled Contrastive Learning
by: Acharya, Anish, et al.
Published: (2022)
by: Acharya, Anish, et al.
Published: (2022)
Deep Generative Models for Offline Policy Learning: Tutorial, Survey, and Perspectives on Future Directions
by: Chen, Jiayu, et al.
Published: (2024)
by: Chen, Jiayu, et al.
Published: (2024)
CAPS: Unifying Attention, Recurrence, and Alignment in Transformer-based Time Series Forecasting
by: Pati, Viresh, et al.
Published: (2026)
by: Pati, Viresh, et al.
Published: (2026)
Flying By ML -- CNN Inversion of Affine Transforms
by: Van Warren, L.
Published: (2023)
by: Van Warren, L.
Published: (2023)
Minion Gated Recurrent Unit for Continual Learning
by: Zyarah, Abdullah M., et al.
Published: (2025)
by: Zyarah, Abdullah M., et al.
Published: (2025)
Interpretable-by-Design Transformers via Architectural Stream Independence
by: Kerce, Clayton, et al.
Published: (2026)
by: Kerce, Clayton, et al.
Published: (2026)
nnterp: A Standardized Interface for Mechanistic Interpretability of Transformers
by: Dumas, Clément
Published: (2025)
by: Dumas, Clément
Published: (2025)
Mechanistic Interpretability for Transformer-based Time Series Classification
by: Kalnāre, Matīss, et al.
Published: (2025)
by: Kalnāre, Matīss, et al.
Published: (2025)
Renaissance of RNNs in Streaming Clinical Time Series: Compact Recurrence Remains Competitive with Transformers
by: Tong, Ran, et al.
Published: (2025)
by: Tong, Ran, et al.
Published: (2025)
Revisiting Glorot Initialization for Long-Range Linear Recurrences
by: Bar, Noga, et al.
Published: (2025)
by: Bar, Noga, et al.
Published: (2025)
Understanding Dynamic Compute Allocation in Recurrent Transformers
by: Moosa, Ibraheem Muhammad, et al.
Published: (2026)
by: Moosa, Ibraheem Muhammad, et al.
Published: (2026)
GARNN: An Interpretable Graph Attentive Recurrent Neural Network for Predicting Blood Glucose Levels via Multivariate Time Series
by: Piao, Chengzhe, et al.
Published: (2024)
by: Piao, Chengzhe, et al.
Published: (2024)
Data-Centric Interpretability for LLM-based Multi-Agent Reinforcement Learning
by: Yan, John, et al.
Published: (2026)
by: Yan, John, et al.
Published: (2026)
Differentiated Thyroid Cancer Recurrence Classification Using Machine Learning Models and Bayesian Neural Networks with Varying Priors: A SHAP-Based Interpretation of the Best Performing Model
by: Kumari, HMNS, et al.
Published: (2025)
by: Kumari, HMNS, et al.
Published: (2025)
Real-Time Recurrent Learning using Trace Units in Reinforcement Learning
by: Elelimy, Esraa, et al.
Published: (2024)
by: Elelimy, Esraa, et al.
Published: (2024)
Attention Mechanism, Max-Affine Partition, and Universal Approximation
by: Liu, Hude, et al.
Published: (2025)
by: Liu, Hude, et al.
Published: (2025)
Max-Affine Spline Insights Into Deep Network Pruning
by: You, Haoran, et al.
Published: (2021)
by: You, Haoran, et al.
Published: (2021)
Efficient and Interpretable Neural Networks Using Complex Lehmer Transform
by: Ataei, Masoud, et al.
Published: (2025)
by: Ataei, Masoud, et al.
Published: (2025)
Towards Mechanistic Interpretability of Graph Transformers via Attention Graphs
by: El, Batu, et al.
Published: (2025)
by: El, Batu, et al.
Published: (2025)
Similar Items
-
Transfer Learning of Surrogate Models: Integrating Domain Warping and Affine Transformations
by: Pan, Shuaiqun, et al.
Published: (2025) -
Likelihood-based Sensor Calibration using Affine Transformation
by: Machhamer, Rüdiger, et al.
Published: (2023) -
nGPT: Normalized Transformer with Representation Learning on the Hypersphere
by: Loshchilov, Ilya, et al.
Published: (2024) -
Recurrent Action Transformer with Memory
by: Cherepanov, Egor, et al.
Published: (2023) -
DeepFilter: A Transformer-style Framework for Accurate and Efficient Process Monitoring
by: Wang, Hao, et al.
Published: (2025)