The Cursive Transformer
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Greydanus, Sam, Wimpee, Zachary |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
When Models Examine Themselves: Vocabulary-Activation Correspondence in Self-Referential Processing
von: Dadfar, Zachary Pedram
Veröffentlicht: (2026)
von: Dadfar, Zachary Pedram
Veröffentlicht: (2026)
Proving that Cryptic Crossword Clue Answers are Correct
von: Andrews, Martin, et al.
Veröffentlicht: (2024)
von: Andrews, Martin, et al.
Veröffentlicht: (2024)
Scaling Down Deep Learning with MNIST-1D
von: Greydanus, Sam, et al.
Veröffentlicht: (2020)
von: Greydanus, Sam, et al.
Veröffentlicht: (2020)
Efficient Benchmarking Is Just Feature Selection and Multiple Regression
von: Bowyer, Sam, et al.
Veröffentlicht: (2026)
von: Bowyer, Sam, et al.
Veröffentlicht: (2026)
Towards Detecting Contextual Real-Time Toxicity for In-Game Chat
von: Yang, Zachary, et al.
Veröffentlicht: (2023)
von: Yang, Zachary, et al.
Veröffentlicht: (2023)
TPTT: Transforming Pretrained Transformers into Titans
von: Furfaro, Fabien
Veröffentlicht: (2025)
von: Furfaro, Fabien
Veröffentlicht: (2025)
LLM-Select: Feature Selection with Large Language Models
von: Jeong, Daniel P., et al.
Veröffentlicht: (2024)
von: Jeong, Daniel P., et al.
Veröffentlicht: (2024)
DINT Transformer
von: Cang, Yueyang, et al.
Veröffentlicht: (2025)
von: Cang, Yueyang, et al.
Veröffentlicht: (2025)
Personalized Language Modeling from Personalized Human Feedback
von: Li, Xinyu, et al.
Veröffentlicht: (2024)
von: Li, Xinyu, et al.
Veröffentlicht: (2024)
Birdie: Advancing State Space Models with Reward-Driven Objectives and Curricula
von: Blouir, Sam, et al.
Veröffentlicht: (2024)
von: Blouir, Sam, et al.
Veröffentlicht: (2024)
The Belief State Transformer
von: Hu, Edward S., et al.
Veröffentlicht: (2024)
von: Hu, Edward S., et al.
Veröffentlicht: (2024)
Three-Phase Transformer
von: Ayyash, Mohammad R. Abu
Veröffentlicht: (2026)
von: Ayyash, Mohammad R. Abu
Veröffentlicht: (2026)
Medical Adaptation of Large Language and Vision-Language Models: Are We Making Progress?
von: Jeong, Daniel P., et al.
Veröffentlicht: (2024)
von: Jeong, Daniel P., et al.
Veröffentlicht: (2024)
Evaluating the Factuality of Zero-shot Summarizers Across Varied Domains
von: Ramprasad, Sanjana, et al.
Veröffentlicht: (2024)
von: Ramprasad, Sanjana, et al.
Veröffentlicht: (2024)
Revisiting the Scaling Properties of Downstream Metrics in Large Language Model Training
von: Krajewski, Jakub, et al.
Veröffentlicht: (2025)
von: Krajewski, Jakub, et al.
Veröffentlicht: (2025)
VeRO: An Evaluation Harness for Agents to Optimize Agents
von: Ursekar, Varun, et al.
Veröffentlicht: (2026)
von: Ursekar, Varun, et al.
Veröffentlicht: (2026)
Selective Attention Improves Transformer
von: Leviathan, Yaniv, et al.
Veröffentlicht: (2024)
von: Leviathan, Yaniv, et al.
Veröffentlicht: (2024)
Fast Byte Latent Transformer
von: Kallini, Julie, et al.
Veröffentlicht: (2026)
von: Kallini, Julie, et al.
Veröffentlicht: (2026)
Algorithmic Capabilities of Random Transformers
von: Zhong, Ziqian, et al.
Veröffentlicht: (2024)
von: Zhong, Ziqian, et al.
Veröffentlicht: (2024)
Transformers Struggle to Learn to Search
von: Saparov, Abulhair, et al.
Veröffentlicht: (2024)
von: Saparov, Abulhair, et al.
Veröffentlicht: (2024)
An Evolved Universal Transformer Memory
von: Cetin, Edoardo, et al.
Veröffentlicht: (2024)
von: Cetin, Edoardo, et al.
Veröffentlicht: (2024)
On the Ability of Transformers to Verify Plans
von: Sarrof, Yash, et al.
Veröffentlicht: (2026)
von: Sarrof, Yash, et al.
Veröffentlicht: (2026)
Your Transformer is Secretly Linear
von: Razzhigaev, Anton, et al.
Veröffentlicht: (2024)
von: Razzhigaev, Anton, et al.
Veröffentlicht: (2024)
TransformLLM: Adapting Large Language Models via LLM-Transformed Reading Comprehension Text
von: Arbel, Iftach, et al.
Veröffentlicht: (2024)
von: Arbel, Iftach, et al.
Veröffentlicht: (2024)
On the Spatial Structure of Mixture-of-Experts in Transformers
von: Bershatsky, Daniel, et al.
Veröffentlicht: (2025)
von: Bershatsky, Daniel, et al.
Veröffentlicht: (2025)
Adaptive Computation Pruning for the Forgetting Transformer
von: Lin, Zhixuan, et al.
Veröffentlicht: (2025)
von: Lin, Zhixuan, et al.
Veröffentlicht: (2025)
Strategic Fusion Optimizes Transformer Compression
von: Rahman, Md Shoaibur
Veröffentlicht: (2025)
von: Rahman, Md Shoaibur
Veröffentlicht: (2025)
Transformer-Squared: Self-adaptive LLMs
von: Sun, Qi, et al.
Veröffentlicht: (2025)
von: Sun, Qi, et al.
Veröffentlicht: (2025)
The Role of Sparsity for Length Generalization in Transformers
von: Golowich, Noah, et al.
Veröffentlicht: (2025)
von: Golowich, Noah, et al.
Veröffentlicht: (2025)
An evolutionary perspective on modes of learning in Transformers
von: Ku, Alexander Y., et al.
Veröffentlicht: (2025)
von: Ku, Alexander Y., et al.
Veröffentlicht: (2025)
When Can Transformers Count to n?
von: Yehudai, Gilad, et al.
Veröffentlicht: (2024)
von: Yehudai, Gilad, et al.
Veröffentlicht: (2024)
Transformer Circuit Faithfulness Metrics are not Robust
von: Miller, Joseph, et al.
Veröffentlicht: (2024)
von: Miller, Joseph, et al.
Veröffentlicht: (2024)
Representing Rule-based Chatbots with Transformers
von: Friedman, Dan, et al.
Veröffentlicht: (2024)
von: Friedman, Dan, et al.
Veröffentlicht: (2024)
ALTA: Compiler-Based Analysis of Transformers
von: Shaw, Peter, et al.
Veröffentlicht: (2024)
von: Shaw, Peter, et al.
Veröffentlicht: (2024)
The Geometric Anatomy of Capability Acquisition in Transformers
von: Billa, Jayadev
Veröffentlicht: (2026)
von: Billa, Jayadev
Veröffentlicht: (2026)
Towards Infinite-Long Prefix in Transformer
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
Momentum Streams for Optimizer-Inspired Transformers
von: Gai, Jingchu, et al.
Veröffentlicht: (2026)
von: Gai, Jingchu, et al.
Veröffentlicht: (2026)
Does Transformer Interpretability Transfer to RNNs?
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2024)
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2024)
STAT: Shrinking Transformers After Training
von: Flynn, Megan, et al.
Veröffentlicht: (2024)
von: Flynn, Megan, et al.
Veröffentlicht: (2024)
The Limited Impact of Medical Adaptation of Large Language and Vision-Language Models
von: Jeong, Daniel P., et al.
Veröffentlicht: (2024)
von: Jeong, Daniel P., et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
When Models Examine Themselves: Vocabulary-Activation Correspondence in Self-Referential Processing
von: Dadfar, Zachary Pedram
Veröffentlicht: (2026) -
Proving that Cryptic Crossword Clue Answers are Correct
von: Andrews, Martin, et al.
Veröffentlicht: (2024) -
Scaling Down Deep Learning with MNIST-1D
von: Greydanus, Sam, et al.
Veröffentlicht: (2020) -
Efficient Benchmarking Is Just Feature Selection and Multiple Regression
von: Bowyer, Sam, et al.
Veröffentlicht: (2026) -
Towards Detecting Contextual Real-Time Toxicity for In-Game Chat
von: Yang, Zachary, et al.
Veröffentlicht: (2023)