Mapping the Edge of Chaos: Fractal-Like Boundaries in The Trainability of Decoder-Only Transformer Models
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Torkamandi, Bahman |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Trainable and Explainable Simplicial Map Neural Networks
von: Paluzo-Hidalgo, Eduardo, et al.
Veröffentlicht: (2023)
von: Paluzo-Hidalgo, Eduardo, et al.
Veröffentlicht: (2023)
Indoor Localization using Compact, Telemetry-Agnostic, Transfer-Learning Enabled Decoder-Only Transformer
von: Bhatia, Nayan Sanjay, et al.
Veröffentlicht: (2025)
von: Bhatia, Nayan Sanjay, et al.
Veröffentlicht: (2025)
SynEHRgy: Synthesizing Mixed-Type Structured Electronic Health Records using Decoder-Only Transformers
von: Karami, Hojjat, et al.
Veröffentlicht: (2024)
von: Karami, Hojjat, et al.
Veröffentlicht: (2024)
Trained Persistent Memory for Frozen Decoder-Only LLMs
von: Jeong, Hong
Veröffentlicht: (2026)
von: Jeong, Hong
Veröffentlicht: (2026)
Progressive Inference: Explaining Decoder-Only Sequence Classification Models Using Intermediate Predictions
von: Kariyappa, Sanjay, et al.
Veröffentlicht: (2024)
von: Kariyappa, Sanjay, et al.
Veröffentlicht: (2024)
Projected Compression: Trainable Projection for Efficient Transformer Compression
von: Stefaniak, Maciej, et al.
Veröffentlicht: (2025)
von: Stefaniak, Maciej, et al.
Veröffentlicht: (2025)
LauraTSE: Target Speaker Extraction using Auto-Regressive Decoder-Only Language Models
von: Tang, Beilong, et al.
Veröffentlicht: (2025)
von: Tang, Beilong, et al.
Veröffentlicht: (2025)
On the Trainability of Masked Diffusion Language Models via Blockwise Locality
von: Wang, Yuxiang, et al.
Veröffentlicht: (2026)
von: Wang, Yuxiang, et al.
Veröffentlicht: (2026)
RewriteNets: End-to-End Trainable String-Rewriting for Generative Sequence Modeling
von: Vejendla, Harshil
Veröffentlicht: (2026)
von: Vejendla, Harshil
Veröffentlicht: (2026)
Scale-Consistent State-Space Dynamics via Fractal of Stationary Transformations
von: Yu, Geunhyeok, et al.
Veröffentlicht: (2026)
von: Yu, Geunhyeok, et al.
Veröffentlicht: (2026)
Fractal Language Modelling by Universal Sequence Maps (USM)
von: Almeida, Jonas S, et al.
Veröffentlicht: (2025)
von: Almeida, Jonas S, et al.
Veröffentlicht: (2025)
A Unified Noise-Curvature View of Loss of Trainability
von: Baveja, Gunbir Singh, et al.
Veröffentlicht: (2025)
von: Baveja, Gunbir Singh, et al.
Veröffentlicht: (2025)
When Bias Meets Trainability: Connecting Theories of Initialization
von: Bassi, Alberto, et al.
Veröffentlicht: (2025)
von: Bassi, Alberto, et al.
Veröffentlicht: (2025)
SageBwd: A Trainable Low-bit Attention
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
Natively Trainable Sparse Attention for Hierarchical Point Cloud Datasets
von: Lapautre, Nicolas, et al.
Veröffentlicht: (2025)
von: Lapautre, Nicolas, et al.
Veröffentlicht: (2025)
Playing the Lottery With Concave Regularizers for Sparse Trainable Neural Networks
von: Fracastoro, Giulia, et al.
Veröffentlicht: (2025)
von: Fracastoro, Giulia, et al.
Veröffentlicht: (2025)
Scalable Decision Focused Learning via Online Trainable Surrogates
von: Signorelli, Gaetano, et al.
Veröffentlicht: (2025)
von: Signorelli, Gaetano, et al.
Veröffentlicht: (2025)
A KL Lens on Quantization: Fast, Forward-Only Sensitivity for Mixed-Precision SSM-Transformer Models
von: Kong, Jason, et al.
Veröffentlicht: (2026)
von: Kong, Jason, et al.
Veröffentlicht: (2026)
Are Graph Transformers Necessary? Efficient Long-Range Message Passing with Fractal Nodes in MPNNs
von: Choi, Jeongwhan, et al.
Veröffentlicht: (2025)
von: Choi, Jeongwhan, et al.
Veröffentlicht: (2025)
A GREAT Architecture for Edge-Based Graph Problems Like TSP
von: Lischka, Attila, et al.
Veröffentlicht: (2024)
von: Lischka, Attila, et al.
Veröffentlicht: (2024)
Calibration and Transformation-Free Weight-Only LLMs Quantization via Dynamic Grouping
von: Zheng, Xinzhe, et al.
Veröffentlicht: (2025)
von: Zheng, Xinzhe, et al.
Veröffentlicht: (2025)
Outcome-Based RL Provably Leads Transformers to Reason, but Only With the Right Data
von: Ran-Milo, Yuval, et al.
Veröffentlicht: (2026)
von: Ran-Milo, Yuval, et al.
Veröffentlicht: (2026)
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference
von: Gong, Ping, et al.
Veröffentlicht: (2025)
von: Gong, Ping, et al.
Veröffentlicht: (2025)
Trainable Dynamic Mask Sparse Attention
von: Shi, Jingze, et al.
Veröffentlicht: (2025)
von: Shi, Jingze, et al.
Veröffentlicht: (2025)
Fractal Landscapes in Policy Optimization
von: Wang, Tao, et al.
Veröffentlicht: (2023)
von: Wang, Tao, et al.
Veröffentlicht: (2023)
Conjugate Learning Theory: Uncovering the Mechanisms of Trainability and Generalization in Deep Neural Networks
von: Qi, Binchuan
Veröffentlicht: (2026)
von: Qi, Binchuan
Veröffentlicht: (2026)
Semi-Supervised Online Learning on the Edge by Transforming Knowledge from Teacher Models
von: Xue, Jiabin
Veröffentlicht: (2025)
von: Xue, Jiabin
Veröffentlicht: (2025)
DBConformer: Dual-Branch Convolutional Transformer for EEG Decoding
von: Wang, Ziwei, et al.
Veröffentlicht: (2025)
von: Wang, Ziwei, et al.
Veröffentlicht: (2025)
Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps
von: Chuang, Yung-Sung, et al.
Veröffentlicht: (2024)
von: Chuang, Yung-Sung, et al.
Veröffentlicht: (2024)
Partner Modelling Emerges in Recurrent Agents (But Only When It Matters)
von: Mon-Williams, Ruaridh, et al.
Veröffentlicht: (2025)
von: Mon-Williams, Ruaridh, et al.
Veröffentlicht: (2025)
T-TAME: Trainable Attention Mechanism for Explaining Convolutional Networks and Vision Transformers
von: Ntrougkas, Mariano V., et al.
Veröffentlicht: (2024)
von: Ntrougkas, Mariano V., et al.
Veröffentlicht: (2024)
Uncovering Graph Reasoning in Decoder-only Transformers with Circuit Tracing
von: Dai, Xinnan, et al.
Veröffentlicht: (2025)
von: Dai, Xinnan, et al.
Veröffentlicht: (2025)
On The Potential of The Fractal Geometry and The CNNs Ability to Encode it
von: Zini, Julia El, et al.
Veröffentlicht: (2024)
von: Zini, Julia El, et al.
Veröffentlicht: (2024)
Catch-Only-One: Non-Transferable Examples for Model-Specific Authorization
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
Emergence of Minimal Circuits for Indirect Object Identification in Attention-Only Transformers
von: Adhikari, Rabin
Veröffentlicht: (2025)
von: Adhikari, Rabin
Veröffentlicht: (2025)
Weightless Neural Networks for Continuously Trainable Personalized Recommendation Systems
von: Latif, Rafayel, et al.
Veröffentlicht: (2025)
von: Latif, Rafayel, et al.
Veröffentlicht: (2025)
StagFormer: Time Staggering Transformer Decoding for RunningLayers In Parallel
von: Cutler, Dylan, et al.
Veröffentlicht: (2025)
von: Cutler, Dylan, et al.
Veröffentlicht: (2025)
Conformal Sparsification for Bandwidth-Efficient Edge-Cloud Speculative Decoding
von: Bhattacharjee, Payel, et al.
Veröffentlicht: (2025)
von: Bhattacharjee, Payel, et al.
Veröffentlicht: (2025)
BaNEL: Exploration Posteriors for Generative Modeling Using Only Negative Rewards
von: Lee, Sangyun, et al.
Veröffentlicht: (2025)
von: Lee, Sangyun, et al.
Veröffentlicht: (2025)
Topology Only Pre-Training: Towards Generalised Multi-Domain Graph Models
von: Davies, Alex O., et al.
Veröffentlicht: (2023)
von: Davies, Alex O., et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Trainable and Explainable Simplicial Map Neural Networks
von: Paluzo-Hidalgo, Eduardo, et al.
Veröffentlicht: (2023) -
Indoor Localization using Compact, Telemetry-Agnostic, Transfer-Learning Enabled Decoder-Only Transformer
von: Bhatia, Nayan Sanjay, et al.
Veröffentlicht: (2025) -
SynEHRgy: Synthesizing Mixed-Type Structured Electronic Health Records using Decoder-Only Transformers
von: Karami, Hojjat, et al.
Veröffentlicht: (2024) -
Trained Persistent Memory for Frozen Decoder-Only LLMs
von: Jeong, Hong
Veröffentlicht: (2026) -
Progressive Inference: Explaining Decoder-Only Sequence Classification Models Using Intermediate Predictions
von: Kariyappa, Sanjay, et al.
Veröffentlicht: (2024)