A completely uniform transformer for parity
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Kozachinskiy, Alexander, Steifer, Tomasz |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Simple online learning with consistent oracle
par: Kozachinskiy, Alexander, et autres
Publié: (2023)
par: Kozachinskiy, Alexander, et autres
Publié: (2023)
Ehrenfeucht-Haussler Rank and Chain of Thought
par: Barceló, Pablo, et autres
Publié: (2025)
par: Barceló, Pablo, et autres
Publié: (2025)
Parity, Sensitivity, and Transformers
par: Kozachinskiy, Alexander, et autres
Publié: (2026)
par: Kozachinskiy, Alexander, et autres
Publié: (2026)
Optimal bounds for dissatisfaction in perpetual voting
par: Kozachinskiy, Alexander, et autres
Publié: (2024)
par: Kozachinskiy, Alexander, et autres
Publié: (2024)
Lower bounds on transformers with infinite precision
par: Kozachinskiy, Alexander
Publié: (2024)
par: Kozachinskiy, Alexander
Publié: (2024)
Strassen Attention, Split VC Dimension and Compositionality in Transformers
par: Kozachinskiy, Alexander, et autres
Publié: (2025)
par: Kozachinskiy, Alexander, et autres
Publié: (2025)
Effective Littlestone Dimension
par: Rose, Valentino Delle, et autres
Publié: (2024)
par: Rose, Valentino Delle, et autres
Publié: (2024)
Provably Shorter Scratchpads in Hybrid DeltaNet-Attention Decoders
par: Steifer, Tomasz
Publié: (2026)
par: Steifer, Tomasz
Publié: (2026)
Explaining k-Nearest Neighbors: Abductive and Counterfactual Explanations
par: Barceló, Pablo, et autres
Publié: (2025)
par: Barceló, Pablo, et autres
Publié: (2025)
Message Passing on the Edge: Towards Scalable and Expressive GNNs
par: Barceló, Pablo, et autres
Publié: (2025)
par: Barceló, Pablo, et autres
Publié: (2025)
Continuity and Isolation Lead to Doubts or Dilemmas in Large Language Models
par: Pasten, Hector, et autres
Publié: (2025)
par: Pasten, Hector, et autres
Publié: (2025)
Language Generation: Complexity Barriers and Implications for Learning
par: Arenas, Marcelo, et autres
Publié: (2025)
par: Arenas, Marcelo, et autres
Publié: (2025)
Computable universal online learning
par: Kalociński, Dariusz, et autres
Publié: (2025)
par: Kalociński, Dariusz, et autres
Publié: (2025)
HKAN: Hierarchical Kolmogorov-Arnold Network without Backpropagation
par: Dudek, Grzegorz, et autres
Publié: (2025)
par: Dudek, Grzegorz, et autres
Publié: (2025)
Carrying over algorithm in transformers
par: Kruthoff, Jorrit
Publié: (2024)
par: Kruthoff, Jorrit
Publié: (2024)
Advancing time series completion via RFAMoE and MDFF
par: Zhang, Ci, et autres
Publié: (2025)
par: Zhang, Ci, et autres
Publié: (2025)
Small transformer architectures for task switching
par: Gros, Claudius
Publié: (2025)
par: Gros, Claudius
Publié: (2025)
NLI:Non-uniform Linear Interpolation Approximation of Nonlinear Operations for Efficient LLMs Inference
par: Yu, Jiangyong, et autres
Publié: (2026)
par: Yu, Jiangyong, et autres
Publié: (2026)
Weight-sparse transformers have interpretable circuits
par: Gao, Leo, et autres
Publié: (2025)
par: Gao, Leo, et autres
Publié: (2025)
Step-resolved data attribution for looped transformers
par: Kaissis, Georgios, et autres
Publié: (2026)
par: Kaissis, Georgios, et autres
Publié: (2026)
Learning the greatest common divisor: explaining transformer predictions
par: Charton, François
Publié: (2023)
par: Charton, François
Publié: (2023)
Pretrained battery transformer (PBT): A foundation model for universal battery life prediction
par: Tan, Ruifeng, et autres
Publié: (2025)
par: Tan, Ruifeng, et autres
Publié: (2025)
Dockformer: A transformer-based molecular docking paradigm for large-scale virtual screening
par: Yang, Zhangfan, et autres
Publié: (2024)
par: Yang, Zhangfan, et autres
Publié: (2024)
Arnold: a generalist muscle transformer policy
par: Chiappa, Alberto Silvio, et autres
Publié: (2025)
par: Chiappa, Alberto Silvio, et autres
Publié: (2025)
Don't be lazy: CompleteP enables compute-efficient deep transformers
par: Dey, Nolan, et autres
Publié: (2025)
par: Dey, Nolan, et autres
Publié: (2025)
Financial time series augmentation using transformer based GAN architecture
par: Podobiński, Andrzej, et autres
Publié: (2026)
par: Podobiński, Andrzej, et autres
Publié: (2026)
Static and multivariate-temporal attentive fusion transformer for readmission risk prediction
par: Sun, Zhe, et autres
Publié: (2024)
par: Sun, Zhe, et autres
Publié: (2024)
DPRM: A Plug-in Doob h transform-induced Token-Ordering Module for Diffusion Language Models
par: Bu, Dake, et autres
Publié: (2026)
par: Bu, Dake, et autres
Publié: (2026)
Fair-FLIP: Fair Deepfake Detection with Fairness-Oriented Final Layer Input Prioritising
par: Szandala, Tomasz, et autres
Publié: (2025)
par: Szandala, Tomasz, et autres
Publié: (2025)
One Shot vs. Iterative: Rethinking Pruning Strategies for Model Compression
par: Janusz, Mikołaj, et autres
Publié: (2025)
par: Janusz, Mikołaj, et autres
Publié: (2025)
1000 Layer Networks for Self-Supervised RL: Scaling Depth Can Enable New Goal-Reaching Capabilities
par: Wang, Kevin, et autres
Publié: (2025)
par: Wang, Kevin, et autres
Publié: (2025)
SONG: Self-Organizing Neural Graphs
par: Struski, Łukasz, et autres
Publié: (2021)
par: Struski, Łukasz, et autres
Publié: (2021)
Decomposition-based multi-scale transformer framework for time series anomaly detection
par: Zhang, Wenxin, et autres
Publié: (2025)
par: Zhang, Wenxin, et autres
Publié: (2025)
$σ$-PCA: a building block for neural learning of identifiable linear transformations
par: Kanavati, Fahdi, et autres
Publié: (2023)
par: Kanavati, Fahdi, et autres
Publié: (2023)
Structural Positional Encoding for knowledge integration in transformer-based medical process monitoring
par: Irwin, Christopher, et autres
Publié: (2024)
par: Irwin, Christopher, et autres
Publié: (2024)
A federated learning framework with knowledge graph and temporal transformer for early sepsis prediction in multi-center ICUs
par: Chang, Yue, et autres
Publié: (2026)
par: Chang, Yue, et autres
Publié: (2026)
A standard transformer and attention with linear biases for molecular conformer generation
par: Gurev, Viatcheslav, et autres
Publié: (2025)
par: Gurev, Viatcheslav, et autres
Publié: (2025)
Fast and scalable retrosynthetic planning with a transformer neural network and speculative beam search
par: Andronov, Mikhail, et autres
Publié: (2025)
par: Andronov, Mikhail, et autres
Publié: (2025)
Beyond 2:4: exploring V:N:M sparsity for efficient transformer inference on GPUs
par: Zhao, Kang, et autres
Publié: (2024)
par: Zhao, Kang, et autres
Publié: (2024)
Towards modeling evolving longitudinal health trajectories with a transformer-based deep learning model
par: Moen, Hans, et autres
Publié: (2024)
par: Moen, Hans, et autres
Publié: (2024)
Documents similaires
-
Simple online learning with consistent oracle
par: Kozachinskiy, Alexander, et autres
Publié: (2023) -
Ehrenfeucht-Haussler Rank and Chain of Thought
par: Barceló, Pablo, et autres
Publié: (2025) -
Parity, Sensitivity, and Transformers
par: Kozachinskiy, Alexander, et autres
Publié: (2026) -
Optimal bounds for dissatisfaction in perpetual voting
par: Kozachinskiy, Alexander, et autres
Publié: (2024) -
Lower bounds on transformers with infinite precision
par: Kozachinskiy, Alexander
Publié: (2024)