Saved in:
| Main Authors: | Biegun, Kai, Dolga, Rares, Cunningham, Jake, Barber, David |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2407.07239 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unifying Linear-Time Attention via Latent Probabilistic Modelling
by: Dolga, Rares, et al.
Published: (2024)
by: Dolga, Rares, et al.
Published: (2024)
Incremental Sequence Classification with Temporal Consistency
by: Maystre, Lucas, et al.
Published: (2025)
by: Maystre, Lucas, et al.
Published: (2025)
wd1: Weighted Policy Optimization for Reasoning in Diffusion Language Models
by: Tang, Xiaohang, et al.
Published: (2025)
by: Tang, Xiaohang, et al.
Published: (2025)
Reparameterized Multi-Resolution Convolutions for Long Sequence Modelling
by: Cunningham, Harry Jake, et al.
Published: (2024)
by: Cunningham, Harry Jake, et al.
Published: (2024)
When Embedding Models Meet: Procrustes Bounds and Applications
by: Maystre, Lucas, et al.
Published: (2025)
by: Maystre, Lucas, et al.
Published: (2025)
Lasso and Partially-Rotated Designs
by: Buhai, Rares-Darius
Published: (2025)
by: Buhai, Rares-Darius
Published: (2025)
From Characters to Tokens: Dynamic Grouping with Hierarchical BPE
by: Dolga, Rares, et al.
Published: (2025)
by: Dolga, Rares, et al.
Published: (2025)
Maximum Likelihood Learning of Latent Dynamics Without Reconstruction
by: Hromadka, Samo, et al.
Published: (2025)
by: Hromadka, Samo, et al.
Published: (2025)
QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
by: Ashkboos, Saleh, et al.
Published: (2024)
by: Ashkboos, Saleh, et al.
Published: (2024)
Attention as an RNN
by: Feng, Leo, et al.
Published: (2024)
by: Feng, Leo, et al.
Published: (2024)
SmoothRot: Combining Channel-Wise Scaling and Rotation for Quantization-Friendly LLMs
by: Czakó, Patrik, et al.
Published: (2025)
by: Czakó, Patrik, et al.
Published: (2025)
DiRotQ: Rotation-Aware Quantization for 4-bit Diffusion Transformers
by: Sharify, Sayeh, et al.
Published: (2026)
by: Sharify, Sayeh, et al.
Published: (2026)
PGN: The RNN's New Successor is Effective for Long-Range Time Series Forecasting
by: Jia, Yuxin, et al.
Published: (2024)
by: Jia, Yuxin, et al.
Published: (2024)
RotMoLE: Enhancing Mixture of Low-Rank Experts through Rotational Gating Mechanism
by: Sun, Mengyang, et al.
Published: (2026)
by: Sun, Mengyang, et al.
Published: (2026)
OptRot: Mitigating Weight Outliers via Data-Free Rotations for Post-Training Quantization
by: Gadhikar, Advait, et al.
Published: (2025)
by: Gadhikar, Advait, et al.
Published: (2025)
The $\mathbf{Y}$-Combinator for LLMs: Solving Long-Context Rot with $λ$-Calculus
by: Roy, Amartya, et al.
Published: (2026)
by: Roy, Amartya, et al.
Published: (2026)
Nonparametric estimation of conditional probability distributions using a generative approach based on conditional push-forward neural networks
by: Franco, Nicola Rares, et al.
Published: (2025)
by: Franco, Nicola Rares, et al.
Published: (2025)
ISMRNN: An Implicitly Segmented RNN Method with Mamba for Long-Term Time Series Forecasting
by: Zhao, GaoXiang, et al.
Published: (2024)
by: Zhao, GaoXiang, et al.
Published: (2024)
Code Revert Prediction with Graph Neural Networks: A Case Study at J.P. Morgan Chase
by: Pei, Yulong, et al.
Published: (2024)
by: Pei, Yulong, et al.
Published: (2024)
RNN-Based Models for Predicting Seizure Onset in Epileptic Patients
by: Mounagurusamy, Mathan Kumar, et al.
Published: (2024)
by: Mounagurusamy, Mathan Kumar, et al.
Published: (2024)
Speed of Light Exact Greedy Decoding for RNN-T Speech Recognition Models on GPU
by: Galvez, Daniel, et al.
Published: (2024)
by: Galvez, Daniel, et al.
Published: (2024)
Building a stable classifier with the inflated argmax
by: Soloff, Jake A., et al.
Published: (2024)
by: Soloff, Jake A., et al.
Published: (2024)
Bagging Provides Assumption-free Stability
by: Soloff, Jake A., et al.
Published: (2023)
by: Soloff, Jake A., et al.
Published: (2023)
A practical existence theorem for reduced order models based on convolutional autoencoders
by: Franco, Nicola Rares, et al.
Published: (2024)
by: Franco, Nicola Rares, et al.
Published: (2024)
FedRot-LoRA: Mitigating Rotational Misalignment in Federated LoRA
by: Zhang, Haoran, et al.
Published: (2026)
by: Zhang, Haoran, et al.
Published: (2026)
Stabilizing RNN Gradients through Pre-training
by: Herranz-Celotti, Luca, et al.
Published: (2023)
by: Herranz-Celotti, Luca, et al.
Published: (2023)
LongVQ: Long Sequence Modeling with Vector Quantization on Structured Memory
by: Liu, Zicheng, et al.
Published: (2024)
by: Liu, Zicheng, et al.
Published: (2024)
Bridging Rested and Restless Bandits with Graph-Triggering: Rising and Rotting
by: Genalti, Gianmarco, et al.
Published: (2024)
by: Genalti, Gianmarco, et al.
Published: (2024)
RNN(p) for Power Consumption Forecasting
by: Baviera, Roberto, et al.
Published: (2022)
by: Baviera, Roberto, et al.
Published: (2022)
Activity Sparsity Complements Weight Sparsity for Efficient RNN Inference
by: Mukherji, Rishav, et al.
Published: (2023)
by: Mukherji, Rishav, et al.
Published: (2023)
Tighter Truncated Rectangular Prism Approximation for RNN Robustness Verification
by: Lin, Xingqi, et al.
Published: (2025)
by: Lin, Xingqi, et al.
Published: (2025)
CoRe: Coherency Regularization for Hierarchical Time Series
by: Cristian, Rares, et al.
Published: (2025)
by: Cristian, Rares, et al.
Published: (2025)
Efficient End-to-End Learning for Decision-Making: A Meta-Optimization Approach
by: Cristian, Rares, et al.
Published: (2025)
by: Cristian, Rares, et al.
Published: (2025)
Aligning Learning and Endogenous Decision-Making
by: Cristian, Rares, et al.
Published: (2025)
by: Cristian, Rares, et al.
Published: (2025)
Attention Augmented GNN RNN-Attention Models for Advanced Cybersecurity Intrusion Detection
by: Biradar, Jayant, et al.
Published: (2025)
by: Biradar, Jayant, et al.
Published: (2025)
SMR: State Memory Replay for Long Sequence Modeling
by: Qi, Biqing, et al.
Published: (2024)
by: Qi, Biqing, et al.
Published: (2024)
CAB: Comprehensive Attention Benchmarking on Long Sequence Modeling
by: Zhang, Jun, et al.
Published: (2022)
by: Zhang, Jun, et al.
Published: (2022)
Poolformer: Recurrent Networks with Pooling for Long-Sequence Modeling
by: Fernández, Daniel Gallo
Published: (2025)
by: Fernández, Daniel Gallo
Published: (2025)
Grouped Sequency-arranged Rotation: Optimizing Rotation Transformation for Quantization for Free
by: Choi, Euntae, et al.
Published: (2025)
by: Choi, Euntae, et al.
Published: (2025)
HadamRNN: Binary and Sparse Ternary Orthogonal RNNs
by: Foucault, Armand, et al.
Published: (2025)
by: Foucault, Armand, et al.
Published: (2025)
Similar Items
-
Unifying Linear-Time Attention via Latent Probabilistic Modelling
by: Dolga, Rares, et al.
Published: (2024) -
Incremental Sequence Classification with Temporal Consistency
by: Maystre, Lucas, et al.
Published: (2025) -
wd1: Weighted Policy Optimization for Reasoning in Diffusion Language Models
by: Tang, Xiaohang, et al.
Published: (2025) -
Reparameterized Multi-Resolution Convolutions for Long Sequence Modelling
by: Cunningham, Harry Jake, et al.
Published: (2024) -
When Embedding Models Meet: Procrustes Bounds and Applications
by: Maystre, Lucas, et al.
Published: (2025)