Towards Understanding Self-Pretraining for Sequence Classification
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Coser, Omar, Zollo, Loredana, Soda, Paolo, Orvieto, Antonio |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Deep Learning for Human Locomotion Analysis in Lower-Limb Exoskeletons: A Comparative Study
von: Coser, Omar, et al.
Veröffentlicht: (2025)
von: Coser, Omar, et al.
Veröffentlicht: (2025)
SPARSE Data, Rich Results: Few-Shot Semi-Supervised Learning via Class-Conditioned Image Translation
von: Manni, Guido, et al.
Veröffentlicht: (2025)
von: Manni, Guido, et al.
Veröffentlicht: (2025)
Adam Simplified: Bias Correction Debunked
von: Laing, Sam, et al.
Veröffentlicht: (2025)
von: Laing, Sam, et al.
Veröffentlicht: (2025)
Revisiting associative recall in modern recurrent models
von: Okpekpe, Destiny, et al.
Veröffentlicht: (2025)
von: Okpekpe, Destiny, et al.
Veröffentlicht: (2025)
Design Principles for Sequence Models via Coefficient Dynamics
von: Sieber, Jerome, et al.
Veröffentlicht: (2025)
von: Sieber, Jerome, et al.
Veröffentlicht: (2025)
NIMBA: Towards Robust and Principled Processing of Point Clouds With SSMs
von: Köprücü, Nursena, et al.
Veröffentlicht: (2024)
von: Köprücü, Nursena, et al.
Veröffentlicht: (2024)
An Adaptive Stochastic Gradient Method with Non-negative Gauss-Newton Stepsizes
von: Orvieto, Antonio, et al.
Veröffentlicht: (2024)
von: Orvieto, Antonio, et al.
Veröffentlicht: (2024)
In Search of Adam's Secret Sauce
von: Orvieto, Antonio, et al.
Veröffentlicht: (2025)
von: Orvieto, Antonio, et al.
Veröffentlicht: (2025)
Recurrent neural networks: vanishing and exploding gradients are not the end of the story
von: Zucchet, Nicolas, et al.
Veröffentlicht: (2024)
von: Zucchet, Nicolas, et al.
Veröffentlicht: (2024)
Explaining Grokking in Transformers through the Lens of Inductive Bias
von: Singh, Jaisidh, et al.
Veröffentlicht: (2026)
von: Singh, Jaisidh, et al.
Veröffentlicht: (2026)
Universal Dynamics of Warmup Stable Decay: understanding WSD beyond Transformers
von: Belloni, Annalisa, et al.
Veröffentlicht: (2026)
von: Belloni, Annalisa, et al.
Veröffentlicht: (2026)
Improved state mixing in higher-order and block diagonal linear recurrent networks
von: Dubinin, Igor, et al.
Veröffentlicht: (2026)
von: Dubinin, Igor, et al.
Veröffentlicht: (2026)
An Uncertainty Principle for Linear Recurrent Neural Networks
von: François, Alexandre, et al.
Veröffentlicht: (2025)
von: François, Alexandre, et al.
Veröffentlicht: (2025)
When, Where and Why to Average Weights?
von: Ajroldi, Niccolò, et al.
Veröffentlicht: (2025)
von: Ajroldi, Niccolò, et al.
Veröffentlicht: (2025)
Not Another Imputation Method: A Transformer-based Model for Missing Values in Tabular Datasets
von: Caruso, Camillo Maria, et al.
Veröffentlicht: (2024)
von: Caruso, Camillo Maria, et al.
Veröffentlicht: (2024)
Understanding Differential Transformer Unchains Pretrained Self-Attentions
von: Kong, Chaerin, et al.
Veröffentlicht: (2025)
von: Kong, Chaerin, et al.
Veröffentlicht: (2025)
GASP: Guided Asymmetric Self-Play For Coding LLMs
von: Jana, Swadesh, et al.
Veröffentlicht: (2026)
von: Jana, Swadesh, et al.
Veröffentlicht: (2026)
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling
von: Srećković, Teodora, et al.
Veröffentlicht: (2025)
von: Srećković, Teodora, et al.
Veröffentlicht: (2025)
Towards a Foundation Purchasing Model: Pretrained Generative Autoregression on Transaction Sequences
von: Skalski, Piotr, et al.
Veröffentlicht: (2024)
von: Skalski, Piotr, et al.
Veröffentlicht: (2024)
Hybrid Quantum Neural Network for Multivariate Clinical Time Series Forecasting
von: Iele, Irene, et al.
Veröffentlicht: (2026)
von: Iele, Irene, et al.
Veröffentlicht: (2026)
Confidence Calibration in Vision-Language-Action Models
von: Zollo, Thomas P, et al.
Veröffentlicht: (2025)
von: Zollo, Thomas P, et al.
Veröffentlicht: (2025)
MARIA: a Multimodal Transformer Model for Incomplete Healthcare Data
von: Caruso, Camillo Maria, et al.
Veröffentlicht: (2024)
von: Caruso, Camillo Maria, et al.
Veröffentlicht: (2024)
Unsupervised Confidence Calibration for Reasoning LLMs from a Single Generation
von: Zollo, Thomas, et al.
Veröffentlicht: (2026)
von: Zollo, Thomas, et al.
Veröffentlicht: (2026)
Geometric Inductive Biases of Deep Networks: The Role of Data and Architecture
von: Movahedi, Sajad, et al.
Veröffentlicht: (2024)
von: Movahedi, Sajad, et al.
Veröffentlicht: (2024)
Super Consistency of Neural Network Landscapes and Learning Rate Transfer
von: Noci, Lorenzo, et al.
Veröffentlicht: (2024)
von: Noci, Lorenzo, et al.
Veröffentlicht: (2024)
On the low-shot transferability of [V]-Mamba
von: Misra, Diganta, et al.
Veröffentlicht: (2024)
von: Misra, Diganta, et al.
Veröffentlicht: (2024)
Fixed-Point RNNs: Interpolating from Diagonal to Dense
von: Movahedi, Sajad, et al.
Veröffentlicht: (2025)
von: Movahedi, Sajad, et al.
Veröffentlicht: (2025)
Tell Me What To Learn: Generalizing Neural Memory to be Controllable in Natural Language
von: Bennett, Max S., et al.
Veröffentlicht: (2026)
von: Bennett, Max S., et al.
Veröffentlicht: (2026)
Resilient Vision-Tabular Multimodal Learning under Modality Missingness
von: Caruso, Camillo Maria, et al.
Veröffentlicht: (2026)
von: Caruso, Camillo Maria, et al.
Veröffentlicht: (2026)
A Deep Learning Approach for Overall Survival Prediction in Lung Cancer with Missing Values
von: Caruso, Camillo Maria, et al.
Veröffentlicht: (2023)
von: Caruso, Camillo Maria, et al.
Veröffentlicht: (2023)
Understanding the differences in Foundation Models: Attention, State Space Models, and Recurrent Neural Networks
von: Sieber, Jerome, et al.
Veröffentlicht: (2024)
von: Sieber, Jerome, et al.
Veröffentlicht: (2024)
In Search of Lost DNA Sequence Pretraining
von: Tang, Zhijiang, et al.
Veröffentlicht: (2026)
von: Tang, Zhijiang, et al.
Veröffentlicht: (2026)
Can you Finetune your Binoculars? Embedding Text Watermarks into the Weights of Large Language Models
von: Elhassan, Fay, et al.
Veröffentlicht: (2025)
von: Elhassan, Fay, et al.
Veröffentlicht: (2025)
Loss Landscape Characterization of Neural Networks without Over-Parametrization
von: Islamov, Rustem, et al.
Veröffentlicht: (2024)
von: Islamov, Rustem, et al.
Veröffentlicht: (2024)
Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size
von: Islamov, Rustem, et al.
Veröffentlicht: (2025)
von: Islamov, Rustem, et al.
Veröffentlicht: (2025)
Test-Time Warmup for Multimodal Large Language Models
von: Rajaneesh, Nikita, et al.
Veröffentlicht: (2025)
von: Rajaneesh, Nikita, et al.
Veröffentlicht: (2025)
ReMIA: a Powerful and Efficient Alternative to Membership Inference Attacks against Synthetic Data Generators
von: Scassola, Davide, et al.
Veröffentlicht: (2026)
von: Scassola, Davide, et al.
Veröffentlicht: (2026)
Muown: Row-Norm Control for Muon Optimization
von: Lion, Kai, et al.
Veröffentlicht: (2026)
von: Lion, Kai, et al.
Veröffentlicht: (2026)
Recurrent Distance Filtering for Graph Representation Learning
von: Ding, Yuhui, et al.
Veröffentlicht: (2023)
von: Ding, Yuhui, et al.
Veröffentlicht: (2023)
MATNet: Multi-Level Fusion Transformer-Based Model for Day-Ahead PV Generation Forecasting
von: Tortora, Matteo, et al.
Veröffentlicht: (2023)
von: Tortora, Matteo, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Deep Learning for Human Locomotion Analysis in Lower-Limb Exoskeletons: A Comparative Study
von: Coser, Omar, et al.
Veröffentlicht: (2025) -
SPARSE Data, Rich Results: Few-Shot Semi-Supervised Learning via Class-Conditioned Image Translation
von: Manni, Guido, et al.
Veröffentlicht: (2025) -
Adam Simplified: Bias Correction Debunked
von: Laing, Sam, et al.
Veröffentlicht: (2025) -
Revisiting associative recall in modern recurrent models
von: Okpekpe, Destiny, et al.
Veröffentlicht: (2025) -
Design Principles for Sequence Models via Coefficient Dynamics
von: Sieber, Jerome, et al.
Veröffentlicht: (2025)