The Curious Case of In-Training Compression of State Space Models
Fuente:
arXiv
Saved in:
| Main Authors: | Chahine, Makram, Nazari, Philipp, Rus, Daniela, Rusch, T. Konstantin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Correcting Gradient-Based Circuit Localization via Interaction-Aware Backpropagation
by: Edin, Joakim, et al.
Published: (2025)
by: Edin, Joakim, et al.
Published: (2025)
Causal Dimensionality of Transformer Representations: Measurement, Scaling, and Layer Structure
by: Sarkar, Nilesh, et al.
Published: (2026)
by: Sarkar, Nilesh, et al.
Published: (2026)
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
Improving Efficiency of Sampling-based Motion Planning via Message-Passing Monte Carlo
by: Chahine, Makram, et al.
Published: (2024)
by: Chahine, Makram, et al.
Published: (2024)
Large Language Models Report Subjective Experience Under Self-Referential Processing
by: Berg, Cameron, et al.
Published: (2025)
by: Berg, Cameron, et al.
Published: (2025)
Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form QA
by: Badshah, Sher, et al.
Published: (2024)
by: Badshah, Sher, et al.
Published: (2024)
Correcting Stochastic Update Bias in Preconditioned Language Model Optimizers
by: Nayak, Nikhil, et al.
Published: (2026)
by: Nayak, Nikhil, et al.
Published: (2026)
DeepPersona: A Generative Engine for Scaling Deep Synthetic Personas
by: Wang, Zhen, et al.
Published: (2025)
by: Wang, Zhen, et al.
Published: (2025)
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
by: Basu, Abhinaba
Published: (2026)
by: Basu, Abhinaba
Published: (2026)
JAM: Controllable and Responsible Text Generation via Causal Reasoning and Latent Vector Manipulation
by: Huang, Yingbing, et al.
Published: (2025)
by: Huang, Yingbing, et al.
Published: (2025)
Stealth edits to large language models
by: Sutton, Oliver J., et al.
Published: (2024)
by: Sutton, Oliver J., et al.
Published: (2024)
Scalify: scale propagation for efficient low-precision LLM training
by: Balança, Paul, et al.
Published: (2024)
by: Balança, Paul, et al.
Published: (2024)
Extending $μ$P: Spectral Conditions for Feature Learning Across Optimizers
by: Gupta, Akshita, et al.
Published: (2026)
by: Gupta, Akshita, et al.
Published: (2026)
MAcPNN: Mutual Assisted Learning on Data Streams with Temporal Dependence
by: Giannini, Federico, et al.
Published: (2026)
by: Giannini, Federico, et al.
Published: (2026)
ParaRNN: Unlocking Parallel Training of Nonlinear RNNs for Large Language Models
by: Danieli, Federico, et al.
Published: (2025)
by: Danieli, Federico, et al.
Published: (2025)
Adaptive Latent-Space Constraints in Personalized Federated Learning
by: Ayromlou, Sana, et al.
Published: (2025)
by: Ayromlou, Sana, et al.
Published: (2025)
Thinking Machines: Mathematical Reasoning in the Age of LLMs
by: Asperti, Andrea, et al.
Published: (2025)
by: Asperti, Andrea, et al.
Published: (2025)
NeuronSpark: A Spiking Neural Network Language Model with Selective State Space Dynamics
by: Tang, Zhengzheng
Published: (2026)
by: Tang, Zhengzheng
Published: (2026)
Complex-Valued Phase-Coherent Transformer
by: Hioki, Leona
Published: (2026)
by: Hioki, Leona
Published: (2026)
How Pruning Reshapes Features: Sparse Autoencoder Analysis of Weight-Pruned Language Models
by: Borobia, Hector, et al.
Published: (2026)
by: Borobia, Hector, et al.
Published: (2026)
Diverse LLMs or Diverse Question Interpretations? That is the Ensembling Question
by: Rosales, Rafael, et al.
Published: (2025)
by: Rosales, Rafael, et al.
Published: (2025)
Are Transformers with One Layer Self-Attention Using Low-Rank Weight Matrices Universal Approximators?
by: Kajitsuka, Tokio, et al.
Published: (2023)
by: Kajitsuka, Tokio, et al.
Published: (2023)
Geometric Evolution Maps: Extracting Stable Concept Probes from Transformer Residual Streams
by: Henry, James
Published: (2026)
by: Henry, James
Published: (2026)
mHC-SSM: Manifold-Constrained Hyper-Connections for State Space Language Models with Stream-Specialized Adapters
by: Mutlu, Abdulvahap, et al.
Published: (2026)
by: Mutlu, Abdulvahap, et al.
Published: (2026)
ProactBench: Beyond What The User Asked For
by: Harfi, Sepehr, et al.
Published: (2026)
by: Harfi, Sepehr, et al.
Published: (2026)
The Origins of Representation Manifolds in Large Language Models
by: Modell, Alexander, et al.
Published: (2025)
by: Modell, Alexander, et al.
Published: (2025)
The Concept Allocation Zone: Tracking How Concepts Form Across Transformer Depth
by: Henry, James
Published: (2026)
by: Henry, James
Published: (2026)
A Practical Guide to Streaming Continual Learning
by: Cossu, Andrea, et al.
Published: (2026)
by: Cossu, Andrea, et al.
Published: (2026)
Don't Look Back in Anger: MAGIC Net for Streaming Continual Learning with Temporal Dependence
by: Giannini, Federico, et al.
Published: (2026)
by: Giannini, Federico, et al.
Published: (2026)
cPNN: Continuous Progressive Neural Networks for Evolving Streaming Time Series
by: Giannini, Federico, et al.
Published: (2026)
by: Giannini, Federico, et al.
Published: (2026)
CogniLoad: A Synthetic Natural Language Reasoning Benchmark With Tunable Length, Intrinsic Difficulty, and Distractor Density
by: Kaiser, Daniel, et al.
Published: (2025)
by: Kaiser, Daniel, et al.
Published: (2025)
Navigating WebAI: Training Agents to Complete Web Tasks with Large Language Models and Reinforcement Learning
by: Thil, Lucas-Andreï, et al.
Published: (2024)
by: Thil, Lucas-Andreï, et al.
Published: (2024)
Representing LLMs in Prompt Semantic Task Space
by: Kashani, Idan, et al.
Published: (2025)
by: Kashani, Idan, et al.
Published: (2025)
Mechanistic Analysis of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
by: Imanov, Olaf Yunus Laitinen
Published: (2026)
by: Imanov, Olaf Yunus Laitinen
Published: (2026)
Can Agentic AI Match the Performance of Human Data Scientists?
by: Luo, An, et al.
Published: (2025)
by: Luo, An, et al.
Published: (2025)
AgentDS Technical Report: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science
by: Luo, An, et al.
Published: (2026)
by: Luo, An, et al.
Published: (2026)
Strategic Doctrine Language Models (sdLM): A Learning-System Framework for Doctrinal Consistency and Geopolitical Forecasting
by: Imanov, Olaf Yunus Laitinen, et al.
Published: (2026)
by: Imanov, Olaf Yunus Laitinen, et al.
Published: (2026)
Social Learning through Interactions with Other Agents: A Survey
by: Hillier, Dylan, et al.
Published: (2024)
by: Hillier, Dylan, et al.
Published: (2024)
GradES: Significantly Faster Training in Transformers with Gradient-Based Early Stopping
by: Wen, Qifu, et al.
Published: (2025)
by: Wen, Qifu, et al.
Published: (2025)
Probing for Representation Manifolds in Superposition
by: Modell, Alexander
Published: (2026)
by: Modell, Alexander
Published: (2026)
Similar Items
-
Correcting Gradient-Based Circuit Localization via Interaction-Aware Backpropagation
by: Edin, Joakim, et al.
Published: (2025) -
Causal Dimensionality of Transformer Representations: Measurement, Scaling, and Layer Structure
by: Sarkar, Nilesh, et al.
Published: (2026) -
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
by: Guo, Dongxin, et al.
Published: (2026) -
Improving Efficiency of Sampling-based Motion Planning via Message-Passing Monte Carlo
by: Chahine, Makram, et al.
Published: (2024) -
Large Language Models Report Subjective Experience Under Self-Referential Processing
by: Berg, Cameron, et al.
Published: (2025)