Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Peihao, Cai, Ruisi, Wang, Yuehao, Zhu, Jiajun, Srivastava, Pragya, Wang, Zhangyang, Li, Pan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Rethinking Addressing in Language Models via Contexualized Equivariant Positional Encoding
por: Zhu, Jiajun, et al.
Publicado: (2025)
por: Zhu, Jiajun, et al.
Publicado: (2025)
$\nabla$-Reasoner: LLM Reasoning via Test-Time Gradient Descent in Latent Space
por: Wang, Peihao, et al.
Publicado: (2026)
por: Wang, Peihao, et al.
Publicado: (2026)
Why Neural Network Can Discover Symbolic Structures with Gradient-based Training: An Algebraic and Geometric Foundation for Neurosymbolic Reasoning
por: Wang, Peihao, et al.
Publicado: (2025)
por: Wang, Peihao, et al.
Publicado: (2025)
Position: Weight Space Should Be a First-Class Generative AI Modality
por: Wang, Zhangyang, et al.
Publicado: (2026)
por: Wang, Zhangyang, et al.
Publicado: (2026)
When Do Graph Foundation Models Transfer? A Data-Centric Theory
por: Zhu, Jiajun, et al.
Publicado: (2026)
por: Zhu, Jiajun, et al.
Publicado: (2026)
Generalization Error Analysis for Sparse Mixture-of-Experts: A Preliminary Study
por: Zhao, Jinze, et al.
Publicado: (2024)
por: Zhao, Jinze, et al.
Publicado: (2024)
Oscillation Inversion: Understand the structure of Large Flow Model through the Lens of Inversion Method
por: Zheng, Yan, et al.
Publicado: (2024)
por: Zheng, Yan, et al.
Publicado: (2024)
Polynomial Width is Sufficient for Set Representation with High-dimensional Features
por: Wang, Peihao, et al.
Publicado: (2023)
por: Wang, Peihao, et al.
Publicado: (2023)
Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design
por: Cai, Ruisi, et al.
Publicado: (2024)
por: Cai, Ruisi, et al.
Publicado: (2024)
LoCoCo: Dropping In Convolutions for Long Context Compression
por: Cai, Ruisi, et al.
Publicado: (2024)
por: Cai, Ruisi, et al.
Publicado: (2024)
Graph-KV: Breaking Sequence via Injecting Structural Biases into Large Language Models
por: Wang, Haoyu, et al.
Publicado: (2025)
por: Wang, Haoyu, et al.
Publicado: (2025)
Beyond Test-Time Memory: State-Space Optimal Control for LLM Reasoning
por: Wang, Peihao, et al.
Publicado: (2026)
por: Wang, Peihao, et al.
Publicado: (2026)
Sparse Mixture-of-Experts for Compositional Generalization: Empirical Evidence and Theoretical Foundations of Optimal Sparsity
por: Zhao, Jinze, et al.
Publicado: (2024)
por: Zhao, Jinze, et al.
Publicado: (2024)
Mamba-Based Graph Convolutional Networks: Tackling Over-smoothing with Selective State Space
por: He, Xin, et al.
Publicado: (2025)
por: He, Xin, et al.
Publicado: (2025)
Flextron: Many-in-One Flexible Large Language Model
por: Cai, Ruisi, et al.
Publicado: (2024)
por: Cai, Ruisi, et al.
Publicado: (2024)
Meta ControlNet: Enhancing Task Adaptation via Meta Learning
por: Yang, Junjie, et al.
Publicado: (2023)
por: Yang, Junjie, et al.
Publicado: (2023)
Model-GLUE: Democratized LLM Scaling for A Large Model Zoo in the Wild
por: Zhao, Xinyu, et al.
Publicado: (2024)
por: Zhao, Xinyu, et al.
Publicado: (2024)
FlexGS: Train Once, Deploy Everywhere with Many-in-One Flexible 3D Gaussian Splatting
por: Liu, Hengyu, et al.
Publicado: (2025)
por: Liu, Hengyu, et al.
Publicado: (2025)
A Tool Bottleneck Framework for Clinically-Informed and Interpretable Medical Image Understanding
por: Liu, Christina, et al.
Publicado: (2025)
por: Liu, Christina, et al.
Publicado: (2025)
Dual Mamba for Node-Specific Representation Learning: Tackling Over-Smoothing with Selective State Space Modeling
por: He, Xin, et al.
Publicado: (2025)
por: He, Xin, et al.
Publicado: (2025)
Recency-Weighted Temporally-Segmented Ensemble for Time-Series Modeling
por: Johnsen, Pål V., et al.
Publicado: (2024)
por: Johnsen, Pål V., et al.
Publicado: (2024)
Revisiting Counterfactual Regression through the Lens of Gromov-Wasserstein Information Bottleneck
por: Yang, Hao, et al.
Publicado: (2024)
por: Yang, Hao, et al.
Publicado: (2024)
Towards Understanding Sensitive and Decisive Patterns in Explainable AI: A Case Study of Model Interpretation in Geometric Deep Learning
por: Zhu, Jiajun, et al.
Publicado: (2024)
por: Zhu, Jiajun, et al.
Publicado: (2024)
An Analysis of Concept Bottleneck Models: Measuring, Understanding, and Mitigating the Impact of Noisy Annotations
por: Park, Seonghwan, et al.
Publicado: (2025)
por: Park, Seonghwan, et al.
Publicado: (2025)
Interpreting and Steering State-Space Models via Activation Subspace Bottlenecks
por: Mohan, Vamshi Sunku, et al.
Publicado: (2026)
por: Mohan, Vamshi Sunku, et al.
Publicado: (2026)
Understanding and Mitigating Miscalibration in Prompt Tuning for Vision-Language Models
por: Wang, Shuoyuan, et al.
Publicado: (2024)
por: Wang, Shuoyuan, et al.
Publicado: (2024)
An Analysis of Attention via the Lens of Exchangeability and Latent Variable Models
por: Zhang, Yufeng, et al.
Publicado: (2022)
por: Zhang, Yufeng, et al.
Publicado: (2022)
Generalization of Graph Neural Networks through the Lens of Homomorphism
por: Li, Shouheng, et al.
Publicado: (2024)
por: Li, Shouheng, et al.
Publicado: (2024)
Attention Saturation and Gradient Suppression at Inflection Layers: Diagnosing and Mitigating Bottlenecks in Transformer Adaptation
por: Zixian, Wang
Publicado: (2025)
por: Zixian, Wang
Publicado: (2025)
Enhancing Knowledge Tracing through Leakage-Free and Recency-Aware Embeddings
por: Badran, Yahya, et al.
Publicado: (2025)
por: Badran, Yahya, et al.
Publicado: (2025)
COMBA: Cross Batch Aggregation for Learning Large Graphs with Context Gating State Space Models
por: Shen, Jiajun, et al.
Publicado: (2026)
por: Shen, Jiajun, et al.
Publicado: (2026)
Understanding Biases in ChatGPT-based Recommender Systems: Provider Fairness, Temporal Stability, and Recency
por: Deldjoo, Yashar
Publicado: (2024)
por: Deldjoo, Yashar
Publicado: (2024)
Evaluating Time-Series Training Dataset through Lens of Spectrum in Deep State Space Models
por: Kanai, Sekitoshi, et al.
Publicado: (2024)
por: Kanai, Sekitoshi, et al.
Publicado: (2024)
Demystifying the Recency Heuristic in Temporal-Difference Learning
por: Daley, Brett, et al.
Publicado: (2024)
por: Daley, Brett, et al.
Publicado: (2024)
Measuring Recency Bias In Sequential Recommendation Systems
por: Oh, Jeonglyul, et al.
Publicado: (2024)
por: Oh, Jeonglyul, et al.
Publicado: (2024)
Online Distribution Shift Detection via Recency Prediction
por: Luo, Rachel, et al.
Publicado: (2022)
por: Luo, Rachel, et al.
Publicado: (2022)
MambaTS: Improved Selective State Space Models for Long-term Time Series Forecasting
por: Cai, Xiuding, et al.
Publicado: (2024)
por: Cai, Xiuding, et al.
Publicado: (2024)
Topic Identification in LLM Input-Output Pairs through the Lens of Information Bottleneck
por: Halperin, Igor
Publicado: (2025)
por: Halperin, Igor
Publicado: (2025)
Controllable Concept Bottleneck Models
por: Lin, Hongbin, et al.
Publicado: (2026)
por: Lin, Hongbin, et al.
Publicado: (2026)
How Universal Polynomial Bases Enhance Spectral Graph Neural Networks: Heterophily, Over-smoothing, and Over-squashing
por: Huang, Keke, et al.
Publicado: (2024)
por: Huang, Keke, et al.
Publicado: (2024)
Ejemplares similares
-
Rethinking Addressing in Language Models via Contexualized Equivariant Positional Encoding
por: Zhu, Jiajun, et al.
Publicado: (2025) -
$\nabla$-Reasoner: LLM Reasoning via Test-Time Gradient Descent in Latent Space
por: Wang, Peihao, et al.
Publicado: (2026) -
Why Neural Network Can Discover Symbolic Structures with Gradient-based Training: An Algebraic and Geometric Foundation for Neurosymbolic Reasoning
por: Wang, Peihao, et al.
Publicado: (2025) -
Position: Weight Space Should Be a First-Class Generative AI Modality
por: Wang, Zhangyang, et al.
Publicado: (2026) -
When Do Graph Foundation Models Transfer? A Data-Centric Theory
por: Zhu, Jiajun, et al.
Publicado: (2026)