Understanding Transformer Architecture through Continuous Dynamics: A Partial Differential Equation Perspective
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Yukun, Zhou, Xueqing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Continuous-Time Attention: PDE-Guided Mechanisms for Long-Sequence Transformers
di: Zhang, Yukun, et al.
Pubblicazione: (2025)
di: Zhang, Yukun, et al.
Pubblicazione: (2025)
Where to Add PDE Diffusion in Transformers
di: Zhang, Yukun, et al.
Pubblicazione: (2025)
di: Zhang, Yukun, et al.
Pubblicazione: (2025)
Partial Information Decomposition for Data Interpretability and Feature Selection
di: Westphal, Charles, et al.
Pubblicazione: (2024)
di: Westphal, Charles, et al.
Pubblicazione: (2024)
Broadcast Channel Cooperative Gain: An Operational Interpretation of Partial Information Decomposition
di: Tian, Chao, et al.
Pubblicazione: (2025)
di: Tian, Chao, et al.
Pubblicazione: (2025)
LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws
di: Ouyang, Xu, et al.
Pubblicazione: (2026)
di: Ouyang, Xu, et al.
Pubblicazione: (2026)
Continual Learning-Aided Super-Resolution Scheme for Channel Reconstruction and Generalization in OFDM Systems
di: Chen, Jianqiao, et al.
Pubblicazione: (2025)
di: Chen, Jianqiao, et al.
Pubblicazione: (2025)
A Deep Latent Factor Graph Clustering with Fairness-Utility Trade-off Perspective
di: Ghodsi, Siamak, et al.
Pubblicazione: (2025)
di: Ghodsi, Siamak, et al.
Pubblicazione: (2025)
Fairness Overfitting in Machine Learning: An Information-Theoretic Perspective
di: Laakom, Firas, et al.
Pubblicazione: (2025)
di: Laakom, Firas, et al.
Pubblicazione: (2025)
Accelerating Error Correction Code Transformers
di: Levy, Matan, et al.
Pubblicazione: (2024)
di: Levy, Matan, et al.
Pubblicazione: (2024)
Uncertainty Quantification and Data Efficiency in AI: An Information-Theoretic Perspective
di: Simeone, Osvaldo, et al.
Pubblicazione: (2025)
di: Simeone, Osvaldo, et al.
Pubblicazione: (2025)
Context Channel Capacity: An Information-Theoretic Framework for Understanding Catastrophic Forgetting
di: Cheng, Ran
Pubblicazione: (2026)
di: Cheng, Ran
Pubblicazione: (2026)
Compression via Pre-trained Transformers: A Study on Byte-Level Multimodal Data
di: Heurtel-Depeiges, David, et al.
Pubblicazione: (2024)
di: Heurtel-Depeiges, David, et al.
Pubblicazione: (2024)
Understanding Action Effects through Instrumental Empowerment in Multi-Agent Reinforcement Learning
di: Selmonaj, Ardian, et al.
Pubblicazione: (2025)
di: Selmonaj, Ardian, et al.
Pubblicazione: (2025)
Fixed-Budget Differentially Private Best Arm Identification
di: Chen, Zhirui, et al.
Pubblicazione: (2024)
di: Chen, Zhirui, et al.
Pubblicazione: (2024)
Simple Convergence Proof of Adam From a Sign-like Descent Perspective
di: Peng, Hanyang, et al.
Pubblicazione: (2025)
di: Peng, Hanyang, et al.
Pubblicazione: (2025)
Hybrid Mamba-Transformer Decoder for Error-Correcting Codes
di: Cohen, Shy-el, et al.
Pubblicazione: (2025)
di: Cohen, Shy-el, et al.
Pubblicazione: (2025)
Understanding LLM Behaviors via Compression: Data Generation, Knowledge Acquisition and Scaling Laws
di: Pan, Zhixuan, et al.
Pubblicazione: (2025)
di: Pan, Zhixuan, et al.
Pubblicazione: (2025)
Bottlenecked Transformers: Periodic KV Cache Consolidation for Generalised Reasoning
di: Oomerjee, Adnan, et al.
Pubblicazione: (2025)
di: Oomerjee, Adnan, et al.
Pubblicazione: (2025)
Learning During Detection: Continual Learning for Neural OFDM Receivers via DMRS
di: Obeed, Mohanad, et al.
Pubblicazione: (2026)
di: Obeed, Mohanad, et al.
Pubblicazione: (2026)
In-Context Learning for MIMO Equalization Using Transformer-Based Sequence Models
di: Zecchin, Matteo, et al.
Pubblicazione: (2023)
di: Zecchin, Matteo, et al.
Pubblicazione: (2023)
Unraveling Text Generation in LLMs: A Stochastic Differential Equation Approach
di: Zhang, Yukun
Pubblicazione: (2024)
di: Zhang, Yukun
Pubblicazione: (2024)
Generative Model-Aided Continual Learning for CSI Feedback in FDD mMIMO-OFDM Systems
di: Liu, Guijun, et al.
Pubblicazione: (2025)
di: Liu, Guijun, et al.
Pubblicazione: (2025)
Decision Making in Changing Environments: Robustness, Query-Based Learning, and Differential Privacy
di: Chen, Fan, et al.
Pubblicazione: (2025)
di: Chen, Fan, et al.
Pubblicazione: (2025)
Fast Fourier Transform-Based Spectral and Temporal Gradient Filtering for Differential Privacy
di: Shin, Hyeju, et al.
Pubblicazione: (2025)
di: Shin, Hyeju, et al.
Pubblicazione: (2025)
Generalization Bounds: Perspectives from Information Theory and PAC-Bayes
di: Hellström, Fredrik, et al.
Pubblicazione: (2023)
di: Hellström, Fredrik, et al.
Pubblicazione: (2023)
Dynamic Multi-Network Mining of Tensor Time Series
di: Obata, Kohei, et al.
Pubblicazione: (2024)
di: Obata, Kohei, et al.
Pubblicazione: (2024)
Knowledge Graph-Based Explainable and Generalized Zero-Shot Semantic Communications
di: Zhang, Zhaoyu, et al.
Pubblicazione: (2025)
di: Zhang, Zhaoyu, et al.
Pubblicazione: (2025)
Intent-Aware DRL-Based NOMA Uplink Dynamic Scheduler for IIoT
di: Mostafa, Salwa, et al.
Pubblicazione: (2024)
di: Mostafa, Salwa, et al.
Pubblicazione: (2024)
Dynamic Quality-Latency Aware Routing for LLM Inference in Wireless Edge-Device Networks
di: Bao, Rui, et al.
Pubblicazione: (2025)
di: Bao, Rui, et al.
Pubblicazione: (2025)
A Wireless Foundation Model for Multi-Task Prediction
di: Sheng, Yucheng, et al.
Pubblicazione: (2025)
di: Sheng, Yucheng, et al.
Pubblicazione: (2025)
A General Error-Theoretical Analysis Framework for Constructing Compression Strategies
di: Zhang, Boyang, et al.
Pubblicazione: (2025)
di: Zhang, Boyang, et al.
Pubblicazione: (2025)
Physically Parameterized Differentiable MUSIC for DoA Estimation with Uncalibrated Arrays
di: Chatelier, Baptiste, et al.
Pubblicazione: (2024)
di: Chatelier, Baptiste, et al.
Pubblicazione: (2024)
A General Deep Learning Framework for Wireless Resource Allocation under Discrete Constraints
di: Wang, Yikun, et al.
Pubblicazione: (2026)
di: Wang, Yikun, et al.
Pubblicazione: (2026)
Capacity-Constrained Continual Learning
di: Wen, Zheng, et al.
Pubblicazione: (2025)
di: Wen, Zheng, et al.
Pubblicazione: (2025)
Integrated Sensing-Communication-Computation for Edge Artificial Intelligence
di: Wen, Dingzhu, et al.
Pubblicazione: (2023)
di: Wen, Dingzhu, et al.
Pubblicazione: (2023)
Optimality of Staircase Mechanisms for Vector Queries under Differential Privacy
di: Melbourne, James, et al.
Pubblicazione: (2026)
di: Melbourne, James, et al.
Pubblicazione: (2026)
An Information Theoretic Perspective on Agentic System Design
di: He, Shizhe, et al.
Pubblicazione: (2025)
di: He, Shizhe, et al.
Pubblicazione: (2025)
Estimating Mutual Information between Time Series and Temporal Event Sequences Across Diverse Analysis Tasks
di: Hu, Haoji, et al.
Pubblicazione: (2026)
di: Hu, Haoji, et al.
Pubblicazione: (2026)
MambaJSCC: Adaptive Deep Joint Source-Channel Coding with Generalized State Space Model
di: Wu, Tong, et al.
Pubblicazione: (2024)
di: Wu, Tong, et al.
Pubblicazione: (2024)
On the Non-decoupling of Supervised Fine-tuning and Reinforcement Learning in Post-training
di: Niu, Xueyan, et al.
Pubblicazione: (2026)
di: Niu, Xueyan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Continuous-Time Attention: PDE-Guided Mechanisms for Long-Sequence Transformers
di: Zhang, Yukun, et al.
Pubblicazione: (2025) -
Where to Add PDE Diffusion in Transformers
di: Zhang, Yukun, et al.
Pubblicazione: (2025) -
Partial Information Decomposition for Data Interpretability and Feature Selection
di: Westphal, Charles, et al.
Pubblicazione: (2024) -
Broadcast Channel Cooperative Gain: An Operational Interpretation of Partial Information Decomposition
di: Tian, Chao, et al.
Pubblicazione: (2025) -
LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws
di: Ouyang, Xu, et al.
Pubblicazione: (2026)