A Two-Parameter Weibull Framework for Diagnosing Transformer Weight Distributions
Fuente:
arXiv
Guardado en:
| Autor principal: | Ding, Tiexin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Generalizing to New Dynamical Systems via Frequency Domain Adaptation
por: Qin, Tiexin, et al.
Publicado: (2025)
por: Qin, Tiexin, et al.
Publicado: (2025)
To Augment or Not to Augment? Diagnosing Distributional Symmetry Breaking
por: Lawrence, Hannah, et al.
Publicado: (2025)
por: Lawrence, Hannah, et al.
Publicado: (2025)
Permutation Equivariant Neural Controlled Differential Equations for Dynamic Graph Representation Learning
por: Berndt, Torben, et al.
Publicado: (2025)
por: Berndt, Torben, et al.
Publicado: (2025)
DFReg: A Physics-Inspired Framework for Global Weight Distribution Regularization in Neural Networks
por: Ruggieri, Giovanni
Publicado: (2025)
por: Ruggieri, Giovanni
Publicado: (2025)
Scalable Weibull Graph Attention Autoencoder for Modeling Document Networks
por: Wang, Chaojie, et al.
Publicado: (2024)
por: Wang, Chaojie, et al.
Publicado: (2024)
Allocation of Parameters in Transformers
por: Yu, Ruoxi, et al.
Publicado: (2025)
por: Yu, Ruoxi, et al.
Publicado: (2025)
Learning Dynamic Graph Embeddings with Neural Controlled Differential Equations
por: Qin, Tiexin, et al.
Publicado: (2023)
por: Qin, Tiexin, et al.
Publicado: (2023)
The Offline-Frontier Shift: Diagnosing Distributional Limits in Generative Multi-Objective Optimization
por: Holly, Stephanie, et al.
Publicado: (2026)
por: Holly, Stephanie, et al.
Publicado: (2026)
Parameter-Efficient Checkpoint Merging via Metrics-Weighted Averaging
por: Yu, Shi Jie, et al.
Publicado: (2025)
por: Yu, Shi Jie, et al.
Publicado: (2025)
TFWT: Tabular Feature Weighting with Transformer
por: Zhang, Xinhao, et al.
Publicado: (2024)
por: Zhang, Xinhao, et al.
Publicado: (2024)
Distributed Sign Momentum with Local Steps for Training Transformers
por: Yu, Shuhua, et al.
Publicado: (2024)
por: Yu, Shuhua, et al.
Publicado: (2024)
A network-constrain Weibull AFT model for biomarkers discovery
por: Angelini, Claudia, et al.
Publicado: (2024)
por: Angelini, Claudia, et al.
Publicado: (2024)
A Transformer-Based Approach for Diagnosing Fault Cases in Optical Fiber Amplifiers
por: Schneider, Dominic, et al.
Publicado: (2025)
por: Schneider, Dominic, et al.
Publicado: (2025)
Fast and Robust: Computationally Efficient Covariance Estimation for Sub-Weibull Vectors
por: He, Even
Publicado: (2025)
por: He, Even
Publicado: (2025)
Transformer-based Parameter Estimation in Statistics
por: Yin, Xiaoxin, et al.
Publicado: (2024)
por: Yin, Xiaoxin, et al.
Publicado: (2024)
Adaptive Dual-Weighting Framework for Federated Learning via Out-of-Distribution Detection
por: Ling, Zhiwei, et al.
Publicado: (2026)
por: Ling, Zhiwei, et al.
Publicado: (2026)
A Brain-to-Population Graph Learning Framework for Diagnosing Brain Disorders
por: Liao, Qianqian, et al.
Publicado: (2025)
por: Liao, Qianqian, et al.
Publicado: (2025)
Log Neural Controlled Differential Equations: The Lie Brackets Make a Difference
por: Walker, Benjamin, et al.
Publicado: (2024)
por: Walker, Benjamin, et al.
Publicado: (2024)
Preventing Data Leakage in EEG-Based Survival Prediction: A Two-Stage Embedding and Transformer Framework
por: Zhou, Yixin, et al.
Publicado: (2026)
por: Zhou, Yixin, et al.
Publicado: (2026)
A Unified Geometric Framework for Weighted Contrastive Learning
por: Vock, Raphael, et al.
Publicado: (2026)
por: Vock, Raphael, et al.
Publicado: (2026)
Automated Immunophenotyping Assessment for Diagnosing Childhood Acute Leukemia using Set-Transformers
por: Lygizou, Elpiniki Maria, et al.
Publicado: (2024)
por: Lygizou, Elpiniki Maria, et al.
Publicado: (2024)
Towards Trustworthy Audio Deepfake Detection: A Systematic Framework for Diagnosing and Mitigating Gender Bias
por: Fursule, Aishwarya, et al.
Publicado: (2026)
por: Fursule, Aishwarya, et al.
Publicado: (2026)
Diagnosing Transformers: Illuminating Feature Spaces for Clinical Decision-Making
por: Hsu, Aliyah R., et al.
Publicado: (2023)
por: Hsu, Aliyah R., et al.
Publicado: (2023)
Weight Updates as Activation Shifts: A Principled Framework for Steering
por: Adila, Dyah, et al.
Publicado: (2026)
por: Adila, Dyah, et al.
Publicado: (2026)
PEANuT: Parameter-Efficient Adaptation with Weight-aware Neural Tweakers
por: Zhong, Yibo, et al.
Publicado: (2024)
por: Zhong, Yibo, et al.
Publicado: (2024)
Equivalence of Context and Parameter Updates in Modern Transformer Blocks
por: Goldwaser, Adrian, et al.
Publicado: (2025)
por: Goldwaser, Adrian, et al.
Publicado: (2025)
Transformers Provably Learn Sparse XOR with Polylogarithmic Parameters
por: Han, Yaomengxi, et al.
Publicado: (2025)
por: Han, Yaomengxi, et al.
Publicado: (2025)
Optimal Parameter and Neuron Pruning for Out-of-Distribution Detection
por: Chen, Chao, et al.
Publicado: (2024)
por: Chen, Chao, et al.
Publicado: (2024)
Learning to Weight Parameters for Training Data Attribution
por: Li, Shuangqi, et al.
Publicado: (2025)
por: Li, Shuangqi, et al.
Publicado: (2025)
Attention Saturation and Gradient Suppression at Inflection Layers: Diagnosing and Mitigating Bottlenecks in Transformer Adaptation
por: Zixian, Wang
Publicado: (2025)
por: Zixian, Wang
Publicado: (2025)
Dynamic Weight Grafting: Localizing Finetuned Factual Knowledge in Transformers
por: Nief, Todd, et al.
Publicado: (2025)
por: Nief, Todd, et al.
Publicado: (2025)
Going Beyond the Edge: Distributed Inference of Transformer Models on Ultra-Low-Power Wireless Devices
por: Gräfe, Alexander, et al.
Publicado: (2026)
por: Gräfe, Alexander, et al.
Publicado: (2026)
A Two-Timescale Approach for Wireless Federated Learning with Parameter Freezing and Power Control
por: Ouyang, Jinhao, et al.
Publicado: (2025)
por: Ouyang, Jinhao, et al.
Publicado: (2025)
TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters
por: Wang, Haiyang, et al.
Publicado: (2024)
por: Wang, Haiyang, et al.
Publicado: (2024)
Aggregation Models with Optimal Weights for Distributed Gaussian Processes
por: Chen, Haoyuan, et al.
Publicado: (2024)
por: Chen, Haoyuan, et al.
Publicado: (2024)
LoGAH: Predicting 774-Million-Parameter Transformers using Graph HyperNetworks with 1/100 Parameters
por: Zhou, Xinyu, et al.
Publicado: (2024)
por: Zhou, Xinyu, et al.
Publicado: (2024)
Exact Conversion of In-Context Learning to Model Weights in Linearized-Attention Transformers
por: Chen, Brian K, et al.
Publicado: (2024)
por: Chen, Brian K, et al.
Publicado: (2024)
In-Context Algorithm Emulation in Fixed-Weight Transformers
por: Hu, Jerry Yao-Chieh, et al.
Publicado: (2025)
por: Hu, Jerry Yao-Chieh, et al.
Publicado: (2025)
Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL
por: Miahi, Erfan, et al.
Publicado: (2026)
por: Miahi, Erfan, et al.
Publicado: (2026)
Deep Out-of-Distribution Uncertainty Quantification via Weight Entropy Maximization
por: de Mathelin, Antoine, et al.
Publicado: (2023)
por: de Mathelin, Antoine, et al.
Publicado: (2023)
Ejemplares similares
-
Generalizing to New Dynamical Systems via Frequency Domain Adaptation
por: Qin, Tiexin, et al.
Publicado: (2025) -
To Augment or Not to Augment? Diagnosing Distributional Symmetry Breaking
por: Lawrence, Hannah, et al.
Publicado: (2025) -
Permutation Equivariant Neural Controlled Differential Equations for Dynamic Graph Representation Learning
por: Berndt, Torben, et al.
Publicado: (2025) -
DFReg: A Physics-Inspired Framework for Global Weight Distribution Regularization in Neural Networks
por: Ruggieri, Giovanni
Publicado: (2025) -
Scalable Weibull Graph Attention Autoencoder for Modeling Document Networks
por: Wang, Chaojie, et al.
Publicado: (2024)