Wasserstein Distances, Neuronal Entanglement, and Sparsity
Fuente:
arXiv
Saved in:
| Main Authors: | Sawmya, Shashata, Kong, Linghao, Markov, Ilia, Alistarh, Dan, Shavit, Nir |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Birth of Knowledge: Emergent Features across Time, Space, and Scale in Large Language Models
by: Sawmya, Shashata, et al.
Published: (2025)
by: Sawmya, Shashata, et al.
Published: (2025)
Expand Neurons, Not Parameters
by: Kong, Linghao, et al.
Published: (2025)
by: Kong, Linghao, et al.
Published: (2025)
Panza: Design and Analysis of a Fully-Local Personalized Text Writing Assistant
by: Nicolicioiu, Armand, et al.
Published: (2024)
by: Nicolicioiu, Armand, et al.
Published: (2024)
Cascade Detector Analysis and Application to Biomedical Microscopy
by: Athey, Thomas L., et al.
Published: (2025)
by: Athey, Thomas L., et al.
Published: (2025)
The Unseen Frontier: Pushing the Limits of LLM Sparsity with Surrogate-Free ADMM
by: Lee, Kwanhee, et al.
Published: (2025)
by: Lee, Kwanhee, et al.
Published: (2025)
Towards Combinatorial Interpretability of Neural Computation
by: Adler, Micah, et al.
Published: (2025)
by: Adler, Micah, et al.
Published: (2025)
NeuroADDA: Active Discriminative Domain Adaptation in Connectomic
by: Sawmya, Shashata, et al.
Published: (2025)
by: Sawmya, Shashata, et al.
Published: (2025)
Budgeted Broadcast: An Activity-Dependent Pruning Rule for Neural Network Efficiency
by: Meirovitch, Yaron, et al.
Published: (2025)
by: Meirovitch, Yaron, et al.
Published: (2025)
Time Series Forecasting via Direct Per-Step Probability Distribution Modeling
by: Kong, Linghao, et al.
Published: (2025)
by: Kong, Linghao, et al.
Published: (2025)
Negative Pre-activations Differentiate Syntax
by: Kong, Linghao, et al.
Published: (2025)
by: Kong, Linghao, et al.
Published: (2025)
Private Wasserstein Distance
by: Li, Wenqian, et al.
Published: (2024)
by: Li, Wenqian, et al.
Published: (2024)
Distance-Based Tree-Sliced Wasserstein Distance
by: Tran, Hoang V., et al.
Published: (2025)
by: Tran, Hoang V., et al.
Published: (2025)
Scalable Energy-Based Models via Adversarial Training: Unifying Discrimination and Generation
by: Yin, Xuwang, et al.
Published: (2025)
by: Yin, Xuwang, et al.
Published: (2025)
Learning to Interpret Weight Differences in Language Models
by: Goel, Avichal, et al.
Published: (2025)
by: Goel, Avichal, et al.
Published: (2025)
Spherical Tree-Sliced Wasserstein Distance
by: Tran, Viet-Hoang, et al.
Published: (2025)
by: Tran, Viet-Hoang, et al.
Published: (2025)
Tree-Sliced Wasserstein Distance with Nonlinear Projection
by: Tran, Thanh, et al.
Published: (2025)
by: Tran, Thanh, et al.
Published: (2025)
Forbidden Facts: An Investigation of Competing Objectives in Llama-2
by: Wang, Tony T., et al.
Published: (2023)
by: Wang, Tony T., et al.
Published: (2023)
Tree-Sliced Wasserstein Distance: A Geometric Perspective
by: Tran, Viet-Hoang, et al.
Published: (2024)
by: Tran, Viet-Hoang, et al.
Published: (2024)
MatryoshkaLoRA: Learning Accurate Hierarchical Low-Rank Representations for LLM Fine-Tuning
by: Modoranu, Ionut-Vlad, et al.
Published: (2026)
by: Modoranu, Ionut-Vlad, et al.
Published: (2026)
Measuring Time-Series Dataset Similarity using Wasserstein Distance
by: Chen, Hongjie, et al.
Published: (2025)
by: Chen, Hongjie, et al.
Published: (2025)
"Give Me BF16 or Give Me Death"? Accuracy-Performance Trade-Offs in LLM Quantization
by: Kurtic, Eldar, et al.
Published: (2024)
by: Kurtic, Eldar, et al.
Published: (2024)
RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust Adaptation
by: Nikdan, Mahdi, et al.
Published: (2024)
by: Nikdan, Mahdi, et al.
Published: (2024)
ECO: Quantized Training without Full-Precision Master Weights
by: Nikdan, Mahdi, et al.
Published: (2026)
by: Nikdan, Mahdi, et al.
Published: (2026)
MoGU: Mixture-of-Gaussians with Uncertainty-based Gating for Time Series Forecasting
by: Aviv, Gilad, et al.
Published: (2025)
by: Aviv, Gilad, et al.
Published: (2025)
Quantifying the Pre-training Dividend: Generative versus Latent Self-Supervised Learning for Time Series Foundation Models
by: Major, Noam, et al.
Published: (2026)
by: Major, Noam, et al.
Published: (2026)
Mathador-LM: A Dynamic Benchmark for Mathematical Reasoning on Large Language Models
by: Kurtic, Eldar, et al.
Published: (2024)
by: Kurtic, Eldar, et al.
Published: (2024)
FFT-based Dynamic Subspace Selection for Low-Rank Adaptive Optimization of Large Language Models
by: Modoranu, Ionut-Vlad, et al.
Published: (2025)
by: Modoranu, Ionut-Vlad, et al.
Published: (2025)
Stereographic Spherical Sliced Wasserstein Distances
by: Tran, Huy, et al.
Published: (2024)
by: Tran, Huy, et al.
Published: (2024)
An Interpretable Latency Model for Speculative Decoding in LLM Serving
by: Kong, Linghao, et al.
Published: (2026)
by: Kong, Linghao, et al.
Published: (2026)
LLM Advertisement based on Neuron Auctions
by: Yun, Peiran, et al.
Published: (2026)
by: Yun, Peiran, et al.
Published: (2026)
Improving Decision Sparsity
by: Sun, Yiyang, et al.
Published: (2024)
by: Sun, Yiyang, et al.
Published: (2024)
Homeostasis and Sparsity in Transformer
by: Kotyuzanskiy, Leonid, et al.
Published: (2024)
by: Kotyuzanskiy, Leonid, et al.
Published: (2024)
Polar Sparsity: High Throughput Batched LLM Inferencing with Scalable Contextual Sparsity
by: Shrestha, Susav, et al.
Published: (2025)
by: Shrestha, Susav, et al.
Published: (2025)
Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models
by: Abnar, Samira, et al.
Published: (2025)
by: Abnar, Samira, et al.
Published: (2025)
NeuRel-Attack: Neuron Relearning for Safety Disalignment in Large Language Models
by: Zhou, Yi, et al.
Published: (2025)
by: Zhou, Yi, et al.
Published: (2025)
Sparsity and Out-of-Distribution Generalization
by: Aaronson, Scott, et al.
Published: (2026)
by: Aaronson, Scott, et al.
Published: (2026)
Sparsity and Superposition in Mixture of Experts
by: Chaudhari, Marmik, et al.
Published: (2025)
by: Chaudhari, Marmik, et al.
Published: (2025)
Wasserstein Policy Optimization
by: Pfau, David, et al.
Published: (2025)
by: Pfau, David, et al.
Published: (2025)
Wasserstein Distance Rivals Kullback-Leibler Divergence for Knowledge Distillation
by: Lv, Jiaming, et al.
Published: (2024)
by: Lv, Jiaming, et al.
Published: (2024)
The False Dawn: Reevaluating Google's Reinforcement Learning for Chip Macro Placement
by: Markov, Igor L.
Published: (2023)
by: Markov, Igor L.
Published: (2023)
Similar Items
-
The Birth of Knowledge: Emergent Features across Time, Space, and Scale in Large Language Models
by: Sawmya, Shashata, et al.
Published: (2025) -
Expand Neurons, Not Parameters
by: Kong, Linghao, et al.
Published: (2025) -
Panza: Design and Analysis of a Fully-Local Personalized Text Writing Assistant
by: Nicolicioiu, Armand, et al.
Published: (2024) -
Cascade Detector Analysis and Application to Biomedical Microscopy
by: Athey, Thomas L., et al.
Published: (2025) -
The Unseen Frontier: Pushing the Limits of LLM Sparsity with Surrogate-Free ADMM
by: Lee, Kwanhee, et al.
Published: (2025)