Uncovering Model Processing Strategies with Non-Negative Per-Example Fisher Factorization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Matena, Michael, Raffel, Colin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Fishers for Free? Approximating the Fisher Information Matrix by Recycling the Squared Gradient Accumulator
von: Li, YuXin, et al.
Veröffentlicht: (2025)
von: Li, YuXin, et al.
Veröffentlicht: (2025)
Efficiently Estimating Data Efficiency for Language Model Fine-tuning
von: Je, Gyung Hyun, et al.
Veröffentlicht: (2025)
von: Je, Gyung Hyun, et al.
Veröffentlicht: (2025)
Merging by Matching Models in Task Parameter Subspaces
von: Tam, Derek, et al.
Veröffentlicht: (2023)
von: Tam, Derek, et al.
Veröffentlicht: (2023)
Soft Merging of Experts with Adaptive Routing
von: Muqeeth, Mohammed, et al.
Veröffentlicht: (2023)
von: Muqeeth, Mohammed, et al.
Veröffentlicht: (2023)
AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution
von: Liu, Fengyuan, et al.
Veröffentlicht: (2024)
von: Liu, Fengyuan, et al.
Veröffentlicht: (2024)
Position: The Most Expensive Part of an LLM should be its Training Data
von: Kandpal, Nikhil, et al.
Veröffentlicht: (2025)
von: Kandpal, Nikhil, et al.
Veröffentlicht: (2025)
Learning to Route Among Specialized Experts for Zero-Shot Generalization
von: Muqeeth, Mohammed, et al.
Veröffentlicht: (2024)
von: Muqeeth, Mohammed, et al.
Veröffentlicht: (2024)
FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale
von: Patel, Ajay, et al.
Veröffentlicht: (2026)
von: Patel, Ajay, et al.
Veröffentlicht: (2026)
DataDreamer: A Tool for Synthetic Data Generation and Reproducible LLM Workflows
von: Patel, Ajay, et al.
Veröffentlicht: (2024)
von: Patel, Ajay, et al.
Veröffentlicht: (2024)
The Butterfly Effect: Neural Network Training Trajectories Are Highly Sensitive to Initial Conditions
von: Kwok, Devin, et al.
Veröffentlicht: (2025)
von: Kwok, Devin, et al.
Veröffentlicht: (2025)
Enhancing Training Data Attribution with Representational Optimization
von: Sun, Weiwei, et al.
Veröffentlicht: (2025)
von: Sun, Weiwei, et al.
Veröffentlicht: (2025)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
von: Yadav, Prateek, et al.
Veröffentlicht: (2023)
von: Yadav, Prateek, et al.
Veröffentlicht: (2023)
Model Merging via Data-Free Covariance Estimation
von: Hameed, Marawan Gamal Abdel, et al.
Veröffentlicht: (2026)
von: Hameed, Marawan Gamal Abdel, et al.
Veröffentlicht: (2026)
Realistic Evaluation of Model Merging for Compositional Generalization
von: Tam, Derek, et al.
Veröffentlicht: (2024)
von: Tam, Derek, et al.
Veröffentlicht: (2024)
A PID-Controlled Non-Negative Tensor Factorization Model for Analyzing Missing Data in NILM
von: Shi, DengYu
Veröffentlicht: (2024)
von: Shi, DengYu
Veröffentlicht: (2024)
Controllable Game Level Generation: Assessing the Effect of Negative Examples in GAN Models
von: Bazzaz, Mahsa, et al.
Veröffentlicht: (2024)
von: Bazzaz, Mahsa, et al.
Veröffentlicht: (2024)
Learnable Similarity and Dissimilarity Guided Symmetric Non-Negative Matrix Factorization
von: Lyu, Wenlong, et al.
Veröffentlicht: (2024)
von: Lyu, Wenlong, et al.
Veröffentlicht: (2024)
Generalized Fisher-Weighted SVD: Scalable Kronecker-Factored Fisher Approximation for Compressing Large Language Models
von: Chekalina, Viktoriia, et al.
Veröffentlicht: (2025)
von: Chekalina, Viktoriia, et al.
Veröffentlicht: (2025)
FlashKAT: Understanding and Addressing Performance Bottlenecks in the Kolmogorov-Arnold Transformer
von: Raffel, Matthew, et al.
Veröffentlicht: (2025)
von: Raffel, Matthew, et al.
Veröffentlicht: (2025)
TokSuite: Measuring the Impact of Tokenizer Choice on Language Model Behavior
von: Altıntaş, Gül Sena, et al.
Veröffentlicht: (2025)
von: Altıntaş, Gül Sena, et al.
Veröffentlicht: (2025)
Normalization Layer Per-Example Gradients are Sufficient to Predict Gradient Noise Scale in Transformers
von: Gray, Gavia, et al.
Veröffentlicht: (2024)
von: Gray, Gavia, et al.
Veröffentlicht: (2024)
Wild Bootstrap Inference for Non-Negative Matrix Factorization with Random Effects
von: Satoh, Kenichi
Veröffentlicht: (2026)
von: Satoh, Kenichi
Veröffentlicht: (2026)
BeaverTalk: Oregon State University's IWSLT 2025 Simultaneous Speech Translation System
von: Raffel, Matthew, et al.
Veröffentlicht: (2025)
von: Raffel, Matthew, et al.
Veröffentlicht: (2025)
Simultaneous Masking, Not Prompting Optimization: A Paradigm Shift in Fine-tuning LLMs for Simultaneous Translation
von: Raffel, Matthew, et al.
Veröffentlicht: (2024)
von: Raffel, Matthew, et al.
Veröffentlicht: (2024)
The Target Polish: A New Approach to Outlier-Resistant Non-Negative Matrix Factorization
von: Fogel, Paul, et al.
Veröffentlicht: (2025)
von: Fogel, Paul, et al.
Veröffentlicht: (2025)
Non-Negative Matrix Factorization Using Non-Von Neumann Computers
von: Borle, Ajinkya, et al.
Veröffentlicht: (2025)
von: Borle, Ajinkya, et al.
Veröffentlicht: (2025)
Algorithms for Non-Negative Matrix Factorization on Noisy Data With Negative Values
von: Green, Dylan, et al.
Veröffentlicht: (2023)
von: Green, Dylan, et al.
Veröffentlicht: (2023)
Rethinking Non-Negative Matrix Factorization with Implicit Neural Representations
von: Subramani, Krishna, et al.
Veröffentlicht: (2024)
von: Subramani, Krishna, et al.
Veröffentlicht: (2024)
Dynamic QoS Prediction via a Non-Negative Tensor Snowflake Factorization
von: Xia, YongHui, et al.
Veröffentlicht: (2025)
von: Xia, YongHui, et al.
Veröffentlicht: (2025)
Neural Negative Binomial Regression for Weekly Seismicity Forecasting: Per-Cell Dispersion Estimation and Tail Risk Assessment
von: Igilik, Alim
Veröffentlicht: (2026)
von: Igilik, Alim
Veröffentlicht: (2026)
Catch-Only-One: Non-Transferable Examples for Model-Specific Authorization
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
von: Pan, Bowen, et al.
Veröffentlicht: (2024)
von: Pan, Bowen, et al.
Veröffentlicht: (2024)
MetaCluster: Enabling Deep Compression of Kolmogorov-Arnold Network
von: Raffel, Matthew, et al.
Veröffentlicht: (2025)
von: Raffel, Matthew, et al.
Veröffentlicht: (2025)
LUNE: Efficient LLM Unlearning via LoRA Fine-Tuning with Negative Examples
von: Liu, Yezi, et al.
Veröffentlicht: (2025)
von: Liu, Yezi, et al.
Veröffentlicht: (2025)
Per-Axis Weight Deltas for Frequent Model Updates
von: Kuyumdzhiev, Stefan, et al.
Veröffentlicht: (2025)
von: Kuyumdzhiev, Stefan, et al.
Veröffentlicht: (2025)
Exploiting Non-Negativity in DAG Structure Learning
von: Rey, Samuel, et al.
Veröffentlicht: (2026)
von: Rey, Samuel, et al.
Veröffentlicht: (2026)
Gaussian Process Tilted Nonparametric Density Estimation using Fisher Divergence Score Matching
von: Paisley, John, et al.
Veröffentlicht: (2025)
von: Paisley, John, et al.
Veröffentlicht: (2025)
Stratified Non-Negative Tensor Factorization
von: Sietsema, Alexander, et al.
Veröffentlicht: (2024)
von: Sietsema, Alexander, et al.
Veröffentlicht: (2024)
The Appeal and Reality of Recycling LoRAs with Adaptive Merging
von: Liu, Haokun, et al.
Veröffentlicht: (2026)
von: Liu, Haokun, et al.
Veröffentlicht: (2026)
Annotation Entropy Predicts Per-Example Learning Dynamics in LoRA Fine-Tuning
von: Steele, Brady
Veröffentlicht: (2026)
von: Steele, Brady
Veröffentlicht: (2026)
Ähnliche Einträge
-
Fishers for Free? Approximating the Fisher Information Matrix by Recycling the Squared Gradient Accumulator
von: Li, YuXin, et al.
Veröffentlicht: (2025) -
Efficiently Estimating Data Efficiency for Language Model Fine-tuning
von: Je, Gyung Hyun, et al.
Veröffentlicht: (2025) -
Merging by Matching Models in Task Parameter Subspaces
von: Tam, Derek, et al.
Veröffentlicht: (2023) -
Soft Merging of Experts with Adaptive Routing
von: Muqeeth, Mohammed, et al.
Veröffentlicht: (2023) -
AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution
von: Liu, Fengyuan, et al.
Veröffentlicht: (2024)