GaLLoP: Gradient-based Sparse Learning on Low-Magnitude Parameters
Fuente:
arXiv
Saved in:
| Main Authors: | Choudhary, Anand, Sulaıman, Yasser, Mauch, Lukas, Hacene, Ghouthi Boukli, Cardinaux, Fabien, Bosselut, Antoine |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Residual Connections and the Causal Shift: Uncovering a Structural Misalignment in Transformers
by: Lys, Jonathan, et al.
Published: (2026)
by: Lys, Jonathan, et al.
Published: (2026)
D5P4: Partition Determinantal Point Process for Diversity in Parallel Discrete Diffusion Decoding
by: Lys, Jonathan, et al.
Published: (2026)
by: Lys, Jonathan, et al.
Published: (2026)
A Novel Benchmark for Few-Shot Semantic Segmentation in the Era of Foundation Models
by: Bensaid, Reda, et al.
Published: (2024)
by: Bensaid, Reda, et al.
Published: (2024)
Inner Loop Inference for Pretrained Transformers: Unlocking Latent Capabilities Without Training
by: Lys, Jonathan, et al.
Published: (2026)
by: Lys, Jonathan, et al.
Published: (2026)
LLM meets Vision-Language Models for Zero-Shot One-Class Classification
by: Bendou, Yassir, et al.
Published: (2024)
by: Bendou, Yassir, et al.
Published: (2024)
SKILL: Similarity-aware Knowledge distILLation for Speech Self-Supervised Learning
by: Zampierin, Luca, et al.
Published: (2024)
by: Zampierin, Luca, et al.
Published: (2024)
LLoCO: Learning Long Contexts Offline
by: Tan, Sijun, et al.
Published: (2024)
by: Tan, Sijun, et al.
Published: (2024)
Towards Robust FastSpeech 2 by Modelling Residual Multimodality
by: Kögel, Fabian, et al.
Published: (2023)
by: Kögel, Fabian, et al.
Published: (2023)
PERK: Long-Context Reasoning as Parameter-Efficient Test-Time Learning
by: Chen, Zeming, et al.
Published: (2025)
by: Chen, Zeming, et al.
Published: (2025)
PICLe: Pseudo-Annotations for In-Context Learning in Low-Resource Named Entity Detection
by: Mamooler, Sepideh, et al.
Published: (2024)
by: Mamooler, Sepideh, et al.
Published: (2024)
SAFT: Towards Out-of-Distribution Generalization in Fine-Tuning
by: Nguyen, Bac, et al.
Published: (2024)
by: Nguyen, Bac, et al.
Published: (2024)
ConLID: Supervised Contrastive Learning for Low-Resource Language Identification
by: Foroutan, Negar, et al.
Published: (2025)
by: Foroutan, Negar, et al.
Published: (2025)
RLMEval: Evaluating Research-Level Neural Theorem Proving
by: Poiroux, Auguste, et al.
Published: (2025)
by: Poiroux, Auguste, et al.
Published: (2025)
Crosscoding Through Time: Tracking Emergence & Consolidation Of Linguistic Representations Throughout LLM Pretraining
by: Bayazit, Deniz, et al.
Published: (2025)
by: Bayazit, Deniz, et al.
Published: (2025)
JOBSKAPE: A Framework for Generating Synthetic Job Postings to Enhance Skill Matching
by: Magron, Antoine, et al.
Published: (2024)
by: Magron, Antoine, et al.
Published: (2024)
Let Me Teach You: Pedagogical Foundations of Feedback for Language Models
by: Borges, Beatriz, et al.
Published: (2023)
by: Borges, Beatriz, et al.
Published: (2023)
Making Reasoning Matter: Measuring and Improving Faithfulness of Chain-of-Thought Reasoning
by: Paul, Debjit, et al.
Published: (2024)
by: Paul, Debjit, et al.
Published: (2024)
On multi-token prediction for efficient LLM inference
by: Mehra, Somesh, et al.
Published: (2025)
by: Mehra, Somesh, et al.
Published: (2025)
Diversity Is All You Need for Contrastive Learning: Spectral Bounds on Gradient Magnitudes
by: Ochieng, Peter
Published: (2025)
by: Ochieng, Peter
Published: (2025)
Complex Reasoning over Logical Queries on Commonsense Knowledge Graphs
by: Fang, Tianqing, et al.
Published: (2024)
by: Fang, Tianqing, et al.
Published: (2024)
Rethinking Skill Extraction in the Job Market Domain using Large Language Models
by: Nguyen, Khanh Cao, et al.
Published: (2024)
by: Nguyen, Khanh Cao, et al.
Published: (2024)
AbstRaL: Augmenting LLMs' Reasoning by Reinforcing Abstract Thinking
by: Gao, Silin, et al.
Published: (2025)
by: Gao, Silin, et al.
Published: (2025)
Reliable Evaluation and Benchmarks for Statement Autoformalization
by: Poiroux, Auguste, et al.
Published: (2024)
by: Poiroux, Auguste, et al.
Published: (2024)
LLMs Are In-Context Bandit Reinforcement Learners
by: Monea, Giovanni, et al.
Published: (2024)
by: Monea, Giovanni, et al.
Published: (2024)
The LLM Language Network: A Neuroscientific Approach for Identifying Causally Task-Relevant Units
by: AlKhamissi, Badr, et al.
Published: (2024)
by: AlKhamissi, Badr, et al.
Published: (2024)
Brain-Like Language Processing via a Shallow Untrained Multihead Attention Network
by: AlKhamissi, Badr, et al.
Published: (2024)
by: AlKhamissi, Badr, et al.
Published: (2024)
Do LLMs Game Formalization? Evaluating Faithfulness in Logical Reasoning
by: Kim, Kyuhee, et al.
Published: (2026)
by: Kim, Kyuhee, et al.
Published: (2026)
Positional Fragility in LLMs: How Offset Effects Reshape Our Understanding of Memorization Risks
by: Xu, Yixuan, et al.
Published: (2025)
by: Xu, Yixuan, et al.
Published: (2025)
"Flex Tape Can't Fix That": Bias and Misinformation in Edited Language Models
by: Halevy, Karina, et al.
Published: (2024)
by: Halevy, Karina, et al.
Published: (2024)
Revisiting Multilingual Data Mixtures in Language Model Pretraining
by: Foroutan, Negar, et al.
Published: (2025)
by: Foroutan, Negar, et al.
Published: (2025)
Creativity in AI: Progresses and Challenges
by: Ismayilzada, Mete, et al.
Published: (2024)
by: Ismayilzada, Mete, et al.
Published: (2024)
Tracking the Limits of Knowledge Propagation: How LLMs Fail at Multi-Step Reasoning with Conflicting Knowledge
by: Feng, Yiyang, et al.
Published: (2026)
by: Feng, Yiyang, et al.
Published: (2026)
LSR-Adapt: Ultra-Efficient Parameter Tuning with Matrix Low Separation Rank Kernel Adaptation
by: Li, Xin, et al.
Published: (2025)
by: Li, Xin, et al.
Published: (2025)
What Really Drives Language Learning Success: Talent or Hard Work?
by: Yasser Teimouri
Published: (2025)
by: Yasser Teimouri
Published: (2025)
Instruction-tuning Aligns LLMs to the Human Brain
by: Aw, Khai Loong, et al.
Published: (2023)
by: Aw, Khai Loong, et al.
Published: (2023)
DiffuCOMET: Contextual Commonsense Knowledge Diffusion
by: Gao, Silin, et al.
Published: (2024)
by: Gao, Silin, et al.
Published: (2024)
Discovering Knowledge-Critical Subnetworks in Pretrained Language Models
by: Bayazit, Deniz, et al.
Published: (2023)
by: Bayazit, Deniz, et al.
Published: (2023)
Helpful to a Fault: Measuring Illicit Assistance in Multi-Turn, Multilingual LLM Agents
by: Talokar, Nivya, et al.
Published: (2026)
by: Talokar, Nivya, et al.
Published: (2026)
From Language to Cognition: How LLMs Outgrow the Human Language Network
by: AlKhamissi, Badr, et al.
Published: (2025)
by: AlKhamissi, Badr, et al.
Published: (2025)
Motion Planning for Automata-based Objectives using Efficient Gradient-based Methods
by: Balakrishnan, Anand, et al.
Published: (2024)
by: Balakrishnan, Anand, et al.
Published: (2024)
Similar Items
-
Residual Connections and the Causal Shift: Uncovering a Structural Misalignment in Transformers
by: Lys, Jonathan, et al.
Published: (2026) -
D5P4: Partition Determinantal Point Process for Diversity in Parallel Discrete Diffusion Decoding
by: Lys, Jonathan, et al.
Published: (2026) -
A Novel Benchmark for Few-Shot Semantic Segmentation in the Era of Foundation Models
by: Bensaid, Reda, et al.
Published: (2024) -
Inner Loop Inference for Pretrained Transformers: Unlocking Latent Capabilities Without Training
by: Lys, Jonathan, et al.
Published: (2026) -
LLM meets Vision-Language Models for Zero-Shot One-Class Classification
by: Bendou, Yassir, et al.
Published: (2024)