SAGE: Sign-Adaptive Gradient for Memory-Efficient LLM Optimization
Fuente:
arXiv
Salvato in:
| Autori principali: | Lee, Wooin, Kim, Hyun-Tae |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
di: Basu, Abhinaba
Pubblicazione: (2026)
di: Basu, Abhinaba
Pubblicazione: (2026)
NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution
di: Breneur, Oleksandr Marchenko, et al.
Pubblicazione: (2026)
di: Breneur, Oleksandr Marchenko, et al.
Pubblicazione: (2026)
Rethinking the Multilingual Reasoning Gap with Layer Swap
di: Lasbordes, Maxence, et al.
Pubblicazione: (2026)
di: Lasbordes, Maxence, et al.
Pubblicazione: (2026)
Harnessing non-adversarial robustness in large language models
di: Zhou, Qinghua, et al.
Pubblicazione: (2026)
di: Zhou, Qinghua, et al.
Pubblicazione: (2026)
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
di: Dang, Kieu, et al.
Pubblicazione: (2025)
di: Dang, Kieu, et al.
Pubblicazione: (2025)
Parameter-Efficient Neuroevolution for Diverse LLM Generation: Quality-Diversity Optimization via Prompt Embedding Evolution
di: Guo, Dongxin, et al.
Pubblicazione: (2026)
di: Guo, Dongxin, et al.
Pubblicazione: (2026)
CLMN: Concept based Language Models via Neural Symbolic Reasoning
di: Yang, Yibo
Pubblicazione: (2025)
di: Yang, Yibo
Pubblicazione: (2025)
From Noise to Diversity: Random Embedding Injection in LLM Reasoning
di: Kim, Heejun, et al.
Pubblicazione: (2026)
di: Kim, Heejun, et al.
Pubblicazione: (2026)
How Pruning Reshapes Features: Sparse Autoencoder Analysis of Weight-Pruned Language Models
di: Borobia, Hector, et al.
Pubblicazione: (2026)
di: Borobia, Hector, et al.
Pubblicazione: (2026)
AIPsy-Affect: A Keyword-Free Clinical Stimulus Battery for Mechanistic Interpretability of Emotion in Language Models
di: Keeman, Michael
Pubblicazione: (2026)
di: Keeman, Michael
Pubblicazione: (2026)
The Concept Allocation Zone: Tracking How Concepts Form Across Transformer Depth
di: Henry, James
Pubblicazione: (2026)
di: Henry, James
Pubblicazione: (2026)
Product-of-Experts Training Reduces Dataset Artifacts in Natural Language Inference
di: Mathew, Aby Mammen
Pubblicazione: (2026)
di: Mathew, Aby Mammen
Pubblicazione: (2026)
Enhancing Ultra-Low-Bit Quantization of Large Language Models Through Saliency-Aware Partial Retraining
di: Cao, Deyu, et al.
Pubblicazione: (2025)
di: Cao, Deyu, et al.
Pubblicazione: (2025)
Measuring Alignment-Induced Activation Shifts Correctly: A Template-Controlled Difference-in-Differences Protocol
di: Nakamura, Yuki
Pubblicazione: (2026)
di: Nakamura, Yuki
Pubblicazione: (2026)
ProactBench: Beyond What The User Asked For
di: Harfi, Sepehr, et al.
Pubblicazione: (2026)
di: Harfi, Sepehr, et al.
Pubblicazione: (2026)
ProbeScale: Probing Analysis to Optimize Neural Scaling Laws for Efficient Small Language Model Inference
di: Das, Sourav
Pubblicazione: (2026)
di: Das, Sourav
Pubblicazione: (2026)
Mechanistic Analysis of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
di: Imanov, Olaf Yunus Laitinen
Pubblicazione: (2026)
di: Imanov, Olaf Yunus Laitinen
Pubblicazione: (2026)
Latent Object Permanence: Topological Phase Transitions, Free-Energy Principles, and Renormalization Group Flows in Deep Transformer Manifolds
di: Alpay, Faruk, et al.
Pubblicazione: (2026)
di: Alpay, Faruk, et al.
Pubblicazione: (2026)
Inference acceleration for large language models using "stairs" assisted greedy generation
di: Grigaliūnas, Domas, et al.
Pubblicazione: (2024)
di: Grigaliūnas, Domas, et al.
Pubblicazione: (2024)
Strategic Doctrine Language Models (sdLM): A Learning-System Framework for Doctrinal Consistency and Geopolitical Forecasting
di: Imanov, Olaf Yunus Laitinen, et al.
Pubblicazione: (2026)
di: Imanov, Olaf Yunus Laitinen, et al.
Pubblicazione: (2026)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
di: Pather, Kaviraj, et al.
Pubblicazione: (2025)
di: Pather, Kaviraj, et al.
Pubblicazione: (2025)
Unsolvability Ceiling in Multi-LLM Routing: An Empirical Study of Evaluation Artifacts
di: Garg, Saloni, et al.
Pubblicazione: (2026)
di: Garg, Saloni, et al.
Pubblicazione: (2026)
Causal Dimensionality of Transformer Representations: Measurement, Scaling, and Layer Structure
di: Sarkar, Nilesh, et al.
Pubblicazione: (2026)
di: Sarkar, Nilesh, et al.
Pubblicazione: (2026)
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
di: Guo, Dongxin, et al.
Pubblicazione: (2026)
di: Guo, Dongxin, et al.
Pubblicazione: (2026)
Unpacking Hateful Memes: Presupposed Context and False Claims
di: Cai, Weibin, et al.
Pubblicazione: (2025)
di: Cai, Weibin, et al.
Pubblicazione: (2025)
Sliced-Wasserstein Distribution Alignment Loss Improves the Ultra-Low-Bit Quantization of Large Language Models
di: Cao, Deyu, et al.
Pubblicazione: (2026)
di: Cao, Deyu, et al.
Pubblicazione: (2026)
EvoPref: Multi-Objective Evolutionary Optimization Discovers Diverse LLM Alignments Beyond Gradient Descent
di: Guo, Dongxin, et al.
Pubblicazione: (2026)
di: Guo, Dongxin, et al.
Pubblicazione: (2026)
Understanding Reinforcement Learning for Model Training, and future directions with GRAPE
di: Patel, Rohit
Pubblicazione: (2025)
di: Patel, Rohit
Pubblicazione: (2025)
Do Reasoning Models Enhance Embedding Models?
di: Chan, Wun Yu, et al.
Pubblicazione: (2026)
di: Chan, Wun Yu, et al.
Pubblicazione: (2026)
Research on a hybrid LSTM-CNN-Attention model for text-based web content classification
di: Kuz, Mykola, et al.
Pubblicazione: (2025)
di: Kuz, Mykola, et al.
Pubblicazione: (2025)
Linguistic Collapse: Neural Collapse in (Large) Language Models
di: Wu, Robert, et al.
Pubblicazione: (2024)
di: Wu, Robert, et al.
Pubblicazione: (2024)
optimize_anything: A Universal API for Optimizing any Text Parameter
di: Agrawal, Lakshya A, et al.
Pubblicazione: (2026)
di: Agrawal, Lakshya A, et al.
Pubblicazione: (2026)
HEFT: A Coarse-to-Fine Hierarchy for Enhancing the Efficiency and Accuracy of Language Model Reasoning
di: Hill, Brennen
Pubblicazione: (2025)
di: Hill, Brennen
Pubblicazione: (2025)
MAcPNN: Mutual Assisted Learning on Data Streams with Temporal Dependence
di: Giannini, Federico, et al.
Pubblicazione: (2026)
di: Giannini, Federico, et al.
Pubblicazione: (2026)
mHC-SSM: Manifold-Constrained Hyper-Connections for State Space Language Models with Stream-Specialized Adapters
di: Mutlu, Abdulvahap, et al.
Pubblicazione: (2026)
di: Mutlu, Abdulvahap, et al.
Pubblicazione: (2026)
Extracting Sentence Embeddings from Pretrained Transformer Models
di: Stankevičius, Lukas, et al.
Pubblicazione: (2024)
di: Stankevičius, Lukas, et al.
Pubblicazione: (2024)
Sentiment Analysis of Lithuanian Online Reviews Using Large Language Models
di: Vileikytė, Brigita, et al.
Pubblicazione: (2024)
di: Vileikytė, Brigita, et al.
Pubblicazione: (2024)
ReFactor GNNs: Revisiting Factorisation-based Models from a Message-Passing Perspective
di: Chen, Yihong, et al.
Pubblicazione: (2022)
di: Chen, Yihong, et al.
Pubblicazione: (2022)
How much do LLMs learn from negative examples?
di: Hamdan, Shadi, et al.
Pubblicazione: (2025)
di: Hamdan, Shadi, et al.
Pubblicazione: (2025)
Communicative Agents for Slideshow Storytelling Video Generation based on LLMs
di: Fan, Jingxing, et al.
Pubblicazione: (2025)
di: Fan, Jingxing, et al.
Pubblicazione: (2025)
Documenti analoghi
-
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
di: Basu, Abhinaba
Pubblicazione: (2026) -
NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution
di: Breneur, Oleksandr Marchenko, et al.
Pubblicazione: (2026) -
Rethinking the Multilingual Reasoning Gap with Layer Swap
di: Lasbordes, Maxence, et al.
Pubblicazione: (2026) -
Harnessing non-adversarial robustness in large language models
di: Zhou, Qinghua, et al.
Pubblicazione: (2026) -
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
di: Dang, Kieu, et al.
Pubblicazione: (2025)