Massive Activations in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sun, Mingjie, Chen, Xinlei, Kolter, J. Zico, Liu, Zhuang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Simple and Effective Pruning Approach for Large Language Models
von: Sun, Mingjie, et al.
Veröffentlicht: (2023)
von: Sun, Mingjie, et al.
Veröffentlicht: (2023)
FUSE-ing Language Models: Zero-Shot Adapter Discovery for Prompt Optimization Across Tokenizers
von: Williams, Joshua Nathaniel, et al.
Veröffentlicht: (2024)
von: Williams, Joshua Nathaniel, et al.
Veröffentlicht: (2024)
Idiosyncrasies in Large Language Models
von: Sun, Mingjie, et al.
Veröffentlicht: (2025)
von: Sun, Mingjie, et al.
Veröffentlicht: (2025)
Forcing Diffuse Distributions out of Language Models
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)
Predicting the Performance of Black-box LLMs through Follow-up Queries
von: Sam, Dylan, et al.
Veröffentlicht: (2025)
von: Sam, Dylan, et al.
Veröffentlicht: (2025)
Compressed Sensing for Capability Localization in Large Language Models
von: Bair, Anna, et al.
Veröffentlicht: (2026)
von: Bair, Anna, et al.
Veröffentlicht: (2026)
Context-Parametric Inversion: Why Instruction Finetuning Can Worsen Context Reliance
von: Goyal, Sachin, et al.
Veröffentlicht: (2024)
von: Goyal, Sachin, et al.
Veröffentlicht: (2024)
Mimetic Initialization Helps State Space Models Learn to Recall
von: Trockman, Asher, et al.
Veröffentlicht: (2024)
von: Trockman, Asher, et al.
Veröffentlicht: (2024)
Existing Large Language Model Unlearning Evaluations Are Inconclusive
von: Feng, Zhili, et al.
Veröffentlicht: (2025)
von: Feng, Zhili, et al.
Veröffentlicht: (2025)
Looking beyond the next token
von: Thankaraj, Abitha, et al.
Veröffentlicht: (2025)
von: Thankaraj, Abitha, et al.
Veröffentlicht: (2025)
TOFU: A Task of Fictitious Unlearning for LLMs
von: Maini, Pratyush, et al.
Veröffentlicht: (2024)
von: Maini, Pratyush, et al.
Veröffentlicht: (2024)
Rethinking LLM Memorization through the Lens of Adversarial Compression
von: Schwarzschild, Avi, et al.
Veröffentlicht: (2024)
von: Schwarzschild, Avi, et al.
Veröffentlicht: (2024)
Base Models Look Human To AI Detectors
von: Xu, Yixuan Even, et al.
Veröffentlicht: (2026)
von: Xu, Yixuan Even, et al.
Veröffentlicht: (2026)
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
von: Xu, Yixuan Even, et al.
Veröffentlicht: (2025)
von: Xu, Yixuan Even, et al.
Veröffentlicht: (2025)
T-MARS: Improving Visual Representations by Circumventing Text Feature Learning
von: Maini, Pratyush, et al.
Veröffentlicht: (2023)
von: Maini, Pratyush, et al.
Veröffentlicht: (2023)
AcceleratedLiNGAM: Learning Causal DAGs at the speed of GPUs
von: Akinwande, Victor, et al.
Veröffentlicht: (2024)
von: Akinwande, Victor, et al.
Veröffentlicht: (2024)
Mimetic Initialization of MLPs
von: Trockman, Asher, et al.
Veröffentlicht: (2026)
von: Trockman, Asher, et al.
Veröffentlicht: (2026)
One-Step Diffusion Distillation via Deep Equilibrium Models
von: Geng, Zhengyang, et al.
Veröffentlicht: (2023)
von: Geng, Zhengyang, et al.
Veröffentlicht: (2023)
ActTail: Global Activation Sparsity in Large Language Models
von: Hou, Wenwen, et al.
Veröffentlicht: (2026)
von: Hou, Wenwen, et al.
Veröffentlicht: (2026)
Beyond Size: How Gradients Shape Pruning Decisions in Large Language Models
von: Das, Rocktim Jyoti, et al.
Veröffentlicht: (2023)
von: Das, Rocktim Jyoti, et al.
Veröffentlicht: (2023)
Massive Editing for Large Language Models via Meta Learning
von: Tan, Chenmien, et al.
Veröffentlicht: (2023)
von: Tan, Chenmien, et al.
Veröffentlicht: (2023)
First Activations Matter: Training-Free Methods for Dynamic Activation in Large Language Models
von: Ma, Chi, et al.
Veröffentlicht: (2024)
von: Ma, Chi, et al.
Veröffentlicht: (2024)
Training a Generally Curious Agent
von: Tajwar, Fahim, et al.
Veröffentlicht: (2025)
von: Tajwar, Fahim, et al.
Veröffentlicht: (2025)
Reasoning-preserved Efficient Distillation of Large Language Models via Activation-aware Initialization
von: He, Junlin, et al.
Veröffentlicht: (2026)
von: He, Junlin, et al.
Veröffentlicht: (2026)
Antidistillation Fingerprinting
von: Xu, Yixuan Even, et al.
Veröffentlicht: (2026)
von: Xu, Yixuan Even, et al.
Veröffentlicht: (2026)
Diffusing Differentiable Representations
von: Savani, Yash, et al.
Veröffentlicht: (2024)
von: Savani, Yash, et al.
Veröffentlicht: (2024)
Robust and Scalable Model Editing for Large Language Models
von: Chen, Yingfa, et al.
Veröffentlicht: (2024)
von: Chen, Yingfa, et al.
Veröffentlicht: (2024)
Exploring Scaling Laws for Local SGD in Large Language Model Training
von: He, Qiaozhi, et al.
Veröffentlicht: (2024)
von: He, Qiaozhi, et al.
Veröffentlicht: (2024)
Test-Time Adaptation Induces Stronger Accuracy and Agreement-on-the-Line
von: Kim, Eungyeup, et al.
Veröffentlicht: (2023)
von: Kim, Eungyeup, et al.
Veröffentlicht: (2023)
Sparsing Law: Towards Large Language Models with Greater Activation Sparsity
von: Luo, Yuqi, et al.
Veröffentlicht: (2024)
von: Luo, Yuqi, et al.
Veröffentlicht: (2024)
Hierarchical Orthogonal Residual Spread for Precise Massive Editing in Large Language Models
von: Gu, Xiaojie, et al.
Veröffentlicht: (2026)
von: Gu, Xiaojie, et al.
Veröffentlicht: (2026)
FacLens: Transferable Probe for Foreseeing Non-Factuality in Fact-Seeking Question Answering of Large Language Models
von: Wang, Yanling, et al.
Veröffentlicht: (2024)
von: Wang, Yanling, et al.
Veröffentlicht: (2024)
CoreInfer: Accelerating Large Language Model Inference with Semantics-Inspired Adaptive Sparse Activation
von: Wang, Qinsi, et al.
Veröffentlicht: (2024)
von: Wang, Qinsi, et al.
Veröffentlicht: (2024)
On the Thinking-Language Modeling Gap in Large Language Models
von: Liu, Chenxi, et al.
Veröffentlicht: (2025)
von: Liu, Chenxi, et al.
Veröffentlicht: (2025)
LOLA -- An Open-Source Massively Multilingual Large Language Model
von: Srivastava, Nikit, et al.
Veröffentlicht: (2024)
von: Srivastava, Nikit, et al.
Veröffentlicht: (2024)
Large Language Models as Agents in Two-Player Games
von: Liu, Yang, et al.
Veröffentlicht: (2024)
von: Liu, Yang, et al.
Veröffentlicht: (2024)
Steering Language Models With Activation Engineering
von: Turner, Alexander Matt, et al.
Veröffentlicht: (2023)
von: Turner, Alexander Matt, et al.
Veröffentlicht: (2023)
Q-Sparse: All Large Language Models can be Fully Sparsely-Activated
von: Wang, Hongyu, et al.
Veröffentlicht: (2024)
von: Wang, Hongyu, et al.
Veröffentlicht: (2024)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
Stronger Normalization-Free Transformers
von: Chen, Mingzhi, et al.
Veröffentlicht: (2025)
von: Chen, Mingzhi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Simple and Effective Pruning Approach for Large Language Models
von: Sun, Mingjie, et al.
Veröffentlicht: (2023) -
FUSE-ing Language Models: Zero-Shot Adapter Discovery for Prompt Optimization Across Tokenizers
von: Williams, Joshua Nathaniel, et al.
Veröffentlicht: (2024) -
Idiosyncrasies in Large Language Models
von: Sun, Mingjie, et al.
Veröffentlicht: (2025) -
Forcing Diffuse Distributions out of Language Models
von: Zhang, Yiming, et al.
Veröffentlicht: (2024) -
Predicting the Performance of Black-box LLMs through Follow-up Queries
von: Sam, Dylan, et al.
Veröffentlicht: (2025)