ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Song, Chenyang, Han, Xu, Zhang, Zhengyan, Hu, Shengding, Shi, Xiyu, Li, Kuai, Chen, Chen, Liu, Zhiyuan, Li, Guangli, Yang, Tao, Sun, Maosong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Sparsing Law: Towards Large Language Models with Greater Activation Sparsity
von: Luo, Yuqi, et al.
Veröffentlicht: (2024)
von: Luo, Yuqi, et al.
Veröffentlicht: (2024)
ConPET: Continual Parameter-Efficient Tuning for Large Language Models
von: Song, Chenyang, et al.
Veröffentlicht: (2023)
von: Song, Chenyang, et al.
Veröffentlicht: (2023)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
Learning to Generate Structured Output with Schema Reinforcement Learning
von: Lu, Yaxi, et al.
Veröffentlicht: (2025)
von: Lu, Yaxi, et al.
Veröffentlicht: (2025)
LoRS: Efficient Low-Rank Adaptation for Sparse Large Language Model
von: Hu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Hu, Yuxuan, et al.
Veröffentlicht: (2025)
Low-Resource Court Judgment Summarization for Common Law Systems
von: Liu, Shuaiqi, et al.
Veröffentlicht: (2024)
von: Liu, Shuaiqi, et al.
Veröffentlicht: (2024)
A Library of LLM Intrinsics for Retrieval-Augmented Generation
von: Danilevsky, Marina, et al.
Veröffentlicht: (2025)
von: Danilevsky, Marina, et al.
Veröffentlicht: (2025)
QUAD: Quantization and Parameter-Efficient Tuning of LLM with Activation Decomposition
von: Hu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Hu, Yuxuan, et al.
Veröffentlicht: (2025)
Intrinsic Evaluation of RAG Systems for Deep-Logic Questions
von: Hu, Junyi, et al.
Veröffentlicht: (2024)
von: Hu, Junyi, et al.
Veröffentlicht: (2024)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
Learning Software Bug Reports: A Systematic Literature Review
von: Long, Guoming, et al.
Veröffentlicht: (2025)
von: Long, Guoming, et al.
Veröffentlicht: (2025)
Introducing Brain-like Concepts to Embodied Hand-crafted Dialog Management System
von: Joublin, Frank, et al.
Veröffentlicht: (2024)
von: Joublin, Frank, et al.
Veröffentlicht: (2024)
After Retrieval, Before Generation: Enhancing the Trustworthiness of Large Language Models in Retrieval-Augmented Generation
von: Dai, Xinbang, et al.
Veröffentlicht: (2025)
von: Dai, Xinbang, et al.
Veröffentlicht: (2025)
LLM-Ref: Enhancing Reference Handling in Technical Writing with Large Language Models
von: Fuad, Kazi Ahmed Asif, et al.
Veröffentlicht: (2024)
von: Fuad, Kazi Ahmed Asif, et al.
Veröffentlicht: (2024)
Unstructured Text Enhanced Open-domain Dialogue System: A Systematic Survey
von: Ma, Longxuan, et al.
Veröffentlicht: (2024)
von: Ma, Longxuan, et al.
Veröffentlicht: (2024)
Deciphering Digital Detectives: Understanding LLM Behaviors and Capabilities in Multi-Agent Mystery Games
von: Wu, Dekun, et al.
Veröffentlicht: (2023)
von: Wu, Dekun, et al.
Veröffentlicht: (2023)
$\rm SP^3$: Enhancing Structured Pruning via PCA Projection
von: Hu, Yuxuan, et al.
Veröffentlicht: (2023)
von: Hu, Yuxuan, et al.
Veröffentlicht: (2023)
Refining Packing and Shuffling Strategies for Enhanced Performance in Generative Language Models
von: Chen, Yanbing, et al.
Veröffentlicht: (2024)
von: Chen, Yanbing, et al.
Veröffentlicht: (2024)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
von: Saji, Alan, et al.
Veröffentlicht: (2025)
von: Saji, Alan, et al.
Veröffentlicht: (2025)
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
AMELI: Enhancing Multimodal Entity Linking with Fine-Grained Attributes
von: Yao, Barry Menglong, et al.
Veröffentlicht: (2023)
von: Yao, Barry Menglong, et al.
Veröffentlicht: (2023)
Are LLM Uncertainty and Correctness Encoded by the Same Features? A Functional Dissociation via Sparse Autoencoders
von: Patel, Het, et al.
Veröffentlicht: (2026)
von: Patel, Het, et al.
Veröffentlicht: (2026)
AsyncTLS: Efficient Generative LLM Inference with Asynchronous Two-level Sparse Attention
von: Hu, Yuxuan, et al.
Veröffentlicht: (2026)
von: Hu, Yuxuan, et al.
Veröffentlicht: (2026)
Thought-Like-Pro: Enhancing Reasoning of Large Language Models through Self-Driven Prolog-based Chain-of-Thought
von: Tan, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Tan, Xiaoyu, et al.
Veröffentlicht: (2024)
Introducing Three New Benchmark Datasets for Hierarchical Text Classification
von: Toit, Jaco du, et al.
Veröffentlicht: (2024)
von: Toit, Jaco du, et al.
Veröffentlicht: (2024)
Preventing Safety Drift in Large Language Models via Coupled Weight and Activation Constraints
von: Peng, Songping, et al.
Veröffentlicht: (2026)
von: Peng, Songping, et al.
Veröffentlicht: (2026)
Proactive Agent: Shifting LLM Agents from Reactive Responses to Active Assistance
von: Lu, Yaxi, et al.
Veröffentlicht: (2024)
von: Lu, Yaxi, et al.
Veröffentlicht: (2024)
EVM-QuestBench: An Execution-Grounded Benchmark for Natural-Language Transaction Code Generation
von: Yang, Pei, et al.
Veröffentlicht: (2026)
von: Yang, Pei, et al.
Veröffentlicht: (2026)
OPOR-Bench: Evaluating Large Language Models on Online Public Opinion Report Generation
von: Yu, Jinzheng, et al.
Veröffentlicht: (2025)
von: Yu, Jinzheng, et al.
Veröffentlicht: (2025)
The Unified Cognitive Consciousness Theory for Language Models: Anchoring Semantics, Thresholds of Activation, and Emergent Reasoning
von: Chang, Edward Y., et al.
Veröffentlicht: (2025)
von: Chang, Edward Y., et al.
Veröffentlicht: (2025)
An Investigation of Neuron Activation as a Unified Lens to Explain Chain-of-Thought Eliciting Arithmetic Reasoning of LLMs
von: Rai, Daking, et al.
Veröffentlicht: (2024)
von: Rai, Daking, et al.
Veröffentlicht: (2024)
MALoRA: Mixture of Asymmetric Low-Rank Adaptation for Enhanced Multi-Task Learning
von: Wang, Xujia, et al.
Veröffentlicht: (2024)
von: Wang, Xujia, et al.
Veröffentlicht: (2024)
SecEmb: Sparsity-Aware Secure Federated Learning of On-Device Recommender System with Large Embedding
von: Mai, Peihua, et al.
Veröffentlicht: (2025)
von: Mai, Peihua, et al.
Veröffentlicht: (2025)
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
von: Collado-Montañez, Jaime, et al.
Veröffentlicht: (2025)
von: Collado-Montañez, Jaime, et al.
Veröffentlicht: (2025)
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
von: Smădu, Răzvan-Alexandru, et al.
Veröffentlicht: (2025)
von: Smădu, Răzvan-Alexandru, et al.
Veröffentlicht: (2025)
PromptSAM+: Malware Detection based on Prompt Segment Anything Model
von: Wei, Xingyuan, et al.
Veröffentlicht: (2024)
von: Wei, Xingyuan, et al.
Veröffentlicht: (2024)
Automated Bug Triaging using Instruction-Tuned Large Language Models
von: Kiashemshaki, Kiana, et al.
Veröffentlicht: (2025)
von: Kiashemshaki, Kiana, et al.
Veröffentlicht: (2025)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
von: Souza, Débora, et al.
Veröffentlicht: (2026)
von: Souza, Débora, et al.
Veröffentlicht: (2026)
Merge-Bench: Resolve Merge Conflicts with Large Language Models
von: Schesch, Benedikt, et al.
Veröffentlicht: (2026)
von: Schesch, Benedikt, et al.
Veröffentlicht: (2026)
PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review
von: Tu, Songjun, et al.
Veröffentlicht: (2026)
von: Tu, Songjun, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Sparsing Law: Towards Large Language Models with Greater Activation Sparsity
von: Luo, Yuqi, et al.
Veröffentlicht: (2024) -
ConPET: Continual Parameter-Efficient Tuning for Large Language Models
von: Song, Chenyang, et al.
Veröffentlicht: (2023) -
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025) -
Learning to Generate Structured Output with Schema Reinforcement Learning
von: Lu, Yaxi, et al.
Veröffentlicht: (2025) -
LoRS: Efficient Low-Rank Adaptation for Sparse Large Language Model
von: Hu, Yuxuan, et al.
Veröffentlicht: (2025)