PowLU: An Activation Function for Stable Pre-Training of LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Jiang, Peijie, Feng, Yuqi, Peng, Cunyin, Zhao, Qian, Liu, Jia, Chen, KunLong, Zhang, Zhiqiang, Zhou, Jun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training
di: Tian, Changxin, et al.
Pubblicazione: (2025)
di: Tian, Changxin, et al.
Pubblicazione: (2025)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
di: Saji, Alan, et al.
Pubblicazione: (2025)
di: Saji, Alan, et al.
Pubblicazione: (2025)
Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models
di: Tian, Changxin, et al.
Pubblicazione: (2025)
di: Tian, Changxin, et al.
Pubblicazione: (2025)
Preventing Safety Drift in Large Language Models via Coupled Weight and Activation Constraints
di: Peng, Songping, et al.
Pubblicazione: (2026)
di: Peng, Songping, et al.
Pubblicazione: (2026)
Pre-trained Language Model with Prompts for Temporal Knowledge Graph Completion
di: Xu, Wenjie, et al.
Pubblicazione: (2023)
di: Xu, Wenjie, et al.
Pubblicazione: (2023)
An Investigation of Neuron Activation as a Unified Lens to Explain Chain-of-Thought Eliciting Arithmetic Reasoning of LLMs
di: Rai, Daking, et al.
Pubblicazione: (2024)
di: Rai, Daking, et al.
Pubblicazione: (2024)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
di: Ashuach, Tomer, et al.
Pubblicazione: (2025)
di: Ashuach, Tomer, et al.
Pubblicazione: (2025)
On Initializing Transformers with Pre-trained Embeddings
di: Kim, Ha Young, et al.
Pubblicazione: (2024)
di: Kim, Ha Young, et al.
Pubblicazione: (2024)
Benchmarking the Performance of Pre-trained LLMs across Urdu NLP Tasks
di: Tahir, Munief Hassan, et al.
Pubblicazione: (2024)
di: Tahir, Munief Hassan, et al.
Pubblicazione: (2024)
Training LLMs to Recognize Hedges in Spontaneous Narratives
di: Paige, Amie J., et al.
Pubblicazione: (2024)
di: Paige, Amie J., et al.
Pubblicazione: (2024)
Sparsing Law: Towards Large Language Models with Greater Activation Sparsity
di: Luo, Yuqi, et al.
Pubblicazione: (2024)
di: Luo, Yuqi, et al.
Pubblicazione: (2024)
SEPTQ: A Simple and Effective Post-Training Quantization Paradigm for Large Language Models
di: Liu, Han, et al.
Pubblicazione: (2026)
di: Liu, Han, et al.
Pubblicazione: (2026)
Towards Effective and Efficient Continual Pre-training of Large Language Models
di: Chen, Jie, et al.
Pubblicazione: (2024)
di: Chen, Jie, et al.
Pubblicazione: (2024)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
di: Peters, Sydney, et al.
Pubblicazione: (2025)
di: Peters, Sydney, et al.
Pubblicazione: (2025)
EduGuardBench: A Holistic Benchmark for Evaluating the Pedagogical Fidelity and Adversarial Safety of LLMs as Simulated Teachers
di: Jiang, Yilin, et al.
Pubblicazione: (2025)
di: Jiang, Yilin, et al.
Pubblicazione: (2025)
Evaluating Relational Reasoning in LLMs with REL
di: Fesser, Lukas, et al.
Pubblicazione: (2026)
di: Fesser, Lukas, et al.
Pubblicazione: (2026)
Pre-training data selection for biomedical domain adaptation using journal impact metrics
di: Laï-king, Mathieu, et al.
Pubblicazione: (2024)
di: Laï-king, Mathieu, et al.
Pubblicazione: (2024)
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
di: Simhi, Adi, et al.
Pubblicazione: (2024)
di: Simhi, Adi, et al.
Pubblicazione: (2024)
LLMs and the Human Condition
di: Wallis, Peter
Pubblicazione: (2024)
di: Wallis, Peter
Pubblicazione: (2024)
Automatic Task Detection and Heterogeneous LLM Speculative Decoding
di: Ge, Danying, et al.
Pubblicazione: (2025)
di: Ge, Danying, et al.
Pubblicazione: (2025)
Learning Software Bug Reports: A Systematic Literature Review
di: Long, Guoming, et al.
Pubblicazione: (2025)
di: Long, Guoming, et al.
Pubblicazione: (2025)
The Unified Cognitive Consciousness Theory for Language Models: Anchoring Semantics, Thresholds of Activation, and Emergent Reasoning
di: Chang, Edward Y., et al.
Pubblicazione: (2025)
di: Chang, Edward Y., et al.
Pubblicazione: (2025)
Steer-MoE: Efficient Audio-Language Alignment with a Mixture-of-Experts Steering Module
di: Feng, Ruitao, et al.
Pubblicazione: (2025)
di: Feng, Ruitao, et al.
Pubblicazione: (2025)
LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance
di: Ivanov, Igor
Pubblicazione: (2025)
di: Ivanov, Igor
Pubblicazione: (2025)
LLMs Are Not Scorers: Rethinking MT Evaluation with Generation-Based Methods
di: Cui, Hyang
Pubblicazione: (2025)
di: Cui, Hyang
Pubblicazione: (2025)
How LLMs Are Persuaded: A Few Attention Heads, Rerouted
di: Sun, Xiangkun, et al.
Pubblicazione: (2026)
di: Sun, Xiangkun, et al.
Pubblicazione: (2026)
RAG-Optimized Tibetan Tourism LLMs: Enhancing Accuracy and Personalization
di: Qi, Jinhu, et al.
Pubblicazione: (2024)
di: Qi, Jinhu, et al.
Pubblicazione: (2024)
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges
di: Chehbouni, Khaoula, et al.
Pubblicazione: (2025)
di: Chehbouni, Khaoula, et al.
Pubblicazione: (2025)
Guided Speculative Inference for Efficient Test-Time Alignment of LLMs
di: Geuter, Jonathan, et al.
Pubblicazione: (2025)
di: Geuter, Jonathan, et al.
Pubblicazione: (2025)
MAWARITH: A Dataset and Benchmark for Legal Inheritance Reasoning with LLMs
di: Bouchekif, Abdessalam, et al.
Pubblicazione: (2026)
di: Bouchekif, Abdessalam, et al.
Pubblicazione: (2026)
MIRIAD: Augmenting LLMs with millions of medical query-response pairs
di: Zheng, Qinyue, et al.
Pubblicazione: (2025)
di: Zheng, Qinyue, et al.
Pubblicazione: (2025)
Hardware Co-Design Scaling Laws via Roofline Modelling for On-Device LLMs
di: Sun, Luoyang, et al.
Pubblicazione: (2026)
di: Sun, Luoyang, et al.
Pubblicazione: (2026)
Auditing Meta-Cognitive Hallucinations in Reasoning Large Language Models
di: Lu, Haolang, et al.
Pubblicazione: (2025)
di: Lu, Haolang, et al.
Pubblicazione: (2025)
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
di: Collado-Montañez, Jaime, et al.
Pubblicazione: (2025)
di: Collado-Montañez, Jaime, et al.
Pubblicazione: (2025)
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
di: Smădu, Răzvan-Alexandru, et al.
Pubblicazione: (2025)
di: Smădu, Răzvan-Alexandru, et al.
Pubblicazione: (2025)
Automated Bug Triaging using Instruction-Tuned Large Language Models
di: Kiashemshaki, Kiana, et al.
Pubblicazione: (2025)
di: Kiashemshaki, Kiana, et al.
Pubblicazione: (2025)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
di: Souza, Débora, et al.
Pubblicazione: (2026)
di: Souza, Débora, et al.
Pubblicazione: (2026)
Merge-Bench: Resolve Merge Conflicts with Large Language Models
di: Schesch, Benedikt, et al.
Pubblicazione: (2026)
di: Schesch, Benedikt, et al.
Pubblicazione: (2026)
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
di: Simhi, Adi, et al.
Pubblicazione: (2025)
di: Simhi, Adi, et al.
Pubblicazione: (2025)
Fine-Tuning LLMs with Fine-Grained Human Feedback on Text Spans
di: CH-Wang, Sky, et al.
Pubblicazione: (2025)
di: CH-Wang, Sky, et al.
Pubblicazione: (2025)
Documenti analoghi
-
WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training
di: Tian, Changxin, et al.
Pubblicazione: (2025) -
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
di: Saji, Alan, et al.
Pubblicazione: (2025) -
Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models
di: Tian, Changxin, et al.
Pubblicazione: (2025) -
Preventing Safety Drift in Large Language Models via Coupled Weight and Activation Constraints
di: Peng, Songping, et al.
Pubblicazione: (2026) -
Pre-trained Language Model with Prompts for Temporal Knowledge Graph Completion
di: Xu, Wenjie, et al.
Pubblicazione: (2023)