Initialization of Large Language Models via Reparameterization to Mitigate Loss Spikes
Fuente:
arXiv
Guardado en:
| Autores principales: | Nishida, Kosuke, Nishida, Kyosuke, Saito, Kuniko |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Debiasing Reward Models via Causally Motivated Inference-Time Intervention
por: Shinoda, Kazutoshi, et al.
Publicado: (2026)
por: Shinoda, Kazutoshi, et al.
Publicado: (2026)
Wavelet-based Positional Representation for Long Context
por: Oka, Yui, et al.
Publicado: (2025)
por: Oka, Yui, et al.
Publicado: (2025)
Can LLMs Detect Their Own Hallucinations?
por: Kadotani, Sora, et al.
Publicado: (2025)
por: Kadotani, Sora, et al.
Publicado: (2025)
InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions
por: Tanaka, Ryota, et al.
Publicado: (2024)
por: Tanaka, Ryota, et al.
Publicado: (2024)
Let's Put Ourselves in Sally's Shoes: Shoes-of-Others Prefilling Improves Theory of Mind in Large Language Models
por: Shinoda, Kazutoshi, et al.
Publicado: (2025)
por: Shinoda, Kazutoshi, et al.
Publicado: (2025)
VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents
por: Tanaka, Ryota, et al.
Publicado: (2025)
por: Tanaka, Ryota, et al.
Publicado: (2025)
ToMATO: Verbalizing the Mental States of Role-Playing LLMs for Benchmarking Theory of Mind
por: Shinoda, Kazutoshi, et al.
Publicado: (2025)
por: Shinoda, Kazutoshi, et al.
Publicado: (2025)
Responses Fall Short of Understanding: Revealing the Gap between Internal Representations and Responses in Visual Document Understanding
por: Kawasaki, Haruka, et al.
Publicado: (2026)
por: Kawasaki, Haruka, et al.
Publicado: (2026)
Portable Reward Tuning: Towards Reusable Fine-Tuning across Different Pretrained Models
por: Chijiwa, Daiki, et al.
Publicado: (2025)
por: Chijiwa, Daiki, et al.
Publicado: (2025)
Lossless Vocabulary Reduction for Auto-Regressive Language Models
por: Chijiwa, Daiki, et al.
Publicado: (2025)
por: Chijiwa, Daiki, et al.
Publicado: (2025)
Do Large Vision-Language Models Distinguish between the Actual and Apparent Features of Illusions?
por: Shinozaki, Taiga, et al.
Publicado: (2025)
por: Shinozaki, Taiga, et al.
Publicado: (2025)
A Comparative Study of Demonstration Selection for Practical Large Language Models-based Next POI Prediction
por: Nishida, Ryo, et al.
Publicado: (2026)
por: Nishida, Ryo, et al.
Publicado: (2026)
Long-Tail Crisis in Nearest Neighbor Language Models
por: Nishida, Yuto, et al.
Publicado: (2025)
por: Nishida, Yuto, et al.
Publicado: (2025)
Revisiting Non-Verbatim Memorization in Large Language Models: The Role of Entity Surface Forms
por: Nishida, Yuto, et al.
Publicado: (2026)
por: Nishida, Yuto, et al.
Publicado: (2026)
Instability in Downstream Task Performance During LLM Pretraining
por: Nishida, Yuto, et al.
Publicado: (2025)
por: Nishida, Yuto, et al.
Publicado: (2025)
Explanation Bottleneck Models
por: Yamaguchi, Shin'ya, et al.
Publicado: (2024)
por: Yamaguchi, Shin'ya, et al.
Publicado: (2024)
Retrieving Time-Series Differences Using Natural Language Queries
por: Dohi, Kota, et al.
Publicado: (2025)
por: Dohi, Kota, et al.
Publicado: (2025)
Relative Density Ratio Optimization for Stable and Statistically Consistent Model Alignment
por: Takahashi, Hiroshi, et al.
Publicado: (2026)
por: Takahashi, Hiroshi, et al.
Publicado: (2026)
LAS: Loss-less ANN-SNN Conversion for Fully Spike-Driven Large Language Models
por: Chen, Long, et al.
Publicado: (2025)
por: Chen, Long, et al.
Publicado: (2025)
MiLe Loss: a New Entropy-Weighed Loss for Mitigating the Bias of Learning Difficulties in Large Language Models
por: Su, Zhenpeng, et al.
Publicado: (2023)
por: Su, Zhenpeng, et al.
Publicado: (2023)
Out-of-the-Box Conditional Text Embeddings from Large Language Models
por: Yamada, Kosuke, et al.
Publicado: (2025)
por: Yamada, Kosuke, et al.
Publicado: (2025)
SpikeLLM: Scaling up Spiking Neural Network to Large Language Models via Saliency-based Spiking
por: Xing, Xingrun, et al.
Publicado: (2024)
por: Xing, Xingrun, et al.
Publicado: (2024)
Generating Diverse Translation with Perturbed kNN-MT
por: Nishida, Yuto, et al.
Publicado: (2024)
por: Nishida, Yuto, et al.
Publicado: (2024)
Quantifying and Mitigating Socially Desirable Responding in LLMs: A Desirability-Matched Graded Forced-Choice Psychometric Study
por: Okada, Kensuke, et al.
Publicado: (2026)
por: Okada, Kensuke, et al.
Publicado: (2026)
Hierarchical Text Classification Using Black Box Large Language Models
por: Yoshimura, Kosuke, et al.
Publicado: (2025)
por: Yoshimura, Kosuke, et al.
Publicado: (2025)
Post Persona Alignment for Multi-Session Dialogue Generation
por: Chen, Yi-Pei, et al.
Publicado: (2025)
por: Chen, Yi-Pei, et al.
Publicado: (2025)
Bridging Perception and Language: A Systematic Benchmark for LVLMs' Understanding of Amodal Completion Reports
por: Watahiki, Amane, et al.
Publicado: (2025)
por: Watahiki, Amane, et al.
Publicado: (2025)
A Reparameterized Discrete Diffusion Model for Text Generation
por: Zheng, Lin, et al.
Publicado: (2023)
por: Zheng, Lin, et al.
Publicado: (2023)
Mitigating Prompt-Induced Hallucinations in Large Language Models via Structured Reasoning
por: Hao, Jinbo, et al.
Publicado: (2026)
por: Hao, Jinbo, et al.
Publicado: (2026)
Mitigating Exaggerated Safety in Large Language Models
por: Ray, Ruchira, et al.
Publicado: (2024)
por: Ray, Ruchira, et al.
Publicado: (2024)
Mitigating the Bias of Large Language Model Evaluation
por: Zhou, Hongli, et al.
Publicado: (2024)
por: Zhou, Hongli, et al.
Publicado: (2024)
DiffNator: Generating Structured Explanations of Time-Series Differences
por: Dohi, Kota, et al.
Publicado: (2025)
por: Dohi, Kota, et al.
Publicado: (2025)
SpikeLM: Towards General Spike-Driven Language Modeling via Elastic Bi-Spiking Mechanisms
por: Xing, Xingrun, et al.
Publicado: (2024)
por: Xing, Xingrun, et al.
Publicado: (2024)
Spike No More: Stabilizing the Pre-training of Large Language Models
por: Takase, Sho, et al.
Publicado: (2023)
por: Takase, Sho, et al.
Publicado: (2023)
Investigating the Feasibility of Mitigating Potential Copyright Infringement via Large Language Model Unlearning
por: Dou, Guangyao
Publicado: (2024)
por: Dou, Guangyao
Publicado: (2024)
Mitigating Hallucinations of Large Language Models in Medical Information Extraction via Contrastive Decoding
por: Xu, Derong, et al.
Publicado: (2024)
por: Xu, Derong, et al.
Publicado: (2024)
Reparameterized LLM Training via Orthogonal Equivalence Transformation
por: Qiu, Zeju, et al.
Publicado: (2025)
por: Qiu, Zeju, et al.
Publicado: (2025)
Recent Trends in Personalized Dialogue Generation: A Review of Datasets, Methodologies, and Evaluations
por: Chen, Yi-Pei, et al.
Publicado: (2024)
por: Chen, Yi-Pei, et al.
Publicado: (2024)
Mitigating Large Language Model Hallucination with Faithful Finetuning
por: Hu, Minda, et al.
Publicado: (2024)
por: Hu, Minda, et al.
Publicado: (2024)
Exploring and Mitigating Fawning Hallucinations in Large Language Models
por: Shangguan, Zixuan, et al.
Publicado: (2025)
por: Shangguan, Zixuan, et al.
Publicado: (2025)
Ejemplares similares
-
Debiasing Reward Models via Causally Motivated Inference-Time Intervention
por: Shinoda, Kazutoshi, et al.
Publicado: (2026) -
Wavelet-based Positional Representation for Long Context
por: Oka, Yui, et al.
Publicado: (2025) -
Can LLMs Detect Their Own Hallucinations?
por: Kadotani, Sora, et al.
Publicado: (2025) -
InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions
por: Tanaka, Ryota, et al.
Publicado: (2024) -
Let's Put Ourselves in Sally's Shoes: Shoes-of-Others Prefilling Improves Theory of Mind in Large Language Models
por: Shinoda, Kazutoshi, et al.
Publicado: (2025)