Beyond Perplexity: A Geometric and Spectral Study of Low-Rank Pre-Training
Fuente:
arXiv
Saved in:
| Main Authors: | Shivagunde, Namrata, Deshpande, Vijeta, Muckatira, Sherin, Rumshisky, Anna |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Pre-Training Analogue of Grokking in Language Models: Tracing Delayed Grammatical Generalization
by: Muckatira, Sherin, et al.
Published: (2026)
by: Muckatira, Sherin, et al.
Published: (2026)
Emergent Abilities in Reduced-Scale Generative Language Models
by: Muckatira, Sherin, et al.
Published: (2024)
by: Muckatira, Sherin, et al.
Published: (2024)
Deconstructing In-Context Learning: Understanding Prompts via Corruption
by: Shivagunde, Namrata, et al.
Published: (2024)
by: Shivagunde, Namrata, et al.
Published: (2024)
Playing with Words, Improving with Rewards: Training Language Models for Creative Association
by: Deshpande, Vijeta, et al.
Published: (2026)
by: Deshpande, Vijeta, et al.
Published: (2026)
Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection
by: Deshpande, Vijeta, et al.
Published: (2026)
by: Deshpande, Vijeta, et al.
Published: (2026)
Low-Rank Quantization-Aware Training for LLMs
by: Bondarenko, Yelysei, et al.
Published: (2024)
by: Bondarenko, Yelysei, et al.
Published: (2024)
Rethinking GSPO: The Perplexity-Entropy Equivalence
by: Liu, Chi
Published: (2025)
by: Liu, Chi
Published: (2025)
AdaPreLoRA: Adafactor Preconditioned Low-Rank Adaptation
by: Liu, Ziyun, et al.
Published: (2026)
by: Liu, Ziyun, et al.
Published: (2026)
Training Acceleration of Low-Rank Decomposed Networks using Sequential Freezing and Rank Quantization
by: Hajimolahoseini, Habib, et al.
Published: (2023)
by: Hajimolahoseini, Habib, et al.
Published: (2023)
Scaling Down to Scale Up: A Guide to Parameter-Efficient Fine-Tuning
by: Lialin, Vladislav, et al.
Published: (2023)
by: Lialin, Vladislav, et al.
Published: (2023)
Perplexity Cannot Always Tell Right from Wrong
by: Veličković, Petar, et al.
Published: (2026)
by: Veličković, Petar, et al.
Published: (2026)
Momentum Point-Perplexity Mechanics in Large Language Models
by: Tomaz, Lorenzo, et al.
Published: (2025)
by: Tomaz, Lorenzo, et al.
Published: (2025)
Spectral Logit Sculpting: Adaptive Low-Rank Logit Transformation for Controlled Text Generation
by: Li, Jin, et al.
Published: (2025)
by: Li, Jin, et al.
Published: (2025)
Analysis of student understanding in short‐answer explanations to concept questions using a human‐centered AI approach
by: Harpreet Auby, et al.
Published: (2025)
by: Harpreet Auby, et al.
Published: (2025)
Training-Free Bayesianization for Low-Rank Adapters of Large Language Models
by: Shi, Haizhou, et al.
Published: (2024)
by: Shi, Haizhou, et al.
Published: (2024)
Inconsistent Tokenizations Cause Language Models to be Perplexed by Japanese Grammar
by: Gambardella, Andrew, et al.
Published: (2025)
by: Gambardella, Andrew, et al.
Published: (2025)
Low-Rank Adaptation for Multilingual Summarization: An Empirical Study
by: Whitehouse, Chenxi, et al.
Published: (2023)
by: Whitehouse, Chenxi, et al.
Published: (2023)
Domain-Specific Quality Estimation for Machine Translation in Low-Resource Scenarios
by: Gurav, Namrata Patil, et al.
Published: (2026)
by: Gurav, Namrata Patil, et al.
Published: (2026)
SePer: Measure Retrieval Utility Through The Lens Of Semantic Perplexity Reduction
by: Dai, Lu, et al.
Published: (2025)
by: Dai, Lu, et al.
Published: (2025)
Alzheimer's Dementia Detection Using Perplexity from Paired Large Language Models
by: Xiao, Yao, et al.
Published: (2025)
by: Xiao, Yao, et al.
Published: (2025)
LoRETTA: Low-Rank Economic Tensor-Train Adaptation for Ultra-Low-Parameter Fine-Tuning of Large Language Models
by: Yang, Yifan, et al.
Published: (2024)
by: Yang, Yifan, et al.
Published: (2024)
Stepwise Perplexity-Guided Refinement for Efficient Chain-of-Thought Reasoning in Large Language Models
by: Cui, Yingqian, et al.
Published: (2025)
by: Cui, Yingqian, et al.
Published: (2025)
MoR: Mixture of Ranks for Low-Rank Adaptation Tuning
by: Tang, Chuanyu, et al.
Published: (2024)
by: Tang, Chuanyu, et al.
Published: (2024)
Finer Parameter Steps for Low-Rank PEFT: A Controlled Study with CP Tensor Adapters
by: Wang, Xinjue, et al.
Published: (2026)
by: Wang, Xinjue, et al.
Published: (2026)
Ensembles of Low-Rank Expert Adapters
by: Li, Yinghao, et al.
Published: (2025)
by: Li, Yinghao, et al.
Published: (2025)
The Expressive Power of Low-Rank Adaptation
by: Zeng, Yuchen, et al.
Published: (2023)
by: Zeng, Yuchen, et al.
Published: (2023)
Reinforcement Learning on Pre-Training Data
by: Li, Siheng, et al.
Published: (2025)
by: Li, Siheng, et al.
Published: (2025)
Batched Low-Rank Adaptation of Foundation Models
by: Wen, Yeming, et al.
Published: (2023)
by: Wen, Yeming, et al.
Published: (2023)
Variational Low-Rank Adaptation Using IVON
by: Cong, Bai, et al.
Published: (2024)
by: Cong, Bai, et al.
Published: (2024)
A3 : an Analytical Low-Rank Approximation Framework for Attention
by: Wong, Jeffrey T. H., et al.
Published: (2025)
by: Wong, Jeffrey T. H., et al.
Published: (2025)
Beyond Temperature: Hyperfitting as a Late-Stage Geometric Expansion
by: Li, Meimingwei, et al.
Published: (2026)
by: Li, Meimingwei, et al.
Published: (2026)
SwitchLoRA: Switched Low-Rank Adaptation Can Learn Full-Rank Information
by: Zhou, Kaiye, et al.
Published: (2024)
by: Zhou, Kaiye, et al.
Published: (2024)
LoRMA: Low-Rank Multiplicative Adaptation for LLMs
by: Bihany, Harsh, et al.
Published: (2025)
by: Bihany, Harsh, et al.
Published: (2025)
LoTR: Low Tensor Rank Weight Adaptation
by: Bershatsky, Daniel, et al.
Published: (2024)
by: Bershatsky, Daniel, et al.
Published: (2024)
AutoLoRA: Automatically Tuning Matrix Ranks in Low-Rank Adaptation Based on Meta Learning
by: Zhang, Ruiyi, et al.
Published: (2024)
by: Zhang, Ruiyi, et al.
Published: (2024)
A Single Linear Layer Yields Task-Adapted Low-Rank Matrices
by: Kim, Hwichan, et al.
Published: (2024)
by: Kim, Hwichan, et al.
Published: (2024)
Not How Many, But Which: Parameter Placement in Low-Rank Adaptation
by: Sehanobish, Arijit, et al.
Published: (2026)
by: Sehanobish, Arijit, et al.
Published: (2026)
Sequences of Logits Reveal the Low Rank Structure of Language Models
by: Golowich, Noah, et al.
Published: (2025)
by: Golowich, Noah, et al.
Published: (2025)
LoRA-Pro: Are Low-Rank Adapters Properly Optimized?
by: Wang, Zhengbo, et al.
Published: (2024)
by: Wang, Zhengbo, et al.
Published: (2024)
GoRA: Gradient-driven Adaptive Low Rank Adaptation
by: He, Haonan, et al.
Published: (2025)
by: He, Haonan, et al.
Published: (2025)
Similar Items
-
A Pre-Training Analogue of Grokking in Language Models: Tracing Delayed Grammatical Generalization
by: Muckatira, Sherin, et al.
Published: (2026) -
Emergent Abilities in Reduced-Scale Generative Language Models
by: Muckatira, Sherin, et al.
Published: (2024) -
Deconstructing In-Context Learning: Understanding Prompts via Corruption
by: Shivagunde, Namrata, et al.
Published: (2024) -
Playing with Words, Improving with Rewards: Training Language Models for Creative Association
by: Deshpande, Vijeta, et al.
Published: (2026) -
Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection
by: Deshpande, Vijeta, et al.
Published: (2026)