Leveraging the true depth of LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | González, Ramón Calvo, Paliotta, Daniele, Pagliardini, Matteo, Jaggi, Martin, Fleuret, François |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted Averaging
by: Pagliardini, Matteo, et al.
Published: (2024)
by: Pagliardini, Matteo, et al.
Published: (2024)
CoTFormer: A Chain-of-Thought Driven Architecture with Budget-Adaptive Computation Cost at Inference
by: Mohtashami, Amirkeivan, et al.
Published: (2023)
by: Mohtashami, Amirkeivan, et al.
Published: (2023)
DoGE: Domain Reweighting with Generalization Estimation
by: Fan, Simin, et al.
Published: (2023)
by: Fan, Simin, et al.
Published: (2023)
Benchmarking Optimizers for Large Language Model Pretraining
by: Semenov, Andrei, et al.
Published: (2025)
by: Semenov, Andrei, et al.
Published: (2025)
Thinking Slow, Fast: Scaling Inference Compute with Distilled Reasoners
by: Paliotta, Daniele, et al.
Published: (2025)
by: Paliotta, Daniele, et al.
Published: (2025)
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
by: Messmer, Bettina, et al.
Published: (2025)
by: Messmer, Bettina, et al.
Published: (2025)
Personalized Collaborative Fine-Tuning for On-Device Large Language Models
by: Wagner, Nicolas, et al.
Published: (2024)
by: Wagner, Nicolas, et al.
Published: (2024)
Can Performant LLMs Be Ethical? Quantifying the Impact of Web Crawling Opt-Outs
by: Fan, Dongyang, et al.
Published: (2025)
by: Fan, Dongyang, et al.
Published: (2025)
On-Device Collaborative Language Modeling via a Mixture of Generalists and Specialists
by: Fan, Dongyang, et al.
Published: (2024)
by: Fan, Dongyang, et al.
Published: (2024)
Mitigating Unintended Memorization with LoRA in Federated Learning for LLMs
by: Bossy, Thierry, et al.
Published: (2025)
by: Bossy, Thierry, et al.
Published: (2025)
Towards an empirical understanding of MoE design choices
by: Fan, Dongyang, et al.
Published: (2024)
by: Fan, Dongyang, et al.
Published: (2024)
Leverage Unlearning to Sanitize LLMs
by: Boutet, Antoine, et al.
Published: (2025)
by: Boutet, Antoine, et al.
Published: (2025)
Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
by: Fan, Dongyang, et al.
Published: (2025)
by: Fan, Dongyang, et al.
Published: (2025)
In-Context Symbolic Regression: Leveraging Large Language Models for Function Discovery
by: Merler, Matteo, et al.
Published: (2024)
by: Merler, Matteo, et al.
Published: (2024)
Meta-Tuning LLMs to Leverage Lexical Knowledge for Generalizable Language Style Understanding
by: Guo, Ruohao, et al.
Published: (2023)
by: Guo, Ruohao, et al.
Published: (2023)
Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning
by: Wang, Xiyao, et al.
Published: (2024)
by: Wang, Xiyao, et al.
Published: (2024)
The ProLiFIC dataset: Leveraging LLMs to Unveil the Italian Lawmaking Process
by: Contestabile, Matilde, et al.
Published: (2025)
by: Contestabile, Matilde, et al.
Published: (2025)
Leveraging Parameter Space Symmetries for Reasoning Skill Transfer in LLMs
by: Horoi, Stefan, et al.
Published: (2025)
by: Horoi, Stefan, et al.
Published: (2025)
Out-of-Distribution Detection by Leveraging Between-Layer Transformation Smoothness
by: Jelenić, Fran, et al.
Published: (2023)
by: Jelenić, Fran, et al.
Published: (2023)
MAP's not dead yet: Uncovering true language model modes by conditioning away degeneracy
by: Yoshida, Davis, et al.
Published: (2023)
by: Yoshida, Davis, et al.
Published: (2023)
Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes
by: Kolawole, Steven, et al.
Published: (2024)
by: Kolawole, Steven, et al.
Published: (2024)
Calibrating LLMs for Text-to-SQL Parsing by Leveraging Sub-clause Frequencies
by: Liu, Terrance, et al.
Published: (2025)
by: Liu, Terrance, et al.
Published: (2025)
LegalLens: Leveraging LLMs for Legal Violation Identification in Unstructured Text
by: Bernsohn, Dor, et al.
Published: (2024)
by: Bernsohn, Dor, et al.
Published: (2024)
Speaking the Same Language: Leveraging LLMs in Standardizing Clinical Data for AI
by: Sett, Arindam, et al.
Published: (2024)
by: Sett, Arindam, et al.
Published: (2024)
Multicalibration for Confidence Scoring in LLMs
by: Detommaso, Gianluca, et al.
Published: (2024)
by: Detommaso, Gianluca, et al.
Published: (2024)
SimulRAG: Simulator-based RAG for Grounding LLMs in Long-form Scientific QA
by: Xu, Haozhou, et al.
Published: (2025)
by: Xu, Haozhou, et al.
Published: (2025)
Attention with Markov: A Framework for Principled Analysis of Transformers via Markov Chains
by: Makkuva, Ashok Vardhan, et al.
Published: (2024)
by: Makkuva, Ashok Vardhan, et al.
Published: (2024)
Trans-Tokenization and Cross-lingual Vocabulary Transfers: Language Adaptation of LLMs for Low-Resource NLP
by: Remy, François, et al.
Published: (2024)
by: Remy, François, et al.
Published: (2024)
Split, Unlearn, Merge: Leveraging Data Attributes for More Effective Unlearning in LLMs
by: Kadhe, Swanand Ravindra, et al.
Published: (2024)
by: Kadhe, Swanand Ravindra, et al.
Published: (2024)
Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation
by: Mahdavi, Sadegh, et al.
Published: (2025)
by: Mahdavi, Sadegh, et al.
Published: (2025)
From Source to Target: Leveraging Transfer Learning for Predictive Process Monitoring in Organizations
by: Weinzierl, Sven, et al.
Published: (2025)
by: Weinzierl, Sven, et al.
Published: (2025)
Understanding Chain-of-Thought in LLMs through Information Theory
by: Ton, Jean-Francois, et al.
Published: (2024)
by: Ton, Jean-Francois, et al.
Published: (2024)
Leveraging large language models for structured information extraction from pathology reports
by: Balasubramanian, Jeya Balaji, et al.
Published: (2025)
by: Balasubramanian, Jeya Balaji, et al.
Published: (2025)
Better as Generators Than Classifiers: Leveraging LLMs and Synthetic Data for Low-Resource Multilingual Classification
by: Pecher, Branislav, et al.
Published: (2026)
by: Pecher, Branislav, et al.
Published: (2026)
Guiding LLMs The Right Way: Fast, Non-Invasive Constrained Generation
by: Beurer-Kellner, Luca, et al.
Published: (2024)
by: Beurer-Kellner, Luca, et al.
Published: (2024)
Leveraging Language Models to Detect Greenwashing
by: Vinella, Avalon, et al.
Published: (2023)
by: Vinella, Avalon, et al.
Published: (2023)
The AdEMAMix Optimizer: Better, Faster, Older
by: Pagliardini, Matteo, et al.
Published: (2024)
by: Pagliardini, Matteo, et al.
Published: (2024)
Evaluating Compact LLMs for Zero-Shot Iberian Language Tasks on End-User Devices
by: Seller, Luís Couto, et al.
Published: (2025)
by: Seller, Luís Couto, et al.
Published: (2025)
Dynamic Vocabulary Pruning in Early-Exit LLMs
by: Vincenti, Jort, et al.
Published: (2024)
by: Vincenti, Jort, et al.
Published: (2024)
Do Recommender Systems Really Leverage Multimodal Content? A Comprehensive Analysis on Multimodal Representations for Recommendation
by: Pomo, Claudio, et al.
Published: (2025)
by: Pomo, Claudio, et al.
Published: (2025)
Similar Items
-
DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted Averaging
by: Pagliardini, Matteo, et al.
Published: (2024) -
CoTFormer: A Chain-of-Thought Driven Architecture with Budget-Adaptive Computation Cost at Inference
by: Mohtashami, Amirkeivan, et al.
Published: (2023) -
DoGE: Domain Reweighting with Generalization Estimation
by: Fan, Simin, et al.
Published: (2023) -
Benchmarking Optimizers for Large Language Model Pretraining
by: Semenov, Andrei, et al.
Published: (2025) -
Thinking Slow, Fast: Scaling Inference Compute with Distilled Reasoners
by: Paliotta, Daniele, et al.
Published: (2025)