What Matters in Linearizing Language Models? A Comparative Study of Architecture, Scale, and Task Adaptation
Fuente:
arXiv
Saved in:
| Main Authors: | Haller, Patrick, Golde, Jonas, Akbik, Alan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sample-Efficient Language Modeling with Linear Attention and Lightweight Enhancements
by: Haller, Patrick, et al.
Published: (2025)
by: Haller, Patrick, et al.
Published: (2025)
What Matters When Building Universal Multilingual Named Entity Recognition Models?
by: Golde, Jonas, et al.
Published: (2026)
by: Golde, Jonas, et al.
Published: (2026)
Fabricator: An Open Source Toolkit for Generating Labeled Training Data with Teacher LLMs
by: Golde, Jonas, et al.
Published: (2023)
by: Golde, Jonas, et al.
Published: (2023)
BabyHGRN: Exploring RNNs for Sample-Efficient Training of Language Models
by: Haller, Patrick, et al.
Published: (2024)
by: Haller, Patrick, et al.
Published: (2024)
PECC: Problem Extraction and Coding Challenges
by: Haller, Patrick, et al.
Published: (2024)
by: Haller, Patrick, et al.
Published: (2024)
FiNERweb: Datasets and Artifacts for Scalable Multilingual Named Entity Recognition
by: Golde, Jonas, et al.
Published: (2025)
by: Golde, Jonas, et al.
Published: (2025)
Repetition over Diversity: High-Signal Data Filtering for Sample-Efficient German Language Modeling
by: Aynetdinov, Ansar, et al.
Published: (2026)
by: Aynetdinov, Ansar, et al.
Published: (2026)
PISA-Bench: The PISA Index as a Multilingual and Multimodal Metric for the Evaluation of Vision-Language Models
by: Haller, Patrick, et al.
Published: (2025)
by: Haller, Patrick, et al.
Published: (2025)
MastermindEval: A Simple But Scalable Reasoning Benchmark
by: Golde, Jonas, et al.
Published: (2025)
by: Golde, Jonas, et al.
Published: (2025)
Large-Scale Label Interpretation Learning for Few-Shot Named Entity Recognition
by: Golde, Jonas, et al.
Published: (2024)
by: Golde, Jonas, et al.
Published: (2024)
Pre-Training Curriculum for Multi-Token Prediction in Language Models
by: Aynetdinov, Ansar, et al.
Published: (2025)
by: Aynetdinov, Ansar, et al.
Published: (2025)
Familiarity: Better Evaluation of Zero-Shot Named Entity Recognition by Quantifying Label Shifts in Synthetic Training Data
by: Golde, Jonas, et al.
Published: (2024)
by: Golde, Jonas, et al.
Published: (2024)
Question Decomposition for Retrieval-Augmented Generation
by: Ammann, Paul J. L., et al.
Published: (2025)
by: Ammann, Paul J. L., et al.
Published: (2025)
A Comparative Study of Task Adaptation Techniques of Large Language Models for Identifying Sustainable Development Goals
by: Cadeddu, Andrea, et al.
Published: (2025)
by: Cadeddu, Andrea, et al.
Published: (2025)
From Data to Knowledge: Evaluating How Efficiently Language Models Learn Facts
by: Christoph, Daniel, et al.
Published: (2025)
by: Christoph, Daniel, et al.
Published: (2025)
NoiseBench: Benchmarking the Impact of Real Label Noise on Named Entity Recognition
by: Merdjanovska, Elena, et al.
Published: (2024)
by: Merdjanovska, Elena, et al.
Published: (2024)
Dissecting Clinical Reasoning in Language Models: A Comparative Study of Prompts and Model Adaptation Strategies
by: Jullien, Mael, et al.
Published: (2025)
by: Jullien, Mael, et al.
Published: (2025)
What Matters for Model Merging at Scale?
by: Yadav, Prateek, et al.
Published: (2024)
by: Yadav, Prateek, et al.
Published: (2024)
TLoRA: Task-aware Low Rank Adaptation of Large Language Models
by: Lin, Weicheng, et al.
Published: (2026)
by: Lin, Weicheng, et al.
Published: (2026)
Efficient Task Adaptation in Large Language Models via Selective Parameter Optimization
by: Wan, Weijie, et al.
Published: (2026)
by: Wan, Weijie, et al.
Published: (2026)
Adaptation of Biomedical and Clinical Pretrained Models to French Long Documents: A Comparative Study
by: Bazoge, Adrien, et al.
Published: (2024)
by: Bazoge, Adrien, et al.
Published: (2024)
Strategic Prompting for Conversational Tasks: A Comparative Analysis of Large Language Models Across Diverse Conversational Tasks
by: Joshi, Ratnesh Kumar, et al.
Published: (2024)
by: Joshi, Ratnesh Kumar, et al.
Published: (2024)
Contextual Attention Modulation: Towards Efficient Multi-Task Adaptation in Large Language Models
by: Pan, Dayan, et al.
Published: (2025)
by: Pan, Dayan, et al.
Published: (2025)
Efficiency at Scale: Investigating the Performance of Diminutive Language Models in Clinical Tasks
by: Taylor, Niall, et al.
Published: (2024)
by: Taylor, Niall, et al.
Published: (2024)
TransformerRanker: A Tool for Efficiently Finding the Best-Suited Language Models for Downstream Classification Tasks
by: Garbas, Lukas, et al.
Published: (2024)
by: Garbas, Lukas, et al.
Published: (2024)
How to Connect Speech Foundation Models and Large Language Models? What Matters and What Does Not
by: Verdini, Francesco, et al.
Published: (2024)
by: Verdini, Francesco, et al.
Published: (2024)
Pay Attention to What Matters
by: Silva, Pedro Luiz, et al.
Published: (2024)
by: Silva, Pedro Luiz, et al.
Published: (2024)
The Markovian Thinker: Architecture-Agnostic Linear Scaling of Reasoning
by: Aghajohari, Milad, et al.
Published: (2025)
by: Aghajohari, Milad, et al.
Published: (2025)
A Comparative Study of Large Language Models and Human Personality Traits
by: Jiaqi, Wang, et al.
Published: (2025)
by: Jiaqi, Wang, et al.
Published: (2025)
Large Language Models-guided Dynamic Adaptation for Temporal Knowledge Graph Reasoning
by: Wang, Jiapu, et al.
Published: (2024)
by: Wang, Jiapu, et al.
Published: (2024)
A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
by: Zhang, Qiyuan, et al.
Published: (2025)
by: Zhang, Qiyuan, et al.
Published: (2025)
An Empirical Study on Cross-lingual Vocabulary Adaptation for Efficient Language Model Inference
by: Yamaguchi, Atsuki, et al.
Published: (2024)
by: Yamaguchi, Atsuki, et al.
Published: (2024)
Low-Resource Dialect Adaptation of Large Language Models: A French Dialect Case-Study
by: Khan, Eeham, et al.
Published: (2025)
by: Khan, Eeham, et al.
Published: (2025)
Evaluating Large Language Models for Abstract Evaluation Tasks: An Empirical Study
by: Liu, Yinuo, et al.
Published: (2026)
by: Liu, Yinuo, et al.
Published: (2026)
Underspecification in Language Modeling Tasks: A Causality-Informed Study of Gendered Pronoun Resolution
by: McMilin, Emily
Published: (2022)
by: McMilin, Emily
Published: (2022)
A General Highly Accurate Online Planning Method Integrating Large Language Models into Nested Rollout Policy Adaptation for Dialogue Tasks
by: Wang, Hui, et al.
Published: (2025)
by: Wang, Hui, et al.
Published: (2025)
Optimizing Language Models for Grammatical Acceptability: A Comparative Study of Fine-Tuning Techniques
by: Ratan, Shobhit, et al.
Published: (2025)
by: Ratan, Shobhit, et al.
Published: (2025)
Fine-tuning Language Models for Recipe Generation: A Comparative Analysis and Benchmark Study
by: Vij, Anneketh, et al.
Published: (2025)
by: Vij, Anneketh, et al.
Published: (2025)
PairBench: Are Vision-Language Models Reliable at Comparing What They See?
by: Feizi, Aarash, et al.
Published: (2025)
by: Feizi, Aarash, et al.
Published: (2025)
Rethinking LLM Language Adaptation: A Case Study on Chinese Mixtral
by: Cui, Yiming, et al.
Published: (2024)
by: Cui, Yiming, et al.
Published: (2024)
Similar Items
-
Sample-Efficient Language Modeling with Linear Attention and Lightweight Enhancements
by: Haller, Patrick, et al.
Published: (2025) -
What Matters When Building Universal Multilingual Named Entity Recognition Models?
by: Golde, Jonas, et al.
Published: (2026) -
Fabricator: An Open Source Toolkit for Generating Labeled Training Data with Teacher LLMs
by: Golde, Jonas, et al.
Published: (2023) -
BabyHGRN: Exploring RNNs for Sample-Efficient Training of Language Models
by: Haller, Patrick, et al.
Published: (2024) -
PECC: Problem Extraction and Coding Challenges
by: Haller, Patrick, et al.
Published: (2024)