Olmo Hybrid: From Theory to Practice and Back
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Merrill, William, Li, Yanhong, Romero, Tyler, Svete, Anej, Costello, Caia, Dasigi, Pradeep, Groeneveld, Dirk, Heineman, David, Kuehl, Bailey, Lambert, Nathan, Li, Chuan, Lo, Kyle, Malik, Saumya, Matusz, DJ, Minixhofer, Benjamin, Morrison, Jacob, Soldaini, Luca, Timbers, Finbarr, Walsh, Pete, Smith, Noah A., Hajishirzi, Hannaneh, Sabharwal, Ashish |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
On the Reasoning Abilities of Masked Diffusion Language Models
par: Svete, Anej, et autres
Publié: (2025)
par: Svete, Anej, et autres
Publié: (2025)
Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don't
par: Svete, Anej, et autres
Publié: (2026)
par: Svete, Anej, et autres
Publié: (2026)
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models
par: Lyu, Xinxi, et autres
Publié: (2024)
par: Lyu, Xinxi, et autres
Publié: (2024)
Olmo 3
par: Olmo, Team, et autres
Publié: (2025)
par: Olmo, Team, et autres
Publié: (2025)
Generalizing Verifiable Instruction Following
par: Pyatkin, Valentina, et autres
Publié: (2025)
par: Pyatkin, Valentina, et autres
Publié: (2025)
Critical Batch Size Revisited: A Simple Empirical Approach to Large-Batch Language Model Training
par: Merrill, William, et autres
Publié: (2025)
par: Merrill, William, et autres
Publié: (2025)
Transformers Can Represent $n$-gram Language Models
par: Svete, Anej, et autres
Publié: (2024)
par: Svete, Anej, et autres
Publié: (2024)
FlexOlmo: Open Language Models for Flexible Data Use
par: Shi, Weijia, et autres
Publié: (2025)
par: Shi, Weijia, et autres
Publié: (2025)
Context-Free Recognition with Transformers
par: Jerad, Selim, et autres
Publié: (2026)
par: Jerad, Selim, et autres
Publié: (2026)
Merge to Learn: Efficiently Adding Skills to Language Models with Model Merging
par: Morrison, Jacob, et autres
Publié: (2024)
par: Morrison, Jacob, et autres
Publié: (2024)
Olmix: A Framework for Data Mixing Throughout LM Development
par: Chen, Mayee F., et autres
Publié: (2026)
par: Chen, Mayee F., et autres
Publié: (2026)
Meta-Reinforcement Learning with Self-Reflection for Agentic Search
par: Xiao, Teng, et autres
Publié: (2026)
par: Xiao, Teng, et autres
Publié: (2026)
ReFIT: Relevance Feedback from a Reranker during Inference
par: Reddy, Revanth Gangi, et autres
Publié: (2023)
par: Reddy, Revanth Gangi, et autres
Publié: (2023)
Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions
par: Wiegreffe, Sarah, et autres
Publié: (2024)
par: Wiegreffe, Sarah, et autres
Publié: (2024)
OLMES: A Standard for Language Model Evaluations
par: Gu, Yuling, et autres
Publié: (2024)
par: Gu, Yuling, et autres
Publié: (2024)
Organize the Web: Constructing Domains Enhances Pre-Training Data Curation
par: Wettig, Alexander, et autres
Publié: (2025)
par: Wettig, Alexander, et autres
Publié: (2025)
RewardBench 2: Advancing Reward Model Evaluation
par: Malik, Saumya, et autres
Publié: (2025)
par: Malik, Saumya, et autres
Publié: (2025)
Establishing Task Scaling Laws via Compute-Efficient Model Ladders
par: Bhagia, Akshita, et autres
Publié: (2024)
par: Bhagia, Akshita, et autres
Publié: (2024)
On Efficiently Representing Regular Languages as RNNs
par: Svete, Anej, et autres
Publié: (2024)
par: Svete, Anej, et autres
Publié: (2024)
Gumbel Counterfactual Generation From Language Models
par: Ravfogel, Shauli, et autres
Publié: (2024)
par: Ravfogel, Shauli, et autres
Publié: (2024)
On the Representational Capacity of Recurrent Neural Language Models
par: Nowak, Franz, et autres
Publié: (2023)
par: Nowak, Franz, et autres
Publié: (2023)
Unique Hard Attention: A Tale of Two Sides
par: Jerad, Selim, et autres
Publié: (2025)
par: Jerad, Selim, et autres
Publié: (2025)
On the Representational Capacity of Neural Language Models with Chain-of-Thought Reasoning
par: Nowak, Franz, et autres
Publié: (2024)
par: Nowak, Franz, et autres
Publié: (2024)
Think, Prune, Train, Improve: Scaling Reasoning without Scaling Models
par: Costello, Caia, et autres
Publié: (2025)
par: Costello, Caia, et autres
Publié: (2025)
A Little Depth Goes a Long Way: The Expressive Power of Log-Depth Transformers
par: Merrill, William, et autres
Publié: (2025)
par: Merrill, William, et autres
Publié: (2025)
Exact Expressive Power of Transformers with Padding
par: Merrill, William, et autres
Publié: (2025)
par: Merrill, William, et autres
Publié: (2025)
The Expressive Power of Transformers with Chain of Thought
par: Merrill, William, et autres
Publié: (2023)
par: Merrill, William, et autres
Publié: (2023)
A Logic for Expressing Log-Precision Transformers
par: Merrill, William, et autres
Publié: (2022)
par: Merrill, William, et autres
Publié: (2022)
TurnWise: The Gap between Single- and Multi-turn Language Model Capabilities
par: Graf, Victoria, et autres
Publié: (2026)
par: Graf, Victoria, et autres
Publié: (2026)
Why Are Linear RNNs More Parallelizable?
par: Merrill, William, et autres
Publié: (2026)
par: Merrill, William, et autres
Publié: (2026)
Lower Bounds on the Expressivity of Recurrent Neural Language Models
par: Svete, Anej, et autres
Publié: (2024)
par: Svete, Anej, et autres
Publié: (2024)
Hybrid Preferences: Learning to Route Instances for Human vs. AI Feedback
par: Miranda, Lester James V., et autres
Publié: (2024)
par: Miranda, Lester James V., et autres
Publié: (2024)
APT: Adaptive Pruning and Tuning Pretrained Language Models for Efficient Training and Inference
par: Zhao, Bowen, et autres
Publié: (2024)
par: Zhao, Bowen, et autres
Publié: (2024)
A Geometric Notion of Causal Probing
par: Guerner, Clément, et autres
Publié: (2023)
par: Guerner, Clément, et autres
Publié: (2023)
Can Transformers Learn $n$-gram Language Models?
par: Svete, Anej, et autres
Publié: (2024)
par: Svete, Anej, et autres
Publié: (2024)
Formal Aspects of Language Modeling
par: Cotterell, Ryan, et autres
Publié: (2023)
par: Cotterell, Ryan, et autres
Publié: (2023)
Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation
par: Heineman, David, et autres
Publié: (2025)
par: Heineman, David, et autres
Publié: (2025)
Paloma: A Benchmark for Evaluating Language Model Fit
par: Magnusson, Ian, et autres
Publié: (2023)
par: Magnusson, Ian, et autres
Publié: (2023)
What's In My Big Data?
par: Elazar, Yanai, et autres
Publié: (2023)
par: Elazar, Yanai, et autres
Publié: (2023)
The Illusion of State in State-Space Models
par: Merrill, William, et autres
Publié: (2024)
par: Merrill, William, et autres
Publié: (2024)
Documents similaires
-
On the Reasoning Abilities of Masked Diffusion Language Models
par: Svete, Anej, et autres
Publié: (2025) -
Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don't
par: Svete, Anej, et autres
Publié: (2026) -
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models
par: Lyu, Xinxi, et autres
Publié: (2024) -
Olmo 3
par: Olmo, Team, et autres
Publié: (2025) -
Generalizing Verifiable Instruction Following
par: Pyatkin, Valentina, et autres
Publié: (2025)