Two Heads Are Better than One: Simulating Large Transformers with Small Ones
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Hantao, Alman, Josh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fundamental Limitations on Subquadratic Alternatives to Transformers
by: Alman, Josh, et al.
Published: (2024)
by: Alman, Josh, et al.
Published: (2024)
Every Bit Counts: A Theoretical Study of Precision-Expressivity Tradeoffs in Quantized Transformers
by: Chakrabarti, Sayak, et al.
Published: (2026)
by: Chakrabarti, Sayak, et al.
Published: (2026)
Poly-attention: a general scheme for higher-order self-attention
by: Chakrabarti, Sayak, et al.
Published: (2026)
by: Chakrabarti, Sayak, et al.
Published: (2026)
What Makes Looped Transformers Perform Better Than Non-Recursive Ones
by: Gong, Zixuan, et al.
Published: (2025)
by: Gong, Zixuan, et al.
Published: (2025)
Two Heads Are Better than One: Model-Weight and Latent-Space Analysis for Federated Learning on Non-iid Data against Poisoning Attacks
by: Lyu, Xingyu, et al.
Published: (2025)
by: Lyu, Xingyu, et al.
Published: (2025)
SOM Directions are Better than One: Multi-Directional Refusal Suppression in Language Models
by: Piras, Giorgio, et al.
Published: (2025)
by: Piras, Giorgio, et al.
Published: (2025)
What One Cannot, Two Can: Two-Layer Transformers Provably Represent Induction Heads on Any-Order Markov Chains
by: Ekbote, Chanakya, et al.
Published: (2025)
by: Ekbote, Chanakya, et al.
Published: (2025)
Two Is Better Than One: Aligned Representation Pairs for Anomaly Detection
by: Ryser, Alain, et al.
Published: (2024)
by: Ryser, Alain, et al.
Published: (2024)
Two Heads are Better than One: Distilling Large Language Model Features Into Small Models with Feature Decomposition and Mixture
by: Fu, Tianhao, et al.
Published: (2025)
by: Fu, Tianhao, et al.
Published: (2025)
Repeat After Me: Transformers are Better than State Space Models at Copying
by: Jelassi, Samy, et al.
Published: (2024)
by: Jelassi, Samy, et al.
Published: (2024)
Two Tickets are Better than One: Fair and Accurate Hiring Under Strategic LLM Manipulations
by: Cohen, Lee, et al.
Published: (2025)
by: Cohen, Lee, et al.
Published: (2025)
Performance Control in Early Exiting to Deploy Large Models at the Same Cost of Smaller Ones
by: Mofakhami, Mehrnaz, et al.
Published: (2024)
by: Mofakhami, Mehrnaz, et al.
Published: (2024)
Large Transformers are Better EEG Learners
by: Wang, Bingxin, et al.
Published: (2023)
by: Wang, Bingxin, et al.
Published: (2023)
Is Data Shapley Not Better than Random in Data Selection? Ask NASH
by: Tian, Xiao, et al.
Published: (2026)
by: Tian, Xiao, et al.
Published: (2026)
First Hallucination Tokens Are Different from Conditional Ones
by: Snel, Jakob, et al.
Published: (2025)
by: Snel, Jakob, et al.
Published: (2025)
Only Large Weights (And Not Skip Connections) Can Prevent the Perils of Rank Collapse
by: Alman, Josh, et al.
Published: (2025)
by: Alman, Josh, et al.
Published: (2025)
OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting
by: Hu, Xing, et al.
Published: (2025)
by: Hu, Xing, et al.
Published: (2025)
The Devil is in the Condition Numbers: Why is GLU Better than non-GLU Structure?
by: Lyu, Xingyu, et al.
Published: (2026)
by: Lyu, Xingyu, et al.
Published: (2026)
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment
by: Zhang, Jiazheng, et al.
Published: (2025)
by: Zhang, Jiazheng, et al.
Published: (2025)
Improving Large Models with Small models: Lower Costs and Better Performance
by: Chen, Dong, et al.
Published: (2024)
by: Chen, Dong, et al.
Published: (2024)
Better than Your Teacher: LLM Agents that learn from Privileged AI Feedback
by: Choudhury, Sanjiban, et al.
Published: (2024)
by: Choudhury, Sanjiban, et al.
Published: (2024)
Do Transformers Understand Ancient Roman Coin Motifs Better than CNNs?
by: Reid, David, et al.
Published: (2026)
by: Reid, David, et al.
Published: (2026)
Two Stones Hit One Bird: Bilevel Positional Encoding for Better Length Extrapolation
by: He, Zhenyu, et al.
Published: (2024)
by: He, Zhenyu, et al.
Published: (2024)
Fast RoPE Attention: Combining the Polynomial Method and Fast Fourier Transform
by: Alman, Josh, et al.
Published: (2025)
by: Alman, Josh, et al.
Published: (2025)
From Small to Large: Generalization Bounds for Transformers on Variable-Size Inputs
by: Alokhina, Anastasiia, et al.
Published: (2025)
by: Alokhina, Anastasiia, et al.
Published: (2025)
Disentangling Tabular Data Towards Better One-Class Anomaly Detection
by: Ye, Jianan, et al.
Published: (2024)
by: Ye, Jianan, et al.
Published: (2024)
Rethinking Data Curation in LLM Training: Online Reweighting Offers Better Generalization than Offline Methods
by: Zhao, Wanru, et al.
Published: (2026)
by: Zhao, Wanru, et al.
Published: (2026)
This Looks Better than That: Better Interpretable Models with ProtoPNeXt
by: Willard, Frank, et al.
Published: (2024)
by: Willard, Frank, et al.
Published: (2024)
Two Heads are Actually Better than One: Towards Better Adversarial Robustness via Transduction and Rejection
by: Palumbo, Nils, et al.
Published: (2023)
by: Palumbo, Nils, et al.
Published: (2023)
Chemical Reaction Networks Learn Better than Spiking Neural Networks
by: Jaffard, Sophie, et al.
Published: (2026)
by: Jaffard, Sophie, et al.
Published: (2026)
Do Transformer World Models Give Better Policy Gradients?
by: Ma, Michel, et al.
Published: (2024)
by: Ma, Michel, et al.
Published: (2024)
Does "Do Differentiable Simulators Give Better Policy Gradients?'' Give Better Policy Gradients?
by: Onoda, Ku, et al.
Published: (2026)
by: Onoda, Ku, et al.
Published: (2026)
One Step Forward and K Steps Back: Better Reasoning with Denoising Recursion Models
by: Cameron, Chris, et al.
Published: (2026)
by: Cameron, Chris, et al.
Published: (2026)
APTQ: Attention-aware Post-Training Mixed-Precision Quantization for Large Language Models
by: Guan, Ziyi, et al.
Published: (2024)
by: Guan, Ziyi, et al.
Published: (2024)
DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads Fusion
by: Chen, Yilong, et al.
Published: (2024)
by: Chen, Yilong, et al.
Published: (2024)
One Model, Two Roles: Emergent Specialization in a Shared Recurrent Transformer
by: Shen, Jucheng, et al.
Published: (2026)
by: Shen, Jucheng, et al.
Published: (2026)
The Gaussian-Head OFL Family: One-Shot Federated Learning from Client Global Statistics
by: Turazza, Fabio, et al.
Published: (2026)
by: Turazza, Fabio, et al.
Published: (2026)
COLA: Cross-city Mobility Transformer for Human Trajectory Simulation
by: Wang, Yu, et al.
Published: (2024)
by: Wang, Yu, et al.
Published: (2024)
Deep Fusion: Capturing Dependencies in Contrastive Learning via Transformer Projection Heads
by: Li, Huanran, et al.
Published: (2024)
by: Li, Huanran, et al.
Published: (2024)
Large-Small Model Collaborative Framework for Federated Continual Learning
by: Yu, Hao, et al.
Published: (2025)
by: Yu, Hao, et al.
Published: (2025)
Similar Items
-
Fundamental Limitations on Subquadratic Alternatives to Transformers
by: Alman, Josh, et al.
Published: (2024) -
Every Bit Counts: A Theoretical Study of Precision-Expressivity Tradeoffs in Quantized Transformers
by: Chakrabarti, Sayak, et al.
Published: (2026) -
Poly-attention: a general scheme for higher-order self-attention
by: Chakrabarti, Sayak, et al.
Published: (2026) -
What Makes Looped Transformers Perform Better Than Non-Recursive Ones
by: Gong, Zixuan, et al.
Published: (2025) -
Two Heads Are Better than One: Model-Weight and Latent-Space Analysis for Federated Learning on Non-iid Data against Poisoning Attacks
by: Lyu, Xingyu, et al.
Published: (2025)