AutoMixer: Checkpoint Artifacts as Automatic Data Mixers
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Chang, Ernie, Li, Yang, Huber, Patrick, Vogeti, Vish, Kant, David, Shi, Yangyang, Chandra, Vikas |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
AutoMix: Automatically Mixing Language Models
par: Aggarwal, Pranjal, et autres
Publié: (2023)
par: Aggarwal, Pranjal, et autres
Publié: (2023)
Self-Vocabularizing Training for Neural Machine Translation
par: Lin, Pin-Jie, et autres
Publié: (2025)
par: Lin, Pin-Jie, et autres
Publié: (2025)
Scaling Parameter-Constrained Language Models with Quality Data
par: Chang, Ernie, et autres
Publié: (2024)
par: Chang, Ernie, et autres
Publié: (2024)
Free Energy Mixer
par: Lu, Jiecheng, et autres
Publié: (2026)
par: Lu, Jiecheng, et autres
Publié: (2026)
Basis Selection: Low-Rank Decomposition of Pretrained Large Language Models for Target Applications
par: Li, Yang, et autres
Publié: (2024)
par: Li, Yang, et autres
Publié: (2024)
Target-Aware Language Modeling via Granular Data Sampling
par: Chang, Ernie, et autres
Publié: (2024)
par: Chang, Ernie, et autres
Publié: (2024)
Masked Mixers for Language Generation and Retrieval
par: Badger, Benjamin L.
Publié: (2024)
par: Badger, Benjamin L.
Publié: (2024)
Cubit: Token Mixer with Kernel Ridge Regression
par: Zheng, Chuanyang, et autres
Publié: (2026)
par: Zheng, Chuanyang, et autres
Publié: (2026)
Structured Recurrent Mixers for Massively Parallelized Sequence Generation
par: Badger, Benjamin L.
Publié: (2026)
par: Badger, Benjamin L.
Publié: (2026)
LLM-Mixer: Multiscale Mixing in LLMs for Time Series Forecasting
par: Kowsher, Md, et autres
Publié: (2024)
par: Kowsher, Md, et autres
Publié: (2024)
ChebMixer: Efficient Graph Representation Learning with MLP Mixer
par: Kui, Xiaoyan, et autres
Publié: (2024)
par: Kui, Xiaoyan, et autres
Publié: (2024)
Toeplitz MLP Mixers are Low Complexity, Information-Rich Sequence Models
par: Badger, Benjamin L., et autres
Publié: (2026)
par: Badger, Benjamin L., et autres
Publié: (2026)
MixerFlow: MLP-Mixer meets Normalising Flows
par: English, Eshant, et autres
Publié: (2023)
par: English, Eshant, et autres
Publié: (2023)
MobileLLM-R1: Exploring the Limits of Sub-Billion Language Model Reasoners with Open Training Recipes
par: Zhao, Changsheng, et autres
Publié: (2025)
par: Zhao, Changsheng, et autres
Publié: (2025)
Code-Mixer Ya Nahi: Novel Approaches to Measuring Multilingual LLMs' Code-Mixing Capabilities
par: Gupta, Ayushman, et autres
Publié: (2024)
par: Gupta, Ayushman, et autres
Publié: (2024)
FFNet: MetaMixer-based Efficient Convolutional Mixer Design
par: Yun, Seokju, et autres
Publié: (2024)
par: Yun, Seokju, et autres
Publié: (2024)
Short Data, Long Context: Distilling Positional Knowledge in Transformers
par: Huber, Patrick, et autres
Publié: (2026)
par: Huber, Patrick, et autres
Publié: (2026)
CoSMoEs: Compact Sparse Mixture of Experts
par: Huber, Patrick, et autres
Publié: (2025)
par: Huber, Patrick, et autres
Publié: (2025)
GroupedMixer: An Entropy Model with Group-wise Token-Mixers for Learned Image Compression
par: Li, Daxin, et autres
Publié: (2024)
par: Li, Daxin, et autres
Publié: (2024)
STS-Mixer: Spatio-Temporal-Spectral Mixer for 4D Point Cloud Video Understanding
par: Li, Wenhao, et autres
Publié: (2026)
par: Li, Wenhao, et autres
Publié: (2026)
iMixer: hierarchical Hopfield network implies an invertible, implicit and iterative MLP-Mixer
par: Ota, Toshihiro, et autres
Publié: (2023)
par: Ota, Toshihiro, et autres
Publié: (2023)
DeGMix: Efficient Multi-Task Dense Prediction with Deformable and Gating Mixer
par: Xu, Yangyang, et autres
Publié: (2023)
par: Xu, Yangyang, et autres
Publié: (2023)
MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases
par: Liu, Zechun, et autres
Publié: (2024)
par: Liu, Zechun, et autres
Publié: (2024)
MixerCSeg: An Efficient Mixer Architecture for Crack Segmentation via Decoupled Mamba Attention
par: Zhao, Zilong, et autres
Publié: (2026)
par: Zhao, Zilong, et autres
Publié: (2026)
MobileMoE: Scaling On-Device Mixture of Experts
par: Chen, Yanbei, et autres
Publié: (2026)
par: Chen, Yanbei, et autres
Publié: (2026)
zkMixer: A Configurable Zero-Knowledge Mixer with Anti-Money Laundering Consensus Protocols
par: Constantinides, Theodoros, et autres
Publié: (2025)
par: Constantinides, Theodoros, et autres
Publié: (2025)
QN-Mixer: A Quasi-Newton MLP-Mixer Model for Sparse-View CT Reconstruction
par: Ayad, Ishak, et autres
Publié: (2024)
par: Ayad, Ishak, et autres
Publié: (2024)
SCHEME: Scalable Channel Mixer for Vision Transformers
par: Sridhar, Deepak, et autres
Publié: (2023)
par: Sridhar, Deepak, et autres
Publié: (2023)
AutoTriton: Automatic Triton Programming with Reinforcement Learning in LLMs
par: Li, Shangzhan, et autres
Publié: (2025)
par: Li, Shangzhan, et autres
Publié: (2025)
AILive Mixer: A Deep Learning based Zero Latency Automatic Music Mixer for Live Music Performances
par: Zurale, Devansh, et autres
Publié: (2026)
par: Zurale, Devansh, et autres
Publié: (2026)
ATOM: Attention Mixer for Efficient Dataset Distillation
par: Khaki, Samir, et autres
Publié: (2024)
par: Khaki, Samir, et autres
Publié: (2024)
PoM: Efficient Image and Video Generation with the Polynomial Mixer
par: Picard, David, et autres
Publié: (2024)
par: Picard, David, et autres
Publié: (2024)
RPRA: Predicting an LLM-Judge for Efficient but Performant Inference
par: Ashley, Dylan R., et autres
Publié: (2026)
par: Ashley, Dylan R., et autres
Publié: (2026)
U-Mixer: An Unet-Mixer Architecture with Stationarity Correction for Time Series Forecasting
par: Ma, Xiang, et autres
Publié: (2024)
par: Ma, Xiang, et autres
Publié: (2024)
Findings of the Third Automatic Minuting (AutoMin) Challenge
par: Shinde, Kartik, et autres
Publié: (2025)
par: Shinde, Kartik, et autres
Publié: (2025)
AutoFAIR : Automatic Data FAIRification via Machine Reading
par: Ma, Tingyan, et autres
Publié: (2024)
par: Ma, Tingyan, et autres
Publié: (2024)
Topological Analysis of Mixer Activities in the Bitcoin Network
par: Zola, Francesco, et autres
Publié: (2025)
par: Zola, Francesco, et autres
Publié: (2025)
AutoLife: Automatic Life Journaling with Smartphones and LLMs
par: Xu, Huatao, et autres
Publié: (2024)
par: Xu, Huatao, et autres
Publié: (2024)
AutoMetrics: Approximate Human Judgements with Automatically Generated Evaluators
par: Ryan, Michael J., et autres
Publié: (2025)
par: Ryan, Michael J., et autres
Publié: (2025)
AutoSpec: An Agentic Framework for Automatically Drafting Patent Specification
par: Shea, Ryan, et autres
Publié: (2025)
par: Shea, Ryan, et autres
Publié: (2025)
Documents similaires
-
AutoMix: Automatically Mixing Language Models
par: Aggarwal, Pranjal, et autres
Publié: (2023) -
Self-Vocabularizing Training for Neural Machine Translation
par: Lin, Pin-Jie, et autres
Publié: (2025) -
Scaling Parameter-Constrained Language Models with Quality Data
par: Chang, Ernie, et autres
Publié: (2024) -
Free Energy Mixer
par: Lu, Jiecheng, et autres
Publié: (2026) -
Basis Selection: Low-Rank Decomposition of Pretrained Large Language Models for Target Applications
par: Li, Yang, et autres
Publié: (2024)