AutoMixer: Checkpoint Artifacts as Automatic Data Mixers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chang, Ernie, Li, Yang, Huber, Patrick, Vogeti, Vish, Kant, David, Shi, Yangyang, Chandra, Vikas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AutoMix: Automatically Mixing Language Models
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2023)
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2023)
Self-Vocabularizing Training for Neural Machine Translation
von: Lin, Pin-Jie, et al.
Veröffentlicht: (2025)
von: Lin, Pin-Jie, et al.
Veröffentlicht: (2025)
Scaling Parameter-Constrained Language Models with Quality Data
von: Chang, Ernie, et al.
Veröffentlicht: (2024)
von: Chang, Ernie, et al.
Veröffentlicht: (2024)
Free Energy Mixer
von: Lu, Jiecheng, et al.
Veröffentlicht: (2026)
von: Lu, Jiecheng, et al.
Veröffentlicht: (2026)
Basis Selection: Low-Rank Decomposition of Pretrained Large Language Models for Target Applications
von: Li, Yang, et al.
Veröffentlicht: (2024)
von: Li, Yang, et al.
Veröffentlicht: (2024)
Target-Aware Language Modeling via Granular Data Sampling
von: Chang, Ernie, et al.
Veröffentlicht: (2024)
von: Chang, Ernie, et al.
Veröffentlicht: (2024)
Masked Mixers for Language Generation and Retrieval
von: Badger, Benjamin L.
Veröffentlicht: (2024)
von: Badger, Benjamin L.
Veröffentlicht: (2024)
Cubit: Token Mixer with Kernel Ridge Regression
von: Zheng, Chuanyang, et al.
Veröffentlicht: (2026)
von: Zheng, Chuanyang, et al.
Veröffentlicht: (2026)
Structured Recurrent Mixers for Massively Parallelized Sequence Generation
von: Badger, Benjamin L.
Veröffentlicht: (2026)
von: Badger, Benjamin L.
Veröffentlicht: (2026)
LLM-Mixer: Multiscale Mixing in LLMs for Time Series Forecasting
von: Kowsher, Md, et al.
Veröffentlicht: (2024)
von: Kowsher, Md, et al.
Veröffentlicht: (2024)
ChebMixer: Efficient Graph Representation Learning with MLP Mixer
von: Kui, Xiaoyan, et al.
Veröffentlicht: (2024)
von: Kui, Xiaoyan, et al.
Veröffentlicht: (2024)
Toeplitz MLP Mixers are Low Complexity, Information-Rich Sequence Models
von: Badger, Benjamin L., et al.
Veröffentlicht: (2026)
von: Badger, Benjamin L., et al.
Veröffentlicht: (2026)
MixerFlow: MLP-Mixer meets Normalising Flows
von: English, Eshant, et al.
Veröffentlicht: (2023)
von: English, Eshant, et al.
Veröffentlicht: (2023)
MobileLLM-R1: Exploring the Limits of Sub-Billion Language Model Reasoners with Open Training Recipes
von: Zhao, Changsheng, et al.
Veröffentlicht: (2025)
von: Zhao, Changsheng, et al.
Veröffentlicht: (2025)
Code-Mixer Ya Nahi: Novel Approaches to Measuring Multilingual LLMs' Code-Mixing Capabilities
von: Gupta, Ayushman, et al.
Veröffentlicht: (2024)
von: Gupta, Ayushman, et al.
Veröffentlicht: (2024)
FFNet: MetaMixer-based Efficient Convolutional Mixer Design
von: Yun, Seokju, et al.
Veröffentlicht: (2024)
von: Yun, Seokju, et al.
Veröffentlicht: (2024)
Short Data, Long Context: Distilling Positional Knowledge in Transformers
von: Huber, Patrick, et al.
Veröffentlicht: (2026)
von: Huber, Patrick, et al.
Veröffentlicht: (2026)
CoSMoEs: Compact Sparse Mixture of Experts
von: Huber, Patrick, et al.
Veröffentlicht: (2025)
von: Huber, Patrick, et al.
Veröffentlicht: (2025)
GroupedMixer: An Entropy Model with Group-wise Token-Mixers for Learned Image Compression
von: Li, Daxin, et al.
Veröffentlicht: (2024)
von: Li, Daxin, et al.
Veröffentlicht: (2024)
STS-Mixer: Spatio-Temporal-Spectral Mixer for 4D Point Cloud Video Understanding
von: Li, Wenhao, et al.
Veröffentlicht: (2026)
von: Li, Wenhao, et al.
Veröffentlicht: (2026)
iMixer: hierarchical Hopfield network implies an invertible, implicit and iterative MLP-Mixer
von: Ota, Toshihiro, et al.
Veröffentlicht: (2023)
von: Ota, Toshihiro, et al.
Veröffentlicht: (2023)
DeGMix: Efficient Multi-Task Dense Prediction with Deformable and Gating Mixer
von: Xu, Yangyang, et al.
Veröffentlicht: (2023)
von: Xu, Yangyang, et al.
Veröffentlicht: (2023)
MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases
von: Liu, Zechun, et al.
Veröffentlicht: (2024)
von: Liu, Zechun, et al.
Veröffentlicht: (2024)
MixerCSeg: An Efficient Mixer Architecture for Crack Segmentation via Decoupled Mamba Attention
von: Zhao, Zilong, et al.
Veröffentlicht: (2026)
von: Zhao, Zilong, et al.
Veröffentlicht: (2026)
MobileMoE: Scaling On-Device Mixture of Experts
von: Chen, Yanbei, et al.
Veröffentlicht: (2026)
von: Chen, Yanbei, et al.
Veröffentlicht: (2026)
zkMixer: A Configurable Zero-Knowledge Mixer with Anti-Money Laundering Consensus Protocols
von: Constantinides, Theodoros, et al.
Veröffentlicht: (2025)
von: Constantinides, Theodoros, et al.
Veröffentlicht: (2025)
QN-Mixer: A Quasi-Newton MLP-Mixer Model for Sparse-View CT Reconstruction
von: Ayad, Ishak, et al.
Veröffentlicht: (2024)
von: Ayad, Ishak, et al.
Veröffentlicht: (2024)
SCHEME: Scalable Channel Mixer for Vision Transformers
von: Sridhar, Deepak, et al.
Veröffentlicht: (2023)
von: Sridhar, Deepak, et al.
Veröffentlicht: (2023)
AutoTriton: Automatic Triton Programming with Reinforcement Learning in LLMs
von: Li, Shangzhan, et al.
Veröffentlicht: (2025)
von: Li, Shangzhan, et al.
Veröffentlicht: (2025)
AILive Mixer: A Deep Learning based Zero Latency Automatic Music Mixer for Live Music Performances
von: Zurale, Devansh, et al.
Veröffentlicht: (2026)
von: Zurale, Devansh, et al.
Veröffentlicht: (2026)
ATOM: Attention Mixer for Efficient Dataset Distillation
von: Khaki, Samir, et al.
Veröffentlicht: (2024)
von: Khaki, Samir, et al.
Veröffentlicht: (2024)
PoM: Efficient Image and Video Generation with the Polynomial Mixer
von: Picard, David, et al.
Veröffentlicht: (2024)
von: Picard, David, et al.
Veröffentlicht: (2024)
RPRA: Predicting an LLM-Judge for Efficient but Performant Inference
von: Ashley, Dylan R., et al.
Veröffentlicht: (2026)
von: Ashley, Dylan R., et al.
Veröffentlicht: (2026)
U-Mixer: An Unet-Mixer Architecture with Stationarity Correction for Time Series Forecasting
von: Ma, Xiang, et al.
Veröffentlicht: (2024)
von: Ma, Xiang, et al.
Veröffentlicht: (2024)
Findings of the Third Automatic Minuting (AutoMin) Challenge
von: Shinde, Kartik, et al.
Veröffentlicht: (2025)
von: Shinde, Kartik, et al.
Veröffentlicht: (2025)
AutoFAIR : Automatic Data FAIRification via Machine Reading
von: Ma, Tingyan, et al.
Veröffentlicht: (2024)
von: Ma, Tingyan, et al.
Veröffentlicht: (2024)
Topological Analysis of Mixer Activities in the Bitcoin Network
von: Zola, Francesco, et al.
Veröffentlicht: (2025)
von: Zola, Francesco, et al.
Veröffentlicht: (2025)
AutoLife: Automatic Life Journaling with Smartphones and LLMs
von: Xu, Huatao, et al.
Veröffentlicht: (2024)
von: Xu, Huatao, et al.
Veröffentlicht: (2024)
AutoMetrics: Approximate Human Judgements with Automatically Generated Evaluators
von: Ryan, Michael J., et al.
Veröffentlicht: (2025)
von: Ryan, Michael J., et al.
Veröffentlicht: (2025)
AutoSpec: An Agentic Framework for Automatically Drafting Patent Specification
von: Shea, Ryan, et al.
Veröffentlicht: (2025)
von: Shea, Ryan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AutoMix: Automatically Mixing Language Models
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2023) -
Self-Vocabularizing Training for Neural Machine Translation
von: Lin, Pin-Jie, et al.
Veröffentlicht: (2025) -
Scaling Parameter-Constrained Language Models with Quality Data
von: Chang, Ernie, et al.
Veröffentlicht: (2024) -
Free Energy Mixer
von: Lu, Jiecheng, et al.
Veröffentlicht: (2026) -
Basis Selection: Low-Rank Decomposition of Pretrained Large Language Models for Target Applications
von: Li, Yang, et al.
Veröffentlicht: (2024)