Gespeichert in:
| Hauptverfasser: | Chang, Ernie, Paltenghi, Matteo, Li, Yang, Lin, Pin-Jie, Zhao, Changsheng, Huber, Patrick, Liu, Zechun, Rabatin, Rastislav, Shi, Yangyang, Chandra, Vikas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2410.03083 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Target-Aware Language Modeling via Granular Data Sampling
von: Chang, Ernie, et al.
Veröffentlicht: (2024)
von: Chang, Ernie, et al.
Veröffentlicht: (2024)
Navigating the Minefield of MT Beam Search in Cascaded Streaming Speech Translation
von: Rabatin, Rastislav, et al.
Veröffentlicht: (2024)
von: Rabatin, Rastislav, et al.
Veröffentlicht: (2024)
Self-Vocabularizing Training for Neural Machine Translation
von: Lin, Pin-Jie, et al.
Veröffentlicht: (2025)
von: Lin, Pin-Jie, et al.
Veröffentlicht: (2025)
Basis Selection: Low-Rank Decomposition of Pretrained Large Language Models for Target Applications
von: Li, Yang, et al.
Veröffentlicht: (2024)
von: Li, Yang, et al.
Veröffentlicht: (2024)
MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases
von: Liu, Zechun, et al.
Veröffentlicht: (2024)
von: Liu, Zechun, et al.
Veröffentlicht: (2024)
AutoMixer: Checkpoint Artifacts as Automatic Data Mixers
von: Chang, Ernie, et al.
Veröffentlicht: (2025)
von: Chang, Ernie, et al.
Veröffentlicht: (2025)
MobileLLM-R1: Exploring the Limits of Sub-Billion Language Model Reasoners with Open Training Recipes
von: Zhao, Changsheng, et al.
Veröffentlicht: (2025)
von: Zhao, Changsheng, et al.
Veröffentlicht: (2025)
MobileMoE: Scaling On-Device Mixture of Experts
von: Chen, Yanbei, et al.
Veröffentlicht: (2026)
von: Chen, Yanbei, et al.
Veröffentlicht: (2026)
RPRA: Predicting an LLM-Judge for Efficient but Performant Inference
von: Ashley, Dylan R., et al.
Veröffentlicht: (2026)
von: Ashley, Dylan R., et al.
Veröffentlicht: (2026)
ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM Quantization
von: Liu, Zechun, et al.
Veröffentlicht: (2025)
von: Liu, Zechun, et al.
Veröffentlicht: (2025)
Breaking Down Power Barriers in On-Device Streaming ASR: Insights and Solutions
von: Li, Yang, et al.
Veröffentlicht: (2024)
von: Li, Yang, et al.
Veröffentlicht: (2024)
SpinQuant: LLM quantization with learned rotations
von: Liu, Zechun, et al.
Veröffentlicht: (2024)
von: Liu, Zechun, et al.
Veröffentlicht: (2024)
Wink: Recovering from Misbehaviors in Coding Agents
von: Nanda, Rahul, et al.
Veröffentlicht: (2026)
von: Nanda, Rahul, et al.
Veröffentlicht: (2026)
Agent-as-a-Judge: Evaluate Agents with Agents
von: Zhuge, Mingchen, et al.
Veröffentlicht: (2024)
von: Zhuge, Mingchen, et al.
Veröffentlicht: (2024)
Non-Monotonic Attention-based Read/Write Policy Learning for Simultaneous Translation
von: Ahmed, Zeeshan, et al.
Veröffentlicht: (2025)
von: Ahmed, Zeeshan, et al.
Veröffentlicht: (2025)
Short Data, Long Context: Distilling Positional Knowledge in Transformers
von: Huber, Patrick, et al.
Veröffentlicht: (2026)
von: Huber, Patrick, et al.
Veröffentlicht: (2026)
CoSMoEs: Compact Sparse Mixture of Experts
von: Huber, Patrick, et al.
Veröffentlicht: (2025)
von: Huber, Patrick, et al.
Veröffentlicht: (2025)
dTRPO: Trajectory Reduction in Policy Optimization of Diffusion Large Language Models
von: Zhang, Wenxuan, et al.
Veröffentlicht: (2026)
von: Zhang, Wenxuan, et al.
Veröffentlicht: (2026)
A Survey on Testing and Analysis of Quantum Software
von: Paltenghi, Matteo, et al.
Veröffentlicht: (2024)
von: Paltenghi, Matteo, et al.
Veröffentlicht: (2024)
QITE: Assembly-Level, Cross-Platform Testing of Quantum Computing Platforms
von: Paltenghi, Matteo, et al.
Veröffentlicht: (2025)
von: Paltenghi, Matteo, et al.
Veröffentlicht: (2025)
Analyzing Quantum Programs with LintQ: A Static Analysis Framework for Qiskit
von: Paltenghi, Matteo, et al.
Veröffentlicht: (2023)
von: Paltenghi, Matteo, et al.
Veröffentlicht: (2023)
Scaling Data-Constrained Language Models
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2023)
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2023)
Folding Attention: Memory and Power Optimization for On-Device Transformer-based Streaming Speech Recognition
von: Li, Yang, et al.
Veröffentlicht: (2023)
von: Li, Yang, et al.
Veröffentlicht: (2023)
EgoAVU: Egocentric Audio-Visual Understanding
von: Seth, Ashish, et al.
Veröffentlicht: (2026)
von: Seth, Ashish, et al.
Veröffentlicht: (2026)
Exploring Audio Hallucination in Egocentric Video Understanding
von: Seth, Ashish, et al.
Veröffentlicht: (2026)
von: Seth, Ashish, et al.
Veröffentlicht: (2026)
REAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production Usage
von: Jha, Smriti, et al.
Veröffentlicht: (2026)
von: Jha, Smriti, et al.
Veröffentlicht: (2026)
Measuring Social Biases in Masked Language Models by Proxy of Prediction Quality
von: Zalkikar, Rahul, et al.
Veröffentlicht: (2024)
von: Zalkikar, Rahul, et al.
Veröffentlicht: (2024)
Neural Computers
von: Zhuge, Mingchen, et al.
Veröffentlicht: (2026)
von: Zhuge, Mingchen, et al.
Veröffentlicht: (2026)
Mixture-of-Supernets: Improving Weight-Sharing Supernet Training with Architecture-Routed Mixture-of-Experts
von: Jawahar, Ganesh, et al.
Veröffentlicht: (2023)
von: Jawahar, Ganesh, et al.
Veröffentlicht: (2023)
Morello: Compiling Fast Neural Networks with Dynamic Programming and Spatial Compression
von: Kaufman, Samuel J., et al.
Veröffentlicht: (2025)
von: Kaufman, Samuel J., et al.
Veröffentlicht: (2025)
Param$Δ$ for Direct Weight Mixing: Post-Train Large Language Model at Zero Cost
von: Cao, Sheng, et al.
Veröffentlicht: (2025)
von: Cao, Sheng, et al.
Veröffentlicht: (2025)
RapidUn: Influence-Driven Parameter Reweighting for Efficient Large Language Model Unlearning
von: Zhao, Guoshenghui, et al.
Veröffentlicht: (2025)
von: Zhao, Guoshenghui, et al.
Veröffentlicht: (2025)
Prescriptive Scaling Laws for Data Constrained Training
von: Lovelace, Justin, et al.
Veröffentlicht: (2026)
von: Lovelace, Justin, et al.
Veröffentlicht: (2026)
SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
von: Liu, Dongyang, et al.
Veröffentlicht: (2024)
von: Liu, Dongyang, et al.
Veröffentlicht: (2024)
Small Vision-Language Models are Smart Compressors for Long Video Understanding
von: Fei, Junjie, et al.
Veröffentlicht: (2026)
von: Fei, Junjie, et al.
Veröffentlicht: (2026)
Exploring the Effectiveness and Consistency of Task Selection in Intermediate-Task Transfer Learning
von: Lin, Pin-Jie, et al.
Veröffentlicht: (2024)
von: Lin, Pin-Jie, et al.
Veröffentlicht: (2024)
CoDa: Constrained Generation based Data Augmentation for Low-Resource NLP
von: Evuru, Chandra Kiran Reddy, et al.
Veröffentlicht: (2024)
von: Evuru, Chandra Kiran Reddy, et al.
Veröffentlicht: (2024)
DepthLM: Metric Depth From Vision Language Models
von: Cai, Zhipeng, et al.
Veröffentlicht: (2025)
von: Cai, Zhipeng, et al.
Veröffentlicht: (2025)
Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data
von: Gu, Shuhao, et al.
Veröffentlicht: (2024)
von: Gu, Shuhao, et al.
Veröffentlicht: (2024)
Quality or Quantity? On Data Scale and Diversity in Adapting Large Language Models for Low-Resource Translation
von: Iyer, Vivek, et al.
Veröffentlicht: (2024)
von: Iyer, Vivek, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Target-Aware Language Modeling via Granular Data Sampling
von: Chang, Ernie, et al.
Veröffentlicht: (2024) -
Navigating the Minefield of MT Beam Search in Cascaded Streaming Speech Translation
von: Rabatin, Rastislav, et al.
Veröffentlicht: (2024) -
Self-Vocabularizing Training for Neural Machine Translation
von: Lin, Pin-Jie, et al.
Veröffentlicht: (2025) -
Basis Selection: Low-Rank Decomposition of Pretrained Large Language Models for Target Applications
von: Li, Yang, et al.
Veröffentlicht: (2024) -
MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases
von: Liu, Zechun, et al.
Veröffentlicht: (2024)