Collage: Light-Weight Low-Precision Strategy for LLM Training
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yu, Tao, Gupta, Gaurav, Gopalswamy, Karthick, Mamidala, Amith, Zhou, Hao, Huynh, Jeffrey, Park, Youngsuk, Diamant, Ron, Deoras, Anoop, Huan, Luke |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
ContextWeaver: Selective and Dependency-Structured Memory Construction for LLM Agents
par: Wu, Yating, et autres
Publié: (2026)
par: Wu, Yating, et autres
Publié: (2026)
Stochastic Rounding for LLM Training: Theory and Practice
par: Ozkara, Kaan, et autres
Publié: (2025)
par: Ozkara, Kaan, et autres
Publié: (2025)
Theoretical Guarantees of Learning Ensembling Strategies with Applications to Time Series Forecasting
par: Hasson, Hilaf, et autres
Publié: (2023)
par: Hasson, Hilaf, et autres
Publié: (2023)
Training LLMs with MXFP4
par: Tseng, Albert, et autres
Publié: (2025)
par: Tseng, Albert, et autres
Publié: (2025)
Lossless Token Sequence Compression via Meta-Tokens
par: Harvill, John, et autres
Publié: (2025)
par: Harvill, John, et autres
Publié: (2025)
Lightweight reranking for language model generations
par: Jain, Siddhartha, et autres
Publié: (2023)
par: Jain, Siddhartha, et autres
Publié: (2023)
Developing a Novel Holistic, Personalized Dementia Risk Prediction Model via Integration of Machine Learning and Network Systems Biology Approaches
par: Mamidala, Srilekha
Publié: (2023)
par: Mamidala, Srilekha
Publié: (2023)
Automated Evaluation of Retrieval-Augmented Language Models with Task-Specific Exam Generation
par: Guinet, Gauthier, et autres
Publié: (2024)
par: Guinet, Gauthier, et autres
Publié: (2024)
Orchestrating the Media Collage
par: Ohler, Jason
Publié: (2009)
par: Ohler, Jason
Publié: (2009)
A Morning Collage.
par: Zingher, Gary
Publié: (1995)
par: Zingher, Gary
Publié: (1995)
Inversion of Magnetic Anomalies Due to 2-D Cylindrical Structures – By an Artificial Neural Network
par: Bhagwan Das Mamidala
Publié: (2025)
par: Bhagwan Das Mamidala
Publié: (2025)
Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs
par: Bian, Song, et autres
Publié: (2025)
par: Bian, Song, et autres
Publié: (2025)
CodeAssistBench (CAB): Dataset & Benchmarking for Multi-turn Chat-Based Code Assistance
par: Kim, Myeongsoo, et autres
Publié: (2025)
par: Kim, Myeongsoo, et autres
Publié: (2025)
Multi-IaC-Eval: Benchmarking Cloud Infrastructure as Code Across Multiple Formats
par: Davidson, Sam, et autres
Publié: (2025)
par: Davidson, Sam, et autres
Publié: (2025)
AI Pose Analysis and Kinematic Profiling of Range-of-Motion Variations in Resistance Training
par: Diamant, Adam
Publié: (2025)
par: Diamant, Adam
Publié: (2025)
Collage collection / Karen S. Chambers
par: Chambers, Karen S
Publié: (1989)
par: Chambers, Karen S
Publié: (1989)
MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation
par: Ekbote, Chanakya, et autres
Publié: (2025)
par: Ekbote, Chanakya, et autres
Publié: (2025)
UTFix: Change Aware Unit Test Repairing using LLM
par: Rahman, Shanto, et autres
Publié: (2025)
par: Rahman, Shanto, et autres
Publié: (2025)
Solar Energetic Particle Events and Radio Bursts
par: Gopalswamy, Nat
Publié: (2024)
par: Gopalswamy, Nat
Publié: (2024)
Flavonoids as Multifunctional Agents: Targeting Multiple Enzymes in Cancer Therapy
par: PASULA, JANAKIRAMULU, et autres
Publié: (2025)
par: PASULA, JANAKIRAMULU, et autres
Publié: (2025)
On the Smallest Size of Internal Collage Systems
par: Migita, Soichiro, et autres
Publié: (2025)
par: Migita, Soichiro, et autres
Publié: (2025)
‘reportless places’: Janet Malcolm and Collage
par: Natalie Ferris
Publié: (2026)
par: Natalie Ferris
Publié: (2026)
Kierkegaard's Romantic Legacy
par: Gupta, Anoop
Publié: (2017)
par: Gupta, Anoop
Publié: (2017)
DBAutoDoc: Automated Discovery and Documentation of Undocumented Database Schemas via Statistical Analysis and Iterative LLM Refinement
par: Nagarajan, Amith, et autres
Publié: (2026)
par: Nagarajan, Amith, et autres
Publié: (2026)
TerraFormer: Automated Infrastructure-as-Code with LLMs Fine-Tuned via Policy-Guided Verifier Feedback
par: Jana, Prithwish, et autres
Publié: (2026)
par: Jana, Prithwish, et autres
Publié: (2026)
MuonBP: Faster Muon via Block-Periodic Orthogonalization
par: Khaled, Ahmed, et autres
Publié: (2025)
par: Khaled, Ahmed, et autres
Publié: (2025)
Image-Space Collage and Packing with Differentiable Rendering
par: Wang, Zhenyu, et autres
Publié: (2024)
par: Wang, Zhenyu, et autres
Publié: (2024)
TritonRL: Training LLMs to Think and Code Triton Without Cheating
par: Woo, Jiin, et autres
Publié: (2025)
par: Woo, Jiin, et autres
Publié: (2025)
Low-Light Image and Video Enhancement: A Comprehensive Survey and Beyond
par: Zheng, Shen, et autres
Publié: (2022)
par: Zheng, Shen, et autres
Publié: (2022)
LeDex: Training LLMs to Better Self-Debug and Explain Code
par: Jiang, Nan, et autres
Publié: (2024)
par: Jiang, Nan, et autres
Publié: (2024)
SINQ: Sinkhorn-Normalized Quantization for Calibration-Free Low-Precision LLM Weights
par: Müller, Lorenz K., et autres
Publié: (2025)
par: Müller, Lorenz K., et autres
Publié: (2025)
ExecTune: Effective Steering of Black-Box LLMs with Guide Models
par: Lingam, Vijay, et autres
Publié: (2026)
par: Lingam, Vijay, et autres
Publié: (2026)
Pre-trained Recommender Systems: A Causal Debiasing Perspective
par: Lin, Ziqian, et autres
Publié: (2023)
par: Lin, Ziqian, et autres
Publié: (2023)
Momentum Based Reward Design for Low Emission Traffic Signal Control
par: Mundane, Chinmay, et autres
Publié: (2026)
par: Mundane, Chinmay, et autres
Publié: (2026)
ProactivePIM: Accelerating Weight-Sharing Embedding Layer with PIM for Scalable Recommendation System
par: Kim, Youngsuk, et autres
Publié: (2024)
par: Kim, Youngsuk, et autres
Publié: (2024)
Gradient Weight-normalized Low-rank Projection for Efficient LLM Training
par: Huang, Jia-Hong, et autres
Publié: (2024)
par: Huang, Jia-Hong, et autres
Publié: (2024)
In-Silico Molecular Docking Studies of Cymbopogon Nardus Compounds against Malarial Targets of Plasmodium falciparum
par: Raghuma, Reddy, et autres
Publié: (2025)
par: Raghuma, Reddy, et autres
Publié: (2025)
Logic-Scaffolding: Personalized Aspect-Instructed Recommendation Explanation Generation using LLMs
par: Rahdari, Behnam, et autres
Publié: (2023)
par: Rahdari, Behnam, et autres
Publié: (2023)
Fewer Truncations Improve Language Modeling
par: Ding, Hantian, et autres
Publié: (2024)
par: Ding, Hantian, et autres
Publié: (2024)
A Light Weight Cryptographic Solution for 6LoWPAN Protocol Stack
par: Khairnar, Sushil, et autres
Publié: (2025)
par: Khairnar, Sushil, et autres
Publié: (2025)
Documents similaires
-
ContextWeaver: Selective and Dependency-Structured Memory Construction for LLM Agents
par: Wu, Yating, et autres
Publié: (2026) -
Stochastic Rounding for LLM Training: Theory and Practice
par: Ozkara, Kaan, et autres
Publié: (2025) -
Theoretical Guarantees of Learning Ensembling Strategies with Applications to Time Series Forecasting
par: Hasson, Hilaf, et autres
Publié: (2023) -
Training LLMs with MXFP4
par: Tseng, Albert, et autres
Publié: (2025) -
Lossless Token Sequence Compression via Meta-Tokens
par: Harvill, John, et autres
Publié: (2025)