Unveiling Downstream Performance Scaling of LLMs: A Clustering-Based Perspective
Fuente:
arXiv
Salvato in:
| Autori principali: | Xu, Chengyin, Chen, Kaiyuan, Li, Xiao, Shen, Ke, Li, Chenggang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Scaling Laws for Predicting Downstream Performance in LLMs
di: Chen, Yangyi, et al.
Pubblicazione: (2024)
di: Chen, Yangyi, et al.
Pubblicazione: (2024)
Quantifying the Importance of Data Alignment in Downstream Model Performance
di: Chawla, Krrish, et al.
Pubblicazione: (2025)
di: Chawla, Krrish, et al.
Pubblicazione: (2025)
Revisiting the Scaling Properties of Downstream Metrics in Large Language Model Training
di: Krajewski, Jakub, et al.
Pubblicazione: (2025)
di: Krajewski, Jakub, et al.
Pubblicazione: (2025)
Task-Informed Anti-Curriculum by Masking Improves Downstream Performance on Text
di: Jarca, Andrei, et al.
Pubblicazione: (2025)
di: Jarca, Andrei, et al.
Pubblicazione: (2025)
Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive?
di: Schaeffer, Rylan, et al.
Pubblicazione: (2024)
di: Schaeffer, Rylan, et al.
Pubblicazione: (2024)
KcMF: A Knowledge-compliant Framework for Schema and Entity Matching with Fine-tuning-free LLMs
di: Xu, Yongqin, et al.
Pubblicazione: (2024)
di: Xu, Yongqin, et al.
Pubblicazione: (2024)
Understanding and Mitigating Bias Inheritance in LLM-based Data Augmentation on Downstream Tasks
di: Li, Miaomiao, et al.
Pubblicazione: (2025)
di: Li, Miaomiao, et al.
Pubblicazione: (2025)
ClusterUCB: Efficient Gradient-Based Data Selection for Targeted Fine-Tuning of LLMs
di: Wang, Zige, et al.
Pubblicazione: (2025)
di: Wang, Zige, et al.
Pubblicazione: (2025)
Principled RL for Diffusion LLMs Emerges from a Sequence-Level Perspective
di: Ou, Jingyang, et al.
Pubblicazione: (2025)
di: Ou, Jingyang, et al.
Pubblicazione: (2025)
Are the Values of LLMs Structurally Aligned with Humans? A Causal Perspective
di: Kang, Yipeng, et al.
Pubblicazione: (2024)
di: Kang, Yipeng, et al.
Pubblicazione: (2024)
Product of Experts with LLMs: Boosting Performance on ARC Is a Matter of Perspective
di: Franzen, Daniel, et al.
Pubblicazione: (2025)
di: Franzen, Daniel, et al.
Pubblicazione: (2025)
What Happened in LLMs Layers when Trained for Fast vs. Slow Thinking: A Gradient Perspective
di: Li, Ming, et al.
Pubblicazione: (2024)
di: Li, Ming, et al.
Pubblicazione: (2024)
How Data Inter-connectivity Shapes LLMs Unlearning: A Structural Unlearning Perspective
di: Qiu, Xinchi, et al.
Pubblicazione: (2024)
di: Qiu, Xinchi, et al.
Pubblicazione: (2024)
A Decomposition Perspective to Long-context Reasoning for LLMs
di: Xiao, Yanling, et al.
Pubblicazione: (2026)
di: Xiao, Yanling, et al.
Pubblicazione: (2026)
Learning to Rewrite Prompts for Bootstrapping LLMs on Downstream Tasks
di: Zhou, Qinhao, et al.
Pubblicazione: (2025)
di: Zhou, Qinhao, et al.
Pubblicazione: (2025)
RAG-Instruct: Boosting LLMs with Diverse Retrieval-Augmented Instructions
di: Liu, Wanlong, et al.
Pubblicazione: (2024)
di: Liu, Wanlong, et al.
Pubblicazione: (2024)
Rethinking the Role of Prompting Strategies in LLM Test-Time Scaling: A Perspective of Probability Theory
di: Liu, Yexiang, et al.
Pubblicazione: (2025)
di: Liu, Yexiang, et al.
Pubblicazione: (2025)
Tactic: Adaptive Sparse Attention with Clustering and Distribution Fitting for Long-Context LLMs
di: Zhu, Kan, et al.
Pubblicazione: (2025)
di: Zhu, Kan, et al.
Pubblicazione: (2025)
Diversity as a Reward: Fine-Tuning LLMs on a Mixture of Domain-Undetermined Data
di: Ling, Zhenqing, et al.
Pubblicazione: (2025)
di: Ling, Zhenqing, et al.
Pubblicazione: (2025)
Configurable Foundation Models: Building LLMs from a Modular Perspective
di: Xiao, Chaojun, et al.
Pubblicazione: (2024)
di: Xiao, Chaojun, et al.
Pubblicazione: (2024)
Reasoning-Based Refinement of Unsupervised Text Clusters with LLMs
di: Islam, Tunazzina
Pubblicazione: (2026)
di: Islam, Tunazzina
Pubblicazione: (2026)
FBI-LLM: Scaling Up Fully Binarized LLMs from Scratch via Autoregressive Distillation
di: Ma, Liqun, et al.
Pubblicazione: (2024)
di: Ma, Liqun, et al.
Pubblicazione: (2024)
DiLoCoX: A Low-Communication Large-Scale Training Framework for Decentralized Cluster
di: Qi, Ji, et al.
Pubblicazione: (2025)
di: Qi, Ji, et al.
Pubblicazione: (2025)
MarkovScale: Towards Optimal Sequential Scaling at Inference Time
di: Wang, Youkang, et al.
Pubblicazione: (2026)
di: Wang, Youkang, et al.
Pubblicazione: (2026)
AutoScale: Scale-Aware Data Mixing for Pre-Training LLMs
di: Kang, Feiyang, et al.
Pubblicazione: (2024)
di: Kang, Feiyang, et al.
Pubblicazione: (2024)
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs
di: Wang, Ganghua, et al.
Pubblicazione: (2025)
di: Wang, Ganghua, et al.
Pubblicazione: (2025)
Multitask Mayhem: Unveiling and Mitigating Safety Gaps in LLMs Fine-tuning
di: Jan, Essa, et al.
Pubblicazione: (2024)
di: Jan, Essa, et al.
Pubblicazione: (2024)
Self-Distillation as a Performance Recovery Mechanism for LLMs: Counteracting Compression and Catastrophic Forgetting
di: Liu, Chi, et al.
Pubblicazione: (2026)
di: Liu, Chi, et al.
Pubblicazione: (2026)
Mission Impossible: A Statistical Perspective on Jailbreaking LLMs
di: Su, Jingtong, et al.
Pubblicazione: (2024)
di: Su, Jingtong, et al.
Pubblicazione: (2024)
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets
di: Chen, Kaiyuan, et al.
Pubblicazione: (2025)
di: Chen, Kaiyuan, et al.
Pubblicazione: (2025)
Instruction Learning Paradigms: A Dual Perspective on White-box and Black-box LLMs
di: Ren, Yanwei, et al.
Pubblicazione: (2025)
di: Ren, Yanwei, et al.
Pubblicazione: (2025)
Graph Learning in the Era of LLMs: A Survey from the Perspective of Data, Models, and Tasks
di: Li, Xunkai, et al.
Pubblicazione: (2024)
di: Li, Xunkai, et al.
Pubblicazione: (2024)
TDBench: A Benchmark for Top-Down Image Understanding with Reliability Analysis of Vision-Language Models
di: Hou, Kaiyuan, et al.
Pubblicazione: (2025)
di: Hou, Kaiyuan, et al.
Pubblicazione: (2025)
VocabTailor: Dynamic Vocabulary Selection for Downstream Tasks in Small Language Models
di: Zhang, Hanling, et al.
Pubblicazione: (2025)
di: Zhang, Hanling, et al.
Pubblicazione: (2025)
ALPINE: Unveiling the Planning Capability of Autoregressive Learning in Language Models
di: Wang, Siwei, et al.
Pubblicazione: (2024)
di: Wang, Siwei, et al.
Pubblicazione: (2024)
Unveiling and Addressing Pseudo Forgetting in Large Language Models
di: Sun, Huashan, et al.
Pubblicazione: (2024)
di: Sun, Huashan, et al.
Pubblicazione: (2024)
Can LLMs Speak For Diverse People? Tuning LLMs via Debate to Generate Controllable Controversial Statements
di: Li, Ming, et al.
Pubblicazione: (2024)
di: Li, Ming, et al.
Pubblicazione: (2024)
Benefits and Pitfalls of Reinforcement Learning for Language Model Planning: A Theoretical Perspective
di: Wang, Siwei, et al.
Pubblicazione: (2025)
di: Wang, Siwei, et al.
Pubblicazione: (2025)
Benchmarking for Domain-Specific LLMs: A Case Study on Academia and Beyond
di: Chen, Rubing, et al.
Pubblicazione: (2025)
di: Chen, Rubing, et al.
Pubblicazione: (2025)
DLO: Dynamic Layer Operation for Efficient Vertical Scaling of LLMs
di: Tan, Zhen, et al.
Pubblicazione: (2024)
di: Tan, Zhen, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Scaling Laws for Predicting Downstream Performance in LLMs
di: Chen, Yangyi, et al.
Pubblicazione: (2024) -
Quantifying the Importance of Data Alignment in Downstream Model Performance
di: Chawla, Krrish, et al.
Pubblicazione: (2025) -
Revisiting the Scaling Properties of Downstream Metrics in Large Language Model Training
di: Krajewski, Jakub, et al.
Pubblicazione: (2025) -
Task-Informed Anti-Curriculum by Masking Improves Downstream Performance on Text
di: Jarca, Andrei, et al.
Pubblicazione: (2025) -
Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive?
di: Schaeffer, Rylan, et al.
Pubblicazione: (2024)