Capability Salience Vector: Fine-grained Alignment of Loss and Capabilities for Downstream Task Scaling Law
Fuente:
arXiv
Saved in:
| Main Authors: | Ge, Qiming, Xing, Shuhao, Gao, Songyang, Zhou, Yunhua, Zou, Yicheng, Zhang, Songyang, Chen, Zhi, Yan, Hang, Zhang, Qi, Guo, Qipeng, Chen, Kai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Are Your LLMs Capable of Stable Reasoning?
by: Liu, Junnan, et al.
Published: (2024)
by: Liu, Junnan, et al.
Published: (2024)
Linear Alignment: A Closed-form Solution for Aligning Human Preferences without Tuning and Feedback
by: Gao, Songyang, et al.
Published: (2024)
by: Gao, Songyang, et al.
Published: (2024)
Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs
by: Qiao, Yuxuan, et al.
Published: (2024)
by: Qiao, Yuxuan, et al.
Published: (2024)
How to Set the Learning Rate for Large-Scale Pre-training?
by: Zhou, Yunhua, et al.
Published: (2026)
by: Zhou, Yunhua, et al.
Published: (2026)
The Fine Line: Navigating Large Language Model Pretraining with Down-streaming Capability Analysis
by: Yang, Chen, et al.
Published: (2024)
by: Yang, Chen, et al.
Published: (2024)
SuperValid: Capability-Aligned OOD Validation for Generalizable Downstream Scaling
by: Sun, Quanen, et al.
Published: (2026)
by: Sun, Quanen, et al.
Published: (2026)
Task-Adaptive Semantic Communications with Controllable Diffusion-based Data Regeneration
by: Guo, Fupei, et al.
Published: (2025)
by: Guo, Fupei, et al.
Published: (2025)
T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step
by: Chen, Zehui, et al.
Published: (2023)
by: Chen, Zehui, et al.
Published: (2023)
Mousse: Rectifying the Geometry of Muon with Curvature-Aware Preconditioning
by: Zhang, Yechen, et al.
Published: (2026)
by: Zhang, Yechen, et al.
Published: (2026)
Towards Effective and Efficient Graph Alignment without Supervision
by: Chen, Songyang, et al.
Published: (2026)
by: Chen, Songyang, et al.
Published: (2026)
Echotune: A Modular Extractor Leveraging the Variable-Length Nature of Speech in ASR Tasks
by: Chen, Sizhou, et al.
Published: (2023)
by: Chen, Sizhou, et al.
Published: (2023)
Pre-Trained Policy Discriminators are General Reward Models
by: Dou, Shihan, et al.
Published: (2025)
by: Dou, Shihan, et al.
Published: (2025)
CombAlign: Enhancing Model Expressiveness in Unsupervised Graph Alignment
by: Chen, Songyang, et al.
Published: (2024)
by: Chen, Songyang, et al.
Published: (2024)
How Real is Your Jailbreak? Fine-grained Jailbreak Evaluation with Anchored Reference
by: Liu, Songyang, et al.
Published: (2026)
by: Liu, Songyang, et al.
Published: (2026)
HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models
by: Que, Haoran, et al.
Published: (2024)
by: Que, Haoran, et al.
Published: (2024)
Inverse-Q*: Token Level Reinforcement Learning for Aligning Large Language Models Without Preference Data
by: Xia, Han, et al.
Published: (2024)
by: Xia, Han, et al.
Published: (2024)
Rectifying LLM Thought from Lens of Optimization
by: Liu, Junnan, et al.
Published: (2025)
by: Liu, Junnan, et al.
Published: (2025)
FoundaBench: Evaluating Chinese Fundamental Knowledge Capabilities of Large Language Models
by: Li, Wei, et al.
Published: (2024)
by: Li, Wei, et al.
Published: (2024)
TACO: Rethinking Semantic Communications with Task Adaptation and Context Embedding
by: Wijesinghe, Achintha, et al.
Published: (2025)
by: Wijesinghe, Achintha, et al.
Published: (2025)
Cross: A Delay Based Congestion Control Method for RTP Media
by: Zhang, Songyang, et al.
Published: (2024)
by: Zhang, Songyang, et al.
Published: (2024)
TiRE-GAN: Task-Incentivized Generative Learning for Radiomap Estimation
by: Zhou, Yueling, et al.
Published: (2024)
by: Zhou, Yueling, et al.
Published: (2024)
LLaST: Improved End-to-end Speech Translation System Leveraged by Large Language Models
by: Chen, Xi, et al.
Published: (2024)
by: Chen, Xi, et al.
Published: (2024)
CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards
by: Zhang, Taolin, et al.
Published: (2025)
by: Zhang, Taolin, et al.
Published: (2025)
TS-SAM: Fine-Tuning Segment-Anything Model for Downstream Tasks
by: Yu, Yang, et al.
Published: (2024)
by: Yu, Yang, et al.
Published: (2024)
Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement
by: Cao, Maosong, et al.
Published: (2025)
by: Cao, Maosong, et al.
Published: (2025)
ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios
by: Ye, Junjie, et al.
Published: (2024)
by: Ye, Junjie, et al.
Published: (2024)
Topologically protected emergent Fermi surface in an Abrikosov vortex lattice
by: Pu, Songyang, et al.
Published: (2024)
by: Pu, Songyang, et al.
Published: (2024)
From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language Models
by: Li, Rongjie, et al.
Published: (2024)
by: Li, Rongjie, et al.
Published: (2024)
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
by: Wang, Chonghua, et al.
Published: (2024)
by: Wang, Chonghua, et al.
Published: (2024)
Task-Driven Semantic Quantization and Imitation Learning for Goal-Oriented Communications
by: Chao, Yu-Chieh, et al.
Published: (2025)
by: Chao, Yu-Chieh, et al.
Published: (2025)
Exploring the MBTI distribution among Chinese undergraduate physics students: the influence of family income on career trajectories
by: Bai, Songyang, et al.
Published: (2024)
by: Bai, Songyang, et al.
Published: (2024)
DualGFL: Federated Learning with a Dual-Level Coalition-Auction Game
by: Chen, Xiaobing, et al.
Published: (2024)
by: Chen, Xiaobing, et al.
Published: (2024)
Scaling Laws Are Unreliable for Downstream Tasks: A Reality Check
by: Lourie, Nicholas, et al.
Published: (2025)
by: Lourie, Nicholas, et al.
Published: (2025)
Scaling Laws for Downstream Task Performance of Large Language Models
by: Isik, Berivan, et al.
Published: (2024)
by: Isik, Berivan, et al.
Published: (2024)
SGTR+: End-to-end Scene Graph Generation with Transformer
by: Li, Rongjie, et al.
Published: (2024)
by: Li, Rongjie, et al.
Published: (2024)
Robust Saliency-Aware Distillation for Few-shot Fine-grained Visual Recognition
by: Liu, Haiqi, et al.
Published: (2023)
by: Liu, Haiqi, et al.
Published: (2023)
AlchemistCoder: Harmonizing and Eliciting Code Capability by Hindsight Tuning on Multi-source Data
by: Song, Zifan, et al.
Published: (2024)
by: Song, Zifan, et al.
Published: (2024)
Saliency-R1: Incentivizing Unified Saliency Reasoning Capability in MLLM with Confidence-Guided Reinforcement Learning
by: Li, Long, et al.
Published: (2025)
by: Li, Long, et al.
Published: (2025)
Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive?
by: Schaeffer, Rylan, et al.
Published: (2024)
by: Schaeffer, Rylan, et al.
Published: (2024)
Scaling Laws for Predicting Downstream Performance in LLMs
by: Chen, Yangyi, et al.
Published: (2024)
by: Chen, Yangyi, et al.
Published: (2024)
Similar Items
-
Are Your LLMs Capable of Stable Reasoning?
by: Liu, Junnan, et al.
Published: (2024) -
Linear Alignment: A Closed-form Solution for Aligning Human Preferences without Tuning and Feedback
by: Gao, Songyang, et al.
Published: (2024) -
Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs
by: Qiao, Yuxuan, et al.
Published: (2024) -
How to Set the Learning Rate for Large-Scale Pre-training?
by: Zhou, Yunhua, et al.
Published: (2026) -
The Fine Line: Navigating Large Language Model Pretraining with Down-streaming Capability Analysis
by: Yang, Chen, et al.
Published: (2024)