Predicting Emergent Abilities with Infinite Resolution Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Shengding, Liu, Xin, Han, Xu, Zhang, Xinrong, He, Chaoqun, Zhao, Weilin, Lin, Yankai, Ding, Ning, Ou, Zebin, Zeng, Guoyang, Liu, Zhiyuan, Sun, Maosong |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unified View of Grokking, Double Descent and Emergent Abilities: A Perspective from Circuits Competition
by: Huang, Yufei, et al.
Published: (2024)
by: Huang, Yufei, et al.
Published: (2024)
Stuffed Mamba: Oversized States Lead to the Inability to Forget
by: Chen, Yingfa, et al.
Published: (2024)
by: Chen, Yingfa, et al.
Published: (2024)
Beyond the Turn-Based Game: Enabling Real-Time Conversations with Duplex Models
by: Zhang, Xinrong, et al.
Published: (2024)
by: Zhang, Xinrong, et al.
Published: (2024)
States Hidden in Hidden States: LLMs Emerge Discrete State Representations Implicitly
by: Chen, Junhao, et al.
Published: (2024)
by: Chen, Junhao, et al.
Published: (2024)
UltraEval: A Lightweight Platform for Flexible and Comprehensive Evaluation for LLMs
by: He, Chaoqun, et al.
Published: (2024)
by: He, Chaoqun, et al.
Published: (2024)
Representation Learning for Natural Language Processing
by: Liu, Zhiyuan, et al.
Published: (2020)
by: Liu, Zhiyuan, et al.
Published: (2020)
Seq1F1B: Efficient Sequence-Level Pipeline Parallelism for Large Language Model Training
by: Sun, Ao, et al.
Published: (2024)
by: Sun, Ao, et al.
Published: (2024)
A Multi-Power Law for Loss Curve Prediction Across Learning Rate Schedules
by: Luo, Kairong, et al.
Published: (2025)
by: Luo, Kairong, et al.
Published: (2025)
$\infty$Bench: Extending Long Context Evaluation Beyond 100K Tokens
by: Zhang, Xinrong, et al.
Published: (2024)
by: Zhang, Xinrong, et al.
Published: (2024)
MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
by: Hu, Shengding, et al.
Published: (2024)
by: Hu, Shengding, et al.
Published: (2024)
Densing Law of LLMs
by: Xiao, Chaojun, et al.
Published: (2024)
by: Xiao, Chaojun, et al.
Published: (2024)
OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
by: He, Chaoqun, et al.
Published: (2024)
by: He, Chaoqun, et al.
Published: (2024)
Ouroboros: Generating Longer Drafts Phrase by Phrase for Faster Speculative Decoding
by: Zhao, Weilin, et al.
Published: (2024)
by: Zhao, Weilin, et al.
Published: (2024)
LEGENT: Open Platform for Embodied Agents
by: Cheng, Zhili, et al.
Published: (2024)
by: Cheng, Zhili, et al.
Published: (2024)
BurstAttention: An Efficient Distributed Attention Framework for Extremely Long Sequences
by: Sun, Ao, et al.
Published: (2024)
by: Sun, Ao, et al.
Published: (2024)
ACDiT: Interpolating Autoregressive Conditional Modeling and Diffusion Transformer
by: Hu, Jinyi, et al.
Published: (2024)
by: Hu, Jinyi, et al.
Published: (2024)
BurstEngine: an Efficient Distributed Framework for Training Transformers on Extremely Long Sequences of over 1M Tokens
by: Sun, Ao, et al.
Published: (2025)
by: Sun, Ao, et al.
Published: (2025)
Rational Decision-Making Agent with Internalized Utility Judgment
by: Ye, Yining, et al.
Published: (2023)
by: Ye, Yining, et al.
Published: (2023)
Mastering Text, Code and Math Simultaneously via Fusing Highly Specialized Language Models
by: Ding, Ning, et al.
Published: (2024)
by: Ding, Ning, et al.
Published: (2024)
Exploring the Benefit of Activation Sparsity in Pre-training
by: Zhang, Zhengyan, et al.
Published: (2024)
by: Zhang, Zhengyan, et al.
Published: (2024)
DecorateLM: Data Engineering through Corpus Rating, Tagging, and Editing with Language Models
by: Zhao, Ranchi, et al.
Published: (2024)
by: Zhao, Ranchi, et al.
Published: (2024)
DebugBench: Evaluating Debugging Capability of Large Language Models
by: Tian, Runchu, et al.
Published: (2024)
by: Tian, Runchu, et al.
Published: (2024)
Learning to Focus: Causal Attention Distillation via Gradient-Guided Token Pruning
by: Guo, Yiju, et al.
Published: (2025)
by: Guo, Yiju, et al.
Published: (2025)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
by: Xiao, Chaojun, et al.
Published: (2024)
by: Xiao, Chaojun, et al.
Published: (2024)
ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models
by: Song, Chenyang, et al.
Published: (2024)
by: Song, Chenyang, et al.
Published: (2024)
Learning to Generate Structured Output with Schema Reinforcement Learning
by: Lu, Yaxi, et al.
Published: (2025)
by: Lu, Yaxi, et al.
Published: (2025)
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity
by: Song, Chenyang, et al.
Published: (2025)
by: Song, Chenyang, et al.
Published: (2025)
ToLeaP: Rethinking Development of Tool Learning with Large Language Models
by: Chen, Haotian, et al.
Published: (2025)
by: Chen, Haotian, et al.
Published: (2025)
Empowering Private Tutoring by Chaining Large Language Models
by: Chen, Yulin, et al.
Published: (2023)
by: Chen, Yulin, et al.
Published: (2023)
VULCAN: Vision-Language-Model Enhanced Multi-Agent Cooperative Navigation for Indoor Fire-Disaster Response
by: Liu, Shengding, et al.
Published: (2026)
by: Liu, Shengding, et al.
Published: (2026)
MiniCPM-V: A GPT-4V Level MLLM on Your Phone
by: Yao, Yuan, et al.
Published: (2024)
by: Yao, Yuan, et al.
Published: (2024)
UltraFeedback: Boosting Language Models with Scaled AI Feedback
by: Cui, Ganqu, et al.
Published: (2023)
by: Cui, Ganqu, et al.
Published: (2023)
Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment
by: Guo, Yiju, et al.
Published: (2024)
by: Guo, Yiju, et al.
Published: (2024)
Investigate-Consolidate-Exploit: A General Strategy for Inter-Task Agent Self-Evolution
by: Qian, Cheng, et al.
Published: (2024)
by: Qian, Cheng, et al.
Published: (2024)
The Overthinker's DIET: Cutting Token Calories with DIfficulty-AwarE Training
by: Chen, Weize, et al.
Published: (2025)
by: Chen, Weize, et al.
Published: (2025)
Matrix Fejér-Riesz type theorem for a union of an interval and a point
by: Sun, Shengding, et al.
Published: (2025)
by: Sun, Shengding, et al.
Published: (2025)
WorkflowLLM: Enhancing Workflow Orchestration Capability of Large Language Models
by: Fan, Shengda, et al.
Published: (2024)
by: Fan, Shengda, et al.
Published: (2024)
Predictable Emergent Abilities of LLMs: Proxy Tasks Are All You Need
by: Zhang, Bo-Wen, et al.
Published: (2024)
by: Zhang, Bo-Wen, et al.
Published: (2024)
CA-LoRA: Adapting Existing LoRA for Compressed LLMs to Enable Efficient Multi-Tasking on Personal Devices
by: Zhao, Weilin, et al.
Published: (2023)
by: Zhao, Weilin, et al.
Published: (2023)
INTERVENOR: Prompting the Coding Ability of Large Language Models with the Interactive Chain of Repair
by: Wang, Hanbin, et al.
Published: (2023)
by: Wang, Hanbin, et al.
Published: (2023)
Similar Items
-
Unified View of Grokking, Double Descent and Emergent Abilities: A Perspective from Circuits Competition
by: Huang, Yufei, et al.
Published: (2024) -
Stuffed Mamba: Oversized States Lead to the Inability to Forget
by: Chen, Yingfa, et al.
Published: (2024) -
Beyond the Turn-Based Game: Enabling Real-Time Conversations with Duplex Models
by: Zhang, Xinrong, et al.
Published: (2024) -
States Hidden in Hidden States: LLMs Emerge Discrete State Representations Implicitly
by: Chen, Junhao, et al.
Published: (2024) -
UltraEval: A Lightweight Platform for Flexible and Comprehensive Evaluation for LLMs
by: He, Chaoqun, et al.
Published: (2024)