APOLLO: An Optimized Training Approach for Long-form Numerical Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sun, Jiashuo, Zhang, Hang, Lin, Chen, Su, Xiangdong, Gong, Yeyun, Guo, Jian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2022
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Optimizing Large Language Model Training Using FP4 Quantization
von: Wang, Ruizhe, et al.
Veröffentlicht: (2025)
von: Wang, Ruizhe, et al.
Veröffentlicht: (2025)
Training-Inference Consistent Segmented Execution for Long-Context LLMs
von: Shang, Xianpeng, et al.
Veröffentlicht: (2026)
von: Shang, Xianpeng, et al.
Veröffentlicht: (2026)
Ensuring Safe and High-Quality Outputs: A Guideline Library Approach for Language Models
von: Luo, Yi, et al.
Veröffentlicht: (2024)
von: Luo, Yi, et al.
Veröffentlicht: (2024)
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
von: Liang, Xiao, et al.
Veröffentlicht: (2025)
von: Liang, Xiao, et al.
Veröffentlicht: (2025)
DynamixSFT: Dynamic Mixture Optimization of Instruction Tuning Collections
von: Shin, Haebin, et al.
Veröffentlicht: (2025)
von: Shin, Haebin, et al.
Veröffentlicht: (2025)
LayerNorm Induces Recency Bias in Transformer Decoders
von: Kim, Junu, et al.
Veröffentlicht: (2025)
von: Kim, Junu, et al.
Veröffentlicht: (2025)
360-LLaMA-Factory: Plug & Play Sequence Parallelism for Long Post-Training
von: Zou, Haosheng, et al.
Veröffentlicht: (2025)
von: Zou, Haosheng, et al.
Veröffentlicht: (2025)
IAPO: Information-Aware Policy Optimization for Token-Efficient Reasoning
von: He, Yinhan, et al.
Veröffentlicht: (2026)
von: He, Yinhan, et al.
Veröffentlicht: (2026)
Reasoning Bias of Next Token Prediction Training
von: Lin, Pengxiao, et al.
Veröffentlicht: (2025)
von: Lin, Pengxiao, et al.
Veröffentlicht: (2025)
Enhancing Chain-of-Thoughts Prompting with Iterative Bootstrapping in Large Language Models
von: Sun, Jiashuo, et al.
Veröffentlicht: (2023)
von: Sun, Jiashuo, et al.
Veröffentlicht: (2023)
LongFlow: Efficient KV Cache Compression for Reasoning Models
von: Su, Yi, et al.
Veröffentlicht: (2026)
von: Su, Yi, et al.
Veröffentlicht: (2026)
Think-on-Graph: Deep and Responsible Reasoning of Large Language Model on Knowledge Graph
von: Sun, Jiashuo, et al.
Veröffentlicht: (2023)
von: Sun, Jiashuo, et al.
Veröffentlicht: (2023)
LazyEviction: Lagged KV Eviction with Attention Pattern Observation for Efficient Long Reasoning
von: Zhang, Haoyue, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyue, et al.
Veröffentlicht: (2025)
LongLaMP: A Benchmark for Personalized Long-form Text Generation
von: Kumar, Ishita, et al.
Veröffentlicht: (2024)
von: Kumar, Ishita, et al.
Veröffentlicht: (2024)
ForesightKV: Optimizing KV Cache Eviction for Reasoning Models by Learning Long-Term Contribution
von: Dong, Zican, et al.
Veröffentlicht: (2026)
von: Dong, Zican, et al.
Veröffentlicht: (2026)
SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
von: Wang, Hanlin, et al.
Veröffentlicht: (2025)
von: Wang, Hanlin, et al.
Veröffentlicht: (2025)
Training Large Reasoning Models Efficiently via Progressive Thought Encoding
von: Zhang, Zeliang, et al.
Veröffentlicht: (2026)
von: Zhang, Zeliang, et al.
Veröffentlicht: (2026)
CoG: Controllable Graph Reasoning via Relational Blueprints and Failure-Aware Refinement over Knowledge Graphs
von: Liu, Yuanxiang, et al.
Veröffentlicht: (2026)
von: Liu, Yuanxiang, et al.
Veröffentlicht: (2026)
Evaluating Numerical Reasoning in Text-to-Image Models
von: Kajić, Ivana, et al.
Veröffentlicht: (2024)
von: Kajić, Ivana, et al.
Veröffentlicht: (2024)
Decoder-Hybrid-Decoder Architecture for Efficient Reasoning with Long Generation
von: Ren, Liliang, et al.
Veröffentlicht: (2025)
von: Ren, Liliang, et al.
Veröffentlicht: (2025)
Mitigating Heterogeneity among Factor Tensors via Lie Group Manifolds for Tensor Decomposition Based Temporal Knowledge Graph Embedding
von: Li, Jiang, et al.
Veröffentlicht: (2024)
von: Li, Jiang, et al.
Veröffentlicht: (2024)
LLM Reasoning as Trajectories: Step-Specific Representation Geometry and Correctness Signals
von: Sun, Lihao, et al.
Veröffentlicht: (2026)
von: Sun, Lihao, et al.
Veröffentlicht: (2026)
GRAM-R$^2$: Self-Training Generative Foundation Reward Models for Reward Reasoning
von: Wang, Chenglong, et al.
Veröffentlicht: (2025)
von: Wang, Chenglong, et al.
Veröffentlicht: (2025)
Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs
von: Zhang, Xuan, et al.
Veröffentlicht: (2024)
von: Zhang, Xuan, et al.
Veröffentlicht: (2024)
Train Long, Think Short: Curriculum Learning for Efficient Reasoning
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2025)
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2025)
Beyond Outliers: A Data-Free Layer-wise Mixed-Precision Quantization Approach Driven by Numerical and Structural Dual-Sensitivity
von: Zhang, Hengyuan, et al.
Veröffentlicht: (2026)
von: Zhang, Hengyuan, et al.
Veröffentlicht: (2026)
PromptCoT 2.0: Scaling Prompt Synthesis for Large Language Model Reasoning
von: Zhao, Xueliang, et al.
Veröffentlicht: (2025)
von: Zhao, Xueliang, et al.
Veröffentlicht: (2025)
Evolving LLMs' Self-Refinement Capability via Synergistic Training-Inference Optimization
von: Zeng, Yongcheng, et al.
Veröffentlicht: (2025)
von: Zeng, Yongcheng, et al.
Veröffentlicht: (2025)
Distribution-Aligned Sequence Distillation for Superior Long-CoT Reasoning
von: Yan, Shaotian, et al.
Veröffentlicht: (2026)
von: Yan, Shaotian, et al.
Veröffentlicht: (2026)
How to Train Long-Context Language Models (Effectively)
von: Gao, Tianyu, et al.
Veröffentlicht: (2024)
von: Gao, Tianyu, et al.
Veröffentlicht: (2024)
Genius: A Generalizable and Purely Unsupervised Self-Training Framework For Advanced Reasoning
von: Xu, Fangzhi, et al.
Veröffentlicht: (2025)
von: Xu, Fangzhi, et al.
Veröffentlicht: (2025)
Dynamic Thinking-Token Selection for Efficient Reasoning in Large Reasoning Models
von: Guo, Zhenyuan, et al.
Veröffentlicht: (2026)
von: Guo, Zhenyuan, et al.
Veröffentlicht: (2026)
Optimizing Anytime Reasoning via Budget Relative Policy Optimization
von: Qi, Penghui, et al.
Veröffentlicht: (2025)
von: Qi, Penghui, et al.
Veröffentlicht: (2025)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
von: Xiao, Chaojun, et al.
Veröffentlicht: (2024)
von: Xiao, Chaojun, et al.
Veröffentlicht: (2024)
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
von: Lai, Xin, et al.
Veröffentlicht: (2024)
von: Lai, Xin, et al.
Veröffentlicht: (2024)
SparseEval: Efficient Evaluation of Large Language Models by Sparse Optimization
von: Zhang, Taolin, et al.
Veröffentlicht: (2026)
von: Zhang, Taolin, et al.
Veröffentlicht: (2026)
Breaking MLPerf Training: A Case Study on Optimizing BERT
von: Kim, Yongdeok, et al.
Veröffentlicht: (2024)
von: Kim, Yongdeok, et al.
Veröffentlicht: (2024)
A Semantic-based Optimization Approach for Repairing LLMs: Case Study on Code Generation
von: Gu, Jian, et al.
Veröffentlicht: (2025)
von: Gu, Jian, et al.
Veröffentlicht: (2025)
SeMe: Training-Free Language Model Merging via Semantic Alignment
von: Gu, Jian, et al.
Veröffentlicht: (2025)
von: Gu, Jian, et al.
Veröffentlicht: (2025)
Rethinking Weight Tying: Pseudo-Inverse Tying for LM Stable Training and Updates
von: Gu, Jian, et al.
Veröffentlicht: (2026)
von: Gu, Jian, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Optimizing Large Language Model Training Using FP4 Quantization
von: Wang, Ruizhe, et al.
Veröffentlicht: (2025) -
Training-Inference Consistent Segmented Execution for Long-Context LLMs
von: Shang, Xianpeng, et al.
Veröffentlicht: (2026) -
Ensuring Safe and High-Quality Outputs: A Guideline Library Approach for Language Models
von: Luo, Yi, et al.
Veröffentlicht: (2024) -
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
von: Liang, Xiao, et al.
Veröffentlicht: (2025) -
DynamixSFT: Dynamic Mixture Optimization of Instruction Tuning Collections
von: Shin, Haebin, et al.
Veröffentlicht: (2025)