Beyond Fixed Length: Bucket Pre-training is All You Need
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Qing, Peng, Qiyao, Liu, Hongtao, Liu, Kai, Qin, Bing, Liu, Ting |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Length Extrapolation of Transformers: A Survey from the Perspective of Positional Encoding
von: Zhao, Liang, et al.
Veröffentlicht: (2023)
von: Zhao, Liang, et al.
Veröffentlicht: (2023)
Review-LLM: Harnessing Large Language Models for Personalized Review Generation
von: Peng, Qiyao, et al.
Veröffentlicht: (2024)
von: Peng, Qiyao, et al.
Veröffentlicht: (2024)
Tensor Product Attention Is All You Need
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
Meaningful Learning: Enhancing Abstract Reasoning in Large Language Models via Generic Fact Guidance
von: Xiong, Kai, et al.
Veröffentlicht: (2024)
von: Xiong, Kai, et al.
Veröffentlicht: (2024)
Is Translation All You Need? A Study on Solving Multilingual Tasks with Large Language Models
von: Liu, Chaoqun, et al.
Veröffentlicht: (2024)
von: Liu, Chaoqun, et al.
Veröffentlicht: (2024)
Not All Tokens Are What You Need In Thinking
von: Yuan, Hang, et al.
Veröffentlicht: (2025)
von: Yuan, Hang, et al.
Veröffentlicht: (2025)
Predictable Emergent Abilities of LLMs: Proxy Tasks Are All You Need
von: Zhang, Bo-Wen, et al.
Veröffentlicht: (2024)
von: Zhang, Bo-Wen, et al.
Veröffentlicht: (2024)
Answer is All You Need: Instruction-following Text Embedding via Answering the Question
von: Peng, Letian, et al.
Veröffentlicht: (2024)
von: Peng, Letian, et al.
Veröffentlicht: (2024)
Attention Smoothing Is All You Need For Unlearning
von: Zade, Saleh Zare, et al.
Veröffentlicht: (2026)
von: Zade, Saleh Zare, et al.
Veröffentlicht: (2026)
More Agents Is All You Need
von: Li, Junyou, et al.
Veröffentlicht: (2024)
von: Li, Junyou, et al.
Veröffentlicht: (2024)
How Does Sequence Modeling Architecture Influence Base Capabilities of Pre-trained Language Models? Exploring Key Architecture Design Principles to Avoid Base Capabilities Degradation
von: Lu, Xin, et al.
Veröffentlicht: (2025)
von: Lu, Xin, et al.
Veröffentlicht: (2025)
Rho-1: Not All Tokens Are What You Need
von: Lin, Zhenghao, et al.
Veröffentlicht: (2024)
von: Lin, Zhenghao, et al.
Veröffentlicht: (2024)
Selected Languages are All You Need for Cross-lingual Truthfulness Transfer
von: Liu, Weihao, et al.
Veröffentlicht: (2024)
von: Liu, Weihao, et al.
Veröffentlicht: (2024)
Training on the Benchmark Is Not All You Need
von: Ni, Shiwen, et al.
Veröffentlicht: (2024)
von: Ni, Shiwen, et al.
Veröffentlicht: (2024)
Continual Pre-Training is (not) What You Need in Domain Adaption
von: Chen, Pin-Er, et al.
Veröffentlicht: (2025)
von: Chen, Pin-Er, et al.
Veröffentlicht: (2025)
Contrast Is All You Need
von: Kilic, Burak, et al.
Veröffentlicht: (2023)
von: Kilic, Burak, et al.
Veröffentlicht: (2023)
Attention Is All You Need But You Don't Need All Of It For Inference of Large Language Models
von: Tyukin, Georgy, et al.
Veröffentlicht: (2024)
von: Tyukin, Georgy, et al.
Veröffentlicht: (2024)
Not All Documents Are What You Need for Extracting Instruction Tuning Data
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
Probing Language Models for Pre-training Data Detection
von: Liu, Zhenhua, et al.
Veröffentlicht: (2024)
von: Liu, Zhenhua, et al.
Veröffentlicht: (2024)
Increasing the Thinking Budget is Not All You Need
von: Iacobacci, Ignacio, et al.
Veröffentlicht: (2025)
von: Iacobacci, Ignacio, et al.
Veröffentlicht: (2025)
All You Need is One: Capsule Prompt Tuning with a Single Vector
von: Liu, Yiyang, et al.
Veröffentlicht: (2025)
von: Liu, Yiyang, et al.
Veröffentlicht: (2025)
How does Architecture Influence the Base Capabilities of Pre-trained Language Models? A Case Study Based on FFN-Wider and MoE Transformers
von: Lu, Xin, et al.
Veröffentlicht: (2024)
von: Lu, Xin, et al.
Veröffentlicht: (2024)
Rethinking Data Selection at Scale: Random Selection is Almost All You Need
von: Xia, Tingyu, et al.
Veröffentlicht: (2024)
von: Xia, Tingyu, et al.
Veröffentlicht: (2024)
Blending Is All You Need: Cheaper, Better Alternative to Trillion-Parameters LLM
von: Lu, Xiaoding, et al.
Veröffentlicht: (2024)
von: Lu, Xiaoding, et al.
Veröffentlicht: (2024)
Don't Do RAG: When Cache-Augmented Generation is All You Need for Knowledge Tasks
von: Chan, Brian J, et al.
Veröffentlicht: (2024)
von: Chan, Brian J, et al.
Veröffentlicht: (2024)
SecEncoder: Logs are All You Need in Security
von: Bulut, Muhammed Fatih, et al.
Veröffentlicht: (2024)
von: Bulut, Muhammed Fatih, et al.
Veröffentlicht: (2024)
Extending Context Window of Large Language Models from a Distributional Perspective
von: Wu, Yingsheng, et al.
Veröffentlicht: (2024)
von: Wu, Yingsheng, et al.
Veröffentlicht: (2024)
Agents Are All You Need for LLM Unlearning
von: Sanyal, Debdeep, et al.
Veröffentlicht: (2025)
von: Sanyal, Debdeep, et al.
Veröffentlicht: (2025)
Diffusion LMs Can Approximate Optimal Infilling Lengths Implicitly
von: Liu, Hengchang, et al.
Veröffentlicht: (2026)
von: Liu, Hengchang, et al.
Veröffentlicht: (2026)
Beyond Similarity: A Gradient-based Graph Method for Instruction Tuning Data Selection
von: Zhao, Yang, et al.
Veröffentlicht: (2025)
von: Zhao, Yang, et al.
Veröffentlicht: (2025)
Grimoire is All You Need for Enhancing Large Language Models
von: Chen, Ding, et al.
Veröffentlicht: (2024)
von: Chen, Ding, et al.
Veröffentlicht: (2024)
Addition is All You Need for Energy-efficient Language Models
von: Luo, Hongyin, et al.
Veröffentlicht: (2024)
von: Luo, Hongyin, et al.
Veröffentlicht: (2024)
Greed is All You Need: An Evaluation of Tokenizer Inference Methods
von: Uzan, Omri, et al.
Veröffentlicht: (2024)
von: Uzan, Omri, et al.
Veröffentlicht: (2024)
English is Not All You Need: Systematically Exploring the Role of Multilinguality in LLM Post-Training
von: Dhaliwal, Mehak, et al.
Veröffentlicht: (2026)
von: Dhaliwal, Mehak, et al.
Veröffentlicht: (2026)
Verbalized Algorithms: Classical Algorithms are All You Need (Mostly)
von: Lall, Supriya, et al.
Veröffentlicht: (2025)
von: Lall, Supriya, et al.
Veröffentlicht: (2025)
Is Depth All You Need? An Exploration of Iterative Reasoning in LLMs
von: Wu, Zongqian, et al.
Veröffentlicht: (2025)
von: Wu, Zongqian, et al.
Veröffentlicht: (2025)
Pipeline Parallelism is All You Need for Optimized Early-Exit Based Self-Speculative Decoding
von: Li, Ruanjun, et al.
Veröffentlicht: (2025)
von: Li, Ruanjun, et al.
Veröffentlicht: (2025)
On Predicting the Post-training Potential of Pre-trained LLMs
von: Li, Xiaoyuan, et al.
Veröffentlicht: (2026)
von: Li, Xiaoyuan, et al.
Veröffentlicht: (2026)
The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Models
von: Ji, Ke, et al.
Veröffentlicht: (2025)
von: Ji, Ke, et al.
Veröffentlicht: (2025)
COIG-CQIA: Quality is All You Need for Chinese Instruction Fine-tuning
von: Bai, Yuelin, et al.
Veröffentlicht: (2024)
von: Bai, Yuelin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Length Extrapolation of Transformers: A Survey from the Perspective of Positional Encoding
von: Zhao, Liang, et al.
Veröffentlicht: (2023) -
Review-LLM: Harnessing Large Language Models for Personalized Review Generation
von: Peng, Qiyao, et al.
Veröffentlicht: (2024) -
Tensor Product Attention Is All You Need
von: Zhang, Yifan, et al.
Veröffentlicht: (2025) -
Meaningful Learning: Enhancing Abstract Reasoning in Large Language Models via Generic Fact Guidance
von: Xiong, Kai, et al.
Veröffentlicht: (2024) -
Is Translation All You Need? A Study on Solving Multilingual Tasks with Large Language Models
von: Liu, Chaoqun, et al.
Veröffentlicht: (2024)