Beyond Fixed Length: Bucket Pre-training is All You Need
Fuente:
arXiv
Salvato in:
| Autori principali: | Yang, Qing, Peng, Qiyao, Liu, Hongtao, Liu, Kai, Qin, Bing, Liu, Ting |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Length Extrapolation of Transformers: A Survey from the Perspective of Positional Encoding
di: Zhao, Liang, et al.
Pubblicazione: (2023)
di: Zhao, Liang, et al.
Pubblicazione: (2023)
Review-LLM: Harnessing Large Language Models for Personalized Review Generation
di: Peng, Qiyao, et al.
Pubblicazione: (2024)
di: Peng, Qiyao, et al.
Pubblicazione: (2024)
Tensor Product Attention Is All You Need
di: Zhang, Yifan, et al.
Pubblicazione: (2025)
di: Zhang, Yifan, et al.
Pubblicazione: (2025)
Meaningful Learning: Enhancing Abstract Reasoning in Large Language Models via Generic Fact Guidance
di: Xiong, Kai, et al.
Pubblicazione: (2024)
di: Xiong, Kai, et al.
Pubblicazione: (2024)
Is Translation All You Need? A Study on Solving Multilingual Tasks with Large Language Models
di: Liu, Chaoqun, et al.
Pubblicazione: (2024)
di: Liu, Chaoqun, et al.
Pubblicazione: (2024)
Not All Tokens Are What You Need In Thinking
di: Yuan, Hang, et al.
Pubblicazione: (2025)
di: Yuan, Hang, et al.
Pubblicazione: (2025)
Predictable Emergent Abilities of LLMs: Proxy Tasks Are All You Need
di: Zhang, Bo-Wen, et al.
Pubblicazione: (2024)
di: Zhang, Bo-Wen, et al.
Pubblicazione: (2024)
Answer is All You Need: Instruction-following Text Embedding via Answering the Question
di: Peng, Letian, et al.
Pubblicazione: (2024)
di: Peng, Letian, et al.
Pubblicazione: (2024)
Attention Smoothing Is All You Need For Unlearning
di: Zade, Saleh Zare, et al.
Pubblicazione: (2026)
di: Zade, Saleh Zare, et al.
Pubblicazione: (2026)
More Agents Is All You Need
di: Li, Junyou, et al.
Pubblicazione: (2024)
di: Li, Junyou, et al.
Pubblicazione: (2024)
How Does Sequence Modeling Architecture Influence Base Capabilities of Pre-trained Language Models? Exploring Key Architecture Design Principles to Avoid Base Capabilities Degradation
di: Lu, Xin, et al.
Pubblicazione: (2025)
di: Lu, Xin, et al.
Pubblicazione: (2025)
Rho-1: Not All Tokens Are What You Need
di: Lin, Zhenghao, et al.
Pubblicazione: (2024)
di: Lin, Zhenghao, et al.
Pubblicazione: (2024)
Selected Languages are All You Need for Cross-lingual Truthfulness Transfer
di: Liu, Weihao, et al.
Pubblicazione: (2024)
di: Liu, Weihao, et al.
Pubblicazione: (2024)
Training on the Benchmark Is Not All You Need
di: Ni, Shiwen, et al.
Pubblicazione: (2024)
di: Ni, Shiwen, et al.
Pubblicazione: (2024)
Continual Pre-Training is (not) What You Need in Domain Adaption
di: Chen, Pin-Er, et al.
Pubblicazione: (2025)
di: Chen, Pin-Er, et al.
Pubblicazione: (2025)
Contrast Is All You Need
di: Kilic, Burak, et al.
Pubblicazione: (2023)
di: Kilic, Burak, et al.
Pubblicazione: (2023)
Attention Is All You Need But You Don't Need All Of It For Inference of Large Language Models
di: Tyukin, Georgy, et al.
Pubblicazione: (2024)
di: Tyukin, Georgy, et al.
Pubblicazione: (2024)
Not All Documents Are What You Need for Extracting Instruction Tuning Data
di: Zhang, Chi, et al.
Pubblicazione: (2025)
di: Zhang, Chi, et al.
Pubblicazione: (2025)
Probing Language Models for Pre-training Data Detection
di: Liu, Zhenhua, et al.
Pubblicazione: (2024)
di: Liu, Zhenhua, et al.
Pubblicazione: (2024)
Increasing the Thinking Budget is Not All You Need
di: Iacobacci, Ignacio, et al.
Pubblicazione: (2025)
di: Iacobacci, Ignacio, et al.
Pubblicazione: (2025)
All You Need is One: Capsule Prompt Tuning with a Single Vector
di: Liu, Yiyang, et al.
Pubblicazione: (2025)
di: Liu, Yiyang, et al.
Pubblicazione: (2025)
How does Architecture Influence the Base Capabilities of Pre-trained Language Models? A Case Study Based on FFN-Wider and MoE Transformers
di: Lu, Xin, et al.
Pubblicazione: (2024)
di: Lu, Xin, et al.
Pubblicazione: (2024)
Rethinking Data Selection at Scale: Random Selection is Almost All You Need
di: Xia, Tingyu, et al.
Pubblicazione: (2024)
di: Xia, Tingyu, et al.
Pubblicazione: (2024)
Blending Is All You Need: Cheaper, Better Alternative to Trillion-Parameters LLM
di: Lu, Xiaoding, et al.
Pubblicazione: (2024)
di: Lu, Xiaoding, et al.
Pubblicazione: (2024)
Don't Do RAG: When Cache-Augmented Generation is All You Need for Knowledge Tasks
di: Chan, Brian J, et al.
Pubblicazione: (2024)
di: Chan, Brian J, et al.
Pubblicazione: (2024)
SecEncoder: Logs are All You Need in Security
di: Bulut, Muhammed Fatih, et al.
Pubblicazione: (2024)
di: Bulut, Muhammed Fatih, et al.
Pubblicazione: (2024)
Extending Context Window of Large Language Models from a Distributional Perspective
di: Wu, Yingsheng, et al.
Pubblicazione: (2024)
di: Wu, Yingsheng, et al.
Pubblicazione: (2024)
Agents Are All You Need for LLM Unlearning
di: Sanyal, Debdeep, et al.
Pubblicazione: (2025)
di: Sanyal, Debdeep, et al.
Pubblicazione: (2025)
Diffusion LMs Can Approximate Optimal Infilling Lengths Implicitly
di: Liu, Hengchang, et al.
Pubblicazione: (2026)
di: Liu, Hengchang, et al.
Pubblicazione: (2026)
Beyond Similarity: A Gradient-based Graph Method for Instruction Tuning Data Selection
di: Zhao, Yang, et al.
Pubblicazione: (2025)
di: Zhao, Yang, et al.
Pubblicazione: (2025)
Grimoire is All You Need for Enhancing Large Language Models
di: Chen, Ding, et al.
Pubblicazione: (2024)
di: Chen, Ding, et al.
Pubblicazione: (2024)
Addition is All You Need for Energy-efficient Language Models
di: Luo, Hongyin, et al.
Pubblicazione: (2024)
di: Luo, Hongyin, et al.
Pubblicazione: (2024)
Greed is All You Need: An Evaluation of Tokenizer Inference Methods
di: Uzan, Omri, et al.
Pubblicazione: (2024)
di: Uzan, Omri, et al.
Pubblicazione: (2024)
English is Not All You Need: Systematically Exploring the Role of Multilinguality in LLM Post-Training
di: Dhaliwal, Mehak, et al.
Pubblicazione: (2026)
di: Dhaliwal, Mehak, et al.
Pubblicazione: (2026)
Verbalized Algorithms: Classical Algorithms are All You Need (Mostly)
di: Lall, Supriya, et al.
Pubblicazione: (2025)
di: Lall, Supriya, et al.
Pubblicazione: (2025)
Is Depth All You Need? An Exploration of Iterative Reasoning in LLMs
di: Wu, Zongqian, et al.
Pubblicazione: (2025)
di: Wu, Zongqian, et al.
Pubblicazione: (2025)
Pipeline Parallelism is All You Need for Optimized Early-Exit Based Self-Speculative Decoding
di: Li, Ruanjun, et al.
Pubblicazione: (2025)
di: Li, Ruanjun, et al.
Pubblicazione: (2025)
On Predicting the Post-training Potential of Pre-trained LLMs
di: Li, Xiaoyuan, et al.
Pubblicazione: (2026)
di: Li, Xiaoyuan, et al.
Pubblicazione: (2026)
The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Models
di: Ji, Ke, et al.
Pubblicazione: (2025)
di: Ji, Ke, et al.
Pubblicazione: (2025)
COIG-CQIA: Quality is All You Need for Chinese Instruction Fine-tuning
di: Bai, Yuelin, et al.
Pubblicazione: (2024)
di: Bai, Yuelin, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Length Extrapolation of Transformers: A Survey from the Perspective of Positional Encoding
di: Zhao, Liang, et al.
Pubblicazione: (2023) -
Review-LLM: Harnessing Large Language Models for Personalized Review Generation
di: Peng, Qiyao, et al.
Pubblicazione: (2024) -
Tensor Product Attention Is All You Need
di: Zhang, Yifan, et al.
Pubblicazione: (2025) -
Meaningful Learning: Enhancing Abstract Reasoning in Large Language Models via Generic Fact Guidance
di: Xiong, Kai, et al.
Pubblicazione: (2024) -
Is Translation All You Need? A Study on Solving Multilingual Tasks with Large Language Models
di: Liu, Chaoqun, et al.
Pubblicazione: (2024)