Extending Input Contexts of Language Models through Training on Segmented Sequences
Fuente:
arXiv
Guardado en:
| Autores principales: | Karypis, Petros, McAuley, Julian, Karypis, George |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
GNN-RAG: Graph Neural Retrieval for Large Language Model Reasoning
por: Mavromatis, Costas, et al.
Publicado: (2024)
por: Mavromatis, Costas, et al.
Publicado: (2024)
Pack of LLMs: Model Fusion at Test-Time via Perplexity Optimization
por: Mavromatis, Costas, et al.
Publicado: (2024)
por: Mavromatis, Costas, et al.
Publicado: (2024)
SemPool: Simple, robust, and interpretable KG pooling for enhancing language models
por: Mavromatis, Costas, et al.
Publicado: (2024)
por: Mavromatis, Costas, et al.
Publicado: (2024)
Revisiting SMoE Language Models by Evaluating Inefficiencies with Task Specific Expert Pruning
por: Sarkar, Soumajyoti, et al.
Publicado: (2024)
por: Sarkar, Soumajyoti, et al.
Publicado: (2024)
BiasEdit: Debiasing Stereotyped Language Models via Model Editing
por: Xu, Xin, et al.
Publicado: (2025)
por: Xu, Xin, et al.
Publicado: (2025)
Fine-Tuning Language Models on Multiple Datasets for Citation Intention Classification
por: Shui, Zeren, et al.
Publicado: (2024)
por: Shui, Zeren, et al.
Publicado: (2024)
Beyond instruction-conditioning, MoTE: Mixture of Task Experts for Multi-task Embedding Models
por: Romero, Miguel, et al.
Publicado: (2025)
por: Romero, Miguel, et al.
Publicado: (2025)
Parameter-Efficient Tuning Large Language Models for Graph Representation Learning
por: Zhu, Qi, et al.
Publicado: (2024)
por: Zhu, Qi, et al.
Publicado: (2024)
Differentially Private Bias-Term Fine-tuning of Foundation Models
por: Bu, Zhiqi, et al.
Publicado: (2022)
por: Bu, Zhiqi, et al.
Publicado: (2022)
BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
por: Xu, Xin, et al.
Publicado: (2025)
por: Xu, Xin, et al.
Publicado: (2025)
MaxCode: A Max-Reward Reinforcement Learning Framework for Automated Code Optimization
por: Ou, Jiefu, et al.
Publicado: (2026)
por: Ou, Jiefu, et al.
Publicado: (2026)
Understanding Silent Data Corruption in LLM Training
por: Ma, Jeffrey, et al.
Publicado: (2025)
por: Ma, Jeffrey, et al.
Publicado: (2025)
FERA: Uncertainty-Aware Federated Reasoning for Large Language Models
por: Wang, Ruhan, et al.
Publicado: (2026)
por: Wang, Ruhan, et al.
Publicado: (2026)
Learning to Hint for Reinforcement Learning
por: Xia, Yu, et al.
Publicado: (2026)
por: Xia, Yu, et al.
Publicado: (2026)
ProxSparse: Regularized Learning of Semi-Structured Sparsity Masks for Pretrained LLMs
por: Liu, Hongyi, et al.
Publicado: (2025)
por: Liu, Hongyi, et al.
Publicado: (2025)
DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts
por: Gai, Jiading, et al.
Publicado: (2026)
por: Gai, Jiading, et al.
Publicado: (2026)
ThinkRouter: Efficient Reasoning via Routing Thinking between Latent and Discrete Spaces
por: Xu, Xin, et al.
Publicado: (2026)
por: Xu, Xin, et al.
Publicado: (2026)
OCEAN: Offline Chain-of-thought Evaluation and Alignment in Large Language Models
por: Wu, Junda, et al.
Publicado: (2024)
por: Wu, Junda, et al.
Publicado: (2024)
OPERA: Online Data Pruning for Efficient Retrieval Model Adaptation
por: Fang, Haoyang, et al.
Publicado: (2026)
por: Fang, Haoyang, et al.
Publicado: (2026)
When Benchmarks Age: Temporal Misalignment through Large Language Model Factuality Evaluation
por: Jiang, Xunyi, et al.
Publicado: (2025)
por: Jiang, Xunyi, et al.
Publicado: (2025)
How to Train Data-Efficient LLMs
por: Sachdeva, Noveen, et al.
Publicado: (2024)
por: Sachdeva, Noveen, et al.
Publicado: (2024)
Scalable Prompt Routing via Fine-Grained Latent Task Discovery
por: Zhang, Yunyi, et al.
Publicado: (2026)
por: Zhang, Yunyi, et al.
Publicado: (2026)
Federated Large Language Models: Current Progress and Future Directions
por: Yao, Yuhang, et al.
Publicado: (2024)
por: Yao, Yuhang, et al.
Publicado: (2024)
Learning to Generate Answers with Citations via Factual Consistency Models
por: Aly, Rami, et al.
Publicado: (2024)
por: Aly, Rami, et al.
Publicado: (2024)
CoMMIT: Coordinated Multimodal Instruction Tuning
por: Li, Xintong, et al.
Publicado: (2024)
por: Li, Xintong, et al.
Publicado: (2024)
Importance Sampling for Multi-Negative Multimodal Direct Preference Optimization
por: Li, Xintong, et al.
Publicado: (2025)
por: Li, Xintong, et al.
Publicado: (2025)
Self-Updatable Large Language Models by Integrating Context into Model Parameters
por: Wang, Yu, et al.
Publicado: (2024)
por: Wang, Yu, et al.
Publicado: (2024)
CAMELoT: Towards Large Language Models with Training-Free Consolidated Associative Memory
por: He, Zexue, et al.
Publicado: (2024)
por: He, Zexue, et al.
Publicado: (2024)
How to Train Long-Context Language Models (Effectively)
por: Gao, Tianyu, et al.
Publicado: (2024)
por: Gao, Tianyu, et al.
Publicado: (2024)
SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual Scenes
por: Wang, Chuhan, et al.
Publicado: (2026)
por: Wang, Chuhan, et al.
Publicado: (2026)
Training Language Models on Synthetic Edit Sequences Improves Code Synthesis
por: Piterbarg, Ulyana, et al.
Publicado: (2024)
por: Piterbarg, Ulyana, et al.
Publicado: (2024)
Sequence-Level Leakage Risk of Training Data in Large Language Models
por: Tiwari, Trishita, et al.
Publicado: (2024)
por: Tiwari, Trishita, et al.
Publicado: (2024)
Beyond the Limits: A Survey of Techniques to Extend the Context Length in Large Language Models
por: Wang, Xindi, et al.
Publicado: (2024)
por: Wang, Xindi, et al.
Publicado: (2024)
Training-Inference Consistent Segmented Execution for Long-Context LLMs
por: Shang, Xianpeng, et al.
Publicado: (2026)
por: Shang, Xianpeng, et al.
Publicado: (2026)
Baby Scale: Investigating Models Trained on Individual Children's Language Input
por: Feng, Steven Y., et al.
Publicado: (2026)
por: Feng, Steven Y., et al.
Publicado: (2026)
CSyMR: Benchmarking Compositional Music Information Retrieval in Symbolic Music Reasoning
por: Wang, Boyang, et al.
Publicado: (2025)
por: Wang, Boyang, et al.
Publicado: (2025)
Automatic Pair Construction for Contrastive Post-training
por: Xu, Canwen, et al.
Publicado: (2023)
por: Xu, Canwen, et al.
Publicado: (2023)
LongEmbed: Extending Embedding Models for Long Context Retrieval
por: Zhu, Dawei, et al.
Publicado: (2024)
por: Zhu, Dawei, et al.
Publicado: (2024)
Scaling Context, Not Parameters: Training a Compact 7B Language Model for Efficient Long-Context Processing
por: Wu, Chen, et al.
Publicado: (2025)
por: Wu, Chen, et al.
Publicado: (2025)
Multimodal Chain-of-Thought Reasoning in Language Models
por: Zhang, Zhuosheng, et al.
Publicado: (2023)
por: Zhang, Zhuosheng, et al.
Publicado: (2023)
Ejemplares similares
-
GNN-RAG: Graph Neural Retrieval for Large Language Model Reasoning
por: Mavromatis, Costas, et al.
Publicado: (2024) -
Pack of LLMs: Model Fusion at Test-Time via Perplexity Optimization
por: Mavromatis, Costas, et al.
Publicado: (2024) -
SemPool: Simple, robust, and interpretable KG pooling for enhancing language models
por: Mavromatis, Costas, et al.
Publicado: (2024) -
Revisiting SMoE Language Models by Evaluating Inefficiencies with Task Specific Expert Pruning
por: Sarkar, Soumajyoti, et al.
Publicado: (2024) -
BiasEdit: Debiasing Stereotyped Language Models via Model Editing
por: Xu, Xin, et al.
Publicado: (2025)