Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zhaowei, Luo, Lishu, Duan, Haodong, Liu, Weiwei, Wu, Sijin, Luo, Ji, Yan, Shen, Peng, Shuai, Yuan, Sihang, Huang, Chaoyi, Lin, Yi, Song, Yangqiu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can Retrieval Heads See Images? Multimodal Retrieval Heads in Long-Context Vision-Language Models
by: Li, Aaron Branson Cigres, et al.
Published: (2026)
by: Li, Aaron Branson Cigres, et al.
Published: (2026)
MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly
by: Wang, Zhaowei, et al.
Published: (2025)
by: Wang, Zhaowei, et al.
Published: (2025)
From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models
by: Xu, Chejian, et al.
Published: (2025)
by: Xu, Chejian, et al.
Published: (2025)
Training Long-Context LLMs Efficiently via Chunk-wise Optimization
by: Li, Wenhao, et al.
Published: (2025)
by: Li, Wenhao, et al.
Published: (2025)
Data Engineering for Scaling Language Models to 128K Context
by: Fu, Yao, et al.
Published: (2024)
by: Fu, Yao, et al.
Published: (2024)
Scaling Granite Code Models to 128K Context
by: Stallone, Matt, et al.
Published: (2024)
by: Stallone, Matt, et al.
Published: (2024)
How to Train Long-Context Language Models (Effectively)
by: Gao, Tianyu, et al.
Published: (2024)
by: Gao, Tianyu, et al.
Published: (2024)
UltraLLaDA: Scaling the Context Length to 128K for Diffusion Large Language Models
by: He, Guangxin, et al.
Published: (2025)
by: He, Guangxin, et al.
Published: (2025)
EntropyLong: Effective Long-Context Training via Predictive Uncertainty
by: Jia, Junlong, et al.
Published: (2025)
by: Jia, Junlong, et al.
Published: (2025)
NExtLong: Toward Effective Long-Context Training without Long Documents
by: Gao, Chaochen, et al.
Published: (2025)
by: Gao, Chaochen, et al.
Published: (2025)
ContextPilot: Fast Long-Context Inference via Context Reuse
by: Jiang, Yinsicheng, et al.
Published: (2025)
by: Jiang, Yinsicheng, et al.
Published: (2025)
Long Context Pre-Training with Lighthouse Attention
by: Peng, Bowen, et al.
Published: (2026)
by: Peng, Bowen, et al.
Published: (2026)
SpecPV: Improving Self-Speculative Decoding for Long-Context Generation via Partial Verification
by: Tan, Zhendong, et al.
Published: (2025)
by: Tan, Zhendong, et al.
Published: (2025)
LongFly: Long-Horizon UAV Vision-and-Language Navigation with Spatiotemporal Context Integration
by: Jiang, Wen, et al.
Published: (2025)
by: Jiang, Wen, et al.
Published: (2025)
Detecting Legend Items on Historical Maps Using GPT-4o with In-Context Learning
by: Kirsanova, Sofia, et al.
Published: (2025)
by: Kirsanova, Sofia, et al.
Published: (2025)
KNOWCOMP POKEMON Team at DialAM-2024: A Two-Stage Pipeline for Detecting Relations in Dialogical Argument Mining
by: Zheng, Zihao, et al.
Published: (2024)
by: Zheng, Zihao, et al.
Published: (2024)
A Little Goes a Long Way: Efficient Long Context Training and Inference with Partial Contexts
by: Ge, Suyu, et al.
Published: (2024)
by: Ge, Suyu, et al.
Published: (2024)
Beyond Context Limits: Subconscious Threads for Long-Horizon Reasoning
by: Luo, Hongyin, et al.
Published: (2025)
by: Luo, Hongyin, et al.
Published: (2025)
Divide-then-Diagnose: Weaving Clinician-Inspired Contexts for Ultra-Long Capsule Endoscopy Videos
by: Liu, Bowen, et al.
Published: (2026)
by: Liu, Bowen, et al.
Published: (2026)
Global Context Compression with Interleaved Vision-Text Transformation
by: Jiao, Dian, et al.
Published: (2026)
by: Jiao, Dian, et al.
Published: (2026)
DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts
by: Gai, Jiading, et al.
Published: (2026)
by: Gai, Jiading, et al.
Published: (2026)
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
by: Zarch, Hossein Entezari, et al.
Published: (2025)
by: Zarch, Hossein Entezari, et al.
Published: (2025)
$\infty$Bench: Extending Long Context Evaluation Beyond 100K Tokens
by: Zhang, Xinrong, et al.
Published: (2024)
by: Zhang, Xinrong, et al.
Published: (2024)
LongAgent: Scaling Language Models to 128k Context through Multi-Agent Collaboration
by: Zhao, Jun, et al.
Published: (2024)
by: Zhao, Jun, et al.
Published: (2024)
Scaling Context, Not Parameters: Training a Compact 7B Language Model for Efficient Long-Context Processing
by: Wu, Chen, et al.
Published: (2025)
by: Wu, Chen, et al.
Published: (2025)
FltLM: An Intergrated Long-Context Large Language Model for Effective Context Filtering and Understanding
by: Deng, Jingyang, et al.
Published: (2024)
by: Deng, Jingyang, et al.
Published: (2024)
Beyond the Protocol: Unveiling Attack Vectors in the Model Context Protocol (MCP) Ecosystem
by: Song, Hao, et al.
Published: (2025)
by: Song, Hao, et al.
Published: (2025)
KVDrive: A Holistic Multi-Tier KV Cache Management System for Long-Context LLM Inference
by: Lin, Jian, et al.
Published: (2026)
by: Lin, Jian, et al.
Published: (2026)
Learning-to-Context Slope: Evaluating In-Context Learning Effectiveness Beyond Performance Illusions
by: Wang, Dingzriui, et al.
Published: (2025)
by: Wang, Dingzriui, et al.
Published: (2025)
Soft-NBCE: Entropy-Weighted Chunk Fusion for Long-Context
by: Ji, Shihao, et al.
Published: (2026)
by: Ji, Shihao, et al.
Published: (2026)
Gated Sparse Attention: Combining Computational Efficiency with Training Stability for Long-Context Language Models
by: Shen, Alfred, et al.
Published: (2026)
by: Shen, Alfred, et al.
Published: (2026)
Segmental Advantage Estimation: Enhancing PPO for Long-Context LLM Training
by: Gong, Xue, et al.
Published: (2026)
by: Gong, Xue, et al.
Published: (2026)
Training-free Context-adaptive Attention for Efficient Long Context Modeling
by: You, Zeng, et al.
Published: (2025)
by: You, Zeng, et al.
Published: (2025)
Counting-Stars: A Multi-evidence, Position-aware, and Scalable Benchmark for Evaluating Long-Context Large Language Models
by: Song, Mingyang, et al.
Published: (2024)
by: Song, Mingyang, et al.
Published: (2024)
LongAttnComp: Cross-Family Context Compression for Long-Context Reasoning
by: Ji, Mengmeng, et al.
Published: (2026)
by: Ji, Mengmeng, et al.
Published: (2026)
InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference
by: Pan, Xiurui, et al.
Published: (2024)
by: Pan, Xiurui, et al.
Published: (2024)
Beyond Needle(s) in the Embodied Haystack: Environment, Architecture, and Training Considerations for Long Context Reasoning
by: Kim, Bosung, et al.
Published: (2025)
by: Kim, Bosung, et al.
Published: (2025)
ContextLens: Modeling Imperfect Privacy and Safety Context for Legal Compliance
by: Li, Haoran, et al.
Published: (2026)
by: Li, Haoran, et al.
Published: (2026)
Coding Agents are Effective Long-Context Processors
by: Cao, Weili, et al.
Published: (2026)
by: Cao, Weili, et al.
Published: (2026)
SlimPipe: Memory-Thrifty and Efficient Pipeline Parallelism for Long-Context LLM Training
by: Li, Zhouyang, et al.
Published: (2025)
by: Li, Zhouyang, et al.
Published: (2025)
Similar Items
-
Can Retrieval Heads See Images? Multimodal Retrieval Heads in Long-Context Vision-Language Models
by: Li, Aaron Branson Cigres, et al.
Published: (2026) -
MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly
by: Wang, Zhaowei, et al.
Published: (2025) -
From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models
by: Xu, Chejian, et al.
Published: (2025) -
Training Long-Context LLMs Efficiently via Chunk-wise Optimization
by: Li, Wenhao, et al.
Published: (2025) -
Data Engineering for Scaling Language Models to 128K Context
by: Fu, Yao, et al.
Published: (2024)