Context-level Language Modeling by Learning Predictive Context Embeddings
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dai, Beiya, Liu, Yuliang, Xue, Daozheng, Song, Yunchong, Guo, Qipeng, Chen, Kai, Wang, Xinbing, Zhou, Bowen, Lin, Zhouhan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Next Concept Prediction in Discrete Latent Space Leads to Stronger Language Models
von: Liu, Yuliang, et al.
Veröffentlicht: (2026)
von: Liu, Yuliang, et al.
Veröffentlicht: (2026)
KaLM: Knowledge-aligned Autoregressive Language Modeling via Dual-view Knowledge Graph Contrastive Learning
von: Yu, Peng, et al.
Veröffentlicht: (2024)
von: Yu, Peng, et al.
Veröffentlicht: (2024)
Memory Decoder: A Pretrained, Plug-and-Play Memory for Large Language Models
von: Cao, Jiaqi, et al.
Veröffentlicht: (2025)
von: Cao, Jiaqi, et al.
Veröffentlicht: (2025)
MLP Memory: A Retriever-Pretrained Memory for Large Language Models
von: Wei, Rubin, et al.
Veröffentlicht: (2025)
von: Wei, Rubin, et al.
Veröffentlicht: (2025)
AWM: Accurate Weight-Matrix Fingerprint for Large Language Models
von: Zeng, Boyi, et al.
Veröffentlicht: (2025)
von: Zeng, Boyi, et al.
Veröffentlicht: (2025)
Graph Parsing Networks
von: Song, Yunchong, et al.
Veröffentlicht: (2024)
von: Song, Yunchong, et al.
Veröffentlicht: (2024)
GeoGalactica: A Scientific Large Language Model in Geoscience
von: Lin, Zhouhan, et al.
Veröffentlicht: (2023)
von: Lin, Zhouhan, et al.
Veröffentlicht: (2023)
Flow of Spans: Generalizing Language Models to Dynamic Span-Vocabulary via GFlowNets
von: Xue, Bo, et al.
Veröffentlicht: (2026)
von: Xue, Bo, et al.
Veröffentlicht: (2026)
Beyond Real: Imaginary Extension of Rotary Position Embeddings for Long-Context LLMs
von: Liu, Xiaoran, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoran, et al.
Veröffentlicht: (2025)
Critical Data Size of Language Models from a Grokking Perspective
von: Zhu, Xuekai, et al.
Veröffentlicht: (2024)
von: Zhu, Xuekai, et al.
Veröffentlicht: (2024)
Thus Spake Long-Context Large Language Model
von: Liu, Xiaoran, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoran, et al.
Veröffentlicht: (2025)
Identifying Semantic Induction Heads to Understand In-Context Learning
von: Ren, Jie, et al.
Veröffentlicht: (2024)
von: Ren, Jie, et al.
Veröffentlicht: (2024)
HuRef: HUman-REadable Fingerprint for Large Language Models
von: Zeng, Boyi, et al.
Veröffentlicht: (2023)
von: Zeng, Boyi, et al.
Veröffentlicht: (2023)
PonderLM: Pretraining Language Models to Ponder in Continuous Space
von: Zeng, Boyi, et al.
Veröffentlicht: (2025)
von: Zeng, Boyi, et al.
Veröffentlicht: (2025)
FLAME: Empowering Frozen LLMs for Knowledge Graph Completion
von: Xue, Bo, et al.
Veröffentlicht: (2024)
von: Xue, Bo, et al.
Veröffentlicht: (2024)
LongLLaDA: Unlocking Long Context Capabilities in Diffusion LLMs
von: Liu, Xiaoran, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoran, et al.
Veröffentlicht: (2025)
FreqKV: Key-Value Compression in Frequency Domain for Context Window Extension
von: Kai, Jushi, et al.
Veröffentlicht: (2025)
von: Kai, Jushi, et al.
Veröffentlicht: (2025)
ReAttention: Training-Free Infinite Context with Finite Attention Scope
von: Liu, Xiaoran, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoran, et al.
Veröffentlicht: (2024)
Fourier Transformer: Fast Long Range Modeling by Removing Sequence Redundancy with FFT Operator
von: He, Ziwei, et al.
Veröffentlicht: (2023)
von: He, Ziwei, et al.
Veröffentlicht: (2023)
Towards Controlled Table-to-Text Generation with Scientific Reasoning
von: Guo, Zhixin, et al.
Veröffentlicht: (2023)
von: Guo, Zhixin, et al.
Veröffentlicht: (2023)
Cluster-wise Graph Transformer with Dual-granularity Kernelized Attention
von: Huang, Siyuan, et al.
Veröffentlicht: (2024)
von: Huang, Siyuan, et al.
Veröffentlicht: (2024)
ICLEval: Evaluating In-Context Learning Ability of Large Language Models
von: Chen, Wentong, et al.
Veröffentlicht: (2024)
von: Chen, Wentong, et al.
Veröffentlicht: (2024)
Leveraging Grammar Induction for Language Understanding and Generation
von: Kai, Jushi, et al.
Veröffentlicht: (2024)
von: Kai, Jushi, et al.
Veröffentlicht: (2024)
How Do Different Forms of Note‐Taking Affect Second Language Vocabulary Learning?
von: Zhouhan Jin, et al.
Veröffentlicht: (2025)
von: Zhouhan Jin, et al.
Veröffentlicht: (2025)
Iterative Forward Tuning Boosts In-Context Learning in Language Models
von: Yang, Jiaxi, et al.
Veröffentlicht: (2023)
von: Yang, Jiaxi, et al.
Veröffentlicht: (2023)
PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space
von: Zeng, Boyi, et al.
Veröffentlicht: (2025)
von: Zeng, Boyi, et al.
Veröffentlicht: (2025)
In-Context Watermarks for Large Language Models
von: Liu, Yepeng, et al.
Veröffentlicht: (2025)
von: Liu, Yepeng, et al.
Veröffentlicht: (2025)
Label Words as Local Task Vectors in In-Context Learning
von: Zheng, Bowen, et al.
Veröffentlicht: (2024)
von: Zheng, Bowen, et al.
Veröffentlicht: (2024)
LongEmbed: Extending Embedding Models for Long Context Retrieval
von: Zhu, Dawei, et al.
Veröffentlicht: (2024)
von: Zhu, Dawei, et al.
Veröffentlicht: (2024)
PaD: Program-aided Distillation Can Teach Small Models Reasoning Better than Chain-of-thought Fine-tuning
von: Zhu, Xuekai, et al.
Veröffentlicht: (2023)
von: Zhu, Xuekai, et al.
Veröffentlicht: (2023)
What are the Essential Factors in Crafting Effective Long Context Multi-Hop Instruction Datasets? Insights and Best Practices
von: Chen, Zhi, et al.
Veröffentlicht: (2024)
von: Chen, Zhi, et al.
Veröffentlicht: (2024)
Long-Context Language Modeling with Parallel Context Encoding
von: Yen, Howard, et al.
Veröffentlicht: (2024)
von: Yen, Howard, et al.
Veröffentlicht: (2024)
Scaling Context, Not Parameters: Training a Compact 7B Language Model for Efficient Long-Context Processing
von: Wu, Chen, et al.
Veröffentlicht: (2025)
von: Wu, Chen, et al.
Veröffentlicht: (2025)
Reducing Distraction in Long-Context Language Models by Focused Learning
von: Wu, Zijun, et al.
Veröffentlicht: (2024)
von: Wu, Zijun, et al.
Veröffentlicht: (2024)
BGE Landmark Embedding: A Chunking-Free Embedding Method For Retrieval Augmented Long-Context Large Language Models
von: Luo, Kun, et al.
Veröffentlicht: (2024)
von: Luo, Kun, et al.
Veröffentlicht: (2024)
Context as a Tool: Context Management for Long-Horizon SWE-Agents
von: Liu, Shukai, et al.
Veröffentlicht: (2025)
von: Liu, Shukai, et al.
Veröffentlicht: (2025)
Revisiting In-Context Learning with Long Context Language Models
von: Baek, Jinheon, et al.
Veröffentlicht: (2024)
von: Baek, Jinheon, et al.
Veröffentlicht: (2024)
In-Context Former: Lightning-fast Compressing Context for Large Language Model
von: Wang, Xiangfeng, et al.
Veröffentlicht: (2024)
von: Wang, Xiangfeng, et al.
Veröffentlicht: (2024)
Embedding-Based Context-Aware Reranker
von: Yuan, Ye, et al.
Veröffentlicht: (2025)
von: Yuan, Ye, et al.
Veröffentlicht: (2025)
Towards Compressive and Scalable Recurrent Memory
von: Song, Yunchong, et al.
Veröffentlicht: (2026)
von: Song, Yunchong, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Next Concept Prediction in Discrete Latent Space Leads to Stronger Language Models
von: Liu, Yuliang, et al.
Veröffentlicht: (2026) -
KaLM: Knowledge-aligned Autoregressive Language Modeling via Dual-view Knowledge Graph Contrastive Learning
von: Yu, Peng, et al.
Veröffentlicht: (2024) -
Memory Decoder: A Pretrained, Plug-and-Play Memory for Large Language Models
von: Cao, Jiaqi, et al.
Veröffentlicht: (2025) -
MLP Memory: A Retriever-Pretrained Memory for Large Language Models
von: Wei, Rubin, et al.
Veröffentlicht: (2025) -
AWM: Accurate Weight-Matrix Fingerprint for Large Language Models
von: Zeng, Boyi, et al.
Veröffentlicht: (2025)