How Well Can a Long Sequence Model Model Long Sequences? Comparing Architechtural Inductive Biases on Long-Context Abilities
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Huang, Jerry |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?
von: Yang, Wang, et al.
Veröffentlicht: (2025)
von: Yang, Wang, et al.
Veröffentlicht: (2025)
MOM: Memory-Efficient Offloaded Mini-Sequence Inference for Long Context Language Models
von: Zhang, Junyang, et al.
Veröffentlicht: (2025)
von: Zhang, Junyang, et al.
Veröffentlicht: (2025)
Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning
von: Yang, Wang, et al.
Veröffentlicht: (2025)
von: Yang, Wang, et al.
Veröffentlicht: (2025)
The Impossibility Triangle of Long-Context Modeling
von: Zhou, Yan
Veröffentlicht: (2026)
von: Zhou, Yan
Veröffentlicht: (2026)
Star Attention: Efficient LLM Inference over Long Sequences
von: Acharya, Shantanu, et al.
Veröffentlicht: (2024)
von: Acharya, Shantanu, et al.
Veröffentlicht: (2024)
LongSafety: Enhance Safety for Long-Context LLMs
von: Huang, Mianqiu, et al.
Veröffentlicht: (2024)
von: Huang, Mianqiu, et al.
Veröffentlicht: (2024)
Too Long, Didn't Model: Decomposing LLM Long-Context Understanding With Novels
von: Hamilton, Sil, et al.
Veröffentlicht: (2025)
von: Hamilton, Sil, et al.
Veröffentlicht: (2025)
Revisiting In-Context Learning with Long Context Language Models
von: Baek, Jinheon, et al.
Veröffentlicht: (2024)
von: Baek, Jinheon, et al.
Veröffentlicht: (2024)
$π^2$: Structure-Originated Reasoning Data Improves Long-Context Reasoning Ability of Large Language Models
von: Do, Quyet V., et al.
Veröffentlicht: (2026)
von: Do, Quyet V., et al.
Veröffentlicht: (2026)
LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models
von: Chen, Yukang, et al.
Veröffentlicht: (2023)
von: Chen, Yukang, et al.
Veröffentlicht: (2023)
Artificial Hippocampus Networks for Efficient Long-Context Modeling
von: Fang, Yunhao, et al.
Veröffentlicht: (2025)
von: Fang, Yunhao, et al.
Veröffentlicht: (2025)
Do Robot Snakes Dream like Electric Sheep? Investigating the Effects of Architectural Inductive Biases on Hallucination
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
Cost-Optimal Grouped-Query Attention for Long-Context Modeling
von: Chen, Yingfa, et al.
Veröffentlicht: (2025)
von: Chen, Yingfa, et al.
Veröffentlicht: (2025)
Integrating LSTM and BERT for Long-Sequence Data Analysis in Intelligent Tutoring Systems
von: Li, Zhaoxing, et al.
Veröffentlicht: (2024)
von: Li, Zhaoxing, et al.
Veröffentlicht: (2024)
Block-Biased Mamba for Long-Range Sequence Processing
von: Yu, Annan, et al.
Veröffentlicht: (2025)
von: Yu, Annan, et al.
Veröffentlicht: (2025)
Long-Context LLMs Meet RAG: Overcoming Challenges for Long Inputs in RAG
von: Jin, Bowen, et al.
Veröffentlicht: (2024)
von: Jin, Bowen, et al.
Veröffentlicht: (2024)
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
von: Yang, Penghui, et al.
Veröffentlicht: (2025)
von: Yang, Penghui, et al.
Veröffentlicht: (2025)
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA)
von: Dentamaro, Vincenzo
Veröffentlicht: (2025)
von: Dentamaro, Vincenzo
Veröffentlicht: (2025)
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
von: Ye, Jiabo, et al.
Veröffentlicht: (2024)
von: Ye, Jiabo, et al.
Veröffentlicht: (2024)
Stress-Testing Long-Context Language Models with Lifelong ICL and Task Haystack
von: Xu, Xiaoyue, et al.
Veröffentlicht: (2024)
von: Xu, Xiaoyue, et al.
Veröffentlicht: (2024)
ELITR-Bench: A Meeting Assistant Benchmark for Long-Context Language Models
von: Thonet, Thibaut, et al.
Veröffentlicht: (2024)
von: Thonet, Thibaut, et al.
Veröffentlicht: (2024)
Retrieval meets Long Context Large Language Models
von: Xu, Peng, et al.
Veröffentlicht: (2023)
von: Xu, Peng, et al.
Veröffentlicht: (2023)
Sequence-to-Sequence Spanish Pre-trained Language Models
von: Araujo, Vladimir, et al.
Veröffentlicht: (2023)
von: Araujo, Vladimir, et al.
Veröffentlicht: (2023)
LLoCO: Learning Long Contexts Offline
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
Long-form factuality in large language models
von: Wei, Jerry, et al.
Veröffentlicht: (2024)
von: Wei, Jerry, et al.
Veröffentlicht: (2024)
LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation
von: Zhang, Xuan, et al.
Veröffentlicht: (2024)
von: Zhang, Xuan, et al.
Veröffentlicht: (2024)
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
von: MiniCPM Team, et al.
Veröffentlicht: (2026)
von: MiniCPM Team, et al.
Veröffentlicht: (2026)
LaCache: Ladder-Shaped KV Caching for Efficient Long-Context Modeling of Large Language Models
von: Shi, Dachuan, et al.
Veröffentlicht: (2025)
von: Shi, Dachuan, et al.
Veröffentlicht: (2025)
PEEK: Context Map as an Orientation Cache for Long-Context LLM Agents
von: Gu, Zhuohan, et al.
Veröffentlicht: (2026)
von: Gu, Zhuohan, et al.
Veröffentlicht: (2026)
Latent Context Compilation: Distilling Long Context into Compact Portable Memory
von: Li, Zeju, et al.
Veröffentlicht: (2026)
von: Li, Zeju, et al.
Veröffentlicht: (2026)
TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference
von: Dzikanyanga, Gradwell, et al.
Veröffentlicht: (2026)
von: Dzikanyanga, Gradwell, et al.
Veröffentlicht: (2026)
Calibrating Long-form Generations from Large Language Models
von: Huang, Yukun, et al.
Veröffentlicht: (2024)
von: Huang, Yukun, et al.
Veröffentlicht: (2024)
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2025)
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2025)
miniCTX: Neural Theorem Proving with (Long-)Contexts
von: Hu, Jiewen, et al.
Veröffentlicht: (2024)
von: Hu, Jiewen, et al.
Veröffentlicht: (2024)
AcademicEval: Live Long-Context LLM Benchmark
von: Zhang, Haozhen, et al.
Veröffentlicht: (2025)
von: Zhang, Haozhen, et al.
Veröffentlicht: (2025)
SEAL: Scaling to Emphasize Attention for Long-Context Retrieval
von: Lee, Changhun, et al.
Veröffentlicht: (2025)
von: Lee, Changhun, et al.
Veröffentlicht: (2025)
MoBA: Mixture of Block Attention for Long-Context LLMs
von: Lu, Enzhe, et al.
Veröffentlicht: (2025)
von: Lu, Enzhe, et al.
Veröffentlicht: (2025)
Fine-Grained Uncertainty Quantification for Long-Form Language Model Outputs: A Comparative Study
von: Bouchard, Dylan, et al.
Veröffentlicht: (2026)
von: Bouchard, Dylan, et al.
Veröffentlicht: (2026)
AgentFold: Long-Horizon Web Agents with Proactive Context Management
von: Ye, Rui, et al.
Veröffentlicht: (2025)
von: Ye, Rui, et al.
Veröffentlicht: (2025)
LooGLE: Can Long-Context Language Models Understand Long Contexts?
von: Li, Jiaqi, et al.
Veröffentlicht: (2023)
von: Li, Jiaqi, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?
von: Yang, Wang, et al.
Veröffentlicht: (2025) -
MOM: Memory-Efficient Offloaded Mini-Sequence Inference for Long Context Language Models
von: Zhang, Junyang, et al.
Veröffentlicht: (2025) -
Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning
von: Yang, Wang, et al.
Veröffentlicht: (2025) -
The Impossibility Triangle of Long-Context Modeling
von: Zhou, Yan
Veröffentlicht: (2026) -
Star Attention: Efficient LLM Inference over Long Sequences
von: Acharya, Shantanu, et al.
Veröffentlicht: (2024)