Provable Length Generalization in Sequence Prediction via Spectral Filtering
Fuente:
arXiv
Saved in:
| Main Authors: | Marsden, Annie, Dogariu, Evan, Agarwal, Naman, Chen, Xinyi, Suo, Daniel, Hazan, Elad |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FutureFill: Fast Generation from Convolutional Sequence Models
by: Agarwal, Naman, et al.
Published: (2024)
by: Agarwal, Naman, et al.
Published: (2024)
Spectral Filtering for Complex Linear Dynamical Systems
by: Hazan, Elad, et al.
Published: (2026)
by: Hazan, Elad, et al.
Published: (2026)
Spectral State Space Models
by: Agarwal, Naman, et al.
Published: (2023)
by: Agarwal, Naman, et al.
Published: (2023)
Flash STU: Fast Spectral Transform Units
by: Liu, Y. Isabel, et al.
Published: (2024)
by: Liu, Y. Isabel, et al.
Published: (2024)
Universal Sequence Preconditioning
by: Marsden, Annie, et al.
Published: (2025)
by: Marsden, Annie, et al.
Published: (2025)
The Power of Second Order Methods for Sequence Preconditioning
by: Marsden, Annie, et al.
Published: (2026)
by: Marsden, Annie, et al.
Published: (2026)
SFO: Learning PDE Operators via Spectral Filtering
by: Koren, Noam, et al.
Published: (2026)
by: Koren, Noam, et al.
Published: (2026)
Universal Learning of Nonlinear Dynamics
by: Dogariu, Evan, et al.
Published: (2025)
by: Dogariu, Evan, et al.
Published: (2025)
AI Alignment via Incentives and Correction
by: Agarwal, Rohit, et al.
Published: (2026)
by: Agarwal, Rohit, et al.
Published: (2026)
An Empirical Study on Context Length for Open-Domain Dialog Generation
by: Shen, Xinyi, et al.
Published: (2024)
by: Shen, Xinyi, et al.
Published: (2024)
Transformers Can Achieve Length Generalization But Not Robustly
by: Zhou, Yongchao, et al.
Published: (2024)
by: Zhou, Yongchao, et al.
Published: (2024)
Context-Aware Initialization for Reducing Generative Path Length in Diffusion Language Models
by: Miao, Tongyuan, et al.
Published: (2025)
by: Miao, Tongyuan, et al.
Published: (2025)
Research Program: Theory of Learning in Dynamical Systems
by: Hazan, Elad, et al.
Published: (2025)
by: Hazan, Elad, et al.
Published: (2025)
Clip Your Sequences Fairly: Enforcing Length Fairness for Sequence-Level RL
by: Mao, Hanyi, et al.
Published: (2025)
by: Mao, Hanyi, et al.
Published: (2025)
A comprehensive study of on-device NLP applications -- VQA, automated Form filling, Smart Replies for Linguistic Codeswitching
by: Goyal, Naman
Published: (2024)
by: Goyal, Naman
Published: (2024)
BARRED: Synthetic Training of Custom Policy Guardrails via Asymmetric Debate
by: Mazza, Arnon, et al.
Published: (2026)
by: Mazza, Arnon, et al.
Published: (2026)
Rewarding Graph Reasoning Process makes LLMs more Generalized Reasoners
by: Peng, Miao, et al.
Published: (2025)
by: Peng, Miao, et al.
Published: (2025)
IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems
by: Levi, Elad, et al.
Published: (2025)
by: Levi, Elad, et al.
Published: (2025)
The Hidden Game Problem
by: Buzaglo, Gon, et al.
Published: (2025)
by: Buzaglo, Gon, et al.
Published: (2025)
LUMOS: Large User MOdels for User Behavior Prediction
by: Nigam, Dhruv, et al.
Published: (2025)
by: Nigam, Dhruv, et al.
Published: (2025)
Chain of LoRA: Efficient Fine-tuning of Language Models via Residual Learning
by: Xia, Wenhan, et al.
Published: (2024)
by: Xia, Wenhan, et al.
Published: (2024)
Multiscale Byte Language Models -- A Hierarchical Architecture for Causal Million-Length Sequence Modeling
by: Egli, Eric, et al.
Published: (2025)
by: Egli, Eric, et al.
Published: (2025)
The Role of Sparsity for Length Generalization in Transformers
by: Golowich, Noah, et al.
Published: (2025)
by: Golowich, Noah, et al.
Published: (2025)
Debate Helps Weak Judges Reward Stronger Models
by: Elasky, Ethan, et al.
Published: (2026)
by: Elasky, Ethan, et al.
Published: (2026)
RubiConv -- Efficient Boundary-Respecting Convolutions
by: Friso, Linda, et al.
Published: (2026)
by: Friso, Linda, et al.
Published: (2026)
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning
by: Zheng, Congmin, et al.
Published: (2025)
by: Zheng, Congmin, et al.
Published: (2025)
Stable Adaptive Thinking via Advantage Shaping and Length-Aware Gradient Regulation
by: Xu, Zihang, et al.
Published: (2026)
by: Xu, Zihang, et al.
Published: (2026)
Intent-based Prompt Calibration: Enhancing prompt optimization with synthetic boundary cases
by: Levi, Elad, et al.
Published: (2024)
by: Levi, Elad, et al.
Published: (2024)
Transformers Provably Learn Chain-of-Thought Reasoning with Length Generalization
by: Huang, Yu, et al.
Published: (2025)
by: Huang, Yu, et al.
Published: (2025)
Languages are Modalities: Cross-Lingual Alignment via Encoder Injection
by: Agarwal, Rajan, et al.
Published: (2025)
by: Agarwal, Rajan, et al.
Published: (2025)
Provable Scaling Laws for the Test-Time Compute of Large Language Models
by: Chen, Yanxi, et al.
Published: (2024)
by: Chen, Yanxi, et al.
Published: (2024)
Understanding and Improving Length Generalization in Hierarchical Sparse Attention Models
by: Leng, Jiaqi, et al.
Published: (2025)
by: Leng, Jiaqi, et al.
Published: (2025)
Provable Interactive Learning with Hindsight Instruction Feedback
by: Misra, Dipendra, et al.
Published: (2024)
by: Misra, Dipendra, et al.
Published: (2024)
Recursive Introspection: Teaching Language Model Agents How to Self-Improve
by: Qu, Yuxiao, et al.
Published: (2024)
by: Qu, Yuxiao, et al.
Published: (2024)
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA)
by: Dentamaro, Vincenzo
Published: (2025)
by: Dentamaro, Vincenzo
Published: (2025)
On the Expressiveness and Length Generalization of Selective State-Space Models on Regular Languages
by: Terzić, Aleksandar, et al.
Published: (2024)
by: Terzić, Aleksandar, et al.
Published: (2024)
SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths
by: Huang, Kaixuan, et al.
Published: (2024)
by: Huang, Kaixuan, et al.
Published: (2024)
Multiple Abstraction Level Retrieve Augment Generation
by: Zheng, Zheng, et al.
Published: (2025)
by: Zheng, Zheng, et al.
Published: (2025)
Post-training for Efficient Communication via Convention Formation
by: Hua, Yilun, et al.
Published: (2025)
by: Hua, Yilun, et al.
Published: (2025)
On The Truthfulness of 'Surprisingly Likely' Responses of Large Language Models
by: Goel, Naman
Published: (2023)
by: Goel, Naman
Published: (2023)
Similar Items
-
FutureFill: Fast Generation from Convolutional Sequence Models
by: Agarwal, Naman, et al.
Published: (2024) -
Spectral Filtering for Complex Linear Dynamical Systems
by: Hazan, Elad, et al.
Published: (2026) -
Spectral State Space Models
by: Agarwal, Naman, et al.
Published: (2023) -
Flash STU: Fast Spectral Transform Units
by: Liu, Y. Isabel, et al.
Published: (2024) -
Universal Sequence Preconditioning
by: Marsden, Annie, et al.
Published: (2025)