BlockGen: Flexible Blockwise Sequence Modeling with Hybrid Samplers
Fuente:
arXiv
Salvato in:
| Autori principali: | Deschenaux, Justin, Gulcehre, Caglar |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Diffusion Duality, Chapter II: $Ψ$-Samplers
di: Deschenaux, Justin, et al.
Pubblicazione: (2026)
di: Deschenaux, Justin, et al.
Pubblicazione: (2026)
Beyond Autoregression: Fast LLMs via Self-Distillation Through Time
di: Deschenaux, Justin, et al.
Pubblicazione: (2024)
di: Deschenaux, Justin, et al.
Pubblicazione: (2024)
Partition Generative Modeling: Masked Modeling Without Masks
di: Deschenaux, Justin, et al.
Pubblicazione: (2025)
di: Deschenaux, Justin, et al.
Pubblicazione: (2025)
Promises, Outlooks and Challenges of Diffusion Language Modeling
di: Deschenaux, Justin, et al.
Pubblicazione: (2024)
di: Deschenaux, Justin, et al.
Pubblicazione: (2024)
Loopholing Discrete Diffusion: Deterministic Bypass of the Sampling Wall
di: Jo, Mingyu, et al.
Pubblicazione: (2025)
di: Jo, Mingyu, et al.
Pubblicazione: (2025)
One-Step is Enough: Sparse Autoencoders for Text-to-Image Diffusion Models
di: Surkov, Viacheslav, et al.
Pubblicazione: (2024)
di: Surkov, Viacheslav, et al.
Pubblicazione: (2024)
Augmenting Attention with Exponentially Decaying Memory Improves Query-Aware KV Sparsity
di: Wei, Xiuying, et al.
Pubblicazione: (2026)
di: Wei, Xiuying, et al.
Pubblicazione: (2026)
RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference
di: Wei, Xiuying, et al.
Pubblicazione: (2026)
di: Wei, Xiuying, et al.
Pubblicazione: (2026)
In Search for Architectures and Loss Functions in Multi-Objective Reinforcement Learning
di: Terekhov, Mikhail, et al.
Pubblicazione: (2024)
di: Terekhov, Mikhail, et al.
Pubblicazione: (2024)
Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions
di: Matrenok, Simon, et al.
Pubblicazione: (2025)
di: Matrenok, Simon, et al.
Pubblicazione: (2025)
The Role of Deep Learning Regularizations on Actors in Offline RL
di: Tarasov, Denis, et al.
Pubblicazione: (2024)
di: Tarasov, Denis, et al.
Pubblicazione: (2024)
HiPPO-Prophecy: State-Space Models can Provably Learn Dynamical Systems in Context
di: Joseph, Federico Arangath, et al.
Pubblicazione: (2024)
di: Joseph, Federico Arangath, et al.
Pubblicazione: (2024)
BlockCert: Certified Blockwise Extraction of Transformer Mechanisms
di: Andric, Sandro
Pubblicazione: (2025)
di: Andric, Sandro
Pubblicazione: (2025)
No Representation, No Trust: Connecting Representation, Collapse, and Trust Issues in PPO
di: Moalla, Skander, et al.
Pubblicazione: (2024)
di: Moalla, Skander, et al.
Pubblicazione: (2024)
Simple Hierarchical Planning with Diffusion
di: Chen, Chang, et al.
Pubblicazione: (2024)
di: Chen, Chang, et al.
Pubblicazione: (2024)
Control Tax: The Price of Keeping AI in Check
di: Terekhov, Mikhail, et al.
Pubblicazione: (2025)
di: Terekhov, Mikhail, et al.
Pubblicazione: (2025)
Self-Recognition in Language Models
di: Davidson, Tim R., et al.
Pubblicazione: (2024)
di: Davidson, Tim R., et al.
Pubblicazione: (2024)
Universality of Linear Recurrences Followed by Non-linear Projections: Finite-Width Guarantees and Benefits of Complex Eigenvalues
di: Orvieto, Antonio, et al.
Pubblicazione: (2023)
di: Orvieto, Antonio, et al.
Pubblicazione: (2023)
PlanDQ: Hierarchical Plan Orchestration via D-Conductor and Q-Performer
di: Chen, Chang, et al.
Pubblicazione: (2024)
di: Chen, Chang, et al.
Pubblicazione: (2024)
Fleet of Agents: Coordinated Problem Solving with Large Language Models
di: Klein, Lars, et al.
Pubblicazione: (2024)
di: Klein, Lars, et al.
Pubblicazione: (2024)
World Model on Million-Length Video And Language With Blockwise RingAttention
di: Liu, Hao, et al.
Pubblicazione: (2024)
di: Liu, Hao, et al.
Pubblicazione: (2024)
Regret-Optimized Portfolio Enhancement through Deep Reinforcement Learning and Future Looking Rewards
di: Karzanov, Daniil, et al.
Pubblicazione: (2025)
di: Karzanov, Daniil, et al.
Pubblicazione: (2025)
On the Trainability of Masked Diffusion Language Models via Blockwise Locality
di: Wang, Yuxiang, et al.
Pubblicazione: (2026)
di: Wang, Yuxiang, et al.
Pubblicazione: (2026)
The Diffusion Duality
di: Sahoo, Subham Sekhar, et al.
Pubblicazione: (2025)
di: Sahoo, Subham Sekhar, et al.
Pubblicazione: (2025)
CoDe: Blockwise Control for Denoising Diffusion Models
di: Singh, Anuj, et al.
Pubblicazione: (2025)
di: Singh, Anuj, et al.
Pubblicazione: (2025)
Scaling Beyond Masked Diffusion Language Models
di: Sahoo, Subham Sekhar, et al.
Pubblicazione: (2026)
di: Sahoo, Subham Sekhar, et al.
Pubblicazione: (2026)
Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols
di: Terekhov, Mikhail, et al.
Pubblicazione: (2025)
di: Terekhov, Mikhail, et al.
Pubblicazione: (2025)
Corrected Samplers for Discrete Flow Models
di: Wan, Zhengyan, et al.
Pubblicazione: (2026)
di: Wan, Zhengyan, et al.
Pubblicazione: (2026)
Neural Flow Samplers with Shortcut Models
di: Chen, Wuhao, et al.
Pubblicazione: (2025)
di: Chen, Wuhao, et al.
Pubblicazione: (2025)
Exploring and Improving Drafts in Blockwise Parallel Decoding
di: Kim, Taehyeon, et al.
Pubblicazione: (2024)
di: Kim, Taehyeon, et al.
Pubblicazione: (2024)
Blockwise Self-Supervised Learning at Scale
di: Siddiqui, Shoaib Ahmed, et al.
Pubblicazione: (2023)
di: Siddiqui, Shoaib Ahmed, et al.
Pubblicazione: (2023)
Is Your Diffusion Sampler Actually Correct? A Sampler-Centric Evaluation of Discrete Diffusion Language Models
di: Tang, Luhan, et al.
Pubblicazione: (2026)
di: Tang, Luhan, et al.
Pubblicazione: (2026)
BoHA: Blockwise Hadamard Product Adaptation for Parameter-Efficient Fine-Tuning
di: Yu, Feng, et al.
Pubblicazione: (2025)
di: Yu, Feng, et al.
Pubblicazione: (2025)
Blockwise Principal Component Analysis for monotone missing data imputation and dimensionality reduction
di: Do, Tu T., et al.
Pubblicazione: (2023)
di: Do, Tu T., et al.
Pubblicazione: (2023)
GenPlan: Generative Sequence Models as Adaptive Planners
di: Karthikeyan, Akash, et al.
Pubblicazione: (2024)
di: Karthikeyan, Akash, et al.
Pubblicazione: (2024)
Blockwise Missingness meets AI: A Tractable Solution for Semiparametric Inference
di: Xu, Qi, et al.
Pubblicazione: (2025)
di: Xu, Qi, et al.
Pubblicazione: (2025)
Multi-Mixer Models: Flexible Sequence Modeling with Shared Representations
di: Li, Kevin Y., et al.
Pubblicazione: (2026)
di: Li, Kevin Y., et al.
Pubblicazione: (2026)
Learnable Sampler Distillation for Discrete Diffusion Models
di: Fu, Feiyang, et al.
Pubblicazione: (2025)
di: Fu, Feiyang, et al.
Pubblicazione: (2025)
Orchid: Flexible and Data-Dependent Convolution for Sequence Modeling
di: Karami, Mahdi, et al.
Pubblicazione: (2024)
di: Karami, Mahdi, et al.
Pubblicazione: (2024)
Diffusion-PINN Sampler
di: Shi, Zhekun, et al.
Pubblicazione: (2024)
di: Shi, Zhekun, et al.
Pubblicazione: (2024)
Documenti analoghi
-
The Diffusion Duality, Chapter II: $Ψ$-Samplers
di: Deschenaux, Justin, et al.
Pubblicazione: (2026) -
Beyond Autoregression: Fast LLMs via Self-Distillation Through Time
di: Deschenaux, Justin, et al.
Pubblicazione: (2024) -
Partition Generative Modeling: Masked Modeling Without Masks
di: Deschenaux, Justin, et al.
Pubblicazione: (2025) -
Promises, Outlooks and Challenges of Diffusion Language Modeling
di: Deschenaux, Justin, et al.
Pubblicazione: (2024) -
Loopholing Discrete Diffusion: Deterministic Bypass of the Sampling Wall
di: Jo, Mingyu, et al.
Pubblicazione: (2025)