Why Gaussian Diffusion Models Fail on Discrete Data and How to Prevent It?
Fuente:
arXiv
Saved in:
| Main Authors: | Shabalin, Alexander, Elistratov, Simon, Meshchaninov, Viacheslav, Sadrtdinov, Ildus, Vetrov, Dmitry |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Smoothie: Smoothing Diffusion on Token Embeddings for Text Generation
by: Shabalin, Alexander, et al.
Published: (2025)
by: Shabalin, Alexander, et al.
Published: (2025)
Cosmos: Compressed and Smooth Latent Space for Text Diffusion Modeling
by: Meshchaninov, Viacheslav, et al.
Published: (2025)
by: Meshchaninov, Viacheslav, et al.
Published: (2025)
How to Train Your Latent Diffusion Language Model Jointly With the Latent Space
by: Meshchaninov, Viacheslav, et al.
Published: (2026)
by: Meshchaninov, Viacheslav, et al.
Published: (2026)
TEncDM: Understanding the Properties of the Diffusion Model in the Space of Language Model Encodings
by: Shabalin, Alexander, et al.
Published: (2024)
by: Shabalin, Alexander, et al.
Published: (2024)
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training
by: Sadrtdinov, Ildus, et al.
Published: (2025)
by: Sadrtdinov, Ildus, et al.
Published: (2025)
To Stay or Not to Stay in the Pre-train Basin: Insights on Ensembling in Transfer Learning
by: Sadrtdinov, Ildus, et al.
Published: (2023)
by: Sadrtdinov, Ildus, et al.
Published: (2023)
Where Do Large Learning Rates Lead Us?
by: Sadrtdinov, Ildus, et al.
Published: (2024)
by: Sadrtdinov, Ildus, et al.
Published: (2024)
Can Stationary Distributions of Scale-Invariant Neural Networks Be Described by the Thermodynamics of an Ideal Gas?
by: Sadrtdinov, Ildus, et al.
Published: (2025)
by: Sadrtdinov, Ildus, et al.
Published: (2025)
Why Instruction-Based Unlearning Fails in Diffusion Models?
by: Zhang, Zeliang, et al.
Published: (2026)
by: Zhang, Zeliang, et al.
Published: (2026)
Diffusion on language model encodings for protein sequence generation
by: Meshchaninov, Viacheslav, et al.
Published: (2024)
by: Meshchaninov, Viacheslav, et al.
Published: (2024)
Why LLMs Fail at Causal Discovery and How Interventional Agents Escape
by: Roy, Amartya, et al.
Published: (2026)
by: Roy, Amartya, et al.
Published: (2026)
Guided Star-Shaped Masked Diffusion
by: Meshchaninov, Viacheslav, et al.
Published: (2025)
by: Meshchaninov, Viacheslav, et al.
Published: (2025)
Why Latent Actions Fail, and How to Prevent It
by: Lee, Jung Min, et al.
Published: (2026)
by: Lee, Jung Min, et al.
Published: (2026)
Why We Need World Models for AGI: Where LLMs Fail and How World Models May Outperform
by: Alaswad, Feisal, et al.
Published: (2026)
by: Alaswad, Feisal, et al.
Published: (2026)
Larger Language Models Don't Care How You Think: Why Chain-of-Thought Prompting Fails in Subjective Tasks
by: Chochlakis, Georgios, et al.
Published: (2024)
by: Chochlakis, Georgios, et al.
Published: (2024)
Why Chain of Thought Fails in Clinical Text Understanding
by: Wu, Jiageng, et al.
Published: (2025)
by: Wu, Jiageng, et al.
Published: (2025)
Why LoRA Fails to Forget: Regularized Low-Rank Adaptation Against Backdoors in Language Models
by: Luong, Hoang-Chau, et al.
Published: (2026)
by: Luong, Hoang-Chau, et al.
Published: (2026)
Why Retrieval-Augmented Generation Fails: A Graph Perspective
by: Guo, Kai, et al.
Published: (2026)
by: Guo, Kai, et al.
Published: (2026)
Neural Flow Diffusion Models: Learnable Forward Process for Improved Diffusion Modelling
by: Bartosh, Grigory, et al.
Published: (2024)
by: Bartosh, Grigory, et al.
Published: (2024)
The Right Answer, the Wrong Direction: Why Transformers Fail at Counting and How to Fix It
by: Garcia, Gabriel
Published: (2026)
by: Garcia, Gabriel
Published: (2026)
Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution
by: Lou, Aaron, et al.
Published: (2023)
by: Lou, Aaron, et al.
Published: (2023)
Why Supervised Fine-Tuning Fails to Learn: A Systematic Study of Incomplete Learning in Large Language Models
by: Xue, Chao, et al.
Published: (2026)
by: Xue, Chao, et al.
Published: (2026)
Lost in the Noise: How Reasoning Models Fail with Contextual Distractors
by: Lee, Seongyun, et al.
Published: (2026)
by: Lee, Seongyun, et al.
Published: (2026)
False Sense of Security: Why Probing-based Malicious Input Detection Fails to Generalize
by: Wang, Cheng, et al.
Published: (2025)
by: Wang, Cheng, et al.
Published: (2025)
Why Better Cross-Lingual Alignment Fails for Better Cross-Lingual Transfer: Case of Encoders
by: Veitsman, Yana, et al.
Published: (2026)
by: Veitsman, Yana, et al.
Published: (2026)
Rolling the DICE on Idiomaticity: How LLMs Fail to Grasp Context
by: Mi, Maggie, et al.
Published: (2024)
by: Mi, Maggie, et al.
Published: (2024)
Why do Large Language Models Fail in Low-resource Translation? Unraveling the Token Dynamics of Large Language Models for Machine Translation
by: Qian, Shenbin, et al.
Published: (2026)
by: Qian, Shenbin, et al.
Published: (2026)
How LLMs Fail to Support Fact-Checking
by: Proma, Adiba Mahbub, et al.
Published: (2025)
by: Proma, Adiba Mahbub, et al.
Published: (2025)
Why Semantic Entropy Fails: Geometry-Aware and Calibrated Uncertainty for Policy Optimization
by: Zhang, Zheyuan, et al.
Published: (2026)
by: Zhang, Zheyuan, et al.
Published: (2026)
Biasless Language Models Learn Unnaturally: How LLMs Fail to Distinguish the Possible from the Impossible
by: Ziv, Imry, et al.
Published: (2025)
by: Ziv, Imry, et al.
Published: (2025)
More Rounds, More Noise: Why Multi-Turn Review Fails to Improve Cross-Context Verification
by: Tae-Eun, Song
Published: (2026)
by: Tae-Eun, Song
Published: (2026)
Balancing Understanding and Generation in Discrete Diffusion Models
by: Liu, Yue, et al.
Published: (2026)
by: Liu, Yue, et al.
Published: (2026)
Streaming Generation of Co-Speech Gestures via Accelerated Rolling Diffusion
by: Vu, Evgeniia, et al.
Published: (2025)
by: Vu, Evgeniia, et al.
Published: (2025)
Why Do LLM-based Web Agents Fail? A Hierarchical Planning Perspective
by: Aghzal, Mohamed, et al.
Published: (2026)
by: Aghzal, Mohamed, et al.
Published: (2026)
Accommodation and Epistemic Vigilance: A Pragmatic Account of Why LLMs Fail to Challenge Harmful Beliefs
by: Cheng, Myra, et al.
Published: (2026)
by: Cheng, Myra, et al.
Published: (2026)
Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data
by: Ou, Jingyang, et al.
Published: (2024)
by: Ou, Jingyang, et al.
Published: (2024)
Constrained Code Generation with Discrete Diffusion
by: Shao, Lize, et al.
Published: (2026)
by: Shao, Lize, et al.
Published: (2026)
Detection Is Cheap, Routing Is Learned: Why Refusal-Based Alignment Evaluation Fails
by: Frank, Gregory N.
Published: (2026)
by: Frank, Gregory N.
Published: (2026)
Adaptation Odyssey in LLMs: Why Does Additional Pretraining Sometimes Fail to Improve?
by: Öncel, Fırat, et al.
Published: (2024)
by: Öncel, Fırat, et al.
Published: (2024)
The Mirage of Performance Gains: Why Contrastive Decoding Fails to Mitigate Object Hallucinations in MLLMs?
by: Yin, Hao, et al.
Published: (2025)
by: Yin, Hao, et al.
Published: (2025)
Similar Items
-
Smoothie: Smoothing Diffusion on Token Embeddings for Text Generation
by: Shabalin, Alexander, et al.
Published: (2025) -
Cosmos: Compressed and Smooth Latent Space for Text Diffusion Modeling
by: Meshchaninov, Viacheslav, et al.
Published: (2025) -
How to Train Your Latent Diffusion Language Model Jointly With the Latent Space
by: Meshchaninov, Viacheslav, et al.
Published: (2026) -
TEncDM: Understanding the Properties of the Diffusion Model in the Space of Language Model Encodings
by: Shabalin, Alexander, et al.
Published: (2024) -
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training
by: Sadrtdinov, Ildus, et al.
Published: (2025)