A Pseudo-Semantic Loss for Autoregressive Models with Logical Constraints
Fuente:
arXiv
Saved in:
| Main Authors: | Ahmed, Kareem, Chang, Kai-Wei, Broeck, Guy Van den |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enabling Autoregressive Models to Fill In Masked Tokens
by: Israel, Daniel, et al.
Published: (2025)
by: Israel, Daniel, et al.
Published: (2025)
Controllable Generation via Locally Constrained Resampling
by: Ahmed, Kareem, et al.
Published: (2024)
by: Ahmed, Kareem, et al.
Published: (2024)
Prepacking: A Simple Method for Fast Prefilling and Increased Throughput in Large Language Models
by: Zhao, Siyan, et al.
Published: (2024)
by: Zhao, Siyan, et al.
Published: (2024)
Adversarial Tokenization
by: Geh, Renato Lui, et al.
Published: (2025)
by: Geh, Renato Lui, et al.
Published: (2025)
Accelerating Diffusion LLMs via Adaptive Parallel Decoding
by: Israel, Daniel, et al.
Published: (2025)
by: Israel, Daniel, et al.
Published: (2025)
SIMPLE: A Gradient Estimator for $k$-Subset Sampling
by: Ahmed, Kareem, et al.
Published: (2022)
by: Ahmed, Kareem, et al.
Published: (2022)
Where is the signal in tokenization space?
by: Geh, Renato Lui, et al.
Published: (2024)
by: Geh, Renato Lui, et al.
Published: (2024)
Collapsed Inference for Bayesian Deep Learning
by: Zeng, Zhe, et al.
Published: (2023)
by: Zeng, Zhe, et al.
Published: (2023)
On the Relationship Between Monotone and Squared Probabilistic Circuits
by: Wang, Benjie, et al.
Published: (2024)
by: Wang, Benjie, et al.
Published: (2024)
Planner and Executor: Collaboration between Discrete Diffusion And Autoregressive Models in Reasoning
by: Berrayana, Lina, et al.
Published: (2025)
by: Berrayana, Lina, et al.
Published: (2025)
Scaling Up Probabilistic Circuits by Latent Variable Distillation
by: Liu, Anji, et al.
Published: (2022)
by: Liu, Anji, et al.
Published: (2022)
How to Marginalize in Causal Structure Learning?
by: Zhao, William, et al.
Published: (2025)
by: Zhao, William, et al.
Published: (2025)
Advancing Sequential Numerical Prediction in Autoregressive Models
by: Fei, Xiang, et al.
Published: (2025)
by: Fei, Xiang, et al.
Published: (2025)
ReFusion: A Diffusion Large Language Model with Parallel Autoregressive Decoding
by: Li, Jia-Nan, et al.
Published: (2025)
by: Li, Jia-Nan, et al.
Published: (2025)
Continuous Autoregressive Language Models
by: Shao, Chenze, et al.
Published: (2025)
by: Shao, Chenze, et al.
Published: (2025)
ALPINE: Unveiling the Planning Capability of Autoregressive Learning in Language Models
by: Wang, Siwei, et al.
Published: (2024)
by: Wang, Siwei, et al.
Published: (2024)
TRACE Back from the Future: A Probabilistic Reasoning Approach to Controllable Language Generation
by: Weng, Gwen Yidou, et al.
Published: (2025)
by: Weng, Gwen Yidou, et al.
Published: (2025)
A Compositional Atlas for Algebraic Circuits
by: Wang, Benjie, et al.
Published: (2024)
by: Wang, Benjie, et al.
Published: (2024)
A Tractable Inference Perspective of Offline RL
by: Liu, Xuejie, et al.
Published: (2023)
by: Liu, Xuejie, et al.
Published: (2023)
Exploring the Reversal Curse and Other Deductive Logical Reasoning in BERT and GPT-Based Large Language Models
by: Wu, Da, et al.
Published: (2023)
by: Wu, Da, et al.
Published: (2023)
Context Dependence and Reliability in Autoregressive Language Models
by: Sengupta, Poushali, et al.
Published: (2026)
by: Sengupta, Poushali, et al.
Published: (2026)
Projected Autoregression: Autoregressive Language Generation in Continuous State Space
by: Naparstek, Oshri
Published: (2026)
by: Naparstek, Oshri
Published: (2026)
AR-Omni: A Unified Autoregressive Model for Any-to-Any Generation
by: Cheng, Dongjie, et al.
Published: (2026)
by: Cheng, Dongjie, et al.
Published: (2026)
Unveiling and Addressing Pseudo Forgetting in Large Language Models
by: Sun, Huashan, et al.
Published: (2024)
by: Sun, Huashan, et al.
Published: (2024)
Position: The Turing-Completeness of Autoregressive Transformers Relies Heavily on Context Management
by: Cui, Guanyu, et al.
Published: (2026)
by: Cui, Guanyu, et al.
Published: (2026)
DivLogicEval: A Framework for Benchmarking Logical Reasoning Evaluation in Large Language Models
by: Chung, Tsz Ting, et al.
Published: (2025)
by: Chung, Tsz Ting, et al.
Published: (2025)
Restructuring Tractable Probabilistic Circuits
by: Zhang, Honghua, et al.
Published: (2024)
by: Zhang, Honghua, et al.
Published: (2024)
ProbMoE: Differentiable Probabilistic Routing for Mixture-of-Experts
by: Zhao, Heng, et al.
Published: (2026)
by: Zhao, Heng, et al.
Published: (2026)
Efficient-DLM: From Autoregressive to Diffusion Language Models, and Beyond in Speed
by: Fu, Yonggan, et al.
Published: (2025)
by: Fu, Yonggan, et al.
Published: (2025)
Scaling Tractable Probabilistic Circuits: A Systems Perspective
by: Liu, Anji, et al.
Published: (2024)
by: Liu, Anji, et al.
Published: (2024)
Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
by: Kim, Minseo, et al.
Published: (2025)
by: Kim, Minseo, et al.
Published: (2025)
LogicTree: Structured Proof Exploration for Coherent and Rigorous Logical Reasoning with Large Language Models
by: He, Kang, et al.
Published: (2025)
by: He, Kang, et al.
Published: (2025)
Circuit Complexity Bounds for Visual Autoregressive Model
by: Ke, Yekun, et al.
Published: (2025)
by: Ke, Yekun, et al.
Published: (2025)
Improving Autoregressive Training with Dynamic Oracles
by: Yang, Jianing, et al.
Published: (2024)
by: Yang, Jianing, et al.
Published: (2024)
From Instructions to Constraints: Language Model Alignment with Automatic Constraint Verification
by: Wang, Fei, et al.
Published: (2024)
by: Wang, Fei, et al.
Published: (2024)
Escaping Collapse: The Strength of Weak Data for Large Language Model Training
by: Amin, Kareem, et al.
Published: (2025)
by: Amin, Kareem, et al.
Published: (2025)
Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models
by: Wang, Fei, et al.
Published: (2024)
by: Wang, Fei, et al.
Published: (2024)
MixCE: Training Autoregressive Language Models by Mixing Forward and Reverse Cross-Entropies
by: Zhang, Shiyue, et al.
Published: (2023)
by: Zhang, Shiyue, et al.
Published: (2023)
Mechanism Shift During Post-training from Autoregressive to Masked Diffusion Language Models
by: Kong, Injin, et al.
Published: (2026)
by: Kong, Injin, et al.
Published: (2026)
ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
by: Lin, Bill Yuchen, et al.
Published: (2025)
by: Lin, Bill Yuchen, et al.
Published: (2025)
Similar Items
-
Enabling Autoregressive Models to Fill In Masked Tokens
by: Israel, Daniel, et al.
Published: (2025) -
Controllable Generation via Locally Constrained Resampling
by: Ahmed, Kareem, et al.
Published: (2024) -
Prepacking: A Simple Method for Fast Prefilling and Increased Throughput in Large Language Models
by: Zhao, Siyan, et al.
Published: (2024) -
Adversarial Tokenization
by: Geh, Renato Lui, et al.
Published: (2025) -
Accelerating Diffusion LLMs via Adaptive Parallel Decoding
by: Israel, Daniel, et al.
Published: (2025)