Guardado en:
| Autor principal: | Corielli, Francesco |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2605.26711 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
When Is Next-Token Prediction Useful? Marginalization, Ergodicity, Mixture Identifiability, Local Sufficiency, RAG, Tools, and Programming
por: Corielli, Francesco
Publicado: (2026)
por: Corielli, Francesco
Publicado: (2026)
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
por: Dong, Yihe, et al.
Publicado: (2025)
por: Dong, Yihe, et al.
Publicado: (2025)
No Need to Talk: Asynchronous Mixture of Language Models
por: Filippova, Anastasiia, et al.
Publicado: (2024)
por: Filippova, Anastasiia, et al.
Publicado: (2024)
Meta-DiffuB: A Contextualized Sequence-to-Sequence Text Diffusion Model with Meta-Exploration
por: Chuang, Yun-Yen, et al.
Publicado: (2024)
por: Chuang, Yun-Yen, et al.
Publicado: (2024)
Neural Sequence-to-Sequence Modeling with Attention by Leveraging Deep Learning Architectures for Enhanced Contextual Understanding in Abstractive Text Summarization
por: Challagundla, Bhavith Chandra, et al.
Publicado: (2024)
por: Challagundla, Bhavith Chandra, et al.
Publicado: (2024)
Is In-Context Learning Sufficient for Instruction Following in LLMs?
por: Zhao, Hao, et al.
Publicado: (2024)
por: Zhao, Hao, et al.
Publicado: (2024)
MoM: Linear Sequence Modeling with Mixture-of-Memories
por: Du, Jusen, et al.
Publicado: (2025)
por: Du, Jusen, et al.
Publicado: (2025)
Necessary and Sufficient Watermark for Large Language Models
por: Takezawa, Yuki, et al.
Publicado: (2023)
por: Takezawa, Yuki, et al.
Publicado: (2023)
Do We Need Frontier Models to Verify Mathematical Proofs?
por: Naik, Aaditya, et al.
Publicado: (2026)
por: Naik, Aaditya, et al.
Publicado: (2026)
Uncertainty as Feature Gaps: Epistemic Uncertainty Quantification of LLMs in Contextual Question-Answering
por: Bakman, Yavuz, et al.
Publicado: (2025)
por: Bakman, Yavuz, et al.
Publicado: (2025)
ProcessBench: Identifying Process Errors in Mathematical Reasoning
por: Zheng, Chujie, et al.
Publicado: (2024)
por: Zheng, Chujie, et al.
Publicado: (2024)
Outcome-Grounded Advantage Reshaping for Fine-Grained Credit Assignment in Mathematical Reasoning
por: Li, Ziheng, et al.
Publicado: (2026)
por: Li, Ziheng, et al.
Publicado: (2026)
DRA-GRPO: Your GRPO Needs to Know Diverse Reasoning Paths for Mathematical Reasoning
por: Chen, Xiwen, et al.
Publicado: (2025)
por: Chen, Xiwen, et al.
Publicado: (2025)
Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal Verification
por: Liu, Chengwu, et al.
Publicado: (2025)
por: Liu, Chengwu, et al.
Publicado: (2025)
Identifying Factual Inconsistencies in Summaries: Grounding LLM Inference via Task Taxonomy
por: Xu, Liyan, et al.
Publicado: (2024)
por: Xu, Liyan, et al.
Publicado: (2024)
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
por: Sun, Weigao, et al.
Publicado: (2025)
por: Sun, Weigao, et al.
Publicado: (2025)
Is Next Token Prediction Sufficient for GPT? Exploration on Code Logic Comprehension
por: Qi, Mengnan, et al.
Publicado: (2024)
por: Qi, Mengnan, et al.
Publicado: (2024)
Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models
por: Belenki, Lior, et al.
Publicado: (2025)
por: Belenki, Lior, et al.
Publicado: (2025)
A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models
por: Kazemi, Hamid, et al.
Publicado: (2026)
por: Kazemi, Hamid, et al.
Publicado: (2026)
Contextual Text Denoising with Masked Language Models
por: Sun, Yifu, et al.
Publicado: (2019)
por: Sun, Yifu, et al.
Publicado: (2019)
Solving Situation Puzzles with Large Language Model and External Reformulation
por: Li, Kun, et al.
Publicado: (2025)
por: Li, Kun, et al.
Publicado: (2025)
AWARE, Beyond Sentence Boundaries: A Contextual Transformer Framework for Identifying Cultural Capital in STEM Narratives
por: Khan, Khalid Mehtab, et al.
Publicado: (2025)
por: Khan, Khalid Mehtab, et al.
Publicado: (2025)
A Survey on Mixture of Experts in Large Language Models
por: Cai, Weilin, et al.
Publicado: (2024)
por: Cai, Weilin, et al.
Publicado: (2024)
CEMTM: Contextual Embedding-based Multimodal Topic Modeling
por: Abaskohi, Amirhossein, et al.
Publicado: (2025)
por: Abaskohi, Amirhossein, et al.
Publicado: (2025)
$K$-MSHC: Unmasking Minimally Sufficient Head Circuits in Large Language Models with Experiments on Syntactic Classification Tasks
por: Chowdhary, Pratim, et al.
Publicado: (2025)
por: Chowdhary, Pratim, et al.
Publicado: (2025)
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective
por: He, Shenghua, et al.
Publicado: (2025)
por: He, Shenghua, et al.
Publicado: (2025)
On the "Induction Bias" in Sequence Models
por: Ebrahimi, M. Reza, et al.
Publicado: (2026)
por: Ebrahimi, M. Reza, et al.
Publicado: (2026)
A Closer Look into Mixture-of-Experts in Large Language Models
por: Lo, Ka Man, et al.
Publicado: (2024)
por: Lo, Ka Man, et al.
Publicado: (2024)
Estimating Privacy Leakage of Augmented Contextual Knowledge in Language Models
por: Flemings, James, et al.
Publicado: (2024)
por: Flemings, James, et al.
Publicado: (2024)
CATS: Contextually-Aware Thresholding for Sparsity in Large Language Models
por: Lee, Donghyun, et al.
Publicado: (2024)
por: Lee, Donghyun, et al.
Publicado: (2024)
FormalProofBench: Can Models Write Graduate Level Math Proofs That Are Formally Verified?
por: Ravi, Nikil, et al.
Publicado: (2026)
por: Ravi, Nikil, et al.
Publicado: (2026)
A Formal Framework for Uncertainty Analysis of Text Generation with Large Language Models
por: Herbold, Steffen, et al.
Publicado: (2026)
por: Herbold, Steffen, et al.
Publicado: (2026)
Bayesian Mixture of Experts For Large Language Models
por: Dialameh, Maryam, et al.
Publicado: (2025)
por: Dialameh, Maryam, et al.
Publicado: (2025)
Mixture-of-Personas Language Models for Population Simulation
por: Bui, Ngoc, et al.
Publicado: (2025)
por: Bui, Ngoc, et al.
Publicado: (2025)
Learning to Reason without External Rewards
por: Zhao, Xuandong, et al.
Publicado: (2025)
por: Zhao, Xuandong, et al.
Publicado: (2025)
On the Thinking-Language Modeling Gap in Large Language Models
por: Liu, Chenxi, et al.
Publicado: (2025)
por: Liu, Chenxi, et al.
Publicado: (2025)
Objective Metrics for Evaluating Large Language Models Using External Data Sources
por: Du, Haoze, et al.
Publicado: (2025)
por: Du, Haoze, et al.
Publicado: (2025)
Transforming Chatbot Text: A Sequence-to-Sequence Approach
por: Reddy, Natesh, et al.
Publicado: (2025)
por: Reddy, Natesh, et al.
Publicado: (2025)
VEHME: A Vision-Language Model For Evaluating Handwritten Mathematics Expressions
por: Nguyen, Thu Phuong, et al.
Publicado: (2025)
por: Nguyen, Thu Phuong, et al.
Publicado: (2025)
Mixture Compressor for Mixture-of-Experts LLMs Gains More
por: Huang, Wei, et al.
Publicado: (2024)
por: Huang, Wei, et al.
Publicado: (2024)
Ejemplares similares
-
When Is Next-Token Prediction Useful? Marginalization, Ergodicity, Mixture Identifiability, Local Sufficiency, RAG, Tools, and Programming
por: Corielli, Francesco
Publicado: (2026) -
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
por: Dong, Yihe, et al.
Publicado: (2025) -
No Need to Talk: Asynchronous Mixture of Language Models
por: Filippova, Anastasiia, et al.
Publicado: (2024) -
Meta-DiffuB: A Contextualized Sequence-to-Sequence Text Diffusion Model with Meta-Exploration
por: Chuang, Yun-Yen, et al.
Publicado: (2024) -
Neural Sequence-to-Sequence Modeling with Attention by Leveraging Deep Learning Architectures for Enhanced Contextual Understanding in Abstractive Text Summarization
por: Challagundla, Bhavith Chandra, et al.
Publicado: (2024)