What Is the Minimum Architecture for Prolepsis? Early Irrevocable Commitment Across Tasks in Small Transformers
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Jacopin, Éric |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Co-Located Tests, Better AI Code: How Test Syntax Structure Affects Foundation Model Code Generation
von: Jacopin, Éric
Veröffentlicht: (2026)
von: Jacopin, Éric
Veröffentlicht: (2026)
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
von: Ramesh, Shyam Sundhar, et al.
Veröffentlicht: (2026)
von: Ramesh, Shyam Sundhar, et al.
Veröffentlicht: (2026)
The Geometries of Truth Are Orthogonal Across Tasks
von: Azizian, Waiss, et al.
Veröffentlicht: (2025)
von: Azizian, Waiss, et al.
Veröffentlicht: (2025)
Hallucination as Trajectory Commitment: Causal Evidence for Asymmetric Attractor Dynamics in Transformer Generation
von: Akarlar, G. Aytug
Veröffentlicht: (2026)
von: Akarlar, G. Aytug
Veröffentlicht: (2026)
When Does a Language Model Commit? A Finite-Answer Theory of Pre-Verbalization Commitment
von: Zhang, Long, et al.
Veröffentlicht: (2026)
von: Zhang, Long, et al.
Veröffentlicht: (2026)
Hymba: A Hybrid-head Architecture for Small Language Models
von: Dong, Xin, et al.
Veröffentlicht: (2024)
von: Dong, Xin, et al.
Veröffentlicht: (2024)
Residual Stream Duality in Modern Transformer Architectures
von: Zhang, Yifan
Veröffentlicht: (2026)
von: Zhang, Yifan
Veröffentlicht: (2026)
Supernova: Achieving More with Less in Transformer Architectures
von: Tanase, Andrei-Valentin, et al.
Veröffentlicht: (2025)
von: Tanase, Andrei-Valentin, et al.
Veröffentlicht: (2025)
Dependency Parsing Across the Resource Spectrum: Evaluating Architectures on High and Low-Resource Languages
von: Guan, Kevin, et al.
Veröffentlicht: (2026)
von: Guan, Kevin, et al.
Veröffentlicht: (2026)
What Matters in Transformers? Not All Attention is Needed
von: He, Shwai, et al.
Veröffentlicht: (2024)
von: He, Shwai, et al.
Veröffentlicht: (2024)
Gumiho: A Hybrid Architecture to Prioritize Early Tokens in Speculative Decoding
von: Li, Jinze, et al.
Veröffentlicht: (2025)
von: Li, Jinze, et al.
Veröffentlicht: (2025)
Multilingual Prompt Engineering in Large Language Models: A Survey Across NLP Tasks
von: Vatsal, Shubham, et al.
Veröffentlicht: (2025)
von: Vatsal, Shubham, et al.
Veröffentlicht: (2025)
Peri-LN: Revisiting Normalization Layer in the Transformer Architecture
von: Kim, Jeonghoon, et al.
Veröffentlicht: (2025)
von: Kim, Jeonghoon, et al.
Veröffentlicht: (2025)
VocabTailor: Dynamic Vocabulary Selection for Downstream Tasks in Small Language Models
von: Zhang, Hanling, et al.
Veröffentlicht: (2025)
von: Zhang, Hanling, et al.
Veröffentlicht: (2025)
The Dual-Stream Transformer: Channelized Architecture for Interpretable Language Modeling
von: Kerce, J. Clayton, et al.
Veröffentlicht: (2026)
von: Kerce, J. Clayton, et al.
Veröffentlicht: (2026)
Cross-Architecture Transfer Learning for Linear-Cost Inference Transformers
von: Choi, Sehyun
Veröffentlicht: (2024)
von: Choi, Sehyun
Veröffentlicht: (2024)
Uncertainty-Aware Decoding with Minimum Bayes Risk
von: Daheim, Nico, et al.
Veröffentlicht: (2025)
von: Daheim, Nico, et al.
Veröffentlicht: (2025)
Depth-Width tradeoffs in Algorithmic Reasoning of Graph Tasks with Transformers
von: Yehudai, Gilad, et al.
Veröffentlicht: (2025)
von: Yehudai, Gilad, et al.
Veröffentlicht: (2025)
Agreement-Constrained Probabilistic Minimum Bayes Risk Decoding
von: Natsumi, Koki, et al.
Veröffentlicht: (2025)
von: Natsumi, Koki, et al.
Veröffentlicht: (2025)
Decoupling Knowledge and Reasoning in Transformers: A Modular Architecture with Generalized Cross-Attention
von: Guo, Zhenyu, et al.
Veröffentlicht: (2025)
von: Guo, Zhenyu, et al.
Veröffentlicht: (2025)
Multiscale Byte Language Models -- A Hierarchical Architecture for Causal Million-Length Sequence Modeling
von: Egli, Eric, et al.
Veröffentlicht: (2025)
von: Egli, Eric, et al.
Veröffentlicht: (2025)
KUET at StanceNakba Shared Task: StanceMoE: Mixture-of-Experts Architecture for Stance Detection
von: Shafi, Abdullah Al, et al.
Veröffentlicht: (2026)
von: Shafi, Abdullah Al, et al.
Veröffentlicht: (2026)
Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure
von: Cho, Hanseul, et al.
Veröffentlicht: (2024)
von: Cho, Hanseul, et al.
Veröffentlicht: (2024)
What Differentiates Educational Literature? A Multimodal Fusion Approach of Transformers and Computational Linguistics
von: Bird, Jordan J.
Veröffentlicht: (2024)
von: Bird, Jordan J.
Veröffentlicht: (2024)
Pre-training a Transformer-Based Generative Model Using a Small Sepedi Dataset
von: Ramalepe, Simon P., et al.
Veröffentlicht: (2025)
von: Ramalepe, Simon P., et al.
Veröffentlicht: (2025)
Unified Multi-Task Learning & Model Fusion for Efficient Language Model Guardrailing
von: Neill, James O', et al.
Veröffentlicht: (2025)
von: Neill, James O', et al.
Veröffentlicht: (2025)
RadLite: Multi-Task LoRA Fine-Tuning of Small Language Models for CPU-Deployable Radiology AI
von: Gupta, Pankaj, et al.
Veröffentlicht: (2026)
von: Gupta, Pankaj, et al.
Veröffentlicht: (2026)
Detecting Clinical Discrepancies in Health Coaching Agents: A Dual-Stream Memory and Reconciliation Architecture
von: Pugh, Samuel L, et al.
Veröffentlicht: (2026)
von: Pugh, Samuel L, et al.
Veröffentlicht: (2026)
Circuit Complexity Bounds for RoPE-based Transformer Architecture
von: Chen, Bo, et al.
Veröffentlicht: (2024)
von: Chen, Bo, et al.
Veröffentlicht: (2024)
Chasing COMET: Leveraging Minimum Bayes Risk Decoding for Self-Improving Machine Translation
von: Guttmann, Kamil, et al.
Veröffentlicht: (2024)
von: Guttmann, Kamil, et al.
Veröffentlicht: (2024)
Shared Lexical Task Representations Explain Behavioral Variability In LLMs
von: Yang, Zhuonan, et al.
Veröffentlicht: (2026)
von: Yang, Zhuonan, et al.
Veröffentlicht: (2026)
Learning to Rank Chain-of-Thought: Using a Small Model
von: Jiang, Eric Hanchen, et al.
Veröffentlicht: (2025)
von: Jiang, Eric Hanchen, et al.
Veröffentlicht: (2025)
Minimum Bayes Risk Decoding for Error Span Detection in Reference-Free Automatic Machine Translation Evaluation
von: Lyu, Boxuan, et al.
Veröffentlicht: (2025)
von: Lyu, Boxuan, et al.
Veröffentlicht: (2025)
A Transformer-based Neural Architecture Search Method
von: Wang, Shang, et al.
Veröffentlicht: (2025)
von: Wang, Shang, et al.
Veröffentlicht: (2025)
One Jump Is All You Need: Short-Cutting Transformers for Early Exit Prediction with One Jump to Fit All Exit Levels
von: Seshadri, Amrit Diggavi
Veröffentlicht: (2025)
von: Seshadri, Amrit Diggavi
Veröffentlicht: (2025)
Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM Reasoning
von: Parashar, Shubham, et al.
Veröffentlicht: (2025)
von: Parashar, Shubham, et al.
Veröffentlicht: (2025)
Synthetic Lyrics Detection Across Languages and Genres
von: Labrak, Yanis, et al.
Veröffentlicht: (2024)
von: Labrak, Yanis, et al.
Veröffentlicht: (2024)
Scaling Optimal LR Across Token Horizons
von: Bjorck, Johan, et al.
Veröffentlicht: (2024)
von: Bjorck, Johan, et al.
Veröffentlicht: (2024)
Long Chain-of-Thought Reasoning Across Languages
von: Barua, Josh, et al.
Veröffentlicht: (2025)
von: Barua, Josh, et al.
Veröffentlicht: (2025)
1024m at SMM4H 2024: Tasks 3, 5 & 6 -- Ensembles of Transformers and Large Language Models for Medical Text Classification
von: Kadiyala, Ram Mohan Rao, et al.
Veröffentlicht: (2024)
von: Kadiyala, Ram Mohan Rao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Co-Located Tests, Better AI Code: How Test Syntax Structure Affects Foundation Model Code Generation
von: Jacopin, Éric
Veröffentlicht: (2026) -
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
von: Ramesh, Shyam Sundhar, et al.
Veröffentlicht: (2026) -
The Geometries of Truth Are Orthogonal Across Tasks
von: Azizian, Waiss, et al.
Veröffentlicht: (2025) -
Hallucination as Trajectory Commitment: Causal Evidence for Asymmetric Attractor Dynamics in Transformer Generation
von: Akarlar, G. Aytug
Veröffentlicht: (2026) -
When Does a Language Model Commit? A Finite-Answer Theory of Pre-Verbalization Commitment
von: Zhang, Long, et al.
Veröffentlicht: (2026)