Gated Tree Cross-Attention for Checkpoint-Compatible Syntax Injection in Decoder-Only LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Gao, Xinyu, Wang, Shaonan, Ding, Nai |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Bitune: Leveraging Bidirectional Attention to Improve Decoder-Only LLMs
por: Kopiczko, Dawid J., et al.
Publicado: (2024)
por: Kopiczko, Dawid J., et al.
Publicado: (2024)
How Syntax Specialization Emerges in Language Models
por: Duan, Xufeng, et al.
Publicado: (2025)
por: Duan, Xufeng, et al.
Publicado: (2025)
How Do Decoder-Only LLMs Perceive Users? Rethinking Attention Masking for User Representation Learning
por: Yuan, Jiahao, et al.
Publicado: (2026)
por: Yuan, Jiahao, et al.
Publicado: (2026)
ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention
por: Liu, Wenjie, et al.
Publicado: (2026)
por: Liu, Wenjie, et al.
Publicado: (2026)
Memorization and Knowledge Injection in Gated LLMs
por: Pan, Xu, et al.
Publicado: (2025)
por: Pan, Xu, et al.
Publicado: (2025)
Cross-Attention Speculative Decoding
por: Zhong, Wei, et al.
Publicado: (2025)
por: Zhong, Wei, et al.
Publicado: (2025)
DOA: Training-Free Decoder-Only Attention Policy for Long-Form Simultaneous Translation with SpeechLLMs
por: Papi, Sara, et al.
Publicado: (2026)
por: Papi, Sara, et al.
Publicado: (2026)
Decoder-Only LLMs are Better Controllers for Diffusion Models
por: Dong, Ziyi, et al.
Publicado: (2025)
por: Dong, Ziyi, et al.
Publicado: (2025)
Improving In-context Learning of Multilingual Generative Language Models with Cross-lingual Alignment
por: Li, Chong, et al.
Publicado: (2023)
por: Li, Chong, et al.
Publicado: (2023)
Encoder-Decoder or Decoder-Only? Revisiting Encoder-Decoder Large Language Model
por: Zhang, Biao, et al.
Publicado: (2025)
por: Zhang, Biao, et al.
Publicado: (2025)
You Only Cache Once: Decoder-Decoder Architectures for Language Models
por: Sun, Yutao, et al.
Publicado: (2024)
por: Sun, Yutao, et al.
Publicado: (2024)
Don't Fine-Tune, Decode: Syntax Error-Free Tool Use via Constrained Decoding
por: Zhang, Kexun, et al.
Publicado: (2023)
por: Zhang, Kexun, et al.
Publicado: (2023)
Decoding the Multimodal Mind: Generalizable Brain-to-Text Translation via Multimodal Alignment and Adaptive Routing
por: Ye, Chunyu, et al.
Publicado: (2025)
por: Ye, Chunyu, et al.
Publicado: (2025)
A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation
por: Wang, Andrew Z., et al.
Publicado: (2025)
por: Wang, Andrew Z., et al.
Publicado: (2025)
Varying Shades of Wrong: Aligning LLMs with Wrong Answers Only
por: Yao, Jihan, et al.
Publicado: (2024)
por: Yao, Jihan, et al.
Publicado: (2024)
Is More Data Worth the Cost? Dataset Scaling Laws in a Tiny Attention-Only Decoder
por: Wiegand, Götz-Henrik, et al.
Publicado: (2026)
por: Wiegand, Götz-Henrik, et al.
Publicado: (2026)
From Syntax to Emotion: A Mechanistic Analysis of Emotion Inference in LLMs
por: Shu, Bangzhao, et al.
Publicado: (2026)
por: Shu, Bangzhao, et al.
Publicado: (2026)
SASST: Leveraging Syntax-Aware Chunking and LLMs for Simultaneous Speech Translation
por: Yang, Zeyu, et al.
Publicado: (2025)
por: Yang, Zeyu, et al.
Publicado: (2025)
On The Adaptation of Unlimiformer for Decoder-Only Transformers
por: Ahrabian, Kian, et al.
Publicado: (2024)
por: Ahrabian, Kian, et al.
Publicado: (2024)
In-context KV-Cache Eviction for LLMs via Attention-Gate
por: Zeng, Zihao, et al.
Publicado: (2024)
por: Zeng, Zihao, et al.
Publicado: (2024)
Prompt Decorators: A Declarative and Composable Syntax for Reasoning, Formatting, and Control in LLMs
por: Heris, Mostapha Kalami
Publicado: (2025)
por: Heris, Mostapha Kalami
Publicado: (2025)
For-Value: Efficient Forward-Only Data Valuation for finetuning LLMs and VLMs
por: Deng, Wenlong, et al.
Publicado: (2025)
por: Deng, Wenlong, et al.
Publicado: (2025)
EG-MLA: Embedding-Gated Multi-head Latent Attention for Scalable and Efficient LLMs
por: Cai, Zhengge, et al.
Publicado: (2025)
por: Cai, Zhengge, et al.
Publicado: (2025)
Sneaking Syntax into Transformer Language Models with Tree Regularization
por: Nandi, Ananjan, et al.
Publicado: (2024)
por: Nandi, Ananjan, et al.
Publicado: (2024)
Stepwise Reasoning Checkpoint Analysis: A Test Time Scaling Method to Enhance LLMs' Reasoning
por: Wang, Zezhong, et al.
Publicado: (2025)
por: Wang, Zezhong, et al.
Publicado: (2025)
OPT-Tree: Speculative Decoding with Adaptive Draft Tree Structure
por: Wang, Jikai, et al.
Publicado: (2024)
por: Wang, Jikai, et al.
Publicado: (2024)
Gender Disambiguation in Machine Translation: Diagnostic Evaluation in Decoder-Only Architectures
por: Manna, Chiara, et al.
Publicado: (2026)
por: Manna, Chiara, et al.
Publicado: (2026)
LLaMA based Punctuation Restoration With Forward Pass Only Decoding
por: Pang, Yutong, et al.
Publicado: (2024)
por: Pang, Yutong, et al.
Publicado: (2024)
How Powerful are Decoder-Only Transformer Neural Models?
por: Roberts, Jesse
Publicado: (2023)
por: Roberts, Jesse
Publicado: (2023)
You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories
por: Wei, Zhepei, et al.
Publicado: (2026)
por: Wei, Zhepei, et al.
Publicado: (2026)
Evaluation of LLMs on Syntax-Aware Code Fill-in-the-Middle Tasks
por: Gong, Linyuan, et al.
Publicado: (2024)
por: Gong, Linyuan, et al.
Publicado: (2024)
SyntaxShap: Syntax-aware Explainability Method for Text Generation
por: Amara, Kenza, et al.
Publicado: (2024)
por: Amara, Kenza, et al.
Publicado: (2024)
$\text{R}^2\text{R}$: A Route-to-Rerank Post-Training Framework for Multi-Domain Decoder-Only Rerankers
por: Wang, Xinyu, et al.
Publicado: (2025)
por: Wang, Xinyu, et al.
Publicado: (2025)
X-Instruction: Aligning Language Model in Low-resource Languages with Self-curated Cross-lingual Instructions
por: Li, Chong, et al.
Publicado: (2024)
por: Li, Chong, et al.
Publicado: (2024)
Formal Constraints on Dependency Syntax
por: Gómez-Rodríguez, et al.
Publicado: (2026)
por: Gómez-Rodríguez, et al.
Publicado: (2026)
Infusing Prompts with Syntax and Semantics
por: Labate, Anton Bulle, et al.
Publicado: (2024)
por: Labate, Anton Bulle, et al.
Publicado: (2024)
Morphology and Syntax of the Tamil Language
por: Sarveswaran, Kengatharaiyer
Publicado: (2024)
por: Sarveswaran, Kengatharaiyer
Publicado: (2024)
Multilingual Pretraining and Instruction Tuning Improve Cross-Lingual Knowledge Alignment, But Only Shallowly
por: Gao, Changjiang, et al.
Publicado: (2024)
por: Gao, Changjiang, et al.
Publicado: (2024)
Evaluating Small Decoder-Only Language Models for Grammar Correction and Text Simplification
por: Lamelas, Anthony
Publicado: (2026)
por: Lamelas, Anthony
Publicado: (2026)
Gated Slot Attention for Efficient Linear-Time Sequence Modeling
por: Zhang, Yu, et al.
Publicado: (2024)
por: Zhang, Yu, et al.
Publicado: (2024)
Ejemplares similares
-
Bitune: Leveraging Bidirectional Attention to Improve Decoder-Only LLMs
por: Kopiczko, Dawid J., et al.
Publicado: (2024) -
How Syntax Specialization Emerges in Language Models
por: Duan, Xufeng, et al.
Publicado: (2025) -
How Do Decoder-Only LLMs Perceive Users? Rethinking Attention Masking for User Representation Learning
por: Yuan, Jiahao, et al.
Publicado: (2026) -
ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention
por: Liu, Wenjie, et al.
Publicado: (2026) -
Memorization and Knowledge Injection in Gated LLMs
por: Pan, Xu, et al.
Publicado: (2025)