CALYREX: Cross-Attention LaYeR EXtended Transformers for System Prompt Anchoring
Fuente:
arXiv
Salvato in:
| Autore principale: | Lixing, Li |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Dynamically Anchored Prompting for Task-Imbalanced Continual Learning
di: Hong, Chenxing, et al.
Pubblicazione: (2024)
di: Hong, Chenxing, et al.
Pubblicazione: (2024)
ASAP: Attention Sink Anchored Pruning
di: Lee, Jaehyuk, et al.
Pubblicazione: (2026)
di: Lee, Jaehyuk, et al.
Pubblicazione: (2026)
DeepCrossAttention: Supercharging Transformer Residual Connections
di: Heddes, Mike, et al.
Pubblicazione: (2025)
di: Heddes, Mike, et al.
Pubblicazione: (2025)
Selective Prompt Anchoring for Code Generation
di: Tian, Yuan, et al.
Pubblicazione: (2024)
di: Tian, Yuan, et al.
Pubblicazione: (2024)
RFSeek and Ye Shall Find
di: Rotman, Noga H., et al.
Pubblicazione: (2025)
di: Rotman, Noga H., et al.
Pubblicazione: (2025)
Evaluating the Architectural Reasoning Capabilities of LLM Provers via the Obfuscated Natural Number Game
di: Li, Lixing
Pubblicazione: (2026)
di: Li, Lixing
Pubblicazione: (2026)
Online Bandits with (Biased) Offline Data: Adaptive Learning under Distribution Mismatch
di: Cheung, Wang Chi, et al.
Pubblicazione: (2024)
di: Cheung, Wang Chi, et al.
Pubblicazione: (2024)
CVTGAD: Simplified Transformer with Cross-View Attention for Unsupervised Graph-level Anomaly Detection
di: Li, Jindong, et al.
Pubblicazione: (2024)
di: Li, Jindong, et al.
Pubblicazione: (2024)
PACT: Peak-Aware Cross-Attention Graph Transformers for Efficient Storm-Surge Emulation
di: Liu, Zesheng, et al.
Pubblicazione: (2026)
di: Liu, Zesheng, et al.
Pubblicazione: (2026)
CAAT-EHR: Cross-Attentional Autoregressive Transformer for Multimodal Electronic Health Record Embeddings
di: Olaimat, Mohammad Al, et al.
Pubblicazione: (2025)
di: Olaimat, Mohammad Al, et al.
Pubblicazione: (2025)
Tree Cross Attention
di: Feng, Leo, et al.
Pubblicazione: (2023)
di: Feng, Leo, et al.
Pubblicazione: (2023)
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
di: Brandon, William, et al.
Pubblicazione: (2024)
di: Brandon, William, et al.
Pubblicazione: (2024)
S2TX: Cross-Attention Multi-Scale State-Space Transformer for Time Series Forecasting
di: Wu, Zihao, et al.
Pubblicazione: (2025)
di: Wu, Zihao, et al.
Pubblicazione: (2025)
The Effect of Attention Head Count on Transformer Approximation
di: Yu, Penghao, et al.
Pubblicazione: (2025)
di: Yu, Penghao, et al.
Pubblicazione: (2025)
CroSTAta: Cross-State Transition Attention Transformer for Robotic Manipulation
di: Minelli, Giovanni, et al.
Pubblicazione: (2025)
di: Minelli, Giovanni, et al.
Pubblicazione: (2025)
Adaptive Physics Transformer with Fused Global-Local Attention for Subsurface Energy Systems
di: Ju, Xin, et al.
Pubblicazione: (2026)
di: Ju, Xin, et al.
Pubblicazione: (2026)
Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts
di: Liang, Feng, et al.
Pubblicazione: (2025)
di: Liang, Feng, et al.
Pubblicazione: (2025)
Memory Limitations of Prompt Tuning in Transformers
di: Meyer, Maxime, et al.
Pubblicazione: (2025)
di: Meyer, Maxime, et al.
Pubblicazione: (2025)
SFi-Former: Sparse Flow Induced Attention for Graph Transformer
di: Li, Zhonghao, et al.
Pubblicazione: (2025)
di: Li, Zhonghao, et al.
Pubblicazione: (2025)
MMCAformer: Macro-Micro Cross-Attention Transformer for Traffic Speed Prediction with Microscopic Connected Vehicle Driving Behavior
di: Han, Lei, et al.
Pubblicazione: (2026)
di: Han, Lei, et al.
Pubblicazione: (2026)
Efficiently Solving Discounted MDPs with Predictions on Transition Matrices
di: Lyu, Lixing, et al.
Pubblicazione: (2025)
di: Lyu, Lixing, et al.
Pubblicazione: (2025)
A Theoretical Framework for Prompt Engineering: Approximating Smooth Functions with Transformer Prompts
di: Nakada, Ryumei, et al.
Pubblicazione: (2025)
di: Nakada, Ryumei, et al.
Pubblicazione: (2025)
On Zero-Initialized Attention: Optimal Prompt and Gating Factor Estimation
di: Diep, Nghiem T., et al.
Pubblicazione: (2025)
di: Diep, Nghiem T., et al.
Pubblicazione: (2025)
Residual Cross-Attention Transformer-Based Multi-User CSI Feedback with Deep Joint Source-Channel Coding
di: Zhang, Hengwei, et al.
Pubblicazione: (2025)
di: Zhang, Hengwei, et al.
Pubblicazione: (2025)
Quantifying Cross-Attention Interaction in Transformers for Interpreting TCR-pMHC Binding
di: Li, Jiarui, et al.
Pubblicazione: (2025)
di: Li, Jiarui, et al.
Pubblicazione: (2025)
Adaptive Prompt Tuning: Vision Guided Prompt Tuning with Cross-Attention for Fine-Grained Few-Shot Learning
di: Brouwer, Eric, et al.
Pubblicazione: (2024)
di: Brouwer, Eric, et al.
Pubblicazione: (2024)
Preconditioned Attention: Enhancing Efficiency in Transformers
di: Saratchandran, Hemanth
Pubblicazione: (2026)
di: Saratchandran, Hemanth
Pubblicazione: (2026)
Cottention: Linear Transformers With Cosine Attention
di: Mongaras, Gabriel, et al.
Pubblicazione: (2024)
di: Mongaras, Gabriel, et al.
Pubblicazione: (2024)
Transformers with Sparse Attention for Granger Causality
di: Mahesh, Riya, et al.
Pubblicazione: (2024)
di: Mahesh, Riya, et al.
Pubblicazione: (2024)
Transformer Reconstructed with Dynamic Value Attention
di: Wang, Xiaowei
Pubblicazione: (2025)
di: Wang, Xiaowei
Pubblicazione: (2025)
Graph External Attention Enhanced Transformer
di: Liang, Jianqing, et al.
Pubblicazione: (2024)
di: Liang, Jianqing, et al.
Pubblicazione: (2024)
Speaker Style-Aware Phoneme Anchoring for Improved Cross-Lingual Speech Emotion Recognition
di: Upadhyay, Shreya G., et al.
Pubblicazione: (2025)
di: Upadhyay, Shreya G., et al.
Pubblicazione: (2025)
Anchored Langevin Algorithms
di: Gurbuzbalaban, Mert, et al.
Pubblicazione: (2025)
di: Gurbuzbalaban, Mert, et al.
Pubblicazione: (2025)
EfficientECG: Cross-Attention with Feature Fusion for Efficient Electrocardiogram Classification
di: Deng, Hanhui, et al.
Pubblicazione: (2025)
di: Deng, Hanhui, et al.
Pubblicazione: (2025)
Anchored Alignment: Preventing Positional Collapse in Multimodal Recommender Systems
di: Jeong, Yonghun, et al.
Pubblicazione: (2026)
di: Jeong, Yonghun, et al.
Pubblicazione: (2026)
Fusion Matrix Prompt Enhanced Self-Attention Spatial-Temporal Interactive Traffic Forecasting Framework
di: Liu, Mu, et al.
Pubblicazione: (2024)
di: Liu, Mu, et al.
Pubblicazione: (2024)
Molecule Design by Latent Prompt Transformer
di: Kong, Deqian, et al.
Pubblicazione: (2023)
di: Kong, Deqian, et al.
Pubblicazione: (2023)
Molecule Design by Latent Prompt Transformer
di: Kong, Deqian, et al.
Pubblicazione: (2024)
di: Kong, Deqian, et al.
Pubblicazione: (2024)
Masked Diffusion Modeling for Anomaly Detection
di: Zhang, Lixing, et al.
Pubblicazione: (2026)
di: Zhang, Lixing, et al.
Pubblicazione: (2026)
Distributed quasi-Newton robust estimation under differential privacy
di: Wang, Chuhan, et al.
Pubblicazione: (2024)
di: Wang, Chuhan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Dynamically Anchored Prompting for Task-Imbalanced Continual Learning
di: Hong, Chenxing, et al.
Pubblicazione: (2024) -
ASAP: Attention Sink Anchored Pruning
di: Lee, Jaehyuk, et al.
Pubblicazione: (2026) -
DeepCrossAttention: Supercharging Transformer Residual Connections
di: Heddes, Mike, et al.
Pubblicazione: (2025) -
Selective Prompt Anchoring for Code Generation
di: Tian, Yuan, et al.
Pubblicazione: (2024) -
RFSeek and Ye Shall Find
di: Rotman, Noga H., et al.
Pubblicazione: (2025)