On The Adaptation of Unlimiformer for Decoder-Only Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Ahrabian, Kian, Benhaim, Alon, Patra, Barun, Pujara, Jay, Singhal, Saksham, Song, Xia |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Practical Analysis of Human Alignment with *PO
by: Ahrabian, Kian, et al.
Published: (2024)
by: Ahrabian, Kian, et al.
Published: (2024)
Toward Better Temporal Structures for Geopolitical Events Forecasting
by: Ahrabian, Kian, et al.
Published: (2026)
by: Ahrabian, Kian, et al.
Published: (2026)
Scaling Laws for Multilingual Language Models
by: He, Yifei, et al.
Published: (2024)
by: He, Yifei, et al.
Published: (2024)
Escaping the Mode Lottery: Multi-Response Training Improves Language Model Generalization
by: Amin, Hasan, et al.
Published: (2026)
by: Amin, Hasan, et al.
Published: (2026)
Scaling Optimal LR Across Token Horizons
by: Bjorck, Johan, et al.
Published: (2024)
by: Bjorck, Johan, et al.
Published: (2024)
A Systematic Analysis of Base Model Choice for Reward Modeling
by: Ahrabian, Kian, et al.
Published: (2025)
by: Ahrabian, Kian, et al.
Published: (2025)
How Powerful are Decoder-Only Transformer Neural Models?
by: Roberts, Jesse
Published: (2023)
by: Roberts, Jesse
Published: (2023)
The Curious Case of Nonverbal Abstract Reasoning with Multi-Modal Large Language Models
by: Ahrabian, Kian, et al.
Published: (2024)
by: Ahrabian, Kian, et al.
Published: (2024)
MARVEL: Multidimensional Abstraction and Reasoning through Visual Evaluation and Learning
by: Jiang, Yifan, et al.
Published: (2024)
by: Jiang, Yifan, et al.
Published: (2024)
WhisTLE: Deeply Supervised, Text-Only Domain Adaptation for Pretrained Speech Recognition Transformers
by: Pandey, Akshat, et al.
Published: (2025)
by: Pandey, Akshat, et al.
Published: (2025)
Reject Only Critical Tokens: Pivot-Aware Speculative Decoding
by: Ziashahabi, Amir, et al.
Published: (2025)
by: Ziashahabi, Amir, et al.
Published: (2025)
Causal Reasoning Favors Encoders: On The Limits of Decoder-Only Models
by: Roy, Amartya, et al.
Published: (2025)
by: Roy, Amartya, et al.
Published: (2025)
Machine Translation with Large Language Models: Decoder Only vs. Encoder-Decoder
by: M., Abhinav P., et al.
Published: (2024)
by: M., Abhinav P., et al.
Published: (2024)
RED QUEEN: Safeguarding Large Language Models against Concealed Multi-Turn Jailbreaking
by: Jiang, Yifan, et al.
Published: (2024)
by: Jiang, Yifan, et al.
Published: (2024)
Integrating Pre-Trained Language Model with Physical Layer Communications
by: Lee, Ju-Hyung, et al.
Published: (2024)
by: Lee, Ju-Hyung, et al.
Published: (2024)
Is More Data Worth the Cost? Dataset Scaling Laws in a Tiny Attention-Only Decoder
by: Wiegand, Götz-Henrik, et al.
Published: (2026)
by: Wiegand, Götz-Henrik, et al.
Published: (2026)
Small LLMs with Expert Blocks Are Good Enough for Hyperparamter Tuning
by: Naphade, Om, et al.
Published: (2025)
by: Naphade, Om, et al.
Published: (2025)
Decoding Partial Differential Equations: Cross-Modal Adaptation of Decoder-only Models to PDEs
by: García-de-Herreros, Paloma, et al.
Published: (2025)
by: García-de-Herreros, Paloma, et al.
Published: (2025)
REFRAG: Rethinking RAG based Decoding
by: Lin, Xiaoqiang, et al.
Published: (2025)
by: Lin, Xiaoqiang, et al.
Published: (2025)
Revisiting the Robustness of Watermarking to Paraphrasing Attacks
by: Rastogi, Saksham, et al.
Published: (2024)
by: Rastogi, Saksham, et al.
Published: (2024)
FAAST: Forward-Only Associative Learning via Closed-Form Fast Weights for Test-Time Supervised Adaptation
by: Bao, Guangsheng, et al.
Published: (2026)
by: Bao, Guangsheng, et al.
Published: (2026)
Encoder-Decoder Gemma: Improving the Quality-Efficiency Trade-Off via Adaptation
by: Zhang, Biao, et al.
Published: (2025)
by: Zhang, Biao, et al.
Published: (2025)
LayerNorm Induces Recency Bias in Transformer Decoders
by: Kim, Junu, et al.
Published: (2025)
by: Kim, Junu, et al.
Published: (2025)
Dissociating Decodability and Causal Use in Bracket-Sequence Transformers
by: Sharma, Aryan, et al.
Published: (2026)
by: Sharma, Aryan, et al.
Published: (2026)
A Glitch in the Matrix? Locating and Detecting Language Model Grounding with Fakepedia
by: Monea, Giovanni, et al.
Published: (2023)
by: Monea, Giovanni, et al.
Published: (2023)
STAMP Your Content: Proving Dataset Membership via Watermarked Rephrasings
by: Rastogi, Saksham, et al.
Published: (2025)
by: Rastogi, Saksham, et al.
Published: (2025)
Impact of Layer Norm on Memorization and Generalization in Transformers
by: Singhal, Rishi, et al.
Published: (2025)
by: Singhal, Rishi, et al.
Published: (2025)
Faster Transformer Decoding: N-gram Masked Self-Attention
by: Chelba, Ciprian, et al.
Published: (2020)
by: Chelba, Ciprian, et al.
Published: (2020)
Contextually Guided Transformers via Low-Rank Adaptation
by: Zhmoginov, Andrey, et al.
Published: (2025)
by: Zhmoginov, Andrey, et al.
Published: (2025)
BitDelta: Your Fine-Tune May Only Be Worth One Bit
by: Liu, James, et al.
Published: (2024)
by: Liu, James, et al.
Published: (2024)
Trainable Transformer in Transformer
by: Panigrahi, Abhishek, et al.
Published: (2023)
by: Panigrahi, Abhishek, et al.
Published: (2023)
Emergence of Minimal Circuits for Indirect Object Identification in Attention-Only Transformers
by: Adhikari, Rabin
Published: (2025)
by: Adhikari, Rabin
Published: (2025)
Transformers for Program Termination
by: Alon, Yoav, et al.
Published: (2026)
by: Alon, Yoav, et al.
Published: (2026)
A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation
by: Wang, Andrew Z., et al.
Published: (2025)
by: Wang, Andrew Z., et al.
Published: (2025)
Decoding Speculative Decoding
by: Yan, Minghao, et al.
Published: (2024)
by: Yan, Minghao, et al.
Published: (2024)
Sparse Transformer with Local and Seasonal Adaptation for Multivariate Time Series Forecasting
by: Zhang, Yifan, et al.
Published: (2023)
by: Zhang, Yifan, et al.
Published: (2023)
MAAT: Multi-phase Adapter-Aware Targeted Unlearning
by: Yagnik, Suryash, et al.
Published: (2026)
by: Yagnik, Suryash, et al.
Published: (2026)
Decoding-based Regression
by: Song, Xingyou, et al.
Published: (2025)
by: Song, Xingyou, et al.
Published: (2025)
Test-Time Adaptation via Many-Shot Prompting: Benefits, Limits, and Pitfalls
by: Upasani, Shubhangi, et al.
Published: (2026)
by: Upasani, Shubhangi, et al.
Published: (2026)
PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration
by: Wu, Songhao, et al.
Published: (2025)
by: Wu, Songhao, et al.
Published: (2025)
Similar Items
-
A Practical Analysis of Human Alignment with *PO
by: Ahrabian, Kian, et al.
Published: (2024) -
Toward Better Temporal Structures for Geopolitical Events Forecasting
by: Ahrabian, Kian, et al.
Published: (2026) -
Scaling Laws for Multilingual Language Models
by: He, Yifei, et al.
Published: (2024) -
Escaping the Mode Lottery: Multi-Response Training Improves Language Model Generalization
by: Amin, Hasan, et al.
Published: (2026) -
Scaling Optimal LR Across Token Horizons
by: Bjorck, Johan, et al.
Published: (2024)