Trainable Transformer in Transformer
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Panigrahi, Abhishek, Malladi, Sadhika, Xia, Mengzhou, Arora, Sanjeev |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LESS: Selecting Influential Data for Targeted Instruction Tuning
von: Xia, Mengzhou, et al.
Veröffentlicht: (2024)
von: Xia, Mengzhou, et al.
Veröffentlicht: (2024)
On the SDEs and Scaling Rules for Adaptive Gradient Algorithms
von: Malladi, Sadhika, et al.
Veröffentlicht: (2022)
von: Malladi, Sadhika, et al.
Veröffentlicht: (2022)
In Good GRACEs: Principled Teacher Selection for Knowledge Distillation
von: Panigrahi, Abhishek, et al.
Veröffentlicht: (2025)
von: Panigrahi, Abhishek, et al.
Veröffentlicht: (2025)
Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization
von: Razin, Noam, et al.
Veröffentlicht: (2024)
von: Razin, Noam, et al.
Veröffentlicht: (2024)
Fine-Tuning Language Models with Just Forward Passes
von: Malladi, Sadhika, et al.
Veröffentlicht: (2023)
von: Malladi, Sadhika, et al.
Veröffentlicht: (2023)
Provable unlearning in topic modeling and downstream tasks
von: Wei, Stanley, et al.
Veröffentlicht: (2024)
von: Wei, Stanley, et al.
Veröffentlicht: (2024)
Representing Rule-based Chatbots with Transformers
von: Friedman, Dan, et al.
Veröffentlicht: (2024)
von: Friedman, Dan, et al.
Veröffentlicht: (2024)
Progressive distillation induces an implicit curriculum
von: Panigrahi, Abhishek, et al.
Veröffentlicht: (2024)
von: Panigrahi, Abhishek, et al.
Veröffentlicht: (2024)
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?
von: Park, Simon, et al.
Veröffentlicht: (2025)
von: Park, Simon, et al.
Veröffentlicht: (2025)
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
von: Dong, Yihe, et al.
Veröffentlicht: (2025)
von: Dong, Yihe, et al.
Veröffentlicht: (2025)
Projected Compression: Trainable Projection for Efficient Transformer Compression
von: Stefaniak, Maciej, et al.
Veröffentlicht: (2025)
von: Stefaniak, Maciej, et al.
Veröffentlicht: (2025)
On the Power of Context-Enhanced Learning in LLMs
von: Zhu, Xingyu, et al.
Veröffentlicht: (2025)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2025)
Preference Learning Algorithms Do Not Learn Preference Rankings
von: Chen, Angelica, et al.
Veröffentlicht: (2024)
von: Chen, Angelica, et al.
Veröffentlicht: (2024)
SimPO: Simple Preference Optimization with a Reference-Free Reward
von: Meng, Yu, et al.
Veröffentlicht: (2024)
von: Meng, Yu, et al.
Veröffentlicht: (2024)
Selective Neuron Amplification in Transformer Language Models
von: Akhtar, Ryyan, et al.
Veröffentlicht: (2026)
von: Akhtar, Ryyan, et al.
Veröffentlicht: (2026)
Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial?
von: Li, Wenzhe, et al.
Veröffentlicht: (2025)
von: Li, Wenzhe, et al.
Veröffentlicht: (2025)
Lory: Fully Differentiable Mixture-of-Experts for Autoregressive Language Model Pre-training
von: Zhong, Zexuan, et al.
Veröffentlicht: (2024)
von: Zhong, Zexuan, et al.
Veröffentlicht: (2024)
AdaptMI: Adaptive Skill-based In-context Math Instruction for Small Language Models
von: He, Yinghui, et al.
Veröffentlicht: (2025)
von: He, Yinghui, et al.
Veröffentlicht: (2025)
Instruct-SkillMix: A Powerful Pipeline for LLM Instruction Tuning
von: Kaur, Simran, et al.
Veröffentlicht: (2024)
von: Kaur, Simran, et al.
Veröffentlicht: (2024)
The Coverage Principle: How Pre-Training Enables Post-Training
von: Chen, Fan, et al.
Veröffentlicht: (2025)
von: Chen, Fan, et al.
Veröffentlicht: (2025)
CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
von: Wang, Zirui, et al.
Veröffentlicht: (2024)
von: Wang, Zirui, et al.
Veröffentlicht: (2024)
Differential Transformer
von: Ye, Tianzhu, et al.
Veröffentlicht: (2024)
von: Ye, Tianzhu, et al.
Veröffentlicht: (2024)
Skill-Targeted Adaptive Training
von: He, Yinghui, et al.
Veröffentlicht: (2025)
von: He, Yinghui, et al.
Veröffentlicht: (2025)
Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning
von: Xia, Mengzhou, et al.
Veröffentlicht: (2023)
von: Xia, Mengzhou, et al.
Veröffentlicht: (2023)
What is in Your Safe Data? Identifying Benign Data that Breaks Safety
von: He, Luxi, et al.
Veröffentlicht: (2024)
von: He, Luxi, et al.
Veröffentlicht: (2024)
Why is Your Language Model a Poor Implicit Reward Model?
von: Razin, Noam, et al.
Veröffentlicht: (2025)
von: Razin, Noam, et al.
Veröffentlicht: (2025)
Contextual Drag: How Errors in the Context Affect LLM Reasoning
von: Cheng, Yun, et al.
Veröffentlicht: (2026)
von: Cheng, Yun, et al.
Veröffentlicht: (2026)
Learning to Focus: Focal Attention for Selective and Scalable Transformers
von: Ram, Dhananjay, et al.
Veröffentlicht: (2025)
von: Ram, Dhananjay, et al.
Veröffentlicht: (2025)
The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning
von: Zhu, Xinyu, et al.
Veröffentlicht: (2025)
von: Zhu, Xinyu, et al.
Veröffentlicht: (2025)
On The Adaptation of Unlimiformer for Decoder-Only Transformers
von: Ahrabian, Kian, et al.
Veröffentlicht: (2024)
von: Ahrabian, Kian, et al.
Veröffentlicht: (2024)
Wave-PDE Nets: Trainable Wave-Equation Layers as an Alternative to Attention
von: Vejendla, Harshil
Veröffentlicht: (2025)
von: Vejendla, Harshil
Veröffentlicht: (2025)
Efficient Stagewise Pretraining via Progressive Subnetworks
von: Panigrahi, Abhishek, et al.
Veröffentlicht: (2024)
von: Panigrahi, Abhishek, et al.
Veröffentlicht: (2024)
LinkTransformer: A Unified Package for Record Linkage with Transformer Language Models
von: Arora, Abhishek, et al.
Veröffentlicht: (2023)
von: Arora, Abhishek, et al.
Veröffentlicht: (2023)
Trainable Dynamic Mask Sparse Attention
von: Shi, Jingze, et al.
Veröffentlicht: (2025)
von: Shi, Jingze, et al.
Veröffentlicht: (2025)
Can Models Learn Skill Composition from Examples?
von: Zhao, Haoyu, et al.
Veröffentlicht: (2024)
von: Zhao, Haoyu, et al.
Veröffentlicht: (2024)
Hyperloop Transformers
von: Zeitoun, Abbas, et al.
Veröffentlicht: (2026)
von: Zeitoun, Abbas, et al.
Veröffentlicht: (2026)
Rethinking Attention Output Projection: Structured Hadamard Transforms for Efficient Transformers
von: Aggarwal, Shubham, et al.
Veröffentlicht: (2026)
von: Aggarwal, Shubham, et al.
Veröffentlicht: (2026)
Extended Mind Transformers
von: Klett, Phoebe, et al.
Veröffentlicht: (2024)
von: Klett, Phoebe, et al.
Veröffentlicht: (2024)
Transformer See, Transformer Do: Copying as an Intermediate Step in Learning Analogical Reasoning
von: Hellwig, Philipp, et al.
Veröffentlicht: (2026)
von: Hellwig, Philipp, et al.
Veröffentlicht: (2026)
Born a Transformer -- Always a Transformer? On the Effect of Pretraining on Architectural Abilities
von: Jobanputra, Mayank, et al.
Veröffentlicht: (2025)
von: Jobanputra, Mayank, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LESS: Selecting Influential Data for Targeted Instruction Tuning
von: Xia, Mengzhou, et al.
Veröffentlicht: (2024) -
On the SDEs and Scaling Rules for Adaptive Gradient Algorithms
von: Malladi, Sadhika, et al.
Veröffentlicht: (2022) -
In Good GRACEs: Principled Teacher Selection for Knowledge Distillation
von: Panigrahi, Abhishek, et al.
Veröffentlicht: (2025) -
Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization
von: Razin, Noam, et al.
Veröffentlicht: (2024) -
Fine-Tuning Language Models with Just Forward Passes
von: Malladi, Sadhika, et al.
Veröffentlicht: (2023)