Born a Transformer -- Always a Transformer? On the Effect of Pretraining on Architectural Abilities
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jobanputra, Mayank, Veitsman, Yana, Sarrof, Yash, Bakalova, Aleksandra, Demberg, Vera, Pavlick, Ellie, Hahn, Michael |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Expressive Capacity of State Space Models: A Formal Language Perspective
von: Sarrof, Yash, et al.
Veröffentlicht: (2024)
von: Sarrof, Yash, et al.
Veröffentlicht: (2024)
Contextualize-then-Aggregate: Circuits for In-Context Learning in Gemma-2 2B
von: Bakalova, Aleksandra, et al.
Veröffentlicht: (2025)
von: Bakalova, Aleksandra, et al.
Veröffentlicht: (2025)
On the Ability of Transformers to Verify Plans
von: Sarrof, Yash, et al.
Veröffentlicht: (2026)
von: Sarrof, Yash, et al.
Veröffentlicht: (2026)
Barriers to Universal Reasoning With Transformers (And How to Overcome Them)
von: Kraus, Oliver, et al.
Veröffentlicht: (2026)
von: Kraus, Oliver, et al.
Veröffentlicht: (2026)
Can LLMs subtract numbers?
von: Jobanputra, Mayank, et al.
Veröffentlicht: (2025)
von: Jobanputra, Mayank, et al.
Veröffentlicht: (2025)
Discovering Interpretable Algorithms by Decompiling Transformers to RASP
von: Huang, Xinting, et al.
Veröffentlicht: (2026)
von: Huang, Xinting, et al.
Veröffentlicht: (2026)
Circuit Component Reuse Across Tasks in Transformer Language Models
von: Merullo, Jack, et al.
Veröffentlicht: (2023)
von: Merullo, Jack, et al.
Veröffentlicht: (2023)
B-cos LM: Efficiently Transforming Pre-trained Language Models for Improved Explainability
von: Wang, Yifan, et al.
Veröffentlicht: (2025)
von: Wang, Yifan, et al.
Veröffentlicht: (2025)
Does Training on Synthetic Data Make Models Less Robust?
von: Zhang, Lingze, et al.
Veröffentlicht: (2025)
von: Zhang, Lingze, et al.
Veröffentlicht: (2025)
Language Models Implement Simple Word2Vec-style Vector Arithmetic
von: Merullo, Jack, et al.
Veröffentlicht: (2023)
von: Merullo, Jack, et al.
Veröffentlicht: (2023)
Dual Process Learning: Controlling Use of In-Context vs. In-Weights Strategies with Weight Forgetting
von: Anand, Suraj, et al.
Veröffentlicht: (2024)
von: Anand, Suraj, et al.
Veröffentlicht: (2024)
A Formal Framework for Understanding Length Generalization in Transformers
von: Huang, Xinting, et al.
Veröffentlicht: (2024)
von: Huang, Xinting, et al.
Veröffentlicht: (2024)
Transferring Linear Features Across Language Models With Model Stitching
von: Chen, Alan, et al.
Veröffentlicht: (2025)
von: Chen, Alan, et al.
Veröffentlicht: (2025)
ChatGPT vs Human-authored Text: Insights into Controllable Text Summarization and Sentence Style Transfer
von: Liu, Dongqi, et al.
Veröffentlicht: (2023)
von: Liu, Dongqi, et al.
Veröffentlicht: (2023)
RST-LoRA: A Discourse-Aware Low-Rank Adaptation for Long Document Abstractive Summarization
von: Liu, Dongqi, et al.
Veröffentlicht: (2024)
von: Liu, Dongqi, et al.
Veröffentlicht: (2024)
How Do Vision-Language Models Process Conflicting Information Across Modalities?
von: Hua, Tianze, et al.
Veröffentlicht: (2025)
von: Hua, Tianze, et al.
Veröffentlicht: (2025)
Uncovering Intermediate Variables in Transformers using Circuit Probing
von: Lepori, Michael A., et al.
Veröffentlicht: (2023)
von: Lepori, Michael A., et al.
Veröffentlicht: (2023)
Recent Advancements and Challenges of Turkic Central Asian Language Processing
von: Veitsman, Yana, et al.
Veröffentlicht: (2024)
von: Veitsman, Yana, et al.
Veröffentlicht: (2024)
TPTT: Transforming Pretrained Transformers into Titans
von: Furfaro, Fabien
Veröffentlicht: (2025)
von: Furfaro, Fabien
Veröffentlicht: (2025)
Incorporating Distributions of Discourse Structure for Long Document Abstractive Summarization
von: Liu, Dongqi, et al.
Veröffentlicht: (2023)
von: Liu, Dongqi, et al.
Veröffentlicht: (2023)
How Few-Shot Examples Add Up: A Causal Decomposition of Function Vectors in In-Context Learning
von: Wang, Entang, et al.
Veröffentlicht: (2026)
von: Wang, Entang, et al.
Veröffentlicht: (2026)
MUStReason: A Benchmark for Diagnosing Pragmatic Reasoning in Video-LMs for Multimodal Sarcasm Detection
von: Saha, Anisha, et al.
Veröffentlicht: (2025)
von: Saha, Anisha, et al.
Veröffentlicht: (2025)
Emergent Stack Representations in Modeling Counter Languages Using Transformers
von: Tiwari, Utkarsh, et al.
Veröffentlicht: (2025)
von: Tiwari, Utkarsh, et al.
Veröffentlicht: (2025)
Bridging Fairness and Explainability: Can Input-Based Explanations Promote Fairness in Hate Speech Detection?
von: Wang, Yifan, et al.
Veröffentlicht: (2025)
von: Wang, Yifan, et al.
Veröffentlicht: (2025)
Talking Heads: Understanding Inter-layer Communication in Transformer Language Models
von: Merullo, Jack, et al.
Veröffentlicht: (2024)
von: Merullo, Jack, et al.
Veröffentlicht: (2024)
SciNews: From Scholarly Complexities to Public Narratives -- A Dataset for Scientific News Report Generation
von: Liu, Dongqi, et al.
Veröffentlicht: (2024)
von: Liu, Dongqi, et al.
Veröffentlicht: (2024)
Explanatory Summarization with Discourse-Driven Planning
von: Liu, Dongqi, et al.
Veröffentlicht: (2025)
von: Liu, Dongqi, et al.
Veröffentlicht: (2025)
$100K or 100 Days: Trade-offs when Pre-Training with Academic Resources
von: Khandelwal, Apoorv, et al.
Veröffentlicht: (2024)
von: Khandelwal, Apoorv, et al.
Veröffentlicht: (2024)
Pretrained Multilingual Transformers Reveal Quantitative Distance Between Human Languages
von: Zhao, Yue, et al.
Veröffentlicht: (2026)
von: Zhao, Yue, et al.
Veröffentlicht: (2026)
Instilling Inductive Biases with Subnetworks
von: Zhang, Enyan, et al.
Veröffentlicht: (2023)
von: Zhang, Enyan, et al.
Veröffentlicht: (2023)
Is Smaller Always Faster? Tradeoffs in Compressing Self-Supervised Speech Transformers
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2022)
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2022)
Why Better Cross-Lingual Alignment Fails for Better Cross-Lingual Transfer: Case of Encoders
von: Veitsman, Yana, et al.
Veröffentlicht: (2026)
von: Veitsman, Yana, et al.
Veröffentlicht: (2026)
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
von: Brandon, William, et al.
Veröffentlicht: (2024)
von: Brandon, William, et al.
Veröffentlicht: (2024)
Handling and Interpreting Missing Modalities in Patient Clinical Trajectories via Autoregressive Sequence Modeling
von: Wang, Andrew, et al.
Veröffentlicht: (2026)
von: Wang, Andrew, et al.
Veröffentlicht: (2026)
The Curved Spacetime of Transformer Architectures
von: Di Sipio, Riccardo, et al.
Veröffentlicht: (2025)
von: Di Sipio, Riccardo, et al.
Veröffentlicht: (2025)
Pooling Attention: Evaluating Pretrained Transformer Embeddings for Deception Classification
von: Mamtani, Sumit, et al.
Veröffentlicht: (2025)
von: Mamtani, Sumit, et al.
Veröffentlicht: (2025)
Learning Novel Transformer Architecture for Time-series Forecasting
von: Zhang, Juyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Juyuan, et al.
Veröffentlicht: (2025)
Softmax Transformers are Turing-Complete
von: Jiang, Hongjian, et al.
Veröffentlicht: (2025)
von: Jiang, Hongjian, et al.
Veröffentlicht: (2025)
Separations in the Representational Capabilities of Transformers and Recurrent Architectures
von: Bhattamishra, Satwik, et al.
Veröffentlicht: (2024)
von: Bhattamishra, Satwik, et al.
Veröffentlicht: (2024)
Transformer-based Joint Modelling for Automatic Essay Scoring and Off-Topic Detection
von: Das, Sourya Dipta, et al.
Veröffentlicht: (2024)
von: Das, Sourya Dipta, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
The Expressive Capacity of State Space Models: A Formal Language Perspective
von: Sarrof, Yash, et al.
Veröffentlicht: (2024) -
Contextualize-then-Aggregate: Circuits for In-Context Learning in Gemma-2 2B
von: Bakalova, Aleksandra, et al.
Veröffentlicht: (2025) -
On the Ability of Transformers to Verify Plans
von: Sarrof, Yash, et al.
Veröffentlicht: (2026) -
Barriers to Universal Reasoning With Transformers (And How to Overcome Them)
von: Kraus, Oliver, et al.
Veröffentlicht: (2026) -
Can LLMs subtract numbers?
von: Jobanputra, Mayank, et al.
Veröffentlicht: (2025)