Gespeichert in:
| Hauptverfasser: | Fu, Deqing, Chen, Tian-Qi, Jia, Robin, Sharan, Vatsal |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2310.17086 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Convergent Evolution: How Different Language Models Learn Similar Number Representations
von: Fu, Deqing, et al.
Veröffentlicht: (2026)
von: Fu, Deqing, et al.
Veröffentlicht: (2026)
Transformers Learn Low Sensitivity Functions: Investigations and Implications
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2024)
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2024)
Pre-trained Large Language Models Use Fourier Features to Compute Addition
von: Zhou, Tianyi, et al.
Veröffentlicht: (2024)
von: Zhou, Tianyi, et al.
Veröffentlicht: (2024)
FoNE: Precise Single-Token Number Embeddings via Fourier Features
von: Zhou, Tianyi, et al.
Veröffentlicht: (2025)
von: Zhou, Tianyi, et al.
Veröffentlicht: (2025)
Transformers Provably Learn Algorithmic Solutions for Graph Connectivity, But Only with the Right Data
von: Ye, Qilin, et al.
Veröffentlicht: (2025)
von: Ye, Qilin, et al.
Veröffentlicht: (2025)
Latent Concept Disentanglement in Transformer-based Language Models
von: Hong, Guan Zhe, et al.
Veröffentlicht: (2025)
von: Hong, Guan Zhe, et al.
Veröffentlicht: (2025)
Limitations on Accurate, Trusted, Human-level Reasoning
von: Panigrahy, Rina, et al.
Veröffentlicht: (2025)
von: Panigrahy, Rina, et al.
Veröffentlicht: (2025)
On the Robustness of Transformers against Context Hijacking for Linear Classification
von: Li, Tianle, et al.
Veröffentlicht: (2025)
von: Li, Tianle, et al.
Veröffentlicht: (2025)
Ulterior Motives: Detecting Misaligned Reasoning in Continuous Thought Models
von: Ramjee, Sharan
Veröffentlicht: (2026)
von: Ramjee, Sharan
Veröffentlicht: (2026)
Can GPT Redefine Medical Understanding? Evaluating GPT on Biomedical Machine Reading Comprehension
von: Vatsal, Shubham, et al.
Veröffentlicht: (2024)
von: Vatsal, Shubham, et al.
Veröffentlicht: (2024)
DeLLMa: Decision Making Under Uncertainty with Large Language Models
von: Liu, Ollie, et al.
Veröffentlicht: (2024)
von: Liu, Ollie, et al.
Veröffentlicht: (2024)
Transformers Can Achieve Length Generalization But Not Robustly
von: Zhou, Yongchao, et al.
Veröffentlicht: (2024)
von: Zhou, Yongchao, et al.
Veröffentlicht: (2024)
Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
von: Gan, Woody Haosheng, et al.
Veröffentlicht: (2025)
von: Gan, Woody Haosheng, et al.
Veröffentlicht: (2025)
Understanding Contextual Recall in Transformers: How Finetuning Enables In-Context Reasoning over Pretraining Knowledge
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2026)
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2026)
Can GPT Improve the State of Prior Authorization via Guideline Based Automated Question Answering?
von: Vatsal, Shubham, et al.
Veröffentlicht: (2024)
von: Vatsal, Shubham, et al.
Veröffentlicht: (2024)
Emotion Classification in Low and Moderate Resource Languages
von: Tafreshi, Shabnam, et al.
Veröffentlicht: (2024)
von: Tafreshi, Shabnam, et al.
Veröffentlicht: (2024)
Multilingual Prompt Engineering in Large Language Models: A Survey Across NLP Tasks
von: Vatsal, Shubham, et al.
Veröffentlicht: (2025)
von: Vatsal, Shubham, et al.
Veröffentlicht: (2025)
Understanding Emergent In-Context Learning from a Kernel Regression Perspective
von: Han, Chi, et al.
Veröffentlicht: (2023)
von: Han, Chi, et al.
Veröffentlicht: (2023)
EPSVec: Efficient and Private Synthetic Data Generation via Dataset Vectors
von: Banayeeanzade, Amin, et al.
Veröffentlicht: (2026)
von: Banayeeanzade, Amin, et al.
Veröffentlicht: (2026)
Supernova: Achieving More with Less in Transformer Architectures
von: Tanase, Andrei-Valentin, et al.
Veröffentlicht: (2025)
von: Tanase, Andrei-Valentin, et al.
Veröffentlicht: (2025)
Do pretrained Transformers Learn In-Context by Gradient Descent?
von: Shen, Lingfeng, et al.
Veröffentlicht: (2023)
von: Shen, Lingfeng, et al.
Veröffentlicht: (2023)
Linear-Time Demonstration Selection for In-Context Learning via Gradient Estimation
von: Zhang, Ziniu, et al.
Veröffentlicht: (2025)
von: Zhang, Ziniu, et al.
Veröffentlicht: (2025)
Intelligent Learning Rate Distribution to reduce Catastrophic Forgetting in Transformers
von: Kenneweg, Philip, et al.
Veröffentlicht: (2024)
von: Kenneweg, Philip, et al.
Veröffentlicht: (2024)
Meta-Sel: Efficient Demonstration Selection for In-Context Learning via Supervised Meta-Learning
von: Wang, Xubin, et al.
Veröffentlicht: (2026)
von: Wang, Xubin, et al.
Veröffentlicht: (2026)
In-Context Learning with Transformers: Softmax Attention Adapts to Function Lipschitzness
von: Collins, Liam, et al.
Veröffentlicht: (2024)
von: Collins, Liam, et al.
Veröffentlicht: (2024)
Cross-Architecture Transfer Learning for Linear-Cost Inference Transformers
von: Choi, Sehyun
Veröffentlicht: (2024)
von: Choi, Sehyun
Veröffentlicht: (2024)
Second-Order Information Matters: Revisiting Machine Unlearning for Large Language Models
von: Gu, Kang, et al.
Veröffentlicht: (2024)
von: Gu, Kang, et al.
Veröffentlicht: (2024)
Unlocking In-Context Learning for Natural Datasets Beyond Language Modelling
von: Bratulić, Jelena, et al.
Veröffentlicht: (2025)
von: Bratulić, Jelena, et al.
Veröffentlicht: (2025)
Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression
von: Chen, Xingwu, et al.
Veröffentlicht: (2025)
von: Chen, Xingwu, et al.
Veröffentlicht: (2025)
Transformers Handle Endogeneity in In-Context Linear Regression
von: Liang, Haodong, et al.
Veröffentlicht: (2024)
von: Liang, Haodong, et al.
Veröffentlicht: (2024)
Your Transformer is Secretly Linear
von: Razzhigaev, Anton, et al.
Veröffentlicht: (2024)
von: Razzhigaev, Anton, et al.
Veröffentlicht: (2024)
Boosting In-Context Learning in LLMs Through the Lens of Classical Supervised Learning
von: Gundem, Korel, et al.
Veröffentlicht: (2025)
von: Gundem, Korel, et al.
Veröffentlicht: (2025)
Unraveling the Mechanics of Learning-Based Demonstration Selection for In-Context Learning
von: Liu, Hui, et al.
Veröffentlicht: (2024)
von: Liu, Hui, et al.
Veröffentlicht: (2024)
Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI
von: Maiya, Sharan, et al.
Veröffentlicht: (2025)
von: Maiya, Sharan, et al.
Veröffentlicht: (2025)
Promote, Suppress, Iterate: How Language Models Answer One-to-Many Factual Queries
von: Yan, Tianyi Lorena, et al.
Veröffentlicht: (2025)
von: Yan, Tianyi Lorena, et al.
Veröffentlicht: (2025)
On the Rate of Convergence of Kolmogorov-Arnold Network Regression Estimators
von: Liu, Wei, et al.
Veröffentlicht: (2025)
von: Liu, Wei, et al.
Veröffentlicht: (2025)
Resa: Transparent Reasoning Models via SAEs
von: Wang, Shangshang, et al.
Veröffentlicht: (2025)
von: Wang, Shangshang, et al.
Veröffentlicht: (2025)
CoCA: Fusing Position Embedding with Collinear Constrained Attention in Transformers for Long Context Window Extending
von: Zhu, Shiyi, et al.
Veröffentlicht: (2023)
von: Zhu, Shiyi, et al.
Veröffentlicht: (2023)
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
von: MiniCPM Team, et al.
Veröffentlicht: (2026)
von: MiniCPM Team, et al.
Veröffentlicht: (2026)
Function Induction and Task Generalization: An Interpretability Study with Off-by-One Addition
von: Ye, Qinyuan, et al.
Veröffentlicht: (2025)
von: Ye, Qinyuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Convergent Evolution: How Different Language Models Learn Similar Number Representations
von: Fu, Deqing, et al.
Veröffentlicht: (2026) -
Transformers Learn Low Sensitivity Functions: Investigations and Implications
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2024) -
Pre-trained Large Language Models Use Fourier Features to Compute Addition
von: Zhou, Tianyi, et al.
Veröffentlicht: (2024) -
FoNE: Precise Single-Token Number Embeddings via Fourier Features
von: Zhou, Tianyi, et al.
Veröffentlicht: (2025) -
Transformers Provably Learn Algorithmic Solutions for Graph Connectivity, But Only with the Right Data
von: Ye, Qilin, et al.
Veröffentlicht: (2025)