Extended Mind Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Klett, Phoebe, Ahle, Thomas |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Many Minds from One Model: Bayesian-Inspired Transformers for Population Diversity
by: Yang, Diji, et al.
Published: (2025)
by: Yang, Diji, et al.
Published: (2025)
MillStone: How Open-Minded Are LLMs?
by: Triedman, Harold, et al.
Published: (2025)
by: Triedman, Harold, et al.
Published: (2025)
Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models
by: Song, Yuda, et al.
Published: (2024)
by: Song, Yuda, et al.
Published: (2024)
Translating Expert Intuition into Quantifiable Features: Encode Investigator Domain Knowledge via LLM for Enhanced Predictive Analytics
by: Jing, Phoebe, et al.
Published: (2024)
by: Jing, Phoebe, et al.
Published: (2024)
Mind the Gap: Data Rewriting for Stable Off-Policy Supervised Fine-Tuning
by: Zhao, Shiwan, et al.
Published: (2025)
by: Zhao, Shiwan, et al.
Published: (2025)
ShifaMind: A Multiplicative Concept Bottleneck for Interpretable ICD-10 Coding
by: Syed, Mohammed Sameer, et al.
Published: (2026)
by: Syed, Mohammed Sameer, et al.
Published: (2026)
Probabilistic Topic Modelling with Transformer Representations
by: Reuter, Arik, et al.
Published: (2024)
by: Reuter, Arik, et al.
Published: (2024)
An Innovative CGL-MHA Model for Sarcasm Sentiment Recognition Using the MindSpore Framework
by: Qin, Zhenkai, et al.
Published: (2024)
by: Qin, Zhenkai, et al.
Published: (2024)
CoCA: Fusing Position Embedding with Collinear Constrained Attention in Transformers for Long Context Window Extending
by: Zhu, Shiyi, et al.
Published: (2023)
by: Zhu, Shiyi, et al.
Published: (2023)
Trainable Transformer in Transformer
by: Panigrahi, Abhishek, et al.
Published: (2023)
by: Panigrahi, Abhishek, et al.
Published: (2023)
LongEmbed: Extending Embedding Models for Long Context Retrieval
by: Zhu, Dawei, et al.
Published: (2024)
by: Zhu, Dawei, et al.
Published: (2024)
SELF: Self-Extend the Context Length With Logistic Growth Function
by: Dang, Phat Thanh, et al.
Published: (2025)
by: Dang, Phat Thanh, et al.
Published: (2025)
Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers
by: Anagnostidis, Sotiris, et al.
Published: (2023)
by: Anagnostidis, Sotiris, et al.
Published: (2023)
PL-FGSA: A Prompt Learning Framework for Fine-Grained Sentiment Analysis Based on MindSpore
by: Qin, Zhenkai, et al.
Published: (2025)
by: Qin, Zhenkai, et al.
Published: (2025)
RakutenAI-7B: Extending Large Language Models for Japanese
by: Rakuten Group, et al.
Published: (2024)
by: Rakuten Group, et al.
Published: (2024)
Extending Input Contexts of Language Models through Training on Segmented Sequences
by: Karypis, Petros, et al.
Published: (2023)
by: Karypis, Petros, et al.
Published: (2023)
Extending Beacon to Hindi: Cultural Adaptation Drives Cross-Lingual Sycophancy
by: Sattigeri, Sarthak
Published: (2026)
by: Sattigeri, Sarthak
Published: (2026)
ToolExpander: Extending the Frontiers of Tool-Using Reinforcement Learning to Weak LLMs
by: Chen, Fu, et al.
Published: (2025)
by: Chen, Fu, et al.
Published: (2025)
Beyond the Limits: A Survey of Techniques to Extend the Context Length in Large Language Models
by: Wang, Xindi, et al.
Published: (2024)
by: Wang, Xindi, et al.
Published: (2024)
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing
by: Smith, James Seale, et al.
Published: (2025)
by: Smith, James Seale, et al.
Published: (2025)
MediaMind: Revolutionizing Media Monitoring using Agentification
by: Gunduz, Ahmet, et al.
Published: (2025)
by: Gunduz, Ahmet, et al.
Published: (2025)
Automated Meta Prompt Engineering for Alignment with the Theory of Mind
by: Baughman, Aaron, et al.
Published: (2025)
by: Baughman, Aaron, et al.
Published: (2025)
Mini Minds: Exploring Bebeshka and Zlata Baby Models
by: Proskurina, Irina, et al.
Published: (2023)
by: Proskurina, Irina, et al.
Published: (2023)
Scalable Bayesian Learning with posteriors
by: Duffield, Samuel, et al.
Published: (2024)
by: Duffield, Samuel, et al.
Published: (2024)
Differential Transformer
by: Ye, Tianzhu, et al.
Published: (2024)
by: Ye, Tianzhu, et al.
Published: (2024)
Hyperloop Transformers
by: Zeitoun, Abbas, et al.
Published: (2026)
by: Zeitoun, Abbas, et al.
Published: (2026)
Rethinking Attention Output Projection: Structured Hadamard Transforms for Efficient Transformers
by: Aggarwal, Shubham, et al.
Published: (2026)
by: Aggarwal, Shubham, et al.
Published: (2026)
GPT-4o Lacks Core Features of Theory of Mind
by: Muchovej, John, et al.
Published: (2026)
by: Muchovej, John, et al.
Published: (2026)
Transformer See, Transformer Do: Copying as an Intermediate Step in Learning Analogical Reasoning
by: Hellwig, Philipp, et al.
Published: (2026)
by: Hellwig, Philipp, et al.
Published: (2026)
Born a Transformer -- Always a Transformer? On the Effect of Pretraining on Architectural Abilities
by: Jobanputra, Mayank, et al.
Published: (2025)
by: Jobanputra, Mayank, et al.
Published: (2025)
Does Transformer Interpretability Transfer to RNNs?
by: Paulo, Gonçalo, et al.
Published: (2024)
by: Paulo, Gonçalo, et al.
Published: (2024)
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU
by: Lee, Heejun, et al.
Published: (2025)
by: Lee, Heejun, et al.
Published: (2025)
LASER: Attention with Exponential Transformation
by: Duvvuri, Sai Surya, et al.
Published: (2024)
by: Duvvuri, Sai Surya, et al.
Published: (2024)
Understanding Addition and Subtraction in Transformers
by: Quirke, Philip, et al.
Published: (2024)
by: Quirke, Philip, et al.
Published: (2024)
Transformers are Universal In-context Learners
by: Furuya, Takashi, et al.
Published: (2024)
by: Furuya, Takashi, et al.
Published: (2024)
Converting Transformers into DGNNs Form
by: Zhang, Jie, et al.
Published: (2025)
by: Zhang, Jie, et al.
Published: (2025)
On the Geometry of Positional Encodings in Transformers
by: Cirrincione, Giansalvo
Published: (2026)
by: Cirrincione, Giansalvo
Published: (2026)
A Notion of Complexity for Theory of Mind via Discrete World Models
by: Huang, X. Angelo, et al.
Published: (2024)
by: Huang, X. Angelo, et al.
Published: (2024)
Mind the Gap: A Review of Arabic Post-Training Datasets and Their Limitations
by: Alkhowaiter, Mohammed, et al.
Published: (2025)
by: Alkhowaiter, Mohammed, et al.
Published: (2025)
Early Transformers: A study on Efficient Training of Transformer Models through Early-Bird Lottery Tickets
by: Cheekati, Shravan
Published: (2024)
by: Cheekati, Shravan
Published: (2024)
Similar Items
-
Many Minds from One Model: Bayesian-Inspired Transformers for Population Diversity
by: Yang, Diji, et al.
Published: (2025) -
MillStone: How Open-Minded Are LLMs?
by: Triedman, Harold, et al.
Published: (2025) -
Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models
by: Song, Yuda, et al.
Published: (2024) -
Translating Expert Intuition into Quantifiable Features: Encode Investigator Domain Knowledge via LLM for Enhanced Predictive Analytics
by: Jing, Phoebe, et al.
Published: (2024) -
Mind the Gap: Data Rewriting for Stable Off-Policy Supervised Fine-Tuning
by: Zhao, Shiwan, et al.
Published: (2025)