Theoretical Analysis of Positional Encodings in Transformer Models: Impact on Expressiveness and Generalization
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Li, Yin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Understanding the Nature of Generative AI as Threshold Logic in High-Dimensional Space
von: Levin, Ilya
Veröffentlicht: (2026)
von: Levin, Ilya
Veröffentlicht: (2026)
Graded Transformers
von: Shaska Sr, Tony
Veröffentlicht: (2025)
von: Shaska Sr, Tony
Veröffentlicht: (2025)
Tricks and Plug-ins for Gradient Boosting with Transformers
von: Fang, Biyi, et al.
Veröffentlicht: (2025)
von: Fang, Biyi, et al.
Veröffentlicht: (2025)
InhibiDistilbert: Knowledge Distillation for a ReLU and Addition-based Transformer
von: Zhang, Tony, et al.
Veröffentlicht: (2025)
von: Zhang, Tony, et al.
Veröffentlicht: (2025)
Think Thrice Before You Speak: Dual knowledge-enhanced Theory-of-Mind Reasoning for Persuasive Agents
von: Ma, Minghui, et al.
Veröffentlicht: (2026)
von: Ma, Minghui, et al.
Veröffentlicht: (2026)
Binarized Neural Networks Converge Toward Algorithmic Simplicity: Empirical Support for the Learning-as-Compression Hypothesis
von: Sakabe, Eduardo Y., et al.
Veröffentlicht: (2025)
von: Sakabe, Eduardo Y., et al.
Veröffentlicht: (2025)
Modularity in Transformers: Investigating Neuron Separability & Specialization
von: Pochinkov, Nicholas, et al.
Veröffentlicht: (2024)
von: Pochinkov, Nicholas, et al.
Veröffentlicht: (2024)
DeepPersona: A Generative Engine for Scaling Deep Synthetic Personas
von: Wang, Zhen, et al.
Veröffentlicht: (2025)
von: Wang, Zhen, et al.
Veröffentlicht: (2025)
Does Editing Provide Evidence for Localization?
von: Wang, Zihao, et al.
Veröffentlicht: (2025)
von: Wang, Zihao, et al.
Veröffentlicht: (2025)
NeuronSpark: A Spiking Neural Network Language Model with Selective State Space Dynamics
von: Tang, Zhengzheng
Veröffentlicht: (2026)
von: Tang, Zhengzheng
Veröffentlicht: (2026)
Position: Uncertainty Quantification in LLMs is Just Unsupervised Clustering
von: Chen, Tiejin, et al.
Veröffentlicht: (2026)
von: Chen, Tiejin, et al.
Veröffentlicht: (2026)
CLMN: Concept based Language Models via Neural Symbolic Reasoning
von: Yang, Yibo
Veröffentlicht: (2025)
von: Yang, Yibo
Veröffentlicht: (2025)
Harnessing non-adversarial robustness in large language models
von: Zhou, Qinghua, et al.
Veröffentlicht: (2026)
von: Zhou, Qinghua, et al.
Veröffentlicht: (2026)
How Pruning Reshapes Features: Sparse Autoencoder Analysis of Weight-Pruned Language Models
von: Borobia, Hector, et al.
Veröffentlicht: (2026)
von: Borobia, Hector, et al.
Veröffentlicht: (2026)
The Concept Allocation Zone: Tracking How Concepts Form Across Transformer Depth
von: Henry, James
Veröffentlicht: (2026)
von: Henry, James
Veröffentlicht: (2026)
Thinking Machines: Mathematical Reasoning in the Age of LLMs
von: Asperti, Andrea, et al.
Veröffentlicht: (2025)
von: Asperti, Andrea, et al.
Veröffentlicht: (2025)
Geometric Evolution Maps: Extracting Stable Concept Probes from Transformer Residual Streams
von: Henry, James
Veröffentlicht: (2026)
von: Henry, James
Veröffentlicht: (2026)
AIPsy-Affect: A Keyword-Free Clinical Stimulus Battery for Mechanistic Interpretability of Emotion in Language Models
von: Keeman, Michael
Veröffentlicht: (2026)
von: Keeman, Michael
Veröffentlicht: (2026)
Causal Dimensionality of Transformer Representations: Measurement, Scaling, and Layer Structure
von: Sarkar, Nilesh, et al.
Veröffentlicht: (2026)
von: Sarkar, Nilesh, et al.
Veröffentlicht: (2026)
Understanding the Uncertainty of LLM Explanations: A Perspective Based on Reasoning Topology
von: Da, Longchao, et al.
Veröffentlicht: (2025)
von: Da, Longchao, et al.
Veröffentlicht: (2025)
A Practical Guide to Streaming Continual Learning
von: Cossu, Andrea, et al.
Veröffentlicht: (2026)
von: Cossu, Andrea, et al.
Veröffentlicht: (2026)
Don't Look Back in Anger: MAGIC Net for Streaming Continual Learning with Temporal Dependence
von: Giannini, Federico, et al.
Veröffentlicht: (2026)
von: Giannini, Federico, et al.
Veröffentlicht: (2026)
Product-of-Experts Training Reduces Dataset Artifacts in Natural Language Inference
von: Mathew, Aby Mammen
Veröffentlicht: (2026)
von: Mathew, Aby Mammen
Veröffentlicht: (2026)
cPNN: Continuous Progressive Neural Networks for Evolving Streaming Time Series
von: Giannini, Federico, et al.
Veröffentlicht: (2026)
von: Giannini, Federico, et al.
Veröffentlicht: (2026)
Constitution or Collapse? Exploring Constitutional AI with Llama 3-8B
von: Zhang, Xue
Veröffentlicht: (2025)
von: Zhang, Xue
Veröffentlicht: (2025)
AI LLM Proof of Self-Consciousness and User-Specific Attractors
von: Camlin, Jeffrey
Veröffentlicht: (2025)
von: Camlin, Jeffrey
Veröffentlicht: (2025)
ProactBench: Beyond What The User Asked For
von: Harfi, Sepehr, et al.
Veröffentlicht: (2026)
von: Harfi, Sepehr, et al.
Veröffentlicht: (2026)
DRO-InstructZero: Distributionally Robust Prompt Optimization for Large Language Models
von: Li, Yangyang
Veröffentlicht: (2025)
von: Li, Yangyang
Veröffentlicht: (2025)
Manipulating Transformer-Based Models: Controllability, Steerability, and Robust Interventions
von: Alpay, Faruk, et al.
Veröffentlicht: (2025)
von: Alpay, Faruk, et al.
Veröffentlicht: (2025)
Attention Meets Reachability: Structural Equivalence and Efficiency in Grammar-Constrained LLM Decoding
von: Alpay, Faruk, et al.
Veröffentlicht: (2026)
von: Alpay, Faruk, et al.
Veröffentlicht: (2026)
NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles
von: Jia, Xiao
Veröffentlicht: (2026)
von: Jia, Xiao
Veröffentlicht: (2026)
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
Recent Advances in Data-Driven Business Process Management
von: Ackermann, Lars, et al.
Veröffentlicht: (2024)
von: Ackermann, Lars, et al.
Veröffentlicht: (2024)
NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution
von: Breneur, Oleksandr Marchenko, et al.
Veröffentlicht: (2026)
von: Breneur, Oleksandr Marchenko, et al.
Veröffentlicht: (2026)
When Are Two RLHF Objectives the Same?
von: Gaikwad, Madhava
Veröffentlicht: (2025)
von: Gaikwad, Madhava
Veröffentlicht: (2025)
ProbeScale: Probing Analysis to Optimize Neural Scaling Laws for Efficient Small Language Model Inference
von: Das, Sourav
Veröffentlicht: (2026)
von: Das, Sourav
Veröffentlicht: (2026)
Efficient Fine-Tuning Methods for Portuguese Question Answering: A Comparative Study of PEFT on BERTimbau and Exploratory Evaluation of Generative LLMs
von: Nina, Mariela M., et al.
Veröffentlicht: (2026)
von: Nina, Mariela M., et al.
Veröffentlicht: (2026)
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
von: Basu, Abhinaba
Veröffentlicht: (2026)
von: Basu, Abhinaba
Veröffentlicht: (2026)
STACHE: Local Black-Box Explanations for Reinforcement Learning Policies
von: Elashkin, Andrew, et al.
Veröffentlicht: (2025)
von: Elashkin, Andrew, et al.
Veröffentlicht: (2025)
Automated but Atrophied? Student Over-Reliance vs Expert Augmentation of AI in Learning and Cybersecurity
von: Khan, Koffka
Veröffentlicht: (2025)
von: Khan, Koffka
Veröffentlicht: (2025)
Ähnliche Einträge
-
Understanding the Nature of Generative AI as Threshold Logic in High-Dimensional Space
von: Levin, Ilya
Veröffentlicht: (2026) -
Graded Transformers
von: Shaska Sr, Tony
Veröffentlicht: (2025) -
Tricks and Plug-ins for Gradient Boosting with Transformers
von: Fang, Biyi, et al.
Veröffentlicht: (2025) -
InhibiDistilbert: Knowledge Distillation for a ReLU and Addition-based Transformer
von: Zhang, Tony, et al.
Veröffentlicht: (2025) -
Think Thrice Before You Speak: Dual knowledge-enhanced Theory-of-Mind Reasoning for Persuasive Agents
von: Ma, Minghui, et al.
Veröffentlicht: (2026)