Myosotis: structured computation for attention like layer
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Egorov, Evgenii, Ackermann, Hanno, Nagel, Markus, Cai, Hong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Equivariant Machine Learning Decoder for 3D Toric Codes
von: Weissl, Oliver, et al.
Veröffentlicht: (2024)
von: Weissl, Oliver, et al.
Veröffentlicht: (2024)
Ai-Sampler: Adversarial Learning of Markov kernels with involutive maps
von: Egorov, Evgenii, et al.
Veröffentlicht: (2024)
von: Egorov, Evgenii, et al.
Veröffentlicht: (2024)
Weight decay induces low-rank attention layers
von: Kobayashi, Seijin, et al.
Veröffentlicht: (2024)
von: Kobayashi, Seijin, et al.
Veröffentlicht: (2024)
Provably learning a multi-head attention layer
von: Chen, Sitan, et al.
Veröffentlicht: (2024)
von: Chen, Sitan, et al.
Veröffentlicht: (2024)
Deconfounding Imitation Learning with Variational Inference
von: Vuorio, Risto, et al.
Veröffentlicht: (2022)
von: Vuorio, Risto, et al.
Veröffentlicht: (2022)
FPTQuant: Function-Preserving Transforms for LLM Quantization
von: van Breugel, Boris, et al.
Veröffentlicht: (2025)
von: van Breugel, Boris, et al.
Veröffentlicht: (2025)
The Paradox of Stochasticity: Limited Creativity and Computational Decoupling in Temperature-Varied LLM Outputs of Structured Fictional Data
von: Evstafev, Evgenii
Veröffentlicht: (2025)
von: Evstafev, Evgenii
Veröffentlicht: (2025)
Token-by-Token Regeneration and Domain Biases: A Benchmark of LLMs on Advanced Mathematical Problem-Solving
von: Evstafev, Evgenii
Veröffentlicht: (2025)
von: Evstafev, Evgenii
Veröffentlicht: (2025)
Token-Hungry, Yet Precise: DeepSeek R1 Highlights the Need for Multi-Step Reasoning Over Speed in MATH
von: Evstafev, Evgenii
Veröffentlicht: (2025)
von: Evstafev, Evgenii
Veröffentlicht: (2025)
Offline Reinforcement Learning with Domain-Unlabeled Data
von: Nishimori, Soichiro, et al.
Veröffentlicht: (2024)
von: Nishimori, Soichiro, et al.
Veröffentlicht: (2024)
Lower bounds for one-layer transformers that compute parity
von: Hsu, Daniel
Veröffentlicht: (2026)
von: Hsu, Daniel
Veröffentlicht: (2026)
Multi-layer Cross-attention is Provably Optimal for Multi-modal In-context Learning
von: Barnfield, Nicholas, et al.
Veröffentlicht: (2026)
von: Barnfield, Nicholas, et al.
Veröffentlicht: (2026)
Untrained Perceptual Loss for image denoising of line-like structures in MR images
von: Pfaehler, Elisabeth, et al.
Veröffentlicht: (2024)
von: Pfaehler, Elisabeth, et al.
Veröffentlicht: (2024)
Low-Rank Quantization-Aware Training for LLMs
von: Bondarenko, Yelysei, et al.
Veröffentlicht: (2024)
von: Bondarenko, Yelysei, et al.
Veröffentlicht: (2024)
The underlying structures of self-attention: symmetry, directionality, and emergent dynamics in Transformer training
von: Saponati, Matteo, et al.
Veröffentlicht: (2025)
von: Saponati, Matteo, et al.
Veröffentlicht: (2025)
Pruning vs Quantization: Which is Better?
von: Kuzmin, Andrey, et al.
Veröffentlicht: (2023)
von: Kuzmin, Andrey, et al.
Veröffentlicht: (2023)
Dissecting Quantization Error: A Concentration-Alignment Perspective
von: Federici, Marco, et al.
Veröffentlicht: (2026)
von: Federici, Marco, et al.
Veröffentlicht: (2026)
Theory and computation for structured variational inference
von: Sheng, Shunan, et al.
Veröffentlicht: (2025)
von: Sheng, Shunan, et al.
Veröffentlicht: (2025)
Optimizing Humor Generation in Large Language Models: Temperature Configurations and Architectural Trade-offs
von: Evstafev, Evgenii
Veröffentlicht: (2025)
von: Evstafev, Evgenii
Veröffentlicht: (2025)
Easy attention: A simple attention mechanism for temporal predictions with transformers
von: Sanchis-Agudo, Marcial, et al.
Veröffentlicht: (2023)
von: Sanchis-Agudo, Marcial, et al.
Veröffentlicht: (2023)
FP8 Quantization: The Power of the Exponent
von: Kuzmin, Andrey, et al.
Veröffentlicht: (2022)
von: Kuzmin, Andrey, et al.
Veröffentlicht: (2022)
Leech Lattice Vector Quantization for Efficient LLM Compression
von: van der Ouderaa, Tycho F. A., et al.
Veröffentlicht: (2026)
von: van der Ouderaa, Tycho F. A., et al.
Veröffentlicht: (2026)
In-context denoising with one-layer transformers: connections between attention and associative memory retrieval
von: Smart, Matthew, et al.
Veröffentlicht: (2025)
von: Smart, Matthew, et al.
Veröffentlicht: (2025)
Reorganizing attention-space geometry with expressive attention
von: Gros, Claudius
Veröffentlicht: (2024)
von: Gros, Claudius
Veröffentlicht: (2024)
Controlling changes to attention logits
von: Anson, Ben, et al.
Veröffentlicht: (2025)
von: Anson, Ben, et al.
Veröffentlicht: (2025)
LiLMaps: Learnable Implicit Language Maps
von: Kruzhkov, Evgenii, et al.
Veröffentlicht: (2025)
von: Kruzhkov, Evgenii, et al.
Veröffentlicht: (2025)
Flow map learning in nonlinear vector autoregressive models: influence of the feature-library structure on the training error
von: Gross, Markus
Veröffentlicht: (2026)
von: Gross, Markus
Veröffentlicht: (2026)
STaMP: Sequence Transformation and Mixed Precision for Low-Precision Activation Quantization
von: Federici, Marco, et al.
Veröffentlicht: (2025)
von: Federici, Marco, et al.
Veröffentlicht: (2025)
Learning spatially structured open quantum dynamics with regional-attention transformers
von: Du, Dounan, et al.
Veröffentlicht: (2025)
von: Du, Dounan, et al.
Veröffentlicht: (2025)
Kernel-based optimization of measurement operators for quantum reservoir computers
von: Gross, Markus, et al.
Veröffentlicht: (2026)
von: Gross, Markus, et al.
Veröffentlicht: (2026)
Approximation of relation functions and attention mechanisms
von: Altabaa, Awni, et al.
Veröffentlicht: (2024)
von: Altabaa, Awni, et al.
Veröffentlicht: (2024)
Poly-attention: a general scheme for higher-order self-attention
von: Chakrabarti, Sayak, et al.
Veröffentlicht: (2026)
von: Chakrabarti, Sayak, et al.
Veröffentlicht: (2026)
The silence of the weights: a structural pruning strategy for attention-based audio signal architectures with second order metrics
von: Diecidue, Andrea, et al.
Veröffentlicht: (2025)
von: Diecidue, Andrea, et al.
Veröffentlicht: (2025)
Fusion of Multiscale Features Via Centralized Sparse-attention Network for EEG Decoding
von: Cai, Xiangrui, et al.
Veröffentlicht: (2025)
von: Cai, Xiangrui, et al.
Veröffentlicht: (2025)
SynthTree: Co-supervised Local Model Synthesis for Explainable Prediction
von: Kuriabov, Evgenii, et al.
Veröffentlicht: (2024)
von: Kuriabov, Evgenii, et al.
Veröffentlicht: (2024)
Narrowing the Gap between Adversarial and Stochastic MDPs via Policy Optimization
von: Tiapkin, Daniil, et al.
Veröffentlicht: (2024)
von: Tiapkin, Daniil, et al.
Veröffentlicht: (2024)
Offline Reinforcement Learning from Datasets with Structured Non-Stationarity
von: Ackermann, Johannes, et al.
Veröffentlicht: (2024)
von: Ackermann, Johannes, et al.
Veröffentlicht: (2024)
Beyond Rule-based Named Entity Recognition and Relation Extraction for Process Model Generation from Natural Language Text
von: Neuberger, Julian, et al.
Veröffentlicht: (2023)
von: Neuberger, Julian, et al.
Veröffentlicht: (2023)
Fast attention mechanisms: a tale of parallelism
von: Liu, Jingwen, et al.
Veröffentlicht: (2025)
von: Liu, Jingwen, et al.
Veröffentlicht: (2025)
An extension of linear self-attention for in-context learning
von: Hagiwara, Katsuyuki
Veröffentlicht: (2025)
von: Hagiwara, Katsuyuki
Veröffentlicht: (2025)
Ähnliche Einträge
-
Equivariant Machine Learning Decoder for 3D Toric Codes
von: Weissl, Oliver, et al.
Veröffentlicht: (2024) -
Ai-Sampler: Adversarial Learning of Markov kernels with involutive maps
von: Egorov, Evgenii, et al.
Veröffentlicht: (2024) -
Weight decay induces low-rank attention layers
von: Kobayashi, Seijin, et al.
Veröffentlicht: (2024) -
Provably learning a multi-head attention layer
von: Chen, Sitan, et al.
Veröffentlicht: (2024) -
Deconfounding Imitation Learning with Variational Inference
von: Vuorio, Risto, et al.
Veröffentlicht: (2022)