Softmax Linear Attention: Reclaiming Global Competition
Fuente:
arXiv
Salvato in:
| Autori principali: | Xu, Mingwei, Lin, Xuan, Guo, Xinnan, Xu, Wanqing, Cui, Wanyun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Exploring and Reshaping the Weight Distribution in LLM
di: Ye, Chunming, et al.
Pubblicazione: (2025)
di: Ye, Chunming, et al.
Pubblicazione: (2025)
Distractor Injection Attacks on Large Reasoning Models: Characterization and Defense
di: Zhang, Zhehao, et al.
Pubblicazione: (2025)
di: Zhang, Zhehao, et al.
Pubblicazione: (2025)
Analysis on distribution and clustering of weight
di: Ye, Chunming, et al.
Pubblicazione: (2025)
di: Ye, Chunming, et al.
Pubblicazione: (2025)
Investigating on RLHF methodology
di: Kutalev, Alexey, et al.
Pubblicazione: (2024)
di: Kutalev, Alexey, et al.
Pubblicazione: (2024)
CIKT: A Collaborative and Iterative Knowledge Tracing Framework with Large Language Models
di: Li, Runze, et al.
Pubblicazione: (2025)
di: Li, Runze, et al.
Pubblicazione: (2025)
Mixing Times of Glauber Dynamics on Masked Language Models
di: Sana, Suvadip, et al.
Pubblicazione: (2026)
di: Sana, Suvadip, et al.
Pubblicazione: (2026)
ProMedTS: A Self-Supervised, Prompt-Guided Multimodal Approach for Integrating Medical Text and Time Series
di: Niu, Shuai, et al.
Pubblicazione: (2025)
di: Niu, Shuai, et al.
Pubblicazione: (2025)
A Confidence-Diversity Framework for Calibrating AI Judgement in Accessible Qualitative Coding Tasks
di: Zhao, Zhilong, et al.
Pubblicazione: (2025)
di: Zhao, Zhilong, et al.
Pubblicazione: (2025)
Multi-Model Synthetic Training for Mission-Critical Small Language Models
di: Platt, Nolan, et al.
Pubblicazione: (2025)
di: Platt, Nolan, et al.
Pubblicazione: (2025)
PairCFR: Enhancing Model Training on Paired Counterfactually Augmented Data through Contrastive Learning
di: Qiu, Xiaoqi, et al.
Pubblicazione: (2024)
di: Qiu, Xiaoqi, et al.
Pubblicazione: (2024)
How Pruning Reshapes Features: Sparse Autoencoder Analysis of Weight-Pruned Language Models
di: Borobia, Hector, et al.
Pubblicazione: (2026)
di: Borobia, Hector, et al.
Pubblicazione: (2026)
The Concept Allocation Zone: Tracking How Concepts Form Across Transformer Depth
di: Henry, James
Pubblicazione: (2026)
di: Henry, James
Pubblicazione: (2026)
The Geometry of Persona: Disentangling Personality from Reasoning in Large Language Models
di: Wang, Zhixiang
Pubblicazione: (2025)
di: Wang, Zhixiang
Pubblicazione: (2025)
WebCanvas: Benchmarking Web Agents in Online Environments
di: Pan, Yichen, et al.
Pubblicazione: (2024)
di: Pan, Yichen, et al.
Pubblicazione: (2024)
Empowering Tabular Data Preparation with Language Models: Why and How?
di: Chen, Mengshi, et al.
Pubblicazione: (2025)
di: Chen, Mengshi, et al.
Pubblicazione: (2025)
TaylorShift: Shifting the Complexity of Self-Attention from Squared to Linear (and Back) using Taylor-Softmax
di: Nauen, Tobias Christian, et al.
Pubblicazione: (2024)
di: Nauen, Tobias Christian, et al.
Pubblicazione: (2024)
Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining
di: Mitchell, Rupert, et al.
Pubblicazione: (2025)
di: Mitchell, Rupert, et al.
Pubblicazione: (2025)
$\text{Memory}^3$: Language Modeling with Explicit Memory
di: Yang, Hongkang, et al.
Pubblicazione: (2024)
di: Yang, Hongkang, et al.
Pubblicazione: (2024)
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
di: Guo, Dongxin, et al.
Pubblicazione: (2026)
di: Guo, Dongxin, et al.
Pubblicazione: (2026)
Exploring the Effectiveness of Instruction Tuning in Biomedical Language Processing
di: Rohanian, Omid, et al.
Pubblicazione: (2023)
di: Rohanian, Omid, et al.
Pubblicazione: (2023)
ATLAS: Constitution-Conditioned Latent Geometry and Redistribution Across Language Models and Neural Perturbation Data
di: Seneque, Gareth, et al.
Pubblicazione: (2026)
di: Seneque, Gareth, et al.
Pubblicazione: (2026)
Generative AI for Enhancing Active Learning in Education: A Comparative Study of GPT-3.5 and GPT-4 in Crafting Customized Test Questions
di: Rouzegar, Hamdireza, et al.
Pubblicazione: (2024)
di: Rouzegar, Hamdireza, et al.
Pubblicazione: (2024)
Investigating Distributions of Telecom Adapted Sentence Embeddings for Document Retrieval
di: Roychowdhury, Sujoy, et al.
Pubblicazione: (2024)
di: Roychowdhury, Sujoy, et al.
Pubblicazione: (2024)
Clinical information extraction for Low-resource languages with Few-shot learning using Pre-trained language models and Prompting
di: Richter-Pechanski, Phillip, et al.
Pubblicazione: (2024)
di: Richter-Pechanski, Phillip, et al.
Pubblicazione: (2024)
From PDF to RAG-Ready: Evaluating Document Conversion Frameworks for Domain-Specific Question Answering
di: Santos, José Guilherme Marques dos, et al.
Pubblicazione: (2026)
di: Santos, José Guilherme Marques dos, et al.
Pubblicazione: (2026)
KIT-TIP-NLP at MultiPride: Continual Learning with Multilingual Foundation Model
di: HB, Barathi Ganesh, et al.
Pubblicazione: (2026)
di: HB, Barathi Ganesh, et al.
Pubblicazione: (2026)
CSTRL: Context-Driven Sequential Transfer Learning for Abstractive Radiology Report Summarization
di: Naznin, Mst. Fahmida Sultana, et al.
Pubblicazione: (2025)
di: Naznin, Mst. Fahmida Sultana, et al.
Pubblicazione: (2025)
Lightweight Transformers for Clinical Natural Language Processing
di: Rohanian, Omid, et al.
Pubblicazione: (2023)
di: Rohanian, Omid, et al.
Pubblicazione: (2023)
TRUE: A Trustworthy Unified Explanation Framework for Large Language Model Reasoning
di: Yang, Yujiao
Pubblicazione: (2026)
di: Yang, Yujiao
Pubblicazione: (2026)
Fast and Fluent Diffusion Language Models via Convolutional Decoding and Rejective Fine-tuning
di: Seo, Yeongbin, et al.
Pubblicazione: (2025)
di: Seo, Yeongbin, et al.
Pubblicazione: (2025)
Aligning Black-box Language Models with Human Judgments
di: Burg, Gerrit J. J. van den, et al.
Pubblicazione: (2025)
di: Burg, Gerrit J. J. van den, et al.
Pubblicazione: (2025)
Sparse Logit Sampling: Accelerating Knowledge Distillation in LLMs
di: Anshumann, et al.
Pubblicazione: (2025)
di: Anshumann, et al.
Pubblicazione: (2025)
Survey and Evaluation of Converging Architecture in LLMs based on Footsteps of Operations
di: Kim, Seongho, et al.
Pubblicazione: (2024)
di: Kim, Seongho, et al.
Pubblicazione: (2024)
Jina Embeddings 2: 8192-Token General-Purpose Text Embeddings for Long Documents
di: Günther, Michael, et al.
Pubblicazione: (2023)
di: Günther, Michael, et al.
Pubblicazione: (2023)
Bielik v3 Small: Technical Report
di: Ociepa, Krzysztof, et al.
Pubblicazione: (2025)
di: Ociepa, Krzysztof, et al.
Pubblicazione: (2025)
Recursive Training Loops in LLMs: How training data properties modulate distribution shift in generated data?
di: Kovač, Grgur, et al.
Pubblicazione: (2025)
di: Kovač, Grgur, et al.
Pubblicazione: (2025)
Fine-tuning Large Language Models for Entity Matching
di: Steiner, Aaron, et al.
Pubblicazione: (2024)
di: Steiner, Aaron, et al.
Pubblicazione: (2024)
ENIGMA: The Geometry of Reasoning and Alignment in Large-Language Models
di: Seneque, Gareth, et al.
Pubblicazione: (2025)
di: Seneque, Gareth, et al.
Pubblicazione: (2025)
Enhancing Text Classification through LLM-Driven Active Learning and Human Annotation
di: Rouzegar, Hamidreza, et al.
Pubblicazione: (2024)
di: Rouzegar, Hamidreza, et al.
Pubblicazione: (2024)
Revisiting a Pain in the Neck: Semantic Phrase Processing Benchmark for Language Models
di: Liu, Yang, et al.
Pubblicazione: (2024)
di: Liu, Yang, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Exploring and Reshaping the Weight Distribution in LLM
di: Ye, Chunming, et al.
Pubblicazione: (2025) -
Distractor Injection Attacks on Large Reasoning Models: Characterization and Defense
di: Zhang, Zhehao, et al.
Pubblicazione: (2025) -
Analysis on distribution and clustering of weight
di: Ye, Chunming, et al.
Pubblicazione: (2025) -
Investigating on RLHF methodology
di: Kutalev, Alexey, et al.
Pubblicazione: (2024) -
CIKT: A Collaborative and Iterative Knowledge Tracing Framework with Large Language Models
di: Li, Runze, et al.
Pubblicazione: (2025)