Trading Complexity for Expressivity Through Structured Generalized Linear Token Mixing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fagnou, Erwan, Caillon, Paul, Delattre, Blaise, Allauzen, Alexandre |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Chain and Causal Attention for Efficient Entity Tracking
von: Fagnou, Erwan, et al.
Veröffentlicht: (2024)
von: Fagnou, Erwan, et al.
Veröffentlicht: (2024)
Structured-Sparse Attention for Entity Tracking with Subquadratic Sequence Complexity
von: Zhao, Hangyue, et al.
Veröffentlicht: (2026)
von: Zhao, Hangyue, et al.
Veröffentlicht: (2026)
Accelerated Training through Iterative Gradient Propagation Along the Residual Path
von: Fagnou, Erwan, et al.
Veröffentlicht: (2025)
von: Fagnou, Erwan, et al.
Veröffentlicht: (2025)
Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models
von: Shravan, Rohan
Veröffentlicht: (2026)
von: Shravan, Rohan
Veröffentlicht: (2026)
Future Token Prediction -- Causal Language Modelling with Per-Token Semantic State Vector for Multi-Token Prediction
von: Walker, Nicholas
Veröffentlicht: (2024)
von: Walker, Nicholas
Veröffentlicht: (2024)
Structure-Guided Entity Resolution: Fine-Tuning LLMs for Robust Name Matching in Complex Linguistic Contexts
von: Chourasia, Shivam, et al.
Veröffentlicht: (2026)
von: Chourasia, Shivam, et al.
Veröffentlicht: (2026)
Token Erasure as a Footprint of Implicit Vocabulary Items in LLMs
von: Feucht, Sheridan, et al.
Veröffentlicht: (2024)
von: Feucht, Sheridan, et al.
Veröffentlicht: (2024)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023)
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023)
BrahmicTokenizer-131K: An Indic-Capable Drop-In Replacement for o200k_base
von: Shravan, Rohan
Veröffentlicht: (2026)
von: Shravan, Rohan
Veröffentlicht: (2026)
Bridging the Theoretical Gap in Randomized Smoothing
von: Delattre, Blaise, et al.
Veröffentlicht: (2025)
von: Delattre, Blaise, et al.
Veröffentlicht: (2025)
Pre-trained Models Perform the Best When Token Distributions Follow Zipf's Law
von: He, Yanjin, et al.
Veröffentlicht: (2025)
von: He, Yanjin, et al.
Veröffentlicht: (2025)
RWKV-7 "Goose" with Expressive Dynamic State Evolution
von: Peng, Bo, et al.
Veröffentlicht: (2025)
von: Peng, Bo, et al.
Veröffentlicht: (2025)
Forward Only Learning for Orthogonal Neural Networks of any Depth
von: Caillon, Paul, et al.
Veröffentlicht: (2025)
von: Caillon, Paul, et al.
Veröffentlicht: (2025)
Detecting Hallucinations in Large Language Model Generation: A Token Probability Approach
von: Quevedo, Ernesto, et al.
Veröffentlicht: (2024)
von: Quevedo, Ernesto, et al.
Veröffentlicht: (2024)
SpectralLoRA: Is Low-Frequency Structure Sufficient for LoRA Adaptation? A Spectral Analysis of Weight Updates
von: Singh, Rajveer
Veröffentlicht: (2026)
von: Singh, Rajveer
Veröffentlicht: (2026)
ZERA: Zero-init Instruction Evolving Refinement Agent -- From Zero Instructions to Structured Prompts via Principle-based Optimization
von: Yi, Seungyoun, et al.
Veröffentlicht: (2025)
von: Yi, Seungyoun, et al.
Veröffentlicht: (2025)
Geometric Deviation as an Unsupervised Pre-Generation Reliability Signal: Probing LLM Representations for Answerability
von: Du, Yucheng
Veröffentlicht: (2026)
von: Du, Yucheng
Veröffentlicht: (2026)
GenAI Content Detection Task 3: Cross-Domain Machine-Generated Text Detection Challenge
von: Dugan, Liam, et al.
Veröffentlicht: (2025)
von: Dugan, Liam, et al.
Veröffentlicht: (2025)
Decoding-Time Debiasing via Process Reward Models: From Controlled Fill-in to Open-Ended Generation
von: Khan, Muneeb Ur Raheem
Veröffentlicht: (2026)
von: Khan, Muneeb Ur Raheem
Veröffentlicht: (2026)
Large Language Model (LLM) Bias Index -- LLMBI
von: Oketunji, Abiodun Finbarrs, et al.
Veröffentlicht: (2023)
von: Oketunji, Abiodun Finbarrs, et al.
Veröffentlicht: (2023)
Engineering A Large Language Model From Scratch
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2024)
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2024)
Thread Detection and Response Generation using Transformers with Prompt Optimisation
von: T, Kevin Joshua, et al.
Veröffentlicht: (2024)
von: T, Kevin Joshua, et al.
Veröffentlicht: (2024)
Structured Prompt Optimization Meets Reinforcement Learning for Global and Local Interpretability over Complex Text
von: Zhou, Tianyang, et al.
Veröffentlicht: (2026)
von: Zhou, Tianyang, et al.
Veröffentlicht: (2026)
Ouroboros: Dynamic Weight Generation for Recursive Transformers via Input-Conditioned LoRA Modulation
von: Jaber, Jaber, et al.
Veröffentlicht: (2026)
von: Jaber, Jaber, et al.
Veröffentlicht: (2026)
MMSciBench: Benchmarking Language Models on Chinese Multimodal Scientific Problems
von: Ye, Xinwu, et al.
Veröffentlicht: (2025)
von: Ye, Xinwu, et al.
Veröffentlicht: (2025)
Influence-driven Curriculum Learning for Pre-training on Limited Data
von: Schoenegger, Loris, et al.
Veröffentlicht: (2025)
von: Schoenegger, Loris, et al.
Veröffentlicht: (2025)
Towards Latent Diffusion Suitable For Text
von: Midavaine, Nesta, et al.
Veröffentlicht: (2026)
von: Midavaine, Nesta, et al.
Veröffentlicht: (2026)
BLP-2023 Task 2: Sentiment Analysis
von: Hasan, Md. Arid, et al.
Veröffentlicht: (2023)
von: Hasan, Md. Arid, et al.
Veröffentlicht: (2023)
FIM-LoRA: Task-Informative Rank Allocation for LoRA via Calibration-Time Gradient-Variance Estimation
von: Sathyavageeswaran, Ramakrishnan
Veröffentlicht: (2026)
von: Sathyavageeswaran, Ramakrishnan
Veröffentlicht: (2026)
Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
von: Hanna, Michael, et al.
Veröffentlicht: (2024)
von: Hanna, Michael, et al.
Veröffentlicht: (2024)
Integrating Expert Labels into LLM-based Emission Goal Detection: Example Selection vs Automatic Prompt Design
von: Wrzalik, Marco, et al.
Veröffentlicht: (2024)
von: Wrzalik, Marco, et al.
Veröffentlicht: (2024)
Amulet: ReAlignment During Test Time for Personalized Preference Adaptation of LLMs
von: Zhang, Zhaowei, et al.
Veröffentlicht: (2025)
von: Zhang, Zhaowei, et al.
Veröffentlicht: (2025)
Where Should LoRA Go? Component-Type Placement in Hybrid Language Models
von: Borobia, Hector, et al.
Veröffentlicht: (2026)
von: Borobia, Hector, et al.
Veröffentlicht: (2026)
Interpreto: An Explainability Library for Transformers
von: Poché, Antonin, et al.
Veröffentlicht: (2025)
von: Poché, Antonin, et al.
Veröffentlicht: (2025)
RightNow-Arabic-0.5B-Turbo: An Open Sub-1B Arabic Language Model via Vocabulary Injection and Edge-First Deployment
von: Jaber, Jaber, et al.
Veröffentlicht: (2026)
von: Jaber, Jaber, et al.
Veröffentlicht: (2026)
Efficacy of ByT5 in Multilingual Translation of Biblical Texts for Underrepresented Languages
von: Aars, Corinne, et al.
Veröffentlicht: (2024)
von: Aars, Corinne, et al.
Veröffentlicht: (2024)
WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training
von: Tian, Changxin, et al.
Veröffentlicht: (2025)
von: Tian, Changxin, et al.
Veröffentlicht: (2025)
PowLU: An Activation Function for Stable Pre-Training of LLMs
von: Jiang, Peijie, et al.
Veröffentlicht: (2026)
von: Jiang, Peijie, et al.
Veröffentlicht: (2026)
Resonant Context Anchoring: Decoupling Attention Routing and Signal Gain at Inference Time
von: Zhao, Mingkuan, et al.
Veröffentlicht: (2026)
von: Zhao, Mingkuan, et al.
Veröffentlicht: (2026)
Reconstructing Syllable Sequences in Abugida Scripts with Incomplete Inputs
von: Thu, Ye Kyaw, et al.
Veröffentlicht: (2025)
von: Thu, Ye Kyaw, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Chain and Causal Attention for Efficient Entity Tracking
von: Fagnou, Erwan, et al.
Veröffentlicht: (2024) -
Structured-Sparse Attention for Entity Tracking with Subquadratic Sequence Complexity
von: Zhao, Hangyue, et al.
Veröffentlicht: (2026) -
Accelerated Training through Iterative Gradient Propagation Along the Residual Path
von: Fagnou, Erwan, et al.
Veröffentlicht: (2025) -
Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models
von: Shravan, Rohan
Veröffentlicht: (2026) -
Future Token Prediction -- Causal Language Modelling with Per-Token Semantic State Vector for Multi-Token Prediction
von: Walker, Nicholas
Veröffentlicht: (2024)