Dynamic sparsity in tree-structured feed-forward layers at scale
Fuente:
arXiv
Salvato in:
| Autori principali: | Sedghi, Reza, Schiewer, Robin, Subramoney, Anand, Kappel, David |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Utilizing dynamic sparsity on pretrained DETR
di: Sedghi, Reza, et al.
Pubblicazione: (2025)
di: Sedghi, Reza, et al.
Pubblicazione: (2025)
Exploring the limits of Hierarchical World Models in Reinforcement Learning
di: Schiewer, Robin, et al.
Pubblicazione: (2024)
di: Schiewer, Robin, et al.
Pubblicazione: (2024)
State-space models can learn in-context by gradient descent
di: Sushma, Neeraj Mohan, et al.
Pubblicazione: (2024)
di: Sushma, Neeraj Mohan, et al.
Pubblicazione: (2024)
Weight Sparsity Complements Activity Sparsity in Neuromorphic Language Models
di: Mukherji, Rishav, et al.
Pubblicazione: (2024)
di: Mukherji, Rishav, et al.
Pubblicazione: (2024)
Scalable Event-by-event Processing of Neuromorphic Sensory Signals With Deep State-Space Models
di: Schöne, Mark, et al.
Pubblicazione: (2024)
di: Schöne, Mark, et al.
Pubblicazione: (2024)
Probing Length Generalization in Mamba via Image Reconstruction
di: Rathjens, Jan, et al.
Pubblicazione: (2026)
di: Rathjens, Jan, et al.
Pubblicazione: (2026)
AKReF: An argumentative knowledge representation framework for structured argumentation
di: Bhattacharjee, Debarati, et al.
Pubblicazione: (2025)
di: Bhattacharjee, Debarati, et al.
Pubblicazione: (2025)
STREAM: A Universal State-Space Model for Sparse Geometric Data
di: Schöne, Mark, et al.
Pubblicazione: (2024)
di: Schöne, Mark, et al.
Pubblicazione: (2024)
Asynchronous Stochastic Gradient Descent with Decoupled Backpropagation and Layer-Wise Updates
di: Fokam, Cabrel Teguemne, et al.
Pubblicazione: (2024)
di: Fokam, Cabrel Teguemne, et al.
Pubblicazione: (2024)
Exploring and Benchmarking the Planning Capabilities of Large Language Models
di: Bohnet, Bernd, et al.
Pubblicazione: (2024)
di: Bohnet, Bernd, et al.
Pubblicazione: (2024)
Distributional Associations vs In-Context Reasoning: A Study of Feed-forward and Attention Layers
di: Chen, Lei, et al.
Pubblicazione: (2024)
di: Chen, Lei, et al.
Pubblicazione: (2024)
Language Modeling on a SpiNNaker 2 Neuromorphic Chip
di: Nazeer, Khaleelulla Khan, et al.
Pubblicazione: (2023)
di: Nazeer, Khaleelulla Khan, et al.
Pubblicazione: (2023)
Loss-to-Loss Prediction: Scaling Laws for All Datasets
di: Brandfonbrener, David, et al.
Pubblicazione: (2024)
di: Brandfonbrener, David, et al.
Pubblicazione: (2024)
Leviathan: Decoupling Input and Output Representations in Language Models
di: Batley, Reza T., et al.
Pubblicazione: (2026)
di: Batley, Reza T., et al.
Pubblicazione: (2026)
Promote, Suppress, Iterate: How Language Models Answer One-to-Many Factual Queries
di: Yan, Tianyi Lorena, et al.
Pubblicazione: (2025)
di: Yan, Tianyi Lorena, et al.
Pubblicazione: (2025)
Function Induction and Task Generalization: An Interpretability Study with Off-by-One Addition
di: Ye, Qinyuan, et al.
Pubblicazione: (2025)
di: Ye, Qinyuan, et al.
Pubblicazione: (2025)
Rhetorical Questions in LLM Representations: A Linear Probing Study
di: Yao, Louie Hong, et al.
Pubblicazione: (2026)
di: Yao, Louie Hong, et al.
Pubblicazione: (2026)
Neural Isomorphic Fields: A Transformer-based Algebraic Numerical Embedding
di: Sadeghi, Hamidreza, et al.
Pubblicazione: (2026)
di: Sadeghi, Hamidreza, et al.
Pubblicazione: (2026)
Decoupling the "What" and "Where" With Polar Coordinate Positional Embeddings
di: Gopalakrishnan, Anand, et al.
Pubblicazione: (2025)
di: Gopalakrishnan, Anand, et al.
Pubblicazione: (2025)
Toxicity Detection Should Measure Contextual Harm, Not Text-Intrinsic Badness
di: Berezin, Sergei, et al.
Pubblicazione: (2025)
di: Berezin, Sergei, et al.
Pubblicazione: (2025)
s1: Simple test-time scaling
di: Muennighoff, Niklas, et al.
Pubblicazione: (2025)
di: Muennighoff, Niklas, et al.
Pubblicazione: (2025)
Single layer tiny Co$^4$ outpaces GPT-2 and GPT-BERT
di: Zain, Noor Ul, et al.
Pubblicazione: (2025)
di: Zain, Noor Ul, et al.
Pubblicazione: (2025)
ContextFocus: Activation Steering for Contextual Faithfulness in Large Language Models
di: Anand, Nikhil, et al.
Pubblicazione: (2026)
di: Anand, Nikhil, et al.
Pubblicazione: (2026)
Mini-batch Coresets for Memory-efficient Language Model Training on Data Mixtures
di: Nguyen, Dang, et al.
Pubblicazione: (2024)
di: Nguyen, Dang, et al.
Pubblicazione: (2024)
Training event-based neural networks with exact gradients via Differentiable ODE Solving in JAX
di: König, Lukas, et al.
Pubblicazione: (2026)
di: König, Lukas, et al.
Pubblicazione: (2026)
Decomposing Elements of Problem Solving: What "Math" Does RL Teach?
di: Qin, Tian, et al.
Pubblicazione: (2025)
di: Qin, Tian, et al.
Pubblicazione: (2025)
Integrating Time Series into LLMs via Multi-layer Steerable Embedding Fusion for Enhanced Forecasting
di: Chen, Zhuomin, et al.
Pubblicazione: (2025)
di: Chen, Zhuomin, et al.
Pubblicazione: (2025)
A Synthetic Dataset for Personal Attribute Inference
di: Yukhymenko, Hanna, et al.
Pubblicazione: (2024)
di: Yukhymenko, Hanna, et al.
Pubblicazione: (2024)
Transformers Learn to Achieve Second-Order Convergence Rates for In-Context Linear Regression
di: Fu, Deqing, et al.
Pubblicazione: (2023)
di: Fu, Deqing, et al.
Pubblicazione: (2023)
Automatic explanation of the classification of Spanish legal judgments in jurisdiction-dependent law categories with tree estimators
di: González-González, Jaime, et al.
Pubblicazione: (2024)
di: González-González, Jaime, et al.
Pubblicazione: (2024)
LoRA Fine-Tuning Without GPUs: A CPU-Efficient Meta-Generation Framework for LLMs
di: Arabpour, Reza, et al.
Pubblicazione: (2025)
di: Arabpour, Reza, et al.
Pubblicazione: (2025)
Unlocking In-Context Learning for Natural Datasets Beyond Language Modelling
di: Bratulić, Jelena, et al.
Pubblicazione: (2025)
di: Bratulić, Jelena, et al.
Pubblicazione: (2025)
How Do LLMs Persuade? Linear Probes Can Uncover Persuasion Dynamics in Multi-Turn Conversations
di: Jaipersaud, Brandon, et al.
Pubblicazione: (2025)
di: Jaipersaud, Brandon, et al.
Pubblicazione: (2025)
Convergent Evolution: How Different Language Models Learn Similar Number Representations
di: Fu, Deqing, et al.
Pubblicazione: (2026)
di: Fu, Deqing, et al.
Pubblicazione: (2026)
LLM Unlearning Without an Expert Curated Dataset
di: Zhu, Xiaoyuan, et al.
Pubblicazione: (2025)
di: Zhu, Xiaoyuan, et al.
Pubblicazione: (2025)
Adaptive Self-Supervised Learning Strategies for Dynamic On-Device LLM Personalization
di: Mendoza, Rafael, et al.
Pubblicazione: (2024)
di: Mendoza, Rafael, et al.
Pubblicazione: (2024)
Align to Structure: Aligning Large Language Models with Structural Information
di: Kim, Zae Myung, et al.
Pubblicazione: (2025)
di: Kim, Zae Myung, et al.
Pubblicazione: (2025)
Safety and accuracy follow different scaling laws in clinical large language models
di: Wind, Sebastian, et al.
Pubblicazione: (2026)
di: Wind, Sebastian, et al.
Pubblicazione: (2026)
Magic Words or Methodical Work? Challenging Conventional Wisdom in LLM-Based Political Text Annotation
di: McLaren, Lorcan, et al.
Pubblicazione: (2026)
di: McLaren, Lorcan, et al.
Pubblicazione: (2026)
CTBench: A Comprehensive Benchmark for Evaluating Language Model Capabilities in Clinical Trial Design
di: Neehal, Nafis, et al.
Pubblicazione: (2024)
di: Neehal, Nafis, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Utilizing dynamic sparsity on pretrained DETR
di: Sedghi, Reza, et al.
Pubblicazione: (2025) -
Exploring the limits of Hierarchical World Models in Reinforcement Learning
di: Schiewer, Robin, et al.
Pubblicazione: (2024) -
State-space models can learn in-context by gradient descent
di: Sushma, Neeraj Mohan, et al.
Pubblicazione: (2024) -
Weight Sparsity Complements Activity Sparsity in Neuromorphic Language Models
di: Mukherji, Rishav, et al.
Pubblicazione: (2024) -
Scalable Event-by-event Processing of Neuromorphic Sensory Signals With Deep State-Space Models
di: Schöne, Mark, et al.
Pubblicazione: (2024)