BLaST: High Performance Inference and Pretraining using BLock Sparse Transformers
Fuente:
arXiv
Guardado en:
| Autores principales: | Okanovic, Patrik, Deshmukh, Sameer, Kwasniewski, Grzegorz, Zhu, Yi, Fujii, Haruto, Fatima, Sakina, Besta, Maciej, Katayama, Kentaro, Honda, Takumi, Nagasaka, Yusuke, Hoefler, Torsten |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
High Performance Unstructured SpMM Computation Using Tensor Cores
por: Okanovic, Patrik, et al.
Publicado: (2024)
por: Okanovic, Patrik, et al.
Publicado: (2024)
EntryPrune: Neural Network Feature Selection using First Impressions
por: Zimmer, Felix, et al.
Publicado: (2024)
por: Zimmer, Felix, et al.
Publicado: (2024)
Demystifying Higher-Order Graph Neural Networks
por: Besta, Maciej, et al.
Publicado: (2024)
por: Besta, Maciej, et al.
Publicado: (2024)
FoldedHexaTorus: An Inter-Chiplet Interconnect Topology for Chiplet-based Systems using Organic and Glass Substrates
por: Iff, Patrick, et al.
Publicado: (2025)
por: Iff, Patrick, et al.
Publicado: (2025)
Confounder Detection via Treatment Intent: A New Observational Study Design
por: Plecko, Drago, et al.
Publicado: (2026)
por: Plecko, Drago, et al.
Publicado: (2026)
Epidemiology of Large Language Models: A Benchmark for Observational Distribution Knowledge
por: Plecko, Drago, et al.
Publicado: (2025)
por: Plecko, Drago, et al.
Publicado: (2025)
Large Language Model Selection with Limited Annotations
por: Durmazkeser, Yavuz, et al.
Publicado: (2026)
por: Durmazkeser, Yavuz, et al.
Publicado: (2026)
Active Model Selection for Large Language Models
por: Durmazkeser, Yavuz, et al.
Publicado: (2025)
por: Durmazkeser, Yavuz, et al.
Publicado: (2025)
PlaceIT: Placement-based Inter-Chiplet Interconnect Topologies
por: Iff, Patrick, et al.
Publicado: (2025)
por: Iff, Patrick, et al.
Publicado: (2025)
Network Design for Wafer-Scale Systems with Wafer-on-Wafer Hybrid Bonding
por: Iff, Patrick, et al.
Publicado: (2026)
por: Iff, Patrick, et al.
Publicado: (2026)
All models are wrong, some are useful: Model Selection with Limited Labels
por: Okanovic, Patrik, et al.
Publicado: (2024)
por: Okanovic, Patrik, et al.
Publicado: (2024)
Benchmarking Filtered Approximate Nearest Neighbor Search Algorithms on Transformer-based Embedding Vectors
por: Iff, Patrick, et al.
Publicado: (2025)
por: Iff, Patrick, et al.
Publicado: (2025)
Spritz: Path-Aware Load Balancing in Low-Diameter Networks
por: Bonato, Tommaso, et al.
Publicado: (2026)
por: Bonato, Tommaso, et al.
Publicado: (2026)
RapidChiplet: A Toolchain for Rapid Design Space Exploration of Chiplet Architectures
por: Iff, Patrick, et al.
Publicado: (2023)
por: Iff, Patrick, et al.
Publicado: (2023)
Hardware Acceleration for Knowledge Graph Processing: Challenges & Recent Developments
por: Besta, Maciej, et al.
Publicado: (2024)
por: Besta, Maciej, et al.
Publicado: (2024)
Low-Depth Spatial Tree Algorithms
por: Baumann, Yves, et al.
Publicado: (2024)
por: Baumann, Yves, et al.
Publicado: (2024)
PolarStar: Expanding the Scalability Horizon of Diameter-3 Networks
por: Lakhotia, Kartik, et al.
Publicado: (2023)
por: Lakhotia, Kartik, et al.
Publicado: (2023)
HOT: Higher-Order Dynamic Graph Representation Learning with Efficient Transformers
por: Besta, Maciej, et al.
Publicado: (2023)
por: Besta, Maciej, et al.
Publicado: (2023)
Multi-Head RAG: Solving Multi-Aspect Problems with LLMs
por: Besta, Maciej, et al.
Publicado: (2024)
por: Besta, Maciej, et al.
Publicado: (2024)
Memory-Efficient LLM Training with Dynamic Sparsity: From Stability to Practical Scaling
por: Xiao, Qiao, et al.
Publicado: (2026)
por: Xiao, Qiao, et al.
Publicado: (2026)
When Data Is Scarce: Scaling Sparse Language Models with Repeated Training
por: Wu, Boqian, et al.
Publicado: (2026)
por: Wu, Boqian, et al.
Publicado: (2026)
Demystifying Chains, Trees, and Graphs of Thoughts
por: Besta, Maciej, et al.
Publicado: (2024)
por: Besta, Maciej, et al.
Publicado: (2024)
Higher-Order Graph Databases
por: Besta, Maciej, et al.
Publicado: (2025)
por: Besta, Maciej, et al.
Publicado: (2025)
BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding
por: Yuan, Jiayi, et al.
Publicado: (2025)
por: Yuan, Jiayi, et al.
Publicado: (2025)
GraphSeek: Next-Generation Graph Analytics with LLMs
por: Besta, Maciej, et al.
Publicado: (2026)
por: Besta, Maciej, et al.
Publicado: (2026)
Reasoning Language Models: A Blueprint
por: Besta, Maciej, et al.
Publicado: (2025)
por: Besta, Maciej, et al.
Publicado: (2025)
Ab-initio Quantum Transport with the GW Approximation, 42,240 Atoms, and Sustained Exascale Performance
por: Vetsch, Nicolas, et al.
Publicado: (2025)
por: Vetsch, Nicolas, et al.
Publicado: (2025)
Affordable AI Assistants with Knowledge Graph of Thoughts
por: Besta, Maciej, et al.
Publicado: (2025)
por: Besta, Maciej, et al.
Publicado: (2025)
SpComm3D: A Framework for Enabling Sparse Communication in 3D Sparse Kernels
por: Abubaker, Nabil, et al.
Publicado: (2024)
por: Abubaker, Nabil, et al.
Publicado: (2024)
Near-Optimal Sparse Allreduce for Distributed Deep Learning
por: Li, Shigang, et al.
Publicado: (2022)
por: Li, Shigang, et al.
Publicado: (2022)
Chimera: Efficiently Training Large-Scale Neural Networks with Bidirectional Pipelines
por: Li, Shigang, et al.
Publicado: (2021)
por: Li, Shigang, et al.
Publicado: (2021)
Graph of Thoughts: Solving Elaborate Problems with Large Language Models
por: Besta, Maciej, et al.
Publicado: (2023)
por: Besta, Maciej, et al.
Publicado: (2023)
FlakyFix: Using Large Language Models for Predicting Flaky Test Fix Categories and Test Code Repair
por: Fatima, Sakina, et al.
Publicado: (2023)
por: Fatima, Sakina, et al.
Publicado: (2023)
HSTU-BLaIR: Lightweight Contrastive Text Embedding for Generative Recommender
por: Liu, Yijun
Publicado: (2025)
por: Liu, Yijun
Publicado: (2025)
CheckEmbed: Effective Verification of LLM Solutions to Open-Ended Tasks
por: Besta, Maciej, et al.
Publicado: (2024)
por: Besta, Maciej, et al.
Publicado: (2024)
Holonomy preserving transformations of weighted graphs and its application to knot theory
por: Nagasaka, Atsuhide
Publicado: (2025)
por: Nagasaka, Atsuhide
Publicado: (2025)
Computation of geodesics and rhumb lines on spheroid
por: Nagasaka, Naohiko
Publicado: (2013)
por: Nagasaka, Naohiko
Publicado: (2013)
Evaluation of the Automated Labeling Method for Taxonomic Nomenclature Through Prompt-Optimized Large Language Model
por: Inoshita, Keito, et al.
Publicado: (2025)
por: Inoshita, Keito, et al.
Publicado: (2025)
Psychologically Enhanced AI Agents
por: Besta, Maciej, et al.
Publicado: (2025)
por: Besta, Maciej, et al.
Publicado: (2025)
FPsPIN: An FPGA-based Open-Hardware Research Platform for Processing in the Network
por: Schneider, Timo, et al.
Publicado: (2024)
por: Schneider, Timo, et al.
Publicado: (2024)
Ejemplares similares
-
High Performance Unstructured SpMM Computation Using Tensor Cores
por: Okanovic, Patrik, et al.
Publicado: (2024) -
EntryPrune: Neural Network Feature Selection using First Impressions
por: Zimmer, Felix, et al.
Publicado: (2024) -
Demystifying Higher-Order Graph Neural Networks
por: Besta, Maciej, et al.
Publicado: (2024) -
FoldedHexaTorus: An Inter-Chiplet Interconnect Topology for Chiplet-based Systems using Organic and Glass Substrates
por: Iff, Patrick, et al.
Publicado: (2025) -
Confounder Detection via Treatment Intent: A New Observational Study Design
por: Plecko, Drago, et al.
Publicado: (2026)