Matrix-Free 3D SIMP Topology Optimization with Fused Gather-GEMM-Scatter Kernels
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Shaoliang, Wang, Jun, Wang, Yunsheng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mass Matrix Assembly on Tensor Cores for Implicit Particle-In-Cell Methods
by: Pennati, Luca, et al.
Published: (2026)
by: Pennati, Luca, et al.
Published: (2026)
Optimizing the Weather Research and Forecasting Model with OpenMP Offload and Codee
by: Chayanon, et al.
Published: (2024)
by: Chayanon, et al.
Published: (2024)
A GPU-boosted high-performance multi-working condition joint analysis framework for predicting dynamics of textured axial piston pump
by: Yao, Xin, et al.
Published: (2025)
by: Yao, Xin, et al.
Published: (2025)
The Missing Adapter Layer for Research Computing
by: Li, Bowen, et al.
Published: (2026)
by: Li, Bowen, et al.
Published: (2026)
Position: LLM Inference Should Be Evaluated as Energy-to-Token Production
by: Liu, Xiang, et al.
Published: (2026)
by: Liu, Xiang, et al.
Published: (2026)
IterSIMP-σ: Evaluating LLM-Assisted Spatial Interventions in Stress-Aware Topology Optimization
by: Yang, Shaoliang, et al.
Published: (2026)
by: Yang, Shaoliang, et al.
Published: (2026)
Large Language Models as Optimization Controllers: Adaptive Continuation for SIMP Topology Optimization
by: Yang, Shaoliang, et al.
Published: (2026)
by: Yang, Shaoliang, et al.
Published: (2026)
A Matrix-Free Galerkin Multigrid Solver and Failure-Mode Screen for Single-GPU 3D SIMP Linear Systems
by: Yang, Shaoliang, et al.
Published: (2026)
by: Yang, Shaoliang, et al.
Published: (2026)
Reclaiming Idle CPU Cycles on Kubernetes: Sparse-Domain Multiplexing for Concurrent MPI-CFD Simulations
by: Xie, Tianfang
Published: (2026)
by: Xie, Tianfang
Published: (2026)
Smoothed aggregation algebraic multigrid for problems with heterogeneous and anisotropic materials
by: Firmbach, Max, et al.
Published: (2026)
by: Firmbach, Max, et al.
Published: (2026)
From Impermanent Loss to Sustainable Gain: Quantifying Profitability Zones for Liquidity Providers on DEX
by: Melnikov, Ignat, et al.
Published: (2026)
by: Melnikov, Ignat, et al.
Published: (2026)
Characterizing Path-Independent Fees: A Route to Zero Impermanent Loss in CPMMs
by: Voronin, Andrey, et al.
Published: (2026)
by: Voronin, Andrey, et al.
Published: (2026)
Icicle: Scalable Metadata Indexing and Real-Time Monitoring for HPC File Systems
by: Pan, Haochen, et al.
Published: (2026)
by: Pan, Haochen, et al.
Published: (2026)
Reduced and mixed precision turbulent flow simulations using explicit finite difference schemes
by: Siklósi, Bálint, et al.
Published: (2025)
by: Siklósi, Bálint, et al.
Published: (2025)
A Cloud-based Real-time Probabilistic Remaining Useful Life (RUL) Estimation using the Sequential Monte Carlo (SMC) Method
by: Lyathakula, Karthik Reddy, et al.
Published: (2024)
by: Lyathakula, Karthik Reddy, et al.
Published: (2024)
In-Memory Non-Binary LDPC Decoding
by: Ferraz, Oscar, et al.
Published: (2025)
by: Ferraz, Oscar, et al.
Published: (2025)
A parallel implementation of reduced-order modeling of large-scale systems
by: Farcas, Ionut-Gabriel, et al.
Published: (2025)
by: Farcas, Ionut-Gabriel, et al.
Published: (2025)
Making Tax Smart: Feasibility of Distributed Ledger Technology for building tax compliance functionality to Central Bank Digital Currency
by: Louvieris, Panos, et al.
Published: (2024)
by: Louvieris, Panos, et al.
Published: (2024)
GPU-accelerated Linear Algebra for Coupled Solvers in Industrial CFD Applications with OpenFOAM
by: Oliani, Stefano, et al.
Published: (2024)
by: Oliani, Stefano, et al.
Published: (2024)
A GPU-based Compressible Combustion Solver for Applications Exhibiting Disparate Space and Time Scales
by: Carreon, Anthony, et al.
Published: (2025)
by: Carreon, Anthony, et al.
Published: (2025)
Level set-based inverse homogenisation of three-dimensional piezoelectric materials
by: Wegert, Zachary J., et al.
Published: (2024)
by: Wegert, Zachary J., et al.
Published: (2024)
Intertemporal Pricing of Time-Bound Stablecoins: Measuring and Controlling the Liquidity-of-Time Premium
by: Borjigin, Ailiya, et al.
Published: (2025)
by: Borjigin, Ailiya, et al.
Published: (2025)
A Portable Multi-GPU Solver for Collisional Plasmas with Coulombic Interactions
by: Almgren-Bell, James, et al.
Published: (2025)
by: Almgren-Bell, James, et al.
Published: (2025)
A Lock-Free Work-Stealing Algorithm for Bulk Operations
by: Kataru, Raja Sai Nandhan Yadav, et al.
Published: (2026)
by: Kataru, Raja Sai Nandhan Yadav, et al.
Published: (2026)
Towards Robust Blockchain Price Oracle: A Study on Human-Centric Node Selection Strategy and Incentive Mechanism
by: Xian, Youquan, et al.
Published: (2023)
by: Xian, Youquan, et al.
Published: (2023)
Distributed Variational Quantum Algorithm with Many-qubit for Optimization Challenges
by: Kim, Seongmin, et al.
Published: (2025)
by: Kim, Seongmin, et al.
Published: (2025)
Distributed Quantum Optimization for Large-Scale Higher-Order Problems with Dense Interactions
by: Kim, Seongmin, et al.
Published: (2026)
by: Kim, Seongmin, et al.
Published: (2026)
Distributed Quantum Approximate Optimization Algorithm on a Quantum-Centric Supercomputing Architecture
by: Kim, Seongmin, et al.
Published: (2024)
by: Kim, Seongmin, et al.
Published: (2024)
High-performance Effective Scientific Error-bounded Lossy Compression with Auto-tuned Multi-component Interpolation
by: Liu, Jinyang, et al.
Published: (2023)
by: Liu, Jinyang, et al.
Published: (2023)
Characterizing GPU Energy Usage in Exascale-Ready Portable Science Applications
by: Godoy, William F., et al.
Published: (2025)
by: Godoy, William F., et al.
Published: (2025)
QPET: A Versatile and Portable Quantity-of-Interest-Preservation Framework for Error-Bounded Lossy Compression
by: Liu, Jinyang, et al.
Published: (2024)
by: Liu, Jinyang, et al.
Published: (2024)
Tuning of Vectorization Parameters for Molecular Dynamics Simulations in AutoPas
by: Gall, Luis, et al.
Published: (2025)
by: Gall, Luis, et al.
Published: (2025)
Novel Architecture for Distributed Travel Data Integration and Service Provision Using Microservices
by: Barua, Biman, et al.
Published: (2024)
by: Barua, Biman, et al.
Published: (2024)
Comparison of nested geometry treatments within GPU-based Monte Carlo neutron transport simulations of fission reactors
by: Biondo, Elliott, et al.
Published: (2024)
by: Biondo, Elliott, et al.
Published: (2024)
Mojo: MLIR-Based Performance-Portable HPC Science Kernels on GPUs for the Python Ecosystem
by: Godoy, William F., et al.
Published: (2025)
by: Godoy, William F., et al.
Published: (2025)
A Fully GPU-Accelerated Framework for High-Performance Configuration Interaction Selection with Neural Network Quantum States
by: Sun, Daran, et al.
Published: (2026)
by: Sun, Daran, et al.
Published: (2026)
Stencil Computations on Cerebras Wafer-Scale Engine
by: Belli, Elia, et al.
Published: (2026)
by: Belli, Elia, et al.
Published: (2026)
Adaptive and Parallel Multiscale Framework for Modeling Cohesive Failure in Engineering Scale Systems
by: Kim, Sion, et al.
Published: (2024)
by: Kim, Sion, et al.
Published: (2024)
Design, Implementation, and Analysis of Fair Faucets for Blockchain Ecosystems
by: Metin, Serdar
Published: (2025)
by: Metin, Serdar
Published: (2025)
waLBerla-wind: a lattice-Boltzmann-based high-performance flow solver for wind energy applications
by: Schottenhamml, Helen, et al.
Published: (2023)
by: Schottenhamml, Helen, et al.
Published: (2023)
Similar Items
-
Mass Matrix Assembly on Tensor Cores for Implicit Particle-In-Cell Methods
by: Pennati, Luca, et al.
Published: (2026) -
Optimizing the Weather Research and Forecasting Model with OpenMP Offload and Codee
by: Chayanon, et al.
Published: (2024) -
A GPU-boosted high-performance multi-working condition joint analysis framework for predicting dynamics of textured axial piston pump
by: Yao, Xin, et al.
Published: (2025) -
The Missing Adapter Layer for Research Computing
by: Li, Bowen, et al.
Published: (2026) -
Position: LLM Inference Should Be Evaluated as Energy-to-Token Production
by: Liu, Xiang, et al.
Published: (2026)