Gespeichert in:
| Hauptverfasser: | Zhang, Cheng, Cheng, Jianyi, Shumailov, Ilia, Constantinides, George A., Zhao, Yiren |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2310.05079 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LQER: Low-Rank Quantization Error Reconstruction for LLMs
von: Zhang, Cheng, et al.
Veröffentlicht: (2024)
von: Zhang, Cheng, et al.
Veröffentlicht: (2024)
Hardware and Software Platform Inference
von: Zhang, Cheng, et al.
Veröffentlicht: (2024)
von: Zhang, Cheng, et al.
Veröffentlicht: (2024)
AMPLE: Event-Driven Accelerator for Mixed-Precision Inference of Graph Neural Networks
von: Gimenes, Pedro, et al.
Veröffentlicht: (2025)
von: Gimenes, Pedro, et al.
Veröffentlicht: (2025)
Quantamination: Dynamic Quantization Leaks Your Data Across the Batch
von: Foerster, Hanna, et al.
Veröffentlicht: (2026)
von: Foerster, Hanna, et al.
Veröffentlicht: (2026)
QERA: an Analytical Framework for Quantization Error Reconstruction
von: Zhang, Cheng, et al.
Veröffentlicht: (2024)
von: Zhang, Cheng, et al.
Veröffentlicht: (2024)
ImpNet: Imperceptible and blackbox-undetectable backdoors in compiled neural networks
von: Clifford, Eleanor, et al.
Veröffentlicht: (2022)
von: Clifford, Eleanor, et al.
Veröffentlicht: (2022)
NeuraLUT-Assemble: Hardware-aware Assembling of Sub-Neural Networks for Efficient LUT Inference
von: Andronic, Marta, et al.
Veröffentlicht: (2025)
von: Andronic, Marta, et al.
Veröffentlicht: (2025)
A Dataflow Compiler for Efficient LLM Inference using Custom Microscaling Formats
von: Cheng, Jianyi, et al.
Veröffentlicht: (2023)
von: Cheng, Jianyi, et al.
Veröffentlicht: (2023)
Optimised Grouped-Query Attention Mechanism for Transformers
von: Chen, Yuang, et al.
Veröffentlicht: (2024)
von: Chen, Yuang, et al.
Veröffentlicht: (2024)
Architectural Neural Backdoors from First Principles
von: Langford, Harry, et al.
Veröffentlicht: (2024)
von: Langford, Harry, et al.
Veröffentlicht: (2024)
Locking Machine Learning Models into Hardware
von: Clifford, Eleanor, et al.
Veröffentlicht: (2024)
von: Clifford, Eleanor, et al.
Veröffentlicht: (2024)
Unlocking the Global Synergies in Low-Rank Adapters
von: Zhang, Zixi, et al.
Veröffentlicht: (2024)
von: Zhang, Zixi, et al.
Veröffentlicht: (2024)
SEA: Shareable and Explainable Attribution for Query-based Black-box Attacks
von: Gao, Yue, et al.
Veröffentlicht: (2023)
von: Gao, Yue, et al.
Veröffentlicht: (2023)
Beyond Labeling Oracles: What does it mean to steal ML models?
von: Shafran, Avital, et al.
Veröffentlicht: (2023)
von: Shafran, Avital, et al.
Veröffentlicht: (2023)
PolyLUT: Learning Piecewise Polynomials for Ultra-Low Latency FPGA LUT-based Inference
von: Andronic, Marta, et al.
Veröffentlicht: (2023)
von: Andronic, Marta, et al.
Veröffentlicht: (2023)
Machine Learning needs Better Randomness Standards: Randomised Smoothing and PRNG-based attacks
von: Dahiya, Pranav, et al.
Veröffentlicht: (2023)
von: Dahiya, Pranav, et al.
Veröffentlicht: (2023)
The Curse of Recursion: Training on Generated Data Makes Models Forget
von: Shumailov, Ilia, et al.
Veröffentlicht: (2023)
von: Shumailov, Ilia, et al.
Veröffentlicht: (2023)
Architectural Backdoors for Within-Batch Data Stealing and Model Inference Manipulation
von: Küchler, Nicolas, et al.
Veröffentlicht: (2025)
von: Küchler, Nicolas, et al.
Veröffentlicht: (2025)
Reasoning Introduces New Poisoning Attacks Yet Makes Them More Complicated
von: Foerster, Hanna, et al.
Veröffentlicht: (2025)
von: Foerster, Hanna, et al.
Veröffentlicht: (2025)
Fairness Feedback Loops: Training on Synthetic Data Amplifies Bias
von: Wyllie, Sierra, et al.
Veröffentlicht: (2024)
von: Wyllie, Sierra, et al.
Veröffentlicht: (2024)
Buffer Overflow in Mixture of Experts
von: Hayes, Jamie, et al.
Veröffentlicht: (2024)
von: Hayes, Jamie, et al.
Veröffentlicht: (2024)
LLM4DV: Using Large Language Models for Hardware Test Stimuli Generation
von: Zhang, Zixi, et al.
Veröffentlicht: (2023)
von: Zhang, Zixi, et al.
Veröffentlicht: (2023)
Scaling Laws For Mixed Quantization
von: Cao, Zeyu, et al.
Veröffentlicht: (2024)
von: Cao, Zeyu, et al.
Veröffentlicht: (2024)
Ultra-Quantisation: Efficient Embedding Search via 1.58-bit Encodings
von: Connor, Richard, et al.
Veröffentlicht: (2025)
von: Connor, Richard, et al.
Veröffentlicht: (2025)
A3 : an Analytical Low-Rank Approximation Framework for Attention
von: Wong, Jeffrey T. H., et al.
Veröffentlicht: (2025)
von: Wong, Jeffrey T. H., et al.
Veröffentlicht: (2025)
PolyLUT: Ultra-low Latency Polynomial Inference with Hardware-Aware Structured Pruning
von: Andronic, Marta, et al.
Veröffentlicht: (2025)
von: Andronic, Marta, et al.
Veröffentlicht: (2025)
ceLLMate: Sandboxing Browser AI Agents
von: Meng, Luoxi, et al.
Veröffentlicht: (2025)
von: Meng, Luoxi, et al.
Veröffentlicht: (2025)
LO-BCQ: Block Clustered Quantization for 4-bit (W4A4) LLM Inference
von: Elangovan, Reena, et al.
Veröffentlicht: (2025)
von: Elangovan, Reena, et al.
Veröffentlicht: (2025)
Watermarking Needs Input Repetition Masking
von: Khachaturov, David, et al.
Veröffentlicht: (2025)
von: Khachaturov, David, et al.
Veröffentlicht: (2025)
Beyond Slow Signs in High-fidelity Model Extraction
von: Foerster, Hanna, et al.
Veröffentlicht: (2024)
von: Foerster, Hanna, et al.
Veröffentlicht: (2024)
Measuring memorization in RLHF for code completion
von: Pappu, Aneesh, et al.
Veröffentlicht: (2024)
von: Pappu, Aneesh, et al.
Veröffentlicht: (2024)
Stealing User Prompts from Mixture of Experts
von: Yona, Itay, et al.
Veröffentlicht: (2024)
von: Yona, Itay, et al.
Veröffentlicht: (2024)
Trusted Machine Learning Models Unlock Private Inference for Problems Currently Infeasible with Cryptography
von: Shumailov, Ilia, et al.
Veröffentlicht: (2025)
von: Shumailov, Ilia, et al.
Veröffentlicht: (2025)
NeuraLUT: Hiding Neural Network Density in Boolean Synthesizable Functions
von: Andronic, Marta, et al.
Veröffentlicht: (2024)
von: Andronic, Marta, et al.
Veröffentlicht: (2024)
HASS: Hardware-Aware Sparsity Search for Dataflow DNN Accelerator
von: Yu, Zhewen, et al.
Veröffentlicht: (2024)
von: Yu, Zhewen, et al.
Veröffentlicht: (2024)
When Vision Fails: Text Attacks Against ViT and OCR
von: Boucher, Nicholas, et al.
Veröffentlicht: (2023)
von: Boucher, Nicholas, et al.
Veröffentlicht: (2023)
Beyond Laplace and Gaussian: Exploring the Generalized Gaussian Mechanism for Private Machine Learning
von: Rinberg, Roy, et al.
Veröffentlicht: (2025)
von: Rinberg, Roy, et al.
Veröffentlicht: (2025)
Inexact Unlearning Needs More Careful Evaluations to Avoid a False Sense of Privacy
von: Hayes, Jamie, et al.
Veröffentlicht: (2024)
von: Hayes, Jamie, et al.
Veröffentlicht: (2024)
Large Language Models Can Verbatim Reproduce Long Malicious Sequences
von: Lin, Sharon, et al.
Veröffentlicht: (2025)
von: Lin, Sharon, et al.
Veröffentlicht: (2025)
Gradients Look Alike: Sensitivity is Often Overestimated in DP-SGD
von: Thudi, Anvith, et al.
Veröffentlicht: (2023)
von: Thudi, Anvith, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
LQER: Low-Rank Quantization Error Reconstruction for LLMs
von: Zhang, Cheng, et al.
Veröffentlicht: (2024) -
Hardware and Software Platform Inference
von: Zhang, Cheng, et al.
Veröffentlicht: (2024) -
AMPLE: Event-Driven Accelerator for Mixed-Precision Inference of Graph Neural Networks
von: Gimenes, Pedro, et al.
Veröffentlicht: (2025) -
Quantamination: Dynamic Quantization Leaks Your Data Across the Batch
von: Foerster, Hanna, et al.
Veröffentlicht: (2026) -
QERA: an Analytical Framework for Quantization Error Reconstruction
von: Zhang, Cheng, et al.
Veröffentlicht: (2024)