Panza: Design and Analysis of a Fully-Local Personalized Text Writing Assistant
Fuente:
arXiv
Saved in:
| Main Authors: | Nicolicioiu, Armand, Iofinova, Eugenia, Jovanovic, Andrej, Kurtic, Eldar, Nikdan, Mahdi, Panferov, Andrei, Markov, Ilia, Shavit, Nir, Alistarh, Dan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Position: It's Time to Act on the Risk of Efficient Personalized Text Generation
by: Iofinova, Eugenia, et al.
Published: (2025)
by: Iofinova, Eugenia, et al.
Published: (2025)
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling
by: Dadgarnia, Alireza, et al.
Published: (2026)
by: Dadgarnia, Alireza, et al.
Published: (2026)
Behemoth: Benchmarking Unlearning in LLMs Using Fully Synthetic Data
by: Iofinova, Eugenia, et al.
Published: (2026)
by: Iofinova, Eugenia, et al.
Published: (2026)
Mathador-LM: A Dynamic Benchmark for Mathematical Reasoning on Large Language Models
by: Kurtic, Eldar, et al.
Published: (2024)
by: Kurtic, Eldar, et al.
Published: (2024)
Towards Combinatorial Interpretability of Neural Computation
by: Adler, Micah, et al.
Published: (2025)
by: Adler, Micah, et al.
Published: (2025)
ECO: Quantized Training without Full-Precision Master Weights
by: Nikdan, Mahdi, et al.
Published: (2026)
by: Nikdan, Mahdi, et al.
Published: (2026)
RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust Adaptation
by: Nikdan, Mahdi, et al.
Published: (2024)
by: Nikdan, Mahdi, et al.
Published: (2024)
DarwinLM: Evolutionary Structured Pruning of Large Language Models
by: Tang, Shengkun, et al.
Published: (2025)
by: Tang, Shengkun, et al.
Published: (2025)
Wasserstein Distances, Neuronal Entanglement, and Sparsity
by: Sawmya, Shashata, et al.
Published: (2024)
by: Sawmya, Shashata, et al.
Published: (2024)
Efficient Data Selection at Scale via Influence Distillation
by: Nikdan, Mahdi, et al.
Published: (2025)
by: Nikdan, Mahdi, et al.
Published: (2025)
Statistically-Lossless Quantization of Large Language Models
by: Helcig, Michael, et al.
Published: (2026)
by: Helcig, Michael, et al.
Published: (2026)
QuEST: Stable Training of LLMs with 1-Bit Weights and Activations
by: Panferov, Andrei, et al.
Published: (2025)
by: Panferov, Andrei, et al.
Published: (2025)
EvoPress: Accurate Dynamic Model Compression via Evolutionary Search
by: Sieberling, Oliver, et al.
Published: (2024)
by: Sieberling, Oliver, et al.
Published: (2024)
Extreme Compression of Large Language Models via Additive Quantization
by: Egiazarian, Vage, et al.
Published: (2024)
by: Egiazarian, Vage, et al.
Published: (2024)
Unified Scaling Laws for Compressed Representations
by: Panferov, Andrei, et al.
Published: (2025)
by: Panferov, Andrei, et al.
Published: (2025)
Quartet: Native FP4 Training Can Be Optimal for Large Language Models
by: Castro, Roberto L., et al.
Published: (2025)
by: Castro, Roberto L., et al.
Published: (2025)
"Give Me BF16 or Give Me Death"? Accuracy-Performance Trade-Offs in LLM Quantization
by: Kurtic, Eldar, et al.
Published: (2024)
by: Kurtic, Eldar, et al.
Published: (2024)
SPADE: Sparsity-Guided Debugging for Deep Neural Networks
by: Moakhar, Arshia Soltani, et al.
Published: (2023)
by: Moakhar, Arshia Soltani, et al.
Published: (2023)
Neural Redshift: Random Networks are not Random Functions
by: Teney, Damien, et al.
Published: (2024)
by: Teney, Damien, et al.
Published: (2024)
Apertus LLM Family Expansion via Distillation and Quantization
by: Panferov, Andrei, et al.
Published: (2026)
by: Panferov, Andrei, et al.
Published: (2026)
Quartet II: Accurate LLM Pre-Training in NVFP4 by Improved Unbiased Gradient Estimation
by: Panferov, Andrei, et al.
Published: (2026)
by: Panferov, Andrei, et al.
Published: (2026)
Error Feedback Can Accurately Compress Preconditioners
by: Modoranu, Ionut-Vlad, et al.
Published: (2023)
by: Modoranu, Ionut-Vlad, et al.
Published: (2023)
On the Complexity of Neural Computation in Superposition
by: Adler, Micah, et al.
Published: (2024)
by: Adler, Micah, et al.
Published: (2024)
Leveraging LLMs for Context-Aware Implicit Textual and Multimodal Hate Speech Detection
by: Brook, Joshua Wolfe, et al.
Published: (2025)
by: Brook, Joshua Wolfe, et al.
Published: (2025)
Leveraging Open-Source Large Language Models for Native Language Identification
by: Ng, Yee Man, et al.
Published: (2024)
by: Ng, Yee Man, et al.
Published: (2024)
Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization
by: Egiazarian, Vage, et al.
Published: (2025)
by: Egiazarian, Vage, et al.
Published: (2025)
Enabling High-Sparsity Foundational Llama Models with Efficient Pretraining and Deployment
by: Agarwalla, Abhinav, et al.
Published: (2024)
by: Agarwalla, Abhinav, et al.
Published: (2024)
Expand Neurons, Not Parameters
by: Kong, Linghao, et al.
Published: (2025)
by: Kong, Linghao, et al.
Published: (2025)
The Birth of Knowledge: Emergent Features across Time, Space, and Scale in Large Language Models
by: Sawmya, Shashata, et al.
Published: (2025)
by: Sawmya, Shashata, et al.
Published: (2025)
Robust Novelty Detection through Style-Conscious Feature Ranking
by: Smeu, Stefan, et al.
Published: (2023)
by: Smeu, Stefan, et al.
Published: (2023)
Speculative Decoding Speed-of-Light: Optimal Lower Bounds via Branching Random Walks
by: Pankratov, Sergey, et al.
Published: (2025)
by: Pankratov, Sergey, et al.
Published: (2025)
CAGE: Curvature-Aware Gradient Estimation For Accurate Quantization-Aware Training
by: Tabesh, Soroush, et al.
Published: (2025)
by: Tabesh, Soroush, et al.
Published: (2025)
The Constant in HATE: Analyzing Toxicity in Reddit across Topics and Languages
by: Tufa, Wondimagegnhue Tsegaye, et al.
Published: (2024)
by: Tufa, Wondimagegnhue Tsegaye, et al.
Published: (2024)
Grounding Toxicity in Real-World Events across Languages
by: Tufa, Wondimagegnhue Tsegaye, et al.
Published: (2024)
by: Tufa, Wondimagegnhue Tsegaye, et al.
Published: (2024)
Unknown Script: Impact of Script on Cross-Lingual Transfer
by: Tufa, Wondimagegnhue Tsegaye, et al.
Published: (2024)
by: Tufa, Wondimagegnhue Tsegaye, et al.
Published: (2024)
Cascade Detector Analysis and Application to Biomedical Microscopy
by: Athey, Thomas L., et al.
Published: (2025)
by: Athey, Thomas L., et al.
Published: (2025)
Pearl: Personalizing Large Language Model Writing Assistants with Generation-Calibrated Retrievers
by: Mysore, Sheshera, et al.
Published: (2023)
by: Mysore, Sheshera, et al.
Published: (2023)
LLMQ: Efficient Lower-Precision Pretraining for Consumer GPUs
by: Schultheis, Erik, et al.
Published: (2025)
by: Schultheis, Erik, et al.
Published: (2025)
A Design Space for Intelligent and Interactive Writing Assistants
by: Lee, Mina, et al.
Published: (2024)
by: Lee, Mina, et al.
Published: (2024)
HALO: Hadamard-Assisted Lower-Precision Optimization for LLMs
by: Ashkboos, Saleh, et al.
Published: (2025)
by: Ashkboos, Saleh, et al.
Published: (2025)
Similar Items
-
Position: It's Time to Act on the Risk of Efficient Personalized Text Generation
by: Iofinova, Eugenia, et al.
Published: (2025) -
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling
by: Dadgarnia, Alireza, et al.
Published: (2026) -
Behemoth: Benchmarking Unlearning in LLMs Using Fully Synthetic Data
by: Iofinova, Eugenia, et al.
Published: (2026) -
Mathador-LM: A Dynamic Benchmark for Mathematical Reasoning on Large Language Models
by: Kurtic, Eldar, et al.
Published: (2024) -
Towards Combinatorial Interpretability of Neural Computation
by: Adler, Micah, et al.
Published: (2025)