Subword Embedding from Bytes Gains Privacy without Sacrificing Accuracy and Complexity
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Mengjiao, Xu, Jia |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence
por: Herbster, Niklas, et al.
Publicado: (2026)
por: Herbster, Niklas, et al.
Publicado: (2026)
Disrupting Model Merging: A Parameter-Level Defense Without Sacrificing Accuracy
por: Junhao, Wei, et al.
Publicado: (2025)
por: Junhao, Wei, et al.
Publicado: (2025)
Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge
por: Batsuren, Khuyagbaatar, et al.
Publicado: (2024)
por: Batsuren, Khuyagbaatar, et al.
Publicado: (2024)
LGSE: Lexically Grounded Subword Embedding Initialization for Low-Resource Language Adaptation
por: Teklehaymanot, Hailay, et al.
Publicado: (2026)
por: Teklehaymanot, Hailay, et al.
Publicado: (2026)
A Middle Path for On-Premises LLM Deployment: Preserving Privacy Without Sacrificing Model Confidentiality
por: Huang, Hanbo, et al.
Publicado: (2024)
por: Huang, Hanbo, et al.
Publicado: (2024)
Positive-Sum Fairness: Leveraging Demographic Attributes to Achieve Fair AI Outcomes Without Sacrificing Group Gains
por: Belhadj, Samia, et al.
Publicado: (2024)
por: Belhadj, Samia, et al.
Publicado: (2024)
T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings
por: Deiseroth, Björn, et al.
Publicado: (2024)
por: Deiseroth, Björn, et al.
Publicado: (2024)
Token Alignment via Character Matching for Subword Completion
por: Athiwaratkun, Ben, et al.
Publicado: (2024)
por: Athiwaratkun, Ben, et al.
Publicado: (2024)
Privacy and Accuracy Implications of Model Complexity and Integration in Heterogeneous Federated Learning
por: Németh, Gergely Dániel, et al.
Publicado: (2023)
por: Németh, Gergely Dániel, et al.
Publicado: (2023)
Understanding Subword Compositionality of Large Language Models
por: Peng, Qiwei, et al.
Publicado: (2025)
por: Peng, Qiwei, et al.
Publicado: (2025)
PixelBytes: Catching Unified Embedding for Multimodal Generation
por: Furfaro, Fabien
Publicado: (2024)
por: Furfaro, Fabien
Publicado: (2024)
Subwords as Skills: Tokenization for Sparse-Reward Reinforcement Learning
por: Yunis, David, et al.
Publicado: (2023)
por: Yunis, David, et al.
Publicado: (2023)
Team Ryu's Submission to SIGMORPHON 2024 Shared Task on Subword Tokenization
por: Li, Zilong
Publicado: (2024)
por: Li, Zilong
Publicado: (2024)
ECCO: Can We Improve Model-Generated Code Efficiency Without Sacrificing Functional Correctness?
por: Waghjale, Siddhant, et al.
Publicado: (2024)
por: Waghjale, Siddhant, et al.
Publicado: (2024)
Temperature and Persona Shape LLM Agent Consensus With Minimal Accuracy Gains in Qualitative Coding
por: Borchers, Conrad, et al.
Publicado: (2025)
por: Borchers, Conrad, et al.
Publicado: (2025)
Optimal Turkish Subword Strategies at Scale: Systematic Evaluation of Data, Vocabulary, Morphology Interplay
por: Altinok, Duygu
Publicado: (2026)
por: Altinok, Duygu
Publicado: (2026)
Deciphering the Interplay between Attack and Protection Complexity in Privacy-Preserving Federated Learning
por: Zhang, Xiaojin, et al.
Publicado: (2025)
por: Zhang, Xiaojin, et al.
Publicado: (2025)
Symbolic Guardrails for Domain-Specific Agents: Stronger Safety and Security Guarantees Without Sacrificing Utility
por: Hong, Yining, et al.
Publicado: (2026)
por: Hong, Yining, et al.
Publicado: (2026)
Language Complexity and Speech Recognition Accuracy: Orthographic Complexity Hurts, Phonological Complexity Doesn't
por: Taguchi, Chihiro, et al.
Publicado: (2024)
por: Taguchi, Chihiro, et al.
Publicado: (2024)
Pruning Foundation Models for High Accuracy without Retraining
por: Zhao, Pu, et al.
Publicado: (2024)
por: Zhao, Pu, et al.
Publicado: (2024)
MambaNetBurst: Direct Byte-level Network Traffic Classification without Tokenization or Pretraining
por: Kulatilleke, Gayan K., et al.
Publicado: (2026)
por: Kulatilleke, Gayan K., et al.
Publicado: (2026)
IGOT: Information Gain Optimized Tokenizer on Domain Adaptive Pretraining
por: Feng, Dawei, et al.
Publicado: (2024)
por: Feng, Dawei, et al.
Publicado: (2024)
ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model Development
por: Wan, Borui, et al.
Publicado: (2024)
por: Wan, Borui, et al.
Publicado: (2024)
Understanding Privacy Risks of Embeddings Induced by Large Language Models
por: Zhu, Zhihao, et al.
Publicado: (2024)
por: Zhu, Zhihao, et al.
Publicado: (2024)
Concept-Aware Privacy Mechanisms for Defending Embedding Inversion Attacks
por: Tsai, Yu-Che, et al.
Publicado: (2026)
por: Tsai, Yu-Che, et al.
Publicado: (2026)
TokDrift: When LLM Speaks in Subwords but Code Speaks in Grammar
por: Li, Yinxi, et al.
Publicado: (2025)
por: Li, Yinxi, et al.
Publicado: (2025)
Decoding in Geometry: Alleviating Embedding-Space Crowding for Complex Reasoning
por: Yang, Yixin, et al.
Publicado: (2026)
por: Yang, Yixin, et al.
Publicado: (2026)
Learning to Beat ByteRL: Exploitability of Collectible Card Game Agents
por: Haluska, Radovan, et al.
Publicado: (2024)
por: Haluska, Radovan, et al.
Publicado: (2024)
Balancing Fairness, Privacy, and Accuracy: A Multitask Adversarial Framework for Centralized Data-Driven Systems
por: Ekanayake, Imesh, et al.
Publicado: (2026)
por: Ekanayake, Imesh, et al.
Publicado: (2026)
Fast Byte Latent Transformer
por: Kallini, Julie, et al.
Publicado: (2026)
por: Kallini, Julie, et al.
Publicado: (2026)
Quantifying and Defending against Privacy Threats on Federated Knowledge Graph Embedding
por: Hu, Yuke, et al.
Publicado: (2023)
por: Hu, Yuke, et al.
Publicado: (2023)
Multi-resolution Rescored ByteTrack for Video Object Detection on Ultra-low-power Embedded Systems
por: Bompani, Luca, et al.
Publicado: (2024)
por: Bompani, Luca, et al.
Publicado: (2024)
Accuracy-Privacy Trade-off in the Mitigation of Membership Inference Attack in Federated Learning
por: Ahamed, Sayyed Farid, et al.
Publicado: (2024)
por: Ahamed, Sayyed Farid, et al.
Publicado: (2024)
Explainable Knowledge Tracing via Probabilistic Embeddings and Pattern-based Reasoning
por: Wu, Siyu, et al.
Publicado: (2026)
por: Wu, Siyu, et al.
Publicado: (2026)
Complex Ontology Matching with Large Language Model Embeddings
por: Sousa, Guilherme, et al.
Publicado: (2025)
por: Sousa, Guilherme, et al.
Publicado: (2025)
From Bytes to Ideas: Language Modeling with Autoregressive U-Nets
por: Videau, Mathurin, et al.
Publicado: (2025)
por: Videau, Mathurin, et al.
Publicado: (2025)
Brains vs. Bytes: Evaluating LLM Proficiency in Olympiad Mathematics
por: Mahdavi, Hamed, et al.
Publicado: (2025)
por: Mahdavi, Hamed, et al.
Publicado: (2025)
Protein Structure Tokenization via Geometric Byte Pair Encoding
por: Sun, Michael, et al.
Publicado: (2025)
por: Sun, Michael, et al.
Publicado: (2025)
Privacy Artifact ConnecTor (PACT): Embedding Enterprise Artifacts for Compliance AI Agents
por: Fang, Chenhao, et al.
Publicado: (2025)
por: Fang, Chenhao, et al.
Publicado: (2025)
GainRAG: Preference Alignment in Retrieval-Augmented Generation through Gain Signal Synthesis
por: Jiang, Yi, et al.
Publicado: (2025)
por: Jiang, Yi, et al.
Publicado: (2025)
Ejemplares similares
-
Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence
por: Herbster, Niklas, et al.
Publicado: (2026) -
Disrupting Model Merging: A Parameter-Level Defense Without Sacrificing Accuracy
por: Junhao, Wei, et al.
Publicado: (2025) -
Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge
por: Batsuren, Khuyagbaatar, et al.
Publicado: (2024) -
LGSE: Lexically Grounded Subword Embedding Initialization for Low-Resource Language Adaptation
por: Teklehaymanot, Hailay, et al.
Publicado: (2026) -
A Middle Path for On-Premises LLM Deployment: Preserving Privacy Without Sacrificing Model Confidentiality
por: Huang, Hanbo, et al.
Publicado: (2024)