How Different Tokenization Algorithms Impact LLMs and Transformer Models for Binary Code Analysis
Fuente:
arXiv
Guardado en:
| Autores principales: | Mostafa, Ahmed, Nahid, Raisul Arefin, Mulder, Samuel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LLMs Have Rhythm: Fingerprinting Large Language Models Using Inter-Token Times and Network Traffic Analysis
por: Alhazbi, Saeif, et al.
Publicado: (2025)
por: Alhazbi, Saeif, et al.
Publicado: (2025)
Automated Software Vulnerability Static Code Analysis Using Generative Pre-Trained Transformer Models
por: Pelofske, Elijah, et al.
Publicado: (2024)
por: Pelofske, Elijah, et al.
Publicado: (2024)
Less Data, More Security: Advancing Cybersecurity LLMs Specialization via Resource-Efficient Domain-Adaptive Continuous Pre-training with Minimal Tokens
por: Salahuddin, Salahuddin, et al.
Publicado: (2025)
por: Salahuddin, Salahuddin, et al.
Publicado: (2025)
Interpreting the Repeated Token Phenomenon in Large Language Models
por: Yona, Itay, et al.
Publicado: (2025)
por: Yona, Itay, et al.
Publicado: (2025)
MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts?
por: Wahed, Muntasir, et al.
Publicado: (2025)
por: Wahed, Muntasir, et al.
Publicado: (2025)
Gandalf the Red: Adaptive Security for LLMs
por: Pfister, Niklas, et al.
Publicado: (2025)
por: Pfister, Niklas, et al.
Publicado: (2025)
ObfuscaTune: Obfuscated Offsite Fine-tuning and Inference of Proprietary LLMs on Private Datasets
por: Frikha, Ahmed, et al.
Publicado: (2024)
por: Frikha, Ahmed, et al.
Publicado: (2024)
Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization
por: Fang, Zheng, et al.
Publicado: (2026)
por: Fang, Zheng, et al.
Publicado: (2026)
SecureCode: A Production-Grade Multi-Turn Dataset for Training Security-Aware Code Generation Models
por: Thornton, Scott
Publicado: (2025)
por: Thornton, Scott
Publicado: (2025)
Fingerprinting Deep Learning Models via Network Traffic Patterns in Federated Learning
por: Shuvo, Md Nahid Hasan, et al.
Publicado: (2025)
por: Shuvo, Md Nahid Hasan, et al.
Publicado: (2025)
Rethinking How to Evaluate Language Model Jailbreak
por: Cai, Hongyu, et al.
Publicado: (2024)
por: Cai, Hongyu, et al.
Publicado: (2024)
Obliviate: Efficient Unmemorization for Protecting Intellectual Property in Large Language Models
por: Russinovich, Mark, et al.
Publicado: (2025)
por: Russinovich, Mark, et al.
Publicado: (2025)
Time Travel in LLMs: Tracing Data Contamination in Large Language Models
por: Golchin, Shahriar, et al.
Publicado: (2023)
por: Golchin, Shahriar, et al.
Publicado: (2023)
Teach LLMs to Phish: Stealing Private Information from Language Models
por: Panda, Ashwinee, et al.
Publicado: (2024)
por: Panda, Ashwinee, et al.
Publicado: (2024)
A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens
por: Dobre, David, et al.
Publicado: (2025)
por: Dobre, David, et al.
Publicado: (2025)
When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models
por: Wang, Kai, et al.
Publicado: (2025)
por: Wang, Kai, et al.
Publicado: (2025)
LLMs can be Dangerous Reasoners: Analyzing-based Jailbreak Attack on Large Language Models
por: Lin, Shi, et al.
Publicado: (2024)
por: Lin, Shi, et al.
Publicado: (2024)
Can Neural Decompilation Assist Vulnerability Prediction on Binary Code?
por: Cotroneo, D., et al.
Publicado: (2024)
por: Cotroneo, D., et al.
Publicado: (2024)
The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy Risks
por: Chen, Xiaoyi, et al.
Publicado: (2023)
por: Chen, Xiaoyi, et al.
Publicado: (2023)
Jailbreaking LLMs via Calibration
por: Lu, Yuxuan, et al.
Publicado: (2026)
por: Lu, Yuxuan, et al.
Publicado: (2026)
BinaryShield: Cross-Service Threat Intelligence in LLM Services using Privacy-Preserving Fingerprints
por: Gill, Waris, et al.
Publicado: (2025)
por: Gill, Waris, et al.
Publicado: (2025)
Tool Preferences in Agentic LLMs are Unreliable
por: Faghih, Kazem, et al.
Publicado: (2025)
por: Faghih, Kazem, et al.
Publicado: (2025)
Not All Tokens Are Created Equal: Query-Efficient Jailbreak Fuzzing for LLMs
por: Chen, Wenyu, et al.
Publicado: (2026)
por: Chen, Wenyu, et al.
Publicado: (2026)
Shh, don't say that! Domain Certification in LLMs
por: Emde, Cornelius, et al.
Publicado: (2025)
por: Emde, Cornelius, et al.
Publicado: (2025)
Early Signs of Steganographic Capabilities in Frontier LLMs
por: Zolkowski, Artur, et al.
Publicado: (2025)
por: Zolkowski, Artur, et al.
Publicado: (2025)
How Does a Deep Learning Model Architecture Impact Its Privacy? A Comprehensive Study of Privacy Attacks on CNNs and Transformers
por: Zhang, Guangsheng, et al.
Publicado: (2022)
por: Zhang, Guangsheng, et al.
Publicado: (2022)
AutoBaxBuilder: Bootstrapping Code Security Benchmarking
por: von Arx, Tobias, et al.
Publicado: (2025)
por: von Arx, Tobias, et al.
Publicado: (2025)
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
por: Mehrotra, Anay, et al.
Publicado: (2023)
por: Mehrotra, Anay, et al.
Publicado: (2023)
AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
por: Paulus, Anselm, et al.
Publicado: (2024)
por: Paulus, Anselm, et al.
Publicado: (2024)
Bypassing the Safety Training of Open-Source LLMs with Priming Attacks
por: Vega, Jason, et al.
Publicado: (2023)
por: Vega, Jason, et al.
Publicado: (2023)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
por: Chu, Junjie, et al.
Publicado: (2024)
por: Chu, Junjie, et al.
Publicado: (2024)
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
por: Shao, Zedian, et al.
Publicado: (2024)
por: Shao, Zedian, et al.
Publicado: (2024)
HARMONIC: Harnessing LLMs for Tabular Data Synthesis and Privacy Protection
por: Wang, Yuxin, et al.
Publicado: (2024)
por: Wang, Yuxin, et al.
Publicado: (2024)
Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
por: Betley, Jan, et al.
Publicado: (2025)
por: Betley, Jan, et al.
Publicado: (2025)
BaxBench: Can LLMs Generate Correct and Secure Backends?
por: Vero, Mark, et al.
Publicado: (2025)
por: Vero, Mark, et al.
Publicado: (2025)
LLMs can hide text in other text of the same length
por: Norelli, Antonio, et al.
Publicado: (2025)
por: Norelli, Antonio, et al.
Publicado: (2025)
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
por: Rando, Javier, et al.
Publicado: (2024)
por: Rando, Javier, et al.
Publicado: (2024)
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
por: Xu, Xiaoyu, et al.
Publicado: (2025)
por: Xu, Xiaoyu, et al.
Publicado: (2025)
Tell me about yourself: LLMs are aware of their learned behaviors
por: Betley, Jan, et al.
Publicado: (2025)
por: Betley, Jan, et al.
Publicado: (2025)
SequentialBreak: Large Language Models Can be Fooled by Embedding Jailbreak Prompts into Sequential Prompt Chains
por: Saiem, Bijoy Ahmed, et al.
Publicado: (2024)
por: Saiem, Bijoy Ahmed, et al.
Publicado: (2024)
Ejemplares similares
-
LLMs Have Rhythm: Fingerprinting Large Language Models Using Inter-Token Times and Network Traffic Analysis
por: Alhazbi, Saeif, et al.
Publicado: (2025) -
Automated Software Vulnerability Static Code Analysis Using Generative Pre-Trained Transformer Models
por: Pelofske, Elijah, et al.
Publicado: (2024) -
Less Data, More Security: Advancing Cybersecurity LLMs Specialization via Resource-Efficient Domain-Adaptive Continuous Pre-training with Minimal Tokens
por: Salahuddin, Salahuddin, et al.
Publicado: (2025) -
Interpreting the Repeated Token Phenomenon in Large Language Models
por: Yona, Itay, et al.
Publicado: (2025) -
MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts?
por: Wahed, Muntasir, et al.
Publicado: (2025)