Gespeichert in:
| Hauptverfasser: | Nuriyev, Amir, Kulp, Gabriel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2602.04105 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MetaMoE: Diversity-Aware Proxy Selection for Privacy-Preserving Mixture-of-Experts Unification
von: Jiang, Weisen, et al.
Veröffentlicht: (2026)
von: Jiang, Weisen, et al.
Veröffentlicht: (2026)
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework
von: Liang, Zi, et al.
Veröffentlicht: (2025)
von: Liang, Zi, et al.
Veröffentlicht: (2025)
Text2VLM: Adapting Text-Only Datasets to Evaluate Alignment Training in Visual Language Models
von: Downer, Gabriel, et al.
Veröffentlicht: (2025)
von: Downer, Gabriel, et al.
Veröffentlicht: (2025)
LLM Jailbreak Detection for (Almost) Free!
von: Chen, Guorui, et al.
Veröffentlicht: (2025)
von: Chen, Guorui, et al.
Veröffentlicht: (2025)
RouteScan: A Non-Intrusive Approach to Auditing MoE LLMs Safety via Expert Routing Telemetry
von: Lv, Bo, et al.
Veröffentlicht: (2026)
von: Lv, Bo, et al.
Veröffentlicht: (2026)
SecMoE: Communication-Efficient Secure MoE Inference via Select-Then-Compute
von: Shen, Bowen, et al.
Veröffentlicht: (2026)
von: Shen, Bowen, et al.
Veröffentlicht: (2026)
Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics
von: Chrabąszcz, Maciej, et al.
Veröffentlicht: (2026)
von: Chrabąszcz, Maciej, et al.
Veröffentlicht: (2026)
Toward Cybersecurity-Expert Small Language Models
von: Levi, Matan, et al.
Veröffentlicht: (2025)
von: Levi, Matan, et al.
Veröffentlicht: (2025)
GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models
von: Wang, Zilong, et al.
Veröffentlicht: (2025)
von: Wang, Zilong, et al.
Veröffentlicht: (2025)
Attacks against Abstractive Text Summarization Models through Lead Bias and Influence Functions
von: Thota, Poojitha, et al.
Veröffentlicht: (2024)
von: Thota, Poojitha, et al.
Veröffentlicht: (2024)
AEIOU: A Unified Defense Framework against NSFW Prompts in Text-to-Image Models
von: Wang, Yiming, et al.
Veröffentlicht: (2024)
von: Wang, Yiming, et al.
Veröffentlicht: (2024)
EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models
von: Mirbagheri, Mohammad Reza, et al.
Veröffentlicht: (2025)
von: Mirbagheri, Mohammad Reza, et al.
Veröffentlicht: (2025)
Bileve: Securing Text Provenance in Large Language Models Against Spoofing with Bi-level Signature
von: Zhou, Tong, et al.
Veröffentlicht: (2024)
von: Zhou, Tong, et al.
Veröffentlicht: (2024)
StructuralSleight: Automated Jailbreak Attacks on Large Language Models Utilizing Uncommon Text-Organization Structures
von: Li, Bangxin, et al.
Veröffentlicht: (2024)
von: Li, Bangxin, et al.
Veröffentlicht: (2024)
Data Extraction Attacks in Retrieval-Augmented Generation via Backdoors
von: Peng, Yuefeng, et al.
Veröffentlicht: (2024)
von: Peng, Yuefeng, et al.
Veröffentlicht: (2024)
Real-time and Zero-footprint Bag of Synthetic Syllables Algorithm for E-mail Spam Detection Using Subject Line and Short Text Fields
von: Selitskiy, Stanislav
Veröffentlicht: (2025)
von: Selitskiy, Stanislav
Veröffentlicht: (2025)
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures
von: Zhou, Yukai, et al.
Veröffentlicht: (2025)
von: Zhou, Yukai, et al.
Veröffentlicht: (2025)
Adversarial Text Generation with Dynamic Contextual Perturbation
von: Waghela, Hetvi, et al.
Veröffentlicht: (2025)
von: Waghela, Hetvi, et al.
Veröffentlicht: (2025)
SecureGate: Learning When to Reveal PII Safely via Token-Gated Dual-Adapters for Federated LLMs
von: Shaaban, Mohamed, et al.
Veröffentlicht: (2026)
von: Shaaban, Mohamed, et al.
Veröffentlicht: (2026)
Revealing Weaknesses in Text Watermarking Through Self-Information Rewrite Attacks
von: Cheng, Yixin, et al.
Veröffentlicht: (2025)
von: Cheng, Yixin, et al.
Veröffentlicht: (2025)
Robustness Assessment and Enhancement of Text Watermarking for Google's SynthID
von: Han, Xia, et al.
Veröffentlicht: (2025)
von: Han, Xia, et al.
Veröffentlicht: (2025)
Traffic-MoE: A Sparse Foundation Model for Network Traffic Analysis
von: Zhou, Jiajun, et al.
Veröffentlicht: (2026)
von: Zhou, Jiajun, et al.
Veröffentlicht: (2026)
MGTEVAL: An Interactive Platform for Systemtic Evaluation of Machine-Generated Text Detectors
von: Li, Yuanfan, et al.
Veröffentlicht: (2026)
von: Li, Yuanfan, et al.
Veröffentlicht: (2026)
Block-wise Codeword Embedding for Reliable Multi-bit Text Watermarking
von: Kim, Joeun, et al.
Veröffentlicht: (2026)
von: Kim, Joeun, et al.
Veröffentlicht: (2026)
DP-BART for Privatized Text Rewriting under Local Differential Privacy
von: Igamberdiev, Timour, et al.
Veröffentlicht: (2023)
von: Igamberdiev, Timour, et al.
Veröffentlicht: (2023)
Factuality Beyond Coherence: Evaluating LLM Watermarking Methods for Medical Texts
von: Hastuti, Rochana Prih, et al.
Veröffentlicht: (2025)
von: Hastuti, Rochana Prih, et al.
Veröffentlicht: (2025)
HLPD: Aligning LLMs to Human Language Preference for Machine-Revised Text Detection
von: Dai, Fangqi, et al.
Veröffentlicht: (2025)
von: Dai, Fangqi, et al.
Veröffentlicht: (2025)
Beyond Text: Unveiling Privacy Vulnerabilities in Multi-modal Retrieval-Augmented Generation
von: Zhang, Jiankun, et al.
Veröffentlicht: (2025)
von: Zhang, Jiankun, et al.
Veröffentlicht: (2025)
MaskSQL: Safeguarding Privacy for LLM-Based Text-to-SQL via Abstraction
von: Abedini, Sepideh, et al.
Veröffentlicht: (2025)
von: Abedini, Sepideh, et al.
Veröffentlicht: (2025)
TSCheater: Generating High-Quality Tibetan Adversarial Texts via Visual Similarity
von: Cao, Xi, et al.
Veröffentlicht: (2024)
von: Cao, Xi, et al.
Veröffentlicht: (2024)
SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
von: Li, Yucheng, et al.
Veröffentlicht: (2025)
von: Li, Yucheng, et al.
Veröffentlicht: (2025)
BinarySelect to Improve Accessibility of Black-Box Attack Research
von: Ghosh, Shatarupa, et al.
Veröffentlicht: (2024)
von: Ghosh, Shatarupa, et al.
Veröffentlicht: (2024)
GradEscape: A Gradient-Based Evader Against AI-Generated Text Detectors
von: Meng, Wenlong, et al.
Veröffentlicht: (2025)
von: Meng, Wenlong, et al.
Veröffentlicht: (2025)
Is the Digital Forensics and Incident Response Pipeline Ready for Text-Based Threats in LLM Era?
von: Bhandarkar, Avanti, et al.
Veröffentlicht: (2024)
von: Bhandarkar, Avanti, et al.
Veröffentlicht: (2024)
Efficiently and Effectively: A Two-stage Approach to Balance Plaintext and Encrypted Text for Traffic Classification
von: Peng, Wei, et al.
Veröffentlicht: (2024)
von: Peng, Wei, et al.
Veröffentlicht: (2024)
Iron Sharpens Iron: Defending Against Attacks in Machine-Generated Text Detection with Adversarial Training
von: Li, Yuanfan, et al.
Veröffentlicht: (2025)
von: Li, Yuanfan, et al.
Veröffentlicht: (2025)
Text Embedding Inversion Security for Multilingual Language Models
von: Chen, Yiyi, et al.
Veröffentlicht: (2024)
von: Chen, Yiyi, et al.
Veröffentlicht: (2024)
Fight Poison with Poison: Enhancing Robustness in Few-shot Machine-Generated Text Detection with Adversarial Training
von: Duan, Wenjing, et al.
Veröffentlicht: (2026)
von: Duan, Wenjing, et al.
Veröffentlicht: (2026)
Beyond Theoretical Bounds: Empirical Privacy Loss Calibration for Text Rewriting Under Local Differential Privacy
von: Li, Weijun, et al.
Veröffentlicht: (2026)
von: Li, Weijun, et al.
Veröffentlicht: (2026)
A Character-based Diffusion Embedding Algorithm for Enhancing the Generation Quality of Generative Linguistic Steganographic Texts
von: Chen, Yingquan, et al.
Veröffentlicht: (2025)
von: Chen, Yingquan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MetaMoE: Diversity-Aware Proxy Selection for Privacy-Preserving Mixture-of-Experts Unification
von: Jiang, Weisen, et al.
Veröffentlicht: (2026) -
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework
von: Liang, Zi, et al.
Veröffentlicht: (2025) -
Text2VLM: Adapting Text-Only Datasets to Evaluate Alignment Training in Visual Language Models
von: Downer, Gabriel, et al.
Veröffentlicht: (2025) -
LLM Jailbreak Detection for (Almost) Free!
von: Chen, Guorui, et al.
Veröffentlicht: (2025) -
RouteScan: A Non-Intrusive Approach to Auditing MoE LLMs Safety via Expert Routing Telemetry
von: Lv, Bo, et al.
Veröffentlicht: (2026)