The Surprising Universality of LLM Outputs: A Real-Time Verification Primitive
Fuente:
arXiv
Saved in:
| Main Authors: | Bogdan, Alex, de Valois-Franklin, Adrian |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Universal and Context-Independent Triggers for Precise Control of LLM Outputs
by: Liang, Jiashuo, et al.
Published: (2024)
by: Liang, Jiashuo, et al.
Published: (2024)
Localizing Malicious Outputs from CodeLLM
by: Borana, Mayukh, et al.
Published: (2025)
by: Borana, Mayukh, et al.
Published: (2025)
LexiMark: Robust Watermarking via Lexical Substitutions to Enhance Membership Verification of an LLM's Textual Training Data
by: German, Eyal, et al.
Published: (2025)
by: German, Eyal, et al.
Published: (2025)
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments
by: Zhang, Chiyu, et al.
Published: (2026)
by: Zhang, Chiyu, et al.
Published: (2026)
RealVul: Can We Detect Vulnerabilities in Web Applications with LLM?
by: Cao, Di, et al.
Published: (2024)
by: Cao, Di, et al.
Published: (2024)
Time Will Tell: Timing Side Channels via Output Token Count in Large Language Models
by: Zhang, Tianchen, et al.
Published: (2024)
by: Zhang, Tianchen, et al.
Published: (2024)
FakeZero: Real-Time, Privacy-Preserving Misinformation Detection for Facebook and X
by: Essahli, Soufiane, et al.
Published: (2025)
by: Essahli, Soufiane, et al.
Published: (2025)
Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs
by: Xing, Wenpeng, et al.
Published: (2025)
by: Xing, Wenpeng, et al.
Published: (2025)
One Word at a Time: Incremental Completion Decomposition Breaks LLM Safety
by: Arif, Samee, et al.
Published: (2026)
by: Arif, Samee, et al.
Published: (2026)
Adaptive Pre-training Data Detection for Large Language Models via Surprising Tokens
by: Zhang, Anqi, et al.
Published: (2024)
by: Zhang, Anqi, et al.
Published: (2024)
LRCTI: A Large Language Model-Based Framework for Multi-Step Evidence Retrieval and Reasoning in Cyber Threat Intelligence Credibility Verification
by: Tang, Fengxiao, et al.
Published: (2025)
by: Tang, Fengxiao, et al.
Published: (2025)
TimeMark: A Trustworthy Time Watermarking Framework for Exact Generation-Time Recovery from AIGC
by: Che, Shangkun, et al.
Published: (2026)
by: Che, Shangkun, et al.
Published: (2026)
LLM Reinforcement in Context
by: Rivasseau, Thomas
Published: (2025)
by: Rivasseau, Thomas
Published: (2025)
Watermarking LLM Agent Trajectories
by: Meng, Wenlong, et al.
Published: (2026)
by: Meng, Wenlong, et al.
Published: (2026)
BadActs: A Universal Backdoor Defense in the Activation Space
by: Yi, Biao, et al.
Published: (2024)
by: Yi, Biao, et al.
Published: (2024)
Talk is (Not) Cheap: A Taxonomy and Benchmark Coverage Audit for LLM Attacks
by: Iyer, Karthik Raghu, et al.
Published: (2026)
by: Iyer, Karthik Raghu, et al.
Published: (2026)
Proactive defense against LLM Jailbreak
by: Zhao, Weiliang, et al.
Published: (2025)
by: Zhao, Weiliang, et al.
Published: (2025)
FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks
by: Chen, Bocheng, et al.
Published: (2024)
by: Chen, Bocheng, et al.
Published: (2024)
TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts
by: Chu, Hua-Rong, et al.
Published: (2026)
by: Chu, Hua-Rong, et al.
Published: (2026)
Universal Zero-shot Embedding Inversion
by: Zhang, Collin, et al.
Published: (2025)
by: Zhang, Collin, et al.
Published: (2025)
LLM Anonymization Against Agentic Re-Identification
by: Li, Ziwen, et al.
Published: (2026)
by: Li, Ziwen, et al.
Published: (2026)
DAVE: A Policy-Enforcing LLM Spokesperson for Secure Multi-Document Data Sharing
by: Brinkhege, René, et al.
Published: (2026)
by: Brinkhege, René, et al.
Published: (2026)
REMARK-LLM: A Robust and Efficient Watermarking Framework for Generative Large Language Models
by: Zhang, Ruisi, et al.
Published: (2023)
by: Zhang, Ruisi, et al.
Published: (2023)
Chain-of-Lure: A Universal Jailbreak Attack Framework using Unconstrained Synthetic Narratives
by: Chang, Wenhan, et al.
Published: (2025)
by: Chang, Wenhan, et al.
Published: (2025)
PRSA: Prompt Stealing Attacks against Real-World Prompt Services
by: Yang, Yong, et al.
Published: (2024)
by: Yang, Yong, et al.
Published: (2024)
FunFuzz: An LLM-Powered Evolutionary Fuzzing Framework
by: Béjar, Mario Rodríguez, et al.
Published: (2026)
by: Béjar, Mario Rodríguez, et al.
Published: (2026)
On the Hidden Costs of Counterfactual Knowledge Training in LLM Unlearning
by: Ye, Xiaotian, et al.
Published: (2026)
by: Ye, Xiaotian, et al.
Published: (2026)
GLiGuard: Schema-Conditioned Classification for LLM Safeguard
by: Zaratiana, Urchade, et al.
Published: (2026)
by: Zaratiana, Urchade, et al.
Published: (2026)
WorldCup Sampling for Multi-bit LLM Watermarking
by: Wang, Yidan, et al.
Published: (2026)
by: Wang, Yidan, et al.
Published: (2026)
Machine Psychometrics: A Mathematical Psychology of Artificial Intelligence
by: Bogdan, Alex, et al.
Published: (2026)
by: Bogdan, Alex, et al.
Published: (2026)
Security Attacks on LLM-based Code Completion Tools
by: Cheng, Wen, et al.
Published: (2024)
by: Cheng, Wen, et al.
Published: (2024)
Prompt Optimization and Evaluation for LLM Automated Red Teaming
by: Freenor, Michael, et al.
Published: (2025)
by: Freenor, Michael, et al.
Published: (2025)
Confidential Prompting: Privacy-preserving LLM Inference on Cloud
by: Li, Caihua, et al.
Published: (2024)
by: Li, Caihua, et al.
Published: (2024)
Multi-use LLM Watermarking and the False Detection Problem
by: Fu, Zihao, et al.
Published: (2025)
by: Fu, Zihao, et al.
Published: (2025)
Interpretable LLM Guardrails via Sparse Representation Steering
by: He, Zeqing, et al.
Published: (2025)
by: He, Zeqing, et al.
Published: (2025)
Raccoon: Prompt Extraction Benchmark of LLM-Integrated Applications
by: Wang, Junlin, et al.
Published: (2024)
by: Wang, Junlin, et al.
Published: (2024)
How Real is Your Jailbreak? Fine-grained Jailbreak Evaluation with Anchored Reference
by: Liu, Songyang, et al.
Published: (2026)
by: Liu, Songyang, et al.
Published: (2026)
Rethinking LLM Watermark Detection in Black-Box Settings: A Non-Intrusive Third-Party Framework
by: Wang, Zhuoshang, et al.
Published: (2026)
by: Wang, Zhuoshang, et al.
Published: (2026)
CI-Work: Benchmarking Contextual Integrity in Enterprise LLM Agents
by: Fu, Wenjie, et al.
Published: (2026)
by: Fu, Wenjie, et al.
Published: (2026)
Robust LLM Watermarking with Minimal Semantic Distortion for IP Protection
by: Dang, Kieu, et al.
Published: (2026)
by: Dang, Kieu, et al.
Published: (2026)
Similar Items
-
Universal and Context-Independent Triggers for Precise Control of LLM Outputs
by: Liang, Jiashuo, et al.
Published: (2024) -
Localizing Malicious Outputs from CodeLLM
by: Borana, Mayukh, et al.
Published: (2025) -
LexiMark: Robust Watermarking via Lexical Substitutions to Enhance Membership Verification of an LLM's Textual Training Data
by: German, Eyal, et al.
Published: (2025) -
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments
by: Zhang, Chiyu, et al.
Published: (2026) -
RealVul: Can We Detect Vulnerabilities in Web Applications with LLM?
by: Cao, Di, et al.
Published: (2024)