Logits of API-Protected LLMs Leak Proprietary Information
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Finlayson, Matthew, Ren, Xiang, Swayamdipta, Swabha |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Shortcuts Arising from Contrast: Effective and Covert Clean-Label Attacks in Prompt-Based Learning
von: Xie, Xiaopeng, et al.
Veröffentlicht: (2024)
von: Xie, Xiaopeng, et al.
Veröffentlicht: (2024)
Can Watermarked LLMs be Identified by Users via Crafted Prompts?
von: Liu, Aiwei, et al.
Veröffentlicht: (2024)
von: Liu, Aiwei, et al.
Veröffentlicht: (2024)
On Adversarial Examples for Text Classification by Perturbing Latent Representations
von: Sooksatra, Korn, et al.
Veröffentlicht: (2024)
von: Sooksatra, Korn, et al.
Veröffentlicht: (2024)
Dark LLMs: The Growing Threat of Unaligned AI Models
von: Fire, Michael, et al.
Veröffentlicht: (2025)
von: Fire, Michael, et al.
Veröffentlicht: (2025)
MarkDiffusion: An Open-Source Toolkit for Generative Watermarking of Latent Diffusion Models
von: Pan, Leyi, et al.
Veröffentlicht: (2025)
von: Pan, Leyi, et al.
Veröffentlicht: (2025)
Sparse Logit Sampling: Accelerating Knowledge Distillation in LLMs
von: Anshumann, et al.
Veröffentlicht: (2025)
von: Anshumann, et al.
Veröffentlicht: (2025)
Monotonicity as an Architectural Bias for Robust Language Models
von: Cooper, Patrick, et al.
Veröffentlicht: (2026)
von: Cooper, Patrick, et al.
Veröffentlicht: (2026)
MarkLLM: An Open-Source Toolkit for LLM Watermarking
von: Pan, Leyi, et al.
Veröffentlicht: (2024)
von: Pan, Leyi, et al.
Veröffentlicht: (2024)
A Semantic Invariant Robust Watermark for Large Language Models
von: Liu, Aiwei, et al.
Veröffentlicht: (2023)
von: Liu, Aiwei, et al.
Veröffentlicht: (2023)
Scaling Patterns in Adversarial Alignment: Evidence from Multi-LLM Jailbreak Experiments
von: Nathanson, Samuel, et al.
Veröffentlicht: (2025)
von: Nathanson, Samuel, et al.
Veröffentlicht: (2025)
How Few-shot Demonstrations Affect Prompt-based Defenses Against LLM Jailbreak Attacks
von: Wang, Yanshu, et al.
Veröffentlicht: (2026)
von: Wang, Yanshu, et al.
Veröffentlicht: (2026)
BreakFun: Jailbreaking LLMs via Schema Exploitation
von: Oskooei, Amirkia Rafiei, et al.
Veröffentlicht: (2025)
von: Oskooei, Amirkia Rafiei, et al.
Veröffentlicht: (2025)
Tool Receipts, Not Zero-Knowledge Proofs: Practical Hallucination Detection for AI Agents
von: Basu, Abhinaba
Veröffentlicht: (2026)
von: Basu, Abhinaba
Veröffentlicht: (2026)
AlignDP: Hybrid Differential Privacy with Rarity-Aware Protection for LLMs
von: Gaikwad, Madhava
Veröffentlicht: (2025)
von: Gaikwad, Madhava
Veröffentlicht: (2025)
Anonymization-Enhanced Privacy Protection for Mobile GUI Agents: Available but Invisible
von: Zhao, Lepeng, et al.
Veröffentlicht: (2026)
von: Zhao, Lepeng, et al.
Veröffentlicht: (2026)
AVEC: Bootstrapping Privacy for Local LLMs
von: Gaikwad, Madhava
Veröffentlicht: (2025)
von: Gaikwad, Madhava
Veröffentlicht: (2025)
Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs
von: Hill, Brennen, et al.
Veröffentlicht: (2025)
von: Hill, Brennen, et al.
Veröffentlicht: (2025)
Generative AI Models: Opportunities and Risks for Industry and Authorities
von: Alt, Tobias, et al.
Veröffentlicht: (2024)
von: Alt, Tobias, et al.
Veröffentlicht: (2024)
Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks
von: Merves, Tyler H., et al.
Veröffentlicht: (2026)
von: Merves, Tyler H., et al.
Veröffentlicht: (2026)
Pun Unintended: LLMs and the Illusion of Humor Understanding
von: Zangari, Alessandro, et al.
Veröffentlicht: (2025)
von: Zangari, Alessandro, et al.
Veröffentlicht: (2025)
Survey and Evaluation of Converging Architecture in LLMs based on Footsteps of Operations
von: Kim, Seongho, et al.
Veröffentlicht: (2024)
von: Kim, Seongho, et al.
Veröffentlicht: (2024)
SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs
von: Alam, Firoj, et al.
Veröffentlicht: (2025)
von: Alam, Firoj, et al.
Veröffentlicht: (2025)
The Personalization Trap: How User Memory Alters Emotional Reasoning in LLMs
von: Fang, Xi, et al.
Veröffentlicht: (2025)
von: Fang, Xi, et al.
Veröffentlicht: (2025)
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
von: Orgad, Hadas, et al.
Veröffentlicht: (2024)
von: Orgad, Hadas, et al.
Veröffentlicht: (2024)
Task Complexity Matters: An Empirical Study of Reasoning in LLMs for Sentiment Analysis
von: Huang, Donghao, et al.
Veröffentlicht: (2026)
von: Huang, Donghao, et al.
Veröffentlicht: (2026)
MORABLES: A Benchmark for Assessing Abstract Moral Reasoning in LLMs with Fables
von: Marcuzzo, Matteo, et al.
Veröffentlicht: (2025)
von: Marcuzzo, Matteo, et al.
Veröffentlicht: (2025)
On the Validity of Traditional Vulnerability Scoring Systems for Adversarial Attacks against LLMs
von: Bahar, Atmane Ayoub Mansour, et al.
Veröffentlicht: (2024)
von: Bahar, Atmane Ayoub Mansour, et al.
Veröffentlicht: (2024)
Cognitive Workspace: Active Memory Management for LLMs -- An Empirical Study of Functional Infinite Context
von: An, Tao
Veröffentlicht: (2025)
von: An, Tao
Veröffentlicht: (2025)
UnifiedCrawl: Aggregated Common Crawl for Affordable Adaptation of LLMs on Low-Resource Languages
von: Tessema, Bethel Melesse, et al.
Veröffentlicht: (2024)
von: Tessema, Bethel Melesse, et al.
Veröffentlicht: (2024)
Recursive Training Loops in LLMs: How training data properties modulate distribution shift in generated data?
von: Kovač, Grgur, et al.
Veröffentlicht: (2025)
von: Kovač, Grgur, et al.
Veröffentlicht: (2025)
Do LLMs have a Gender (Entropy) Bias?
von: Prabhune, Sonal, et al.
Veröffentlicht: (2025)
von: Prabhune, Sonal, et al.
Veröffentlicht: (2025)
Multi-Model Synthetic Training for Mission-Critical Small Language Models
von: Platt, Nolan, et al.
Veröffentlicht: (2025)
von: Platt, Nolan, et al.
Veröffentlicht: (2025)
Accelerating Suffix Jailbreak attacks with Prefix-Shared KV-cache
von: Wang, Xinhai, et al.
Veröffentlicht: (2026)
von: Wang, Xinhai, et al.
Veröffentlicht: (2026)
LegalGuardian: A Privacy-Preserving Framework for Secure Integration of Large Language Models in Legal Practice
von: Demir, M. Mikail, et al.
Veröffentlicht: (2025)
von: Demir, M. Mikail, et al.
Veröffentlicht: (2025)
Diverse LLMs or Diverse Question Interpretations? That is the Ensembling Question
von: Rosales, Rafael, et al.
Veröffentlicht: (2025)
von: Rosales, Rafael, et al.
Veröffentlicht: (2025)
WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
von: Ali, Zien Sheikh, et al.
Veröffentlicht: (2026)
von: Ali, Zien Sheikh, et al.
Veröffentlicht: (2026)
Do LLMs Truly Understand When a Precedent Is Overruled?
von: Zhang, Li, et al.
Veröffentlicht: (2025)
von: Zhang, Li, et al.
Veröffentlicht: (2025)
Causally Grounded Mechanistic Interpretability for LLMs with Faithful Natural-Language Explanations
von: Mahale, Ajay Pravin
Veröffentlicht: (2026)
von: Mahale, Ajay Pravin
Veröffentlicht: (2026)
Hybrid Gated Flow (HGF): Stabilizing 1.58-bit LLMs via Selective Low-Rank Correction
von: Pizzo, David Alejandro Trejo
Veröffentlicht: (2026)
von: Pizzo, David Alejandro Trejo
Veröffentlicht: (2026)
GSDFuse: Capturing Cognitive Inconsistencies from Multi-Dimensional Weak Signals in Social Media Steganalysis
von: Huang, Kaibo, et al.
Veröffentlicht: (2025)
von: Huang, Kaibo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Shortcuts Arising from Contrast: Effective and Covert Clean-Label Attacks in Prompt-Based Learning
von: Xie, Xiaopeng, et al.
Veröffentlicht: (2024) -
Can Watermarked LLMs be Identified by Users via Crafted Prompts?
von: Liu, Aiwei, et al.
Veröffentlicht: (2024) -
On Adversarial Examples for Text Classification by Perturbing Latent Representations
von: Sooksatra, Korn, et al.
Veröffentlicht: (2024) -
Dark LLMs: The Growing Threat of Unaligned AI Models
von: Fire, Michael, et al.
Veröffentlicht: (2025) -
MarkDiffusion: An Open-Source Toolkit for Generative Watermarking of Latent Diffusion Models
von: Pan, Leyi, et al.
Veröffentlicht: (2025)