DALD: Improving Logits-based Detector without Logits from Black-box LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Zeng, Cong, Tang, Shengkun, Yang, Xianjun, Chen, Yuanzhou, Sun, Yiyou, xu, zhiqiang, Li, Yao, Chen, Haifeng, Cheng, Wei, Xu, Dongkuan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Where's the liability in the Generative Era? Recovery-based Black-Box Detection of AI-Generated Content
di: Bai, Haoyue, et al.
Pubblicazione: (2025)
di: Bai, Haoyue, et al.
Pubblicazione: (2025)
Human Texts Are Outliers: Detecting LLM-generated Texts via Out-of-distribution Detection
di: Zeng, Cong, et al.
Pubblicazione: (2025)
di: Zeng, Cong, et al.
Pubblicazione: (2025)
Knowledge Distillation with Refined Logits
di: Sun, Wujie, et al.
Pubblicazione: (2024)
di: Sun, Wujie, et al.
Pubblicazione: (2024)
Humanizing the Machine: Proxy Attacks to Mislead LLM Detectors
di: Wang, Tianchun, et al.
Pubblicazione: (2024)
di: Wang, Tianchun, et al.
Pubblicazione: (2024)
DFA-RAG: Conversational Semantic Router for Large Language Model with Definite Finite Automaton
di: Sun, Yiyou, et al.
Pubblicazione: (2024)
di: Sun, Yiyou, et al.
Pubblicazione: (2024)
LogitLens4LLMs: Extending Logit Lens Analysis to Modern Large Language Models
di: Wang, Zhenyu
Pubblicazione: (2025)
di: Wang, Zhenyu
Pubblicazione: (2025)
Logits Poisoning Attack in Federated Distillation
di: Tang, Yuhan, et al.
Pubblicazione: (2024)
di: Tang, Yuhan, et al.
Pubblicazione: (2024)
Lost in Overlap: Exploring Logit-based Watermark Collision in LLMs
di: Luo, Yiyang, et al.
Pubblicazione: (2024)
di: Luo, Yiyang, et al.
Pubblicazione: (2024)
Stabilizing Policy Optimization via Logits Convexity
di: Chen, Hongzhan, et al.
Pubblicazione: (2026)
di: Chen, Hongzhan, et al.
Pubblicazione: (2026)
LogitsCoder: Towards Efficient Chain-of-Thought Path Search via Logits Preference Decoding for Code Generation
di: Chen, Jizheng, et al.
Pubblicazione: (2026)
di: Chen, Jizheng, et al.
Pubblicazione: (2026)
On Giant's Shoulders: Effortless Weak to Strong by Dynamic Logits Fusion
di: Fan, Chenghao, et al.
Pubblicazione: (2024)
di: Fan, Chenghao, et al.
Pubblicazione: (2024)
Logit Standardization in Knowledge Distillation
di: Sun, Shangquan, et al.
Pubblicazione: (2024)
di: Sun, Shangquan, et al.
Pubblicazione: (2024)
SLED: Self Logits Evolution Decoding for Improving Factuality in Large Language Models
di: Zhang, Jianyi, et al.
Pubblicazione: (2024)
di: Zhang, Jianyi, et al.
Pubblicazione: (2024)
LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories
di: He, Zirui, et al.
Pubblicazione: (2025)
di: He, Zirui, et al.
Pubblicazione: (2025)
VQ-Logits: Compressing the Output Bottleneck of Large Language Models via Vector Quantized Logits
di: Shao, Jintian, et al.
Pubblicazione: (2025)
di: Shao, Jintian, et al.
Pubblicazione: (2025)
Beyond Hidden-Layer Manipulation: Semantically-Aware Logit Interventions for Debiasing LLMs
di: Xia, Wei
Pubblicazione: (2025)
di: Xia, Wei
Pubblicazione: (2025)
Strategic Deflection: Defending LLMs from Logit Manipulation
di: Rachidy, Yassine, et al.
Pubblicazione: (2025)
di: Rachidy, Yassine, et al.
Pubblicazione: (2025)
Logits of API-Protected LLMs Leak Proprietary Information
di: Finlayson, Matthew, et al.
Pubblicazione: (2024)
di: Finlayson, Matthew, et al.
Pubblicazione: (2024)
Sparse Logit Sampling: Accelerating Knowledge Distillation in LLMs
di: Anshumann, et al.
Pubblicazione: (2025)
di: Anshumann, et al.
Pubblicazione: (2025)
On Logit Weibull Manifold
di: Assandje, Prosper Rosaire Mama, et al.
Pubblicazione: (2025)
di: Assandje, Prosper Rosaire Mama, et al.
Pubblicazione: (2025)
Logits-Based Finetuning
di: Li, Jingyao, et al.
Pubblicazione: (2025)
di: Li, Jingyao, et al.
Pubblicazione: (2025)
On the Estimation of Multinomial Logit and Nested Logit Models: A Conic Optimization Approach
di: Pham, Hoang Giang, et al.
Pubblicazione: (2025)
di: Pham, Hoang Giang, et al.
Pubblicazione: (2025)
LogitDynamics: Reliable ViT Error Detection from Layerwise Logit Trajectories
di: Beigelman, Ido, et al.
Pubblicazione: (2026)
di: Beigelman, Ido, et al.
Pubblicazione: (2026)
Peak-Controlled Logits Poisoning Attack in Federated Distillation
di: Tang, Yuhan, et al.
Pubblicazione: (2024)
di: Tang, Yuhan, et al.
Pubblicazione: (2024)
Top-$nσ$: Not All Logits Are You Need
di: Tang, Chenxia, et al.
Pubblicazione: (2024)
di: Tang, Chenxia, et al.
Pubblicazione: (2024)
AdaDiff: Accelerating Diffusion Models through Step-Wise Adaptive Computation
di: Tang, Shengkun, et al.
Pubblicazione: (2023)
di: Tang, Shengkun, et al.
Pubblicazione: (2023)
Spectral Logit Sculpting: Adaptive Low-Rank Logit Transformation for Controlled Text Generation
di: Li, Jin, et al.
Pubblicazione: (2025)
di: Li, Jin, et al.
Pubblicazione: (2025)
Aligning Logits Generatively for Principled Black-Box Knowledge Distillation
di: Ma, Jing, et al.
Pubblicazione: (2022)
di: Ma, Jing, et al.
Pubblicazione: (2022)
BiGain: Unified Token Compression for Joint Generation and Classification
di: Liu, Jiacheng, et al.
Pubblicazione: (2026)
di: Liu, Jiacheng, et al.
Pubblicazione: (2026)
The Implicit Bias of Logit Regularization
di: Beck, Alon, et al.
Pubblicazione: (2026)
di: Beck, Alon, et al.
Pubblicazione: (2026)
Gradient-Aware Logit Adjustment Loss for Long-tailed Classifier
di: Zhang, Fan, et al.
Pubblicazione: (2024)
di: Zhang, Fan, et al.
Pubblicazione: (2024)
Mitigating Spurious Correlations with Causal Logit Perturbation
di: Zhou, Xiaoling, et al.
Pubblicazione: (2025)
di: Zhou, Xiaoling, et al.
Pubblicazione: (2025)
From Logits to Latents: Contrastive Representation Shaping for LLM Unlearning
di: Tang, Haoran, et al.
Pubblicazione: (2026)
di: Tang, Haoran, et al.
Pubblicazione: (2026)
Logit-based alternatives to two-stage least squares
di: Chetverikov, Denis, et al.
Pubblicazione: (2023)
di: Chetverikov, Denis, et al.
Pubblicazione: (2023)
Last Layer Logits to Logic: Empowering LLMs with Logic-Consistent Structured Knowledge Reasoning
di: Li, Songze, et al.
Pubblicazione: (2025)
di: Li, Songze, et al.
Pubblicazione: (2025)
Towards Cross-Tokenizer Distillation: the Universal Logit Distillation Loss for LLMs
di: Boizard, Nicolas, et al.
Pubblicazione: (2024)
di: Boizard, Nicolas, et al.
Pubblicazione: (2024)
Revisiting Logit Distributions for Reliable Out-of-Distribution Detection
di: Liang, Jiachen, et al.
Pubblicazione: (2025)
di: Liang, Jiachen, et al.
Pubblicazione: (2025)
AMU-Tuning: Effective Logit Bias for CLIP-based Few-shot Learning
di: Tang, Yuwei, et al.
Pubblicazione: (2024)
di: Tang, Yuwei, et al.
Pubblicazione: (2024)
LEAD: Exploring Logit Space Evolution for Model Selection
di: Hu, Zixuan, et al.
Pubblicazione: (2025)
di: Hu, Zixuan, et al.
Pubblicazione: (2025)
Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
di: Li, Yuxi, et al.
Pubblicazione: (2024)
di: Li, Yuxi, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Where's the liability in the Generative Era? Recovery-based Black-Box Detection of AI-Generated Content
di: Bai, Haoyue, et al.
Pubblicazione: (2025) -
Human Texts Are Outliers: Detecting LLM-generated Texts via Out-of-distribution Detection
di: Zeng, Cong, et al.
Pubblicazione: (2025) -
Knowledge Distillation with Refined Logits
di: Sun, Wujie, et al.
Pubblicazione: (2024) -
Humanizing the Machine: Proxy Attacks to Mislead LLM Detectors
di: Wang, Tianchun, et al.
Pubblicazione: (2024) -
DFA-RAG: Conversational Semantic Router for Large Language Model with Definite Finite Automaton
di: Sun, Yiyou, et al.
Pubblicazione: (2024)