Salvato in:
| Autori principali: | Jawad, Hussein, Chenik, Yassine, Brunel, Nicolas J. -B. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2406.02044 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
PSM: Prompt Sensitivity Minimization via LLM-Guided Black-Box Optimization
di: Jawad, Huseein, et al.
Pubblicazione: (2025)
di: Jawad, Huseein, et al.
Pubblicazione: (2025)
ToolFlood: Beyond Selection -- Hiding Valid Tools from LLM Agents via Semantic Covering
di: Jawad, Hussein, et al.
Pubblicazione: (2026)
di: Jawad, Hussein, et al.
Pubblicazione: (2026)
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
di: Mehrotra, Anay, et al.
Pubblicazione: (2023)
di: Mehrotra, Anay, et al.
Pubblicazione: (2023)
Audit Me If You Can: Query-Efficient Active Fairness Auditing of Black-Box LLMs
di: Hartmann, David, et al.
Pubblicazione: (2026)
di: Hartmann, David, et al.
Pubblicazione: (2026)
Predicting the Performance of Black-box LLMs through Follow-up Queries
di: Sam, Dylan, et al.
Pubblicazione: (2025)
di: Sam, Dylan, et al.
Pubblicazione: (2025)
Matryoshka Pilot: Learning to Drive Black-Box LLMs with LLMs
di: Li, Changhao, et al.
Pubblicazione: (2024)
di: Li, Changhao, et al.
Pubblicazione: (2024)
PCS: Perceived Confidence Scoring of Black Box LLMs with Metamorphic Relations
di: Salimian, Sina, et al.
Pubblicazione: (2025)
di: Salimian, Sina, et al.
Pubblicazione: (2025)
SafePassage: High-Fidelity Information Extraction with Black Box LLMs
di: Barrow, Joe, et al.
Pubblicazione: (2025)
di: Barrow, Joe, et al.
Pubblicazione: (2025)
In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
di: Kroeger, Nicholas, et al.
Pubblicazione: (2023)
di: Kroeger, Nicholas, et al.
Pubblicazione: (2023)
Bias Similarity Measurement: A Black-Box Audit of Fairness Across LLMs
di: Jeong, Hyejun, et al.
Pubblicazione: (2024)
di: Jeong, Hyejun, et al.
Pubblicazione: (2024)
How to Train Your Advisor: Steering Black-Box LLMs with Advisor Models
di: Asawa, Parth, et al.
Pubblicazione: (2025)
di: Asawa, Parth, et al.
Pubblicazione: (2025)
FactSelfCheck: Fact-Level Black-Box Hallucination Detection for LLMs
di: Sawczyn, Albert, et al.
Pubblicazione: (2025)
di: Sawczyn, Albert, et al.
Pubblicazione: (2025)
PAL: Proxy-Guided Black-Box Attack on Large Language Models
di: Sitawarin, Chawin, et al.
Pubblicazione: (2024)
di: Sitawarin, Chawin, et al.
Pubblicazione: (2024)
Does It Make Sense to Explain a Black Box With Another Black Box?
di: Delaunay, Julien, et al.
Pubblicazione: (2024)
di: Delaunay, Julien, et al.
Pubblicazione: (2024)
Group Fairness Meets the Black Box: Enabling Fair Algorithms on Closed LLMs via Post-Processing
di: Xian, Ruicheng, et al.
Pubblicazione: (2025)
di: Xian, Ruicheng, et al.
Pubblicazione: (2025)
Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
di: Kaneko, Masahiro, et al.
Pubblicazione: (2025)
di: Kaneko, Masahiro, et al.
Pubblicazione: (2025)
Kov: Transferable and Naturalistic Black-Box LLM Attacks using Markov Decision Processes and Tree Search
di: Moss, Robert J.
Pubblicazione: (2024)
di: Moss, Robert J.
Pubblicazione: (2024)
Universal Response and Emergence of Induction in LLMs
di: Luick, Niclas
Pubblicazione: (2024)
di: Luick, Niclas
Pubblicazione: (2024)
Ten Words Only Still Help: Improving Black-Box AI-Generated Text Detection via Proxy-Guided Efficient Re-Sampling
di: Shi, Yuhui, et al.
Pubblicazione: (2024)
di: Shi, Yuhui, et al.
Pubblicazione: (2024)
Bounded Behavioral Indistinguishability for Black-Box LLM Distillation
di: Hasan, Munawar
Pubblicazione: (2026)
di: Hasan, Munawar
Pubblicazione: (2026)
Pushing the Frontier of Black-Box LVLM Attacks via Fine-Grained Detail Targeting
di: Zhao, Xiaohan, et al.
Pubblicazione: (2026)
di: Zhao, Xiaohan, et al.
Pubblicazione: (2026)
Towards Lightweight Reliability: Using Soft Prompts for Hallucination Mitigation in Large Language Models
di: Siddiqui, S M Tahmid, et al.
Pubblicazione: (2026)
di: Siddiqui, S M Tahmid, et al.
Pubblicazione: (2026)
ACING: Actor-Critic for Instruction Learning in Black-Box LLMs
di: Kharrat, Salma, et al.
Pubblicazione: (2024)
di: Kharrat, Salma, et al.
Pubblicazione: (2024)
An Evaluation of Explanation Methods for Black-Box Detectors of Machine-Generated Text
di: Schoenegger, Loris, et al.
Pubblicazione: (2024)
di: Schoenegger, Loris, et al.
Pubblicazione: (2024)
Hierarchical Text Classification Using Black Box Large Language Models
di: Yoshimura, Kosuke, et al.
Pubblicazione: (2025)
di: Yoshimura, Kosuke, et al.
Pubblicazione: (2025)
SODA: Semi On-Policy Black-Box Distillation for Large Language Models
di: Chen, Xiwen, et al.
Pubblicazione: (2026)
di: Chen, Xiwen, et al.
Pubblicazione: (2026)
Unlocking the Black Box of Latent Reasoning: An Interpretability-Guided Approach to Intervention
di: Chang, Shuochen, et al.
Pubblicazione: (2026)
di: Chang, Shuochen, et al.
Pubblicazione: (2026)
A Watermark for Black-Box Language Models
di: Bahri, Dara, et al.
Pubblicazione: (2024)
di: Bahri, Dara, et al.
Pubblicazione: (2024)
Deep Learning-based Method for Expressing Knowledge Boundary of Black-Box LLM
di: Sheng, Haotian, et al.
Pubblicazione: (2026)
di: Sheng, Haotian, et al.
Pubblicazione: (2026)
Topic Modelling Black Box Optimization
di: Akramov, Roman, et al.
Pubblicazione: (2025)
di: Akramov, Roman, et al.
Pubblicazione: (2025)
You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories
di: Wei, Zhepei, et al.
Pubblicazione: (2026)
di: Wei, Zhepei, et al.
Pubblicazione: (2026)
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs
di: Mendu, Sai Krishna, et al.
Pubblicazione: (2025)
di: Mendu, Sai Krishna, et al.
Pubblicazione: (2025)
Density estimation with LLMs: a geometric investigation of in-context learning trajectories
di: Liu, Toni J. B., et al.
Pubblicazione: (2024)
di: Liu, Toni J. B., et al.
Pubblicazione: (2024)
Training Deliberative Monitors for Black-Box Scheming Detection
di: Sinha, Aditya, et al.
Pubblicazione: (2026)
di: Sinha, Aditya, et al.
Pubblicazione: (2026)
Does Unlearning Truly Unlearn? A Black Box Evaluation of LLM Unlearning Methods
di: Doshi, Jai, et al.
Pubblicazione: (2024)
di: Doshi, Jai, et al.
Pubblicazione: (2024)
TRN-R1-Zero: Text-rich Network Reasoning via LLMs with Reinforcement Learning Only
di: Liu, Yilun, et al.
Pubblicazione: (2026)
di: Liu, Yilun, et al.
Pubblicazione: (2026)
Emergent Response Planning in LLMs
di: Dong, Zhichen, et al.
Pubblicazione: (2025)
di: Dong, Zhichen, et al.
Pubblicazione: (2025)
Towards Modular LLMs by Building and Reusing a Library of LoRAs
di: Ostapenko, Oleksiy, et al.
Pubblicazione: (2024)
di: Ostapenko, Oleksiy, et al.
Pubblicazione: (2024)
Uncertainty Quantification for Language Models: A Suite of Black-Box, White-Box, LLM Judge, and Ensemble Scorers
di: Bouchard, Dylan, et al.
Pubblicazione: (2025)
di: Bouchard, Dylan, et al.
Pubblicazione: (2025)
HYDRA: Model Factorization Framework for Black-Box LLM Personalization
di: Zhuang, Yuchen, et al.
Pubblicazione: (2024)
di: Zhuang, Yuchen, et al.
Pubblicazione: (2024)
Documenti analoghi
-
PSM: Prompt Sensitivity Minimization via LLM-Guided Black-Box Optimization
di: Jawad, Huseein, et al.
Pubblicazione: (2025) -
ToolFlood: Beyond Selection -- Hiding Valid Tools from LLM Agents via Semantic Covering
di: Jawad, Hussein, et al.
Pubblicazione: (2026) -
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
di: Mehrotra, Anay, et al.
Pubblicazione: (2023) -
Audit Me If You Can: Query-Efficient Active Fairness Auditing of Black-Box LLMs
di: Hartmann, David, et al.
Pubblicazione: (2026) -
Predicting the Performance of Black-box LLMs through Follow-up Queries
di: Sam, Dylan, et al.
Pubblicazione: (2025)