Guardado en:
| Autores principales: | Jawad, Hussein, Chenik, Yassine, Brunel, Nicolas J. -B. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2406.02044 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
PSM: Prompt Sensitivity Minimization via LLM-Guided Black-Box Optimization
por: Jawad, Huseein, et al.
Publicado: (2025)
por: Jawad, Huseein, et al.
Publicado: (2025)
ToolFlood: Beyond Selection -- Hiding Valid Tools from LLM Agents via Semantic Covering
por: Jawad, Hussein, et al.
Publicado: (2026)
por: Jawad, Hussein, et al.
Publicado: (2026)
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
por: Mehrotra, Anay, et al.
Publicado: (2023)
por: Mehrotra, Anay, et al.
Publicado: (2023)
Audit Me If You Can: Query-Efficient Active Fairness Auditing of Black-Box LLMs
por: Hartmann, David, et al.
Publicado: (2026)
por: Hartmann, David, et al.
Publicado: (2026)
Predicting the Performance of Black-box LLMs through Follow-up Queries
por: Sam, Dylan, et al.
Publicado: (2025)
por: Sam, Dylan, et al.
Publicado: (2025)
Matryoshka Pilot: Learning to Drive Black-Box LLMs with LLMs
por: Li, Changhao, et al.
Publicado: (2024)
por: Li, Changhao, et al.
Publicado: (2024)
PCS: Perceived Confidence Scoring of Black Box LLMs with Metamorphic Relations
por: Salimian, Sina, et al.
Publicado: (2025)
por: Salimian, Sina, et al.
Publicado: (2025)
SafePassage: High-Fidelity Information Extraction with Black Box LLMs
por: Barrow, Joe, et al.
Publicado: (2025)
por: Barrow, Joe, et al.
Publicado: (2025)
In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
por: Kroeger, Nicholas, et al.
Publicado: (2023)
por: Kroeger, Nicholas, et al.
Publicado: (2023)
Bias Similarity Measurement: A Black-Box Audit of Fairness Across LLMs
por: Jeong, Hyejun, et al.
Publicado: (2024)
por: Jeong, Hyejun, et al.
Publicado: (2024)
How to Train Your Advisor: Steering Black-Box LLMs with Advisor Models
por: Asawa, Parth, et al.
Publicado: (2025)
por: Asawa, Parth, et al.
Publicado: (2025)
FactSelfCheck: Fact-Level Black-Box Hallucination Detection for LLMs
por: Sawczyn, Albert, et al.
Publicado: (2025)
por: Sawczyn, Albert, et al.
Publicado: (2025)
PAL: Proxy-Guided Black-Box Attack on Large Language Models
por: Sitawarin, Chawin, et al.
Publicado: (2024)
por: Sitawarin, Chawin, et al.
Publicado: (2024)
Does It Make Sense to Explain a Black Box With Another Black Box?
por: Delaunay, Julien, et al.
Publicado: (2024)
por: Delaunay, Julien, et al.
Publicado: (2024)
Group Fairness Meets the Black Box: Enabling Fair Algorithms on Closed LLMs via Post-Processing
por: Xian, Ruicheng, et al.
Publicado: (2025)
por: Xian, Ruicheng, et al.
Publicado: (2025)
Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
por: Kaneko, Masahiro, et al.
Publicado: (2025)
por: Kaneko, Masahiro, et al.
Publicado: (2025)
Kov: Transferable and Naturalistic Black-Box LLM Attacks using Markov Decision Processes and Tree Search
por: Moss, Robert J.
Publicado: (2024)
por: Moss, Robert J.
Publicado: (2024)
Universal Response and Emergence of Induction in LLMs
por: Luick, Niclas
Publicado: (2024)
por: Luick, Niclas
Publicado: (2024)
Ten Words Only Still Help: Improving Black-Box AI-Generated Text Detection via Proxy-Guided Efficient Re-Sampling
por: Shi, Yuhui, et al.
Publicado: (2024)
por: Shi, Yuhui, et al.
Publicado: (2024)
Bounded Behavioral Indistinguishability for Black-Box LLM Distillation
por: Hasan, Munawar
Publicado: (2026)
por: Hasan, Munawar
Publicado: (2026)
Pushing the Frontier of Black-Box LVLM Attacks via Fine-Grained Detail Targeting
por: Zhao, Xiaohan, et al.
Publicado: (2026)
por: Zhao, Xiaohan, et al.
Publicado: (2026)
Towards Lightweight Reliability: Using Soft Prompts for Hallucination Mitigation in Large Language Models
por: Siddiqui, S M Tahmid, et al.
Publicado: (2026)
por: Siddiqui, S M Tahmid, et al.
Publicado: (2026)
ACING: Actor-Critic for Instruction Learning in Black-Box LLMs
por: Kharrat, Salma, et al.
Publicado: (2024)
por: Kharrat, Salma, et al.
Publicado: (2024)
An Evaluation of Explanation Methods for Black-Box Detectors of Machine-Generated Text
por: Schoenegger, Loris, et al.
Publicado: (2024)
por: Schoenegger, Loris, et al.
Publicado: (2024)
Hierarchical Text Classification Using Black Box Large Language Models
por: Yoshimura, Kosuke, et al.
Publicado: (2025)
por: Yoshimura, Kosuke, et al.
Publicado: (2025)
SODA: Semi On-Policy Black-Box Distillation for Large Language Models
por: Chen, Xiwen, et al.
Publicado: (2026)
por: Chen, Xiwen, et al.
Publicado: (2026)
Unlocking the Black Box of Latent Reasoning: An Interpretability-Guided Approach to Intervention
por: Chang, Shuochen, et al.
Publicado: (2026)
por: Chang, Shuochen, et al.
Publicado: (2026)
A Watermark for Black-Box Language Models
por: Bahri, Dara, et al.
Publicado: (2024)
por: Bahri, Dara, et al.
Publicado: (2024)
Deep Learning-based Method for Expressing Knowledge Boundary of Black-Box LLM
por: Sheng, Haotian, et al.
Publicado: (2026)
por: Sheng, Haotian, et al.
Publicado: (2026)
Topic Modelling Black Box Optimization
por: Akramov, Roman, et al.
Publicado: (2025)
por: Akramov, Roman, et al.
Publicado: (2025)
You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories
por: Wei, Zhepei, et al.
Publicado: (2026)
por: Wei, Zhepei, et al.
Publicado: (2026)
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs
por: Mendu, Sai Krishna, et al.
Publicado: (2025)
por: Mendu, Sai Krishna, et al.
Publicado: (2025)
Density estimation with LLMs: a geometric investigation of in-context learning trajectories
por: Liu, Toni J. B., et al.
Publicado: (2024)
por: Liu, Toni J. B., et al.
Publicado: (2024)
Training Deliberative Monitors for Black-Box Scheming Detection
por: Sinha, Aditya, et al.
Publicado: (2026)
por: Sinha, Aditya, et al.
Publicado: (2026)
Does Unlearning Truly Unlearn? A Black Box Evaluation of LLM Unlearning Methods
por: Doshi, Jai, et al.
Publicado: (2024)
por: Doshi, Jai, et al.
Publicado: (2024)
TRN-R1-Zero: Text-rich Network Reasoning via LLMs with Reinforcement Learning Only
por: Liu, Yilun, et al.
Publicado: (2026)
por: Liu, Yilun, et al.
Publicado: (2026)
Emergent Response Planning in LLMs
por: Dong, Zhichen, et al.
Publicado: (2025)
por: Dong, Zhichen, et al.
Publicado: (2025)
Towards Modular LLMs by Building and Reusing a Library of LoRAs
por: Ostapenko, Oleksiy, et al.
Publicado: (2024)
por: Ostapenko, Oleksiy, et al.
Publicado: (2024)
Uncertainty Quantification for Language Models: A Suite of Black-Box, White-Box, LLM Judge, and Ensemble Scorers
por: Bouchard, Dylan, et al.
Publicado: (2025)
por: Bouchard, Dylan, et al.
Publicado: (2025)
HYDRA: Model Factorization Framework for Black-Box LLM Personalization
por: Zhuang, Yuchen, et al.
Publicado: (2024)
por: Zhuang, Yuchen, et al.
Publicado: (2024)
Ejemplares similares
-
PSM: Prompt Sensitivity Minimization via LLM-Guided Black-Box Optimization
por: Jawad, Huseein, et al.
Publicado: (2025) -
ToolFlood: Beyond Selection -- Hiding Valid Tools from LLM Agents via Semantic Covering
por: Jawad, Hussein, et al.
Publicado: (2026) -
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
por: Mehrotra, Anay, et al.
Publicado: (2023) -
Audit Me If You Can: Query-Efficient Active Fairness Auditing of Black-Box LLMs
por: Hartmann, David, et al.
Publicado: (2026) -
Predicting the Performance of Black-box LLMs through Follow-up Queries
por: Sam, Dylan, et al.
Publicado: (2025)