You've Changed: Detecting Modification of Black-Box Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Dima, Alden, Foulds, James, Pan, Shimei, Feldman, Philip |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RAGged Edges: The Double-Edged Sword of Retrieval-Augmented Chatbots
by: Feldman, Philip, et al.
Published: (2024)
by: Feldman, Philip, et al.
Published: (2024)
Readme_AI: Dynamic Context Construction for Large Language Models
by: Vyas, Millie, et al.
Published: (2025)
by: Vyas, Millie, et al.
Published: (2025)
Can Generative AI be Egalitarian?
by: Feldman, Philip, et al.
Published: (2025)
by: Feldman, Philip, et al.
Published: (2025)
GenderAlign: An Alignment Dataset for Mitigating Gender Bias in Large Language Models
by: Zhang, Tao, et al.
Published: (2024)
by: Zhang, Tao, et al.
Published: (2024)
Black-Box On-Policy Distillation of Large Language Models
by: Ye, Tianzhu, et al.
Published: (2025)
by: Ye, Tianzhu, et al.
Published: (2025)
Opening the Black Box: A Survey on the Mechanisms of Multi-Step Reasoning in Large Language Models
by: Pan, Liangming, et al.
Published: (2026)
by: Pan, Liangming, et al.
Published: (2026)
SeSE: Black-Box Uncertainty Quantification for Large Language Models Based on Structural Information Theory
by: Zhao, Xingtao, et al.
Published: (2025)
by: Zhao, Xingtao, et al.
Published: (2025)
Inference-Aware Prompt Optimization for Aligning Black-Box Large Language Models
by: Mahmud, Saaduddin, et al.
Published: (2025)
by: Mahmud, Saaduddin, et al.
Published: (2025)
BBox-Adapter: Lightweight Adapting for Black-Box Large Language Models
by: Sun, Haotian, et al.
Published: (2024)
by: Sun, Haotian, et al.
Published: (2024)
Large Language Model Confidence Estimation via Black-Box Access
by: Pedapati, Tejaswini, et al.
Published: (2024)
by: Pedapati, Tejaswini, et al.
Published: (2024)
All in How You Ask for It: Simple Black-Box Method for Jailbreak Attacks
by: Takemoto, Kazuhiro
Published: (2024)
by: Takemoto, Kazuhiro
Published: (2024)
CLOMO: Counterfactual Logical Modification with Large Language Models
by: Huang, Yinya, et al.
Published: (2023)
by: Huang, Yinya, et al.
Published: (2023)
A Unifying Human-Centered AI Fairness Framework
by: Rahman, Munshi Mahbubur, et al.
Published: (2025)
by: Rahman, Munshi Mahbubur, et al.
Published: (2025)
Black-Box Opinion Manipulation Attacks to Retrieval-Augmented Generation of Large Language Models
by: Chen, Zhuo, et al.
Published: (2024)
by: Chen, Zhuo, et al.
Published: (2024)
Moderating Harm: Benchmarking Large Language Models for Cyberbullying Detection in YouTube Comments
by: Muminovic, Amel
Published: (2025)
by: Muminovic, Amel
Published: (2025)
Large Language Models are Biased Because They Are Large Language Models
by: Resnik, Philip
Published: (2024)
by: Resnik, Philip
Published: (2024)
PAL: Proxy-Guided Black-Box Attack on Large Language Models
by: Sitawarin, Chawin, et al.
Published: (2024)
by: Sitawarin, Chawin, et al.
Published: (2024)
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems
by: Geng, Jiayi, et al.
Published: (2025)
by: Geng, Jiayi, et al.
Published: (2025)
Efficient Non-Parametric Uncertainty Quantification for Black-Box Large Language Models and Decision Planning
by: Tsai, Yao-Hung Hubert, et al.
Published: (2024)
by: Tsai, Yao-Hung Hubert, et al.
Published: (2024)
Black-Box Hallucination Detection via Consistency Under the Uncertain Expression
by: Joo, Seongho, et al.
Published: (2025)
by: Joo, Seongho, et al.
Published: (2025)
BlackDAN: A Black-Box Multi-Objective Approach for Effective and Contextual Jailbreaking of Large Language Models
by: Wang, Xinyuan, et al.
Published: (2024)
by: Wang, Xinyuan, et al.
Published: (2024)
Language Models Change Facts Based on the Way You Talk
by: Kearney, Matthew, et al.
Published: (2025)
by: Kearney, Matthew, et al.
Published: (2025)
Black-Box Visual Prompt Engineering for Mitigating Object Hallucination in Large Vision Language Models
by: Woo, Sangmin, et al.
Published: (2025)
by: Woo, Sangmin, et al.
Published: (2025)
"If You're Very Clever, No One Knows You've Used It": The Social Dynamics of Developing Generative AI Literacy in the Workplace
by: Xia, Qing Nancy, et al.
Published: (2026)
by: Xia, Qing Nancy, et al.
Published: (2026)
Verify when Uncertain: Beyond Self-Consistency in Black Box Hallucination Detection
by: Xue, Yihao, et al.
Published: (2025)
by: Xue, Yihao, et al.
Published: (2025)
"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks
by: Wang, Libo
Published: (2024)
by: Wang, Libo
Published: (2024)
Your Inference Request Will Become a Black Box: Confidential Inference for Cloud-based Large Language Models
by: Huang, Chung-ju, et al.
Published: (2026)
by: Huang, Chung-ju, et al.
Published: (2026)
Training Deliberative Monitors for Black-Box Scheming Detection
by: Sinha, Aditya, et al.
Published: (2026)
by: Sinha, Aditya, et al.
Published: (2026)
Uncertainty Quantification for Language Models: A Suite of Black-Box, White-Box, LLM Judge, and Ensemble Scorers
by: Bouchard, Dylan, et al.
Published: (2025)
by: Bouchard, Dylan, et al.
Published: (2025)
High Risk of Political Bias in Black Box Emotion Inference Models
by: Plisiecki, Hubert, et al.
Published: (2024)
by: Plisiecki, Hubert, et al.
Published: (2024)
Learned Hallucination Detection in Black-Box LLMs using Token-level Entropy Production Rate
by: Moslonka, Charles, et al.
Published: (2025)
by: Moslonka, Charles, et al.
Published: (2025)
Language Modeling and Understanding Through Paraphrase Generation and Detection
by: Wahle, Jan Philip
Published: (2026)
by: Wahle, Jan Philip
Published: (2026)
When Does Language Transfer Help? Sequential Fine-Tuning for Cross-Lingual Euphemism Detection
by: Sammartino, Julia, et al.
Published: (2025)
by: Sammartino, Julia, et al.
Published: (2025)
A Survey of Calibration Process for Black-Box LLMs
by: Xie, Liangru, et al.
Published: (2024)
by: Xie, Liangru, et al.
Published: (2024)
Do You Trust Me? Cognitive-Affective Signatures of Trustworthiness in Large Language Models
by: Yeo, Gerard, et al.
Published: (2025)
by: Yeo, Gerard, et al.
Published: (2025)
Think Thrice Before You Act: Progressive Thought Refinement in Large Language Models
by: Du, Chengyu, et al.
Published: (2024)
by: Du, Chengyu, et al.
Published: (2024)
Parameter Efficient Multimodal Instruction Tuning for Romanian Vision Language Models
by: Dima, George-Andrei, et al.
Published: (2025)
by: Dima, George-Andrei, et al.
Published: (2025)
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs
by: von Recum, Alexander, et al.
Published: (2024)
by: von Recum, Alexander, et al.
Published: (2024)
Let's Think Outside the Box: Exploring Leap-of-Thought in Large Language Models with Creative Humor Generation
by: Zhong, Shanshan, et al.
Published: (2023)
by: Zhong, Shanshan, et al.
Published: (2023)
Integrating Large Language Models with Human Expertise for Disease Detection in Electronic Health Records
by: Pan, Jie, et al.
Published: (2025)
by: Pan, Jie, et al.
Published: (2025)
Similar Items
-
RAGged Edges: The Double-Edged Sword of Retrieval-Augmented Chatbots
by: Feldman, Philip, et al.
Published: (2024) -
Readme_AI: Dynamic Context Construction for Large Language Models
by: Vyas, Millie, et al.
Published: (2025) -
Can Generative AI be Egalitarian?
by: Feldman, Philip, et al.
Published: (2025) -
GenderAlign: An Alignment Dataset for Mitigating Gender Bias in Large Language Models
by: Zhang, Tao, et al.
Published: (2024) -
Black-Box On-Policy Distillation of Large Language Models
by: Ye, Tianzhu, et al.
Published: (2025)