PCS: Perceived Confidence Scoring of Black Box LLMs with Metamorphic Relations
Fuente:
arXiv
Salvato in:
| Autori principali: | Salimian, Sina, Uddin, Gias, Raza, Shaina, Leung, Henry |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Bias Testing and Mitigation in Black Box LLMs using Metamorphic Relations
di: Salimian, Sina, et al.
Pubblicazione: (2025)
di: Salimian, Sina, et al.
Pubblicazione: (2025)
Hallucination Detection in Large Language Models with Metamorphic Relations
di: Yang, Borui, et al.
Pubblicazione: (2025)
di: Yang, Borui, et al.
Pubblicazione: (2025)
Multicalibration for Confidence Scoring in LLMs
di: Detommaso, Gianluca, et al.
Pubblicazione: (2024)
di: Detommaso, Gianluca, et al.
Pubblicazione: (2024)
Exploring Bias and Prediction Metrics to Characterise the Fairness of Machine Learning for Equity-Centered Public Health Decision-Making: A Narrative Review
di: Raza, Shaina, et al.
Pubblicazione: (2024)
di: Raza, Shaina, et al.
Pubblicazione: (2024)
Large Language Model Confidence Estimation via Black-Box Access
di: Pedapati, Tejaswini, et al.
Pubblicazione: (2024)
di: Pedapati, Tejaswini, et al.
Pubblicazione: (2024)
SafePassage: High-Fidelity Information Extraction with Black Box LLMs
di: Barrow, Joe, et al.
Pubblicazione: (2025)
di: Barrow, Joe, et al.
Pubblicazione: (2025)
Matryoshka Pilot: Learning to Drive Black-Box LLMs with LLMs
di: Li, Changhao, et al.
Pubblicazione: (2024)
di: Li, Changhao, et al.
Pubblicazione: (2024)
Towards Universal and Black-Box Query-Response Only Attack on LLMs with QROA
di: Jawad, Hussein, et al.
Pubblicazione: (2024)
di: Jawad, Hussein, et al.
Pubblicazione: (2024)
QA-Calibration of Language Model Confidence Scores
di: Manggala, Putra, et al.
Pubblicazione: (2024)
di: Manggala, Putra, et al.
Pubblicazione: (2024)
In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
di: Kroeger, Nicholas, et al.
Pubblicazione: (2023)
di: Kroeger, Nicholas, et al.
Pubblicazione: (2023)
Audit Me If You Can: Query-Efficient Active Fairness Auditing of Black-Box LLMs
di: Hartmann, David, et al.
Pubblicazione: (2026)
di: Hartmann, David, et al.
Pubblicazione: (2026)
SteerConf: Steering LLMs for Confidence Elicitation
di: Zhou, Ziang, et al.
Pubblicazione: (2025)
di: Zhou, Ziang, et al.
Pubblicazione: (2025)
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
di: Mehrotra, Anay, et al.
Pubblicazione: (2023)
di: Mehrotra, Anay, et al.
Pubblicazione: (2023)
How to Train Your Advisor: Steering Black-Box LLMs with Advisor Models
di: Asawa, Parth, et al.
Pubblicazione: (2025)
di: Asawa, Parth, et al.
Pubblicazione: (2025)
FactSelfCheck: Fact-Level Black-Box Hallucination Detection for LLMs
di: Sawczyn, Albert, et al.
Pubblicazione: (2025)
di: Sawczyn, Albert, et al.
Pubblicazione: (2025)
Bias Similarity Measurement: A Black-Box Audit of Fairness Across LLMs
di: Jeong, Hyejun, et al.
Pubblicazione: (2024)
di: Jeong, Hyejun, et al.
Pubblicazione: (2024)
Does It Make Sense to Explain a Black Box With Another Black Box?
di: Delaunay, Julien, et al.
Pubblicazione: (2024)
di: Delaunay, Julien, et al.
Pubblicazione: (2024)
Group Fairness Meets the Black Box: Enabling Fair Algorithms on Closed LLMs via Post-Processing
di: Xian, Ruicheng, et al.
Pubblicazione: (2025)
di: Xian, Ruicheng, et al.
Pubblicazione: (2025)
The Confidence Trap: Gender Bias and Predictive Certainty in LLMs
di: Sabir, Ahmed, et al.
Pubblicazione: (2026)
di: Sabir, Ahmed, et al.
Pubblicazione: (2026)
Factual Confidence of LLMs: on Reliability and Robustness of Current Estimators
di: Mahaut, Matéo, et al.
Pubblicazione: (2024)
di: Mahaut, Matéo, et al.
Pubblicazione: (2024)
Bounded Behavioral Indistinguishability for Black-Box LLM Distillation
di: Hasan, Munawar
Pubblicazione: (2026)
di: Hasan, Munawar
Pubblicazione: (2026)
Does Alignment Tuning Really Break LLMs' Internal Confidence?
di: Oh, Hongseok, et al.
Pubblicazione: (2024)
di: Oh, Hongseok, et al.
Pubblicazione: (2024)
Confidence-Credibility Aware Weighted Ensembles of Small LLMs Outperform Large LLMs in Emotion Detection
di: Elgabry, Menna, et al.
Pubblicazione: (2025)
di: Elgabry, Menna, et al.
Pubblicazione: (2025)
Generating with Confidence: Uncertainty Quantification for Black-box Large Language Models
di: Lin, Zhen, et al.
Pubblicazione: (2023)
di: Lin, Zhen, et al.
Pubblicazione: (2023)
Learning to Route LLMs with Confidence Tokens
di: Chuang, Yu-Neng, et al.
Pubblicazione: (2024)
di: Chuang, Yu-Neng, et al.
Pubblicazione: (2024)
ConfClip: Confidence-Weighted and Clipped Reward for Reinforcement Learning in LLMs
di: Zhang, Bonan, et al.
Pubblicazione: (2025)
di: Zhang, Bonan, et al.
Pubblicazione: (2025)
Hierarchical Text Classification Using Black Box Large Language Models
di: Yoshimura, Kosuke, et al.
Pubblicazione: (2025)
di: Yoshimura, Kosuke, et al.
Pubblicazione: (2025)
An Evaluation of Explanation Methods for Black-Box Detectors of Machine-Generated Text
di: Schoenegger, Loris, et al.
Pubblicazione: (2024)
di: Schoenegger, Loris, et al.
Pubblicazione: (2024)
SODA: Semi On-Policy Black-Box Distillation for Large Language Models
di: Chen, Xiwen, et al.
Pubblicazione: (2026)
di: Chen, Xiwen, et al.
Pubblicazione: (2026)
Unlocking the Black Box of Latent Reasoning: An Interpretability-Guided Approach to Intervention
di: Chang, Shuochen, et al.
Pubblicazione: (2026)
di: Chang, Shuochen, et al.
Pubblicazione: (2026)
Do LLMs Really Forget? Evaluating Unlearning with Knowledge Correlation and Confidence Awareness
di: Wei, Rongzhe, et al.
Pubblicazione: (2025)
di: Wei, Rongzhe, et al.
Pubblicazione: (2025)
Deep Learning-based Method for Expressing Knowledge Boundary of Black-Box LLM
di: Sheng, Haotian, et al.
Pubblicazione: (2026)
di: Sheng, Haotian, et al.
Pubblicazione: (2026)
How do LLMs Compute Verbal Confidence
di: Kumaran, Dharshan, et al.
Pubblicazione: (2026)
di: Kumaran, Dharshan, et al.
Pubblicazione: (2026)
A Watermark for Black-Box Language Models
di: Bahri, Dara, et al.
Pubblicazione: (2024)
di: Bahri, Dara, et al.
Pubblicazione: (2024)
ACING: Actor-Critic for Instruction Learning in Black-Box LLMs
di: Kharrat, Salma, et al.
Pubblicazione: (2024)
di: Kharrat, Salma, et al.
Pubblicazione: (2024)
Does Unlearning Truly Unlearn? A Black Box Evaluation of LLM Unlearning Methods
di: Doshi, Jai, et al.
Pubblicazione: (2024)
di: Doshi, Jai, et al.
Pubblicazione: (2024)
Beyond the Score: Uncertainty-Calibrated LLMs for Automated Essay Assessment
di: Karim, Ahmed, et al.
Pubblicazione: (2025)
di: Karim, Ahmed, et al.
Pubblicazione: (2025)
Enhancing Marker Scoring Accuracy through Ordinal Confidence Modelling in Educational Assessments
di: Chakravarty, Abhirup, et al.
Pubblicazione: (2025)
di: Chakravarty, Abhirup, et al.
Pubblicazione: (2025)
Training Deliberative Monitors for Black-Box Scheming Detection
di: Sinha, Aditya, et al.
Pubblicazione: (2026)
di: Sinha, Aditya, et al.
Pubblicazione: (2026)
MARS: Meaning-Aware Response Scoring for Uncertainty Estimation in Generative LLMs
di: Bakman, Yavuz Faruk, et al.
Pubblicazione: (2024)
di: Bakman, Yavuz Faruk, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Bias Testing and Mitigation in Black Box LLMs using Metamorphic Relations
di: Salimian, Sina, et al.
Pubblicazione: (2025) -
Hallucination Detection in Large Language Models with Metamorphic Relations
di: Yang, Borui, et al.
Pubblicazione: (2025) -
Multicalibration for Confidence Scoring in LLMs
di: Detommaso, Gianluca, et al.
Pubblicazione: (2024) -
Exploring Bias and Prediction Metrics to Characterise the Fairness of Machine Learning for Equity-Centered Public Health Decision-Making: A Narrative Review
di: Raza, Shaina, et al.
Pubblicazione: (2024) -
Large Language Model Confidence Estimation via Black-Box Access
di: Pedapati, Tejaswini, et al.
Pubblicazione: (2024)