Towards Certification of Uncertainty Calibration under Adversarial Attacks
Fuente:
arXiv
Guardado en:
| Autores principales: | Emde, Cornelius, Pinto, Francesco, Lukasiewicz, Thomas, Torr, Philip H. S., Bibi, Adel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Shh, don't say that! Domain Certification in LLMs
por: Emde, Cornelius, et al.
Publicado: (2025)
por: Emde, Cornelius, et al.
Publicado: (2025)
When Do Prompting and Prefix-Tuning Work? A Theory of Capabilities and Limitations
por: Petrov, Aleksandar, et al.
Publicado: (2023)
por: Petrov, Aleksandar, et al.
Publicado: (2023)
Prompting a Pretrained Transformer Can Be a Universal Approximator
por: Petrov, Aleksandar, et al.
Publicado: (2024)
por: Petrov, Aleksandar, et al.
Publicado: (2024)
Efficient Error Certification for Physics-Informed Neural Networks
por: Eiras, Francisco, et al.
Publicado: (2023)
por: Eiras, Francisco, et al.
Publicado: (2023)
Bi-Factorial Preference Optimization: Balancing Safety-Helpfulness in Language Models
por: Zhang, Wenxuan, et al.
Publicado: (2024)
por: Zhang, Wenxuan, et al.
Publicado: (2024)
FORCE: Transferable Visual Jailbreaking Attacks via Feature Over-Reliance CorrEction
por: Lin, Runqi, et al.
Publicado: (2025)
por: Lin, Runqi, et al.
Publicado: (2025)
Beyond Linear Probes: Dynamic Safety Monitoring for Language Models
por: Oldfield, James, et al.
Publicado: (2025)
por: Oldfield, James, et al.
Publicado: (2025)
Detecting LLM Hallucination Through Layer-wise Information Deficiency: Analysis of Ambiguous Prompts and Unanswerable Questions
por: Kim, Hazel, et al.
Publicado: (2024)
por: Kim, Hazel, et al.
Publicado: (2024)
Do as I do (Safely): Mitigating Task-Specific Fine-tuning Risks in Large Language Models
por: Eiras, Francisco, et al.
Publicado: (2024)
por: Eiras, Francisco, et al.
Publicado: (2024)
Towards Interpretable Deep Local Learning with Successive Gradient Reconciliation
por: Yang, Yibo, et al.
Publicado: (2024)
por: Yang, Yibo, et al.
Publicado: (2024)
Universal In-Context Approximation By Prompting Fully Recurrent Models
por: Petrov, Aleksandar, et al.
Publicado: (2024)
por: Petrov, Aleksandar, et al.
Publicado: (2024)
Segment, Select, Correct: A Framework for Weakly-Supervised Referring Segmentation
por: Eiras, Francisco, et al.
Publicado: (2023)
por: Eiras, Francisco, et al.
Publicado: (2023)
MIP against Agent: Malicious Image Patches Hijacking Multimodal OS Agents
por: Aichberger, Lukas, et al.
Publicado: (2025)
por: Aichberger, Lukas, et al.
Publicado: (2025)
Continual Learning on a Diet: Learning from Sparsely Labeled Streams Under Constrained Computation
por: Zhang, Wenxuan, et al.
Publicado: (2024)
por: Zhang, Wenxuan, et al.
Publicado: (2024)
A Stable, Fast, and Fully Automatic Learning Algorithm for Predictive Coding Networks
por: Salvatori, Tommaso, et al.
Publicado: (2022)
por: Salvatori, Tommaso, et al.
Publicado: (2022)
Efficient Lifelong Model Evaluation in an Era of Rapid Progress
por: Prabhu, Ameya, et al.
Publicado: (2024)
por: Prabhu, Ameya, et al.
Publicado: (2024)
From Categories to Classifiers: Name-Only Continual Learning by Exploring the Web
por: Prabhu, Ameya, et al.
Publicado: (2023)
por: Prabhu, Ameya, et al.
Publicado: (2023)
PROSAC: Provably Safe Certification for Machine Learning Models under Adversarial Attacks
por: Feng, Chen, et al.
Publicado: (2024)
por: Feng, Chen, et al.
Publicado: (2024)
Revisiting Uncertainty Estimation and Calibration of Large Language Models
por: Tao, Linwei, et al.
Publicado: (2025)
por: Tao, Linwei, et al.
Publicado: (2025)
Robustness Certificates for Neural Networks against Adversarial Attacks
por: Taheri, Sara, et al.
Publicado: (2025)
por: Taheri, Sara, et al.
Publicado: (2025)
Hard Regularization to Prevent Deep Online Clustering Collapse without Data Augmentation
por: Mahon, Louis, et al.
Publicado: (2023)
por: Mahon, Louis, et al.
Publicado: (2023)
Towards the Training of Deeper Predictive Coding Neural Networks
por: Qi, Chang, et al.
Publicado: (2025)
por: Qi, Chang, et al.
Publicado: (2025)
Mixture of Experts Made Intrinsically Interpretable
por: Yang, Xingyi, et al.
Publicado: (2025)
por: Yang, Xingyi, et al.
Publicado: (2025)
Rethinking Safety in LLM Fine-tuning: An Optimization Perspective
por: Kim, Minseon, et al.
Publicado: (2025)
por: Kim, Minseon, et al.
Publicado: (2025)
On the Robustness of Adversarial Training Against Uncertainty Attacks
por: Ledda, Emanuele, et al.
Publicado: (2024)
por: Ledda, Emanuele, et al.
Publicado: (2024)
On Pretraining Data Diversity for Self-Supervised Learning
por: Hammoud, Hasan Abed Al Kader, et al.
Publicado: (2024)
por: Hammoud, Hasan Abed Al Kader, et al.
Publicado: (2024)
SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?
por: Hammoud, Hasan Abed Al Kader, et al.
Publicado: (2024)
por: Hammoud, Hasan Abed Al Kader, et al.
Publicado: (2024)
Safer by Diffusion, Broken by Context: Diffusion LLM's Safety Blessing and Its Failure Mode
por: He, Zeyuan, et al.
Publicado: (2026)
por: He, Zeyuan, et al.
Publicado: (2026)
No "Zero-Shot" Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model Performance
por: Udandarao, Vishaal, et al.
Publicado: (2024)
por: Udandarao, Vishaal, et al.
Publicado: (2024)
Model Merging and Safety Alignment: One Bad Model Spoils the Bunch
por: Hammoud, Hasan Abed Al Kader, et al.
Publicado: (2024)
por: Hammoud, Hasan Abed Al Kader, et al.
Publicado: (2024)
Calibration Attacks: A Comprehensive Study of Adversarial Attacks on Model Confidence
por: Obadinma, Stephen, et al.
Publicado: (2024)
por: Obadinma, Stephen, et al.
Publicado: (2024)
Measuring Uncertainty Calibration
por: Ciosek, Kamil, et al.
Publicado: (2025)
por: Ciosek, Kamil, et al.
Publicado: (2025)
Exact Certification of Data-Poisoning Attacks Using Mixed-Integer Programming
por: Sosnin, Philip, et al.
Publicado: (2026)
por: Sosnin, Philip, et al.
Publicado: (2026)
Ensuring Calibration Robustness in Split Conformal Prediction Under Adversarial Attacks
por: Qian, Xunlei, et al.
Publicado: (2025)
por: Qian, Xunlei, et al.
Publicado: (2025)
Focus On This, Not That! Steering LLMs with Adaptive Feature Specification
por: Lamb, Tom A., et al.
Publicado: (2024)
por: Lamb, Tom A., et al.
Publicado: (2024)
Label Delay in Online Continual Learning
por: Csaba, Botos, et al.
Publicado: (2023)
por: Csaba, Botos, et al.
Publicado: (2023)
Benchmarking Predictive Coding Networks -- Made Simple
por: Pinchetti, Luca, et al.
Publicado: (2024)
por: Pinchetti, Luca, et al.
Publicado: (2024)
A Survey on Deep Learning Approaches for Tabular Data Generation: Utility, Alignment, Fidelity, Privacy, Diversity, and Beyond
por: Stoian, Mihaela Cătălina, et al.
Publicado: (2025)
por: Stoian, Mihaela Cătălina, et al.
Publicado: (2025)
Improving Semantic Uncertainty Quantification in Language Model Question-Answering via Token-Level Temperature Scaling
por: Lamb, Tom A., et al.
Publicado: (2026)
por: Lamb, Tom A., et al.
Publicado: (2026)
Relationship between Uncertainty in DNNs and Adversarial Attacks
por: Ogonna, Mabel, et al.
Publicado: (2024)
por: Ogonna, Mabel, et al.
Publicado: (2024)
Ejemplares similares
-
Shh, don't say that! Domain Certification in LLMs
por: Emde, Cornelius, et al.
Publicado: (2025) -
When Do Prompting and Prefix-Tuning Work? A Theory of Capabilities and Limitations
por: Petrov, Aleksandar, et al.
Publicado: (2023) -
Prompting a Pretrained Transformer Can Be a Universal Approximator
por: Petrov, Aleksandar, et al.
Publicado: (2024) -
Efficient Error Certification for Physics-Informed Neural Networks
por: Eiras, Francisco, et al.
Publicado: (2023) -
Bi-Factorial Preference Optimization: Balancing Safety-Helpfulness in Language Models
por: Zhang, Wenxuan, et al.
Publicado: (2024)