Towards Realistic Guarantees: A Probabilistic Certificate for SmoothLLM
Fuente:
arXiv
Guardado en:
| Autores principales: | Kumarappan, Adarsh, Mehrotra, Ayushi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Automating Deception: Scalable Multi-Turn LLM Jailbreaks
por: Kumarappan, Adarsh, et al.
Publicado: (2025)
por: Kumarappan, Adarsh, et al.
Publicado: (2025)
SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
por: Robey, Alexander, et al.
Publicado: (2023)
por: Robey, Alexander, et al.
Publicado: (2023)
Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy
por: Kumarappan, Adarsh, et al.
Publicado: (2026)
por: Kumarappan, Adarsh, et al.
Publicado: (2026)
DevBench: A Realistic, Developer-Informed Benchmark for Code Generation Models
por: Kumarappan, Adarsh, et al.
Publicado: (2026)
por: Kumarappan, Adarsh, et al.
Publicado: (2026)
LeanAgent: Lifelong Learning for Formal Theorem Proving
por: Kumarappan, Adarsh, et al.
Publicado: (2024)
por: Kumarappan, Adarsh, et al.
Publicado: (2024)
Probabilistic Stability Guarantees for Feature Attributions
por: Jin, Helen, et al.
Publicado: (2025)
por: Jin, Helen, et al.
Publicado: (2025)
Rigorous Probabilistic Guarantees for Robust Counterfactual Explanations
por: Marzari, Luca, et al.
Publicado: (2024)
por: Marzari, Luca, et al.
Publicado: (2024)
Trust Regions for Explanations via Black-Box Probabilistic Certification
por: Dhurandhar, Amit, et al.
Publicado: (2024)
por: Dhurandhar, Amit, et al.
Publicado: (2024)
No Certificate for Alignment: Two Independent Impossibilities and the Pareto Frontier of Achievable Safety Guarantees
por: Agarwal, Ayushi
Publicado: (2026)
por: Agarwal, Ayushi
Publicado: (2026)
Probabilistic Performance Guarantees for Multi-Task Reinforcement Learning
por: Schnitzer, Yannik, et al.
Publicado: (2026)
por: Schnitzer, Yannik, et al.
Publicado: (2026)
Counterfactual Explanations with Probabilistic Guarantees on their Robustness to Model Change
por: Stępka, Ignacy, et al.
Publicado: (2024)
por: Stępka, Ignacy, et al.
Publicado: (2024)
Reinforcement Learning for Control with Probabilistic Stability Guarantee: A Finite-Sample Approach
por: Han, Minghao, et al.
Publicado: (2026)
por: Han, Minghao, et al.
Publicado: (2026)
Enumerating Safe Regions in Deep Neural Networks with Provable Probabilistic Guarantees
por: Marzari, Luca, et al.
Publicado: (2023)
por: Marzari, Luca, et al.
Publicado: (2023)
Mechanistic Interpretability of Brain-to-Speech Models Across Speech Modes
por: Maghsoudi, Maryam, et al.
Publicado: (2026)
por: Maghsoudi, Maryam, et al.
Publicado: (2026)
Leveraging Approximate Model-based Shielding for Probabilistic Safety Guarantees in Continuous Environments
por: Goodall, Alexander W., et al.
Publicado: (2024)
por: Goodall, Alexander W., et al.
Publicado: (2024)
SimCert: Probabilistic Certification for Behavioral Similarity in Deep Neural Network Compression
por: Li, Jingyang, et al.
Publicado: (2026)
por: Li, Jingyang, et al.
Publicado: (2026)
Improved Guarantees for Heterogeneous Treatment-Effect Estimation via Matrix Completion
por: Mehrotra, Anay, et al.
Publicado: (2026)
por: Mehrotra, Anay, et al.
Publicado: (2026)
Towards Interpretable Hate Speech Detection using Large Language Model-extracted Rationales
por: Nirmal, Ayushi, et al.
Publicado: (2024)
por: Nirmal, Ayushi, et al.
Publicado: (2024)
Anytime-Valid Answer Sufficiency Certificates for LLM Generation via Sequential Information Lift
por: Akter, Sanjeda, et al.
Publicado: (2025)
por: Akter, Sanjeda, et al.
Publicado: (2025)
Toward Maturity-Based Certification of Embodied AI: Quantifying Trustworthiness Through Measurement Mechanisms
por: Darling, Michael C., et al.
Publicado: (2026)
por: Darling, Michael C., et al.
Publicado: (2026)
Rational Tuning of LLM Cascades via Probabilistic Modeling
por: Zellinger, Michael J., et al.
Publicado: (2025)
por: Zellinger, Michael J., et al.
Publicado: (2025)
Smooth InfoMax -- Towards Easier Post-Hoc Interpretability
por: Denoodt, Fabian, et al.
Publicado: (2024)
por: Denoodt, Fabian, et al.
Publicado: (2024)
Robust Counterfactual Explanations for Neural Networks With Probabilistic Guarantees
por: Hamman, Faisal, et al.
Publicado: (2023)
por: Hamman, Faisal, et al.
Publicado: (2023)
How NOT to benchmark your SITE metric: Beyond Static Leaderboards and Towards Realistic Evaluation
por: Singh, Prabhant, et al.
Publicado: (2025)
por: Singh, Prabhant, et al.
Publicado: (2025)
It's Not You, It's Clipping: A Soft Trust-Region via Probability Smoothing for LLM RL
por: Dwyer, Madeleine, et al.
Publicado: (2025)
por: Dwyer, Madeleine, et al.
Publicado: (2025)
Towards a Probabilistic Fusion Approach for Robust Battery Prognostics
por: Alcibar, Jokin, et al.
Publicado: (2024)
por: Alcibar, Jokin, et al.
Publicado: (2024)
Towards Probabilistic Inductive Logic Programming with Neurosymbolic Inference and Relaxation
por: Hillerstrom, Fieke, et al.
Publicado: (2024)
por: Hillerstrom, Fieke, et al.
Publicado: (2024)
GeoLLM-Engine: A Realistic Environment for Building Geospatial Copilots
por: Singh, Simranjit, et al.
Publicado: (2024)
por: Singh, Simranjit, et al.
Publicado: (2024)
FinSearchComp: Towards a Realistic, Expert-Level Evaluation of Financial Search and Reasoning
por: Hu, Liang, et al.
Publicado: (2025)
por: Hu, Liang, et al.
Publicado: (2025)
CluCERT: Certifying LLM Robustness via Clustering-Guided Denoising Smoothing
por: Wang, Zixia, et al.
Publicado: (2025)
por: Wang, Zixia, et al.
Publicado: (2025)
Context-Aware Probabilistic Modeling with LLM for Multimodal Time Series Forecasting
por: Yao, Yueyang, et al.
Publicado: (2025)
por: Yao, Yueyang, et al.
Publicado: (2025)
Pushing Toward the Simplex Vertices: A Simple Remedy for Code Collapse in Smoothed Vector Quantization
por: Morita, Takashi
Publicado: (2025)
por: Morita, Takashi
Publicado: (2025)
Robustly Improving LLM Fairness in Realistic Settings via Interpretability
por: Karvonen, Adam, et al.
Publicado: (2025)
por: Karvonen, Adam, et al.
Publicado: (2025)
Towards One Model for Classical Dimensionality Reduction: A Probabilistic Perspective on UMAP and t-SNE
por: Ravuri, Aditya, et al.
Publicado: (2024)
por: Ravuri, Aditya, et al.
Publicado: (2024)
Lag-Llama: Towards Foundation Models for Probabilistic Time Series Forecasting
por: Rasul, Kashif, et al.
Publicado: (2023)
por: Rasul, Kashif, et al.
Publicado: (2023)
A Testable Certificate for Constant Collapse in Teacher-Guided VAEs
por: Zhang, Zegu, et al.
Publicado: (2026)
por: Zhang, Zegu, et al.
Publicado: (2026)
A Simplex Witness Certificate for Constant Collapse in Variational Autoencoders
por: Zhang, Zegu, et al.
Publicado: (2026)
por: Zhang, Zegu, et al.
Publicado: (2026)
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge
por: Fathullah, Yassir, et al.
Publicado: (2025)
por: Fathullah, Yassir, et al.
Publicado: (2025)
Certificate-Guided Pruning for Stochastic Lipschitz Optimization
por: Shihab, Ibne Farabi, et al.
Publicado: (2026)
por: Shihab, Ibne Farabi, et al.
Publicado: (2026)
Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques
por: Behera, Adarsh Prasad, et al.
Publicado: (2025)
por: Behera, Adarsh Prasad, et al.
Publicado: (2025)
Ejemplares similares
-
Automating Deception: Scalable Multi-Turn LLM Jailbreaks
por: Kumarappan, Adarsh, et al.
Publicado: (2025) -
SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
por: Robey, Alexander, et al.
Publicado: (2023) -
Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy
por: Kumarappan, Adarsh, et al.
Publicado: (2026) -
DevBench: A Realistic, Developer-Informed Benchmark for Code Generation Models
por: Kumarappan, Adarsh, et al.
Publicado: (2026) -
LeanAgent: Lifelong Learning for Formal Theorem Proving
por: Kumarappan, Adarsh, et al.
Publicado: (2024)