Saved in:
| Main Authors: | Kim, Hazel, Lamb, Tom A., Bibi, Adel, Torr, Philip, Gal, Yarin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2412.10246 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MIP against Agent: Malicious Image Patches Hijacking Multimodal OS Agents
by: Aichberger, Lukas, et al.
Published: (2025)
by: Aichberger, Lukas, et al.
Published: (2025)
Universal In-Context Approximation By Prompting Fully Recurrent Models
by: Petrov, Aleksandar, et al.
Published: (2024)
by: Petrov, Aleksandar, et al.
Published: (2024)
When Do Prompting and Prefix-Tuning Work? A Theory of Capabilities and Limitations
by: Petrov, Aleksandar, et al.
Published: (2023)
by: Petrov, Aleksandar, et al.
Published: (2023)
Prompting a Pretrained Transformer Can Be a Universal Approximator
by: Petrov, Aleksandar, et al.
Published: (2024)
by: Petrov, Aleksandar, et al.
Published: (2024)
Fine-tuning can cripple your foundation model; preserving features may be the solution
by: Mukhoti, Jishnu, et al.
Published: (2023)
by: Mukhoti, Jishnu, et al.
Published: (2023)
Beyond Linear Probes: Dynamic Safety Monitoring for Language Models
by: Oldfield, James, et al.
Published: (2025)
by: Oldfield, James, et al.
Published: (2025)
Bi-Factorial Preference Optimization: Balancing Safety-Helpfulness in Language Models
by: Zhang, Wenxuan, et al.
Published: (2024)
by: Zhang, Wenxuan, et al.
Published: (2024)
Towards Certification of Uncertainty Calibration under Adversarial Attacks
by: Emde, Cornelius, et al.
Published: (2024)
by: Emde, Cornelius, et al.
Published: (2024)
Improving Semantic Uncertainty Quantification in Language Model Question-Answering via Token-Level Temperature Scaling
by: Lamb, Tom A., et al.
Published: (2026)
by: Lamb, Tom A., et al.
Published: (2026)
Rethinking Safety in LLM Fine-tuning: An Optimization Perspective
by: Kim, Minseon, et al.
Published: (2025)
by: Kim, Minseon, et al.
Published: (2025)
How Ambiguous Are the Rationales for Natural Language Reasoning? A Simple Approach to Handling Rationale Uncertainty
by: Kim, Hazel H.
Published: (2024)
by: Kim, Hazel H.
Published: (2024)
In-Context Learning Learns Label Relationships but Is Not Conventional Learning
by: Kossen, Jannik, et al.
Published: (2023)
by: Kossen, Jannik, et al.
Published: (2023)
Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs
by: Kossen, Jannik, et al.
Published: (2024)
by: Kossen, Jannik, et al.
Published: (2024)
Do as I do (Safely): Mitigating Task-Specific Fine-tuning Risks in Large Language Models
by: Eiras, Francisco, et al.
Published: (2024)
by: Eiras, Francisco, et al.
Published: (2024)
Segment, Select, Correct: A Framework for Weakly-Supervised Referring Segmentation
by: Eiras, Francisco, et al.
Published: (2023)
by: Eiras, Francisco, et al.
Published: (2023)
Efficient Lifelong Model Evaluation in an Era of Rapid Progress
by: Prabhu, Ameya, et al.
Published: (2024)
by: Prabhu, Ameya, et al.
Published: (2024)
FORCE: Transferable Visual Jailbreaking Attacks via Feature Over-Reliance CorrEction
by: Lin, Runqi, et al.
Published: (2025)
by: Lin, Runqi, et al.
Published: (2025)
Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning
by: Melo, Luckeciano C., et al.
Published: (2025)
by: Melo, Luckeciano C., et al.
Published: (2025)
Continual Learning on a Diet: Learning from Sparsely Labeled Streams Under Constrained Computation
by: Zhang, Wenxuan, et al.
Published: (2024)
by: Zhang, Wenxuan, et al.
Published: (2024)
Efficient Error Certification for Physics-Informed Neural Networks
by: Eiras, Francisco, et al.
Published: (2023)
by: Eiras, Francisco, et al.
Published: (2023)
Safer by Diffusion, Broken by Context: Diffusion LLM's Safety Blessing and Its Failure Mode
by: He, Zeyuan, et al.
Published: (2026)
by: He, Zeyuan, et al.
Published: (2026)
OMNI-LEAK: Orchestrator Multi-Agent Network Induced Data Leakage
by: Naik, Akshat, et al.
Published: (2026)
by: Naik, Akshat, et al.
Published: (2026)
Towards Interpretable Deep Local Learning with Successive Gradient Reconciliation
by: Yang, Yibo, et al.
Published: (2024)
by: Yang, Yibo, et al.
Published: (2024)
Estimating the Hallucination Rate of Generative AI
by: Jesson, Andrew, et al.
Published: (2024)
by: Jesson, Andrew, et al.
Published: (2024)
ToolTweak: An Attack on Tool Selection in LLM-based Agents
by: Sneh, Jonathan, et al.
Published: (2025)
by: Sneh, Jonathan, et al.
Published: (2025)
From Categories to Classifiers: Name-Only Continual Learning by Exploring the Web
by: Prabhu, Ameya, et al.
Published: (2023)
by: Prabhu, Ameya, et al.
Published: (2023)
TreeCut: A Synthetic Unanswerable Math Word Problem Dataset for LLM Hallucination Evaluation
by: Ouyang, Jialin
Published: (2025)
by: Ouyang, Jialin
Published: (2025)
Simple Baselines are Competitive with Code Evolution
by: Gideoni, Yonatan, et al.
Published: (2026)
by: Gideoni, Yonatan, et al.
Published: (2026)
The Benefits and Risks of Transductive Approaches for AI Fairness
by: Razzak, Muhammed, et al.
Published: (2024)
by: Razzak, Muhammed, et al.
Published: (2024)
Focus On This, Not That! Steering LLMs with Adaptive Feature Specification
by: Lamb, Tom A., et al.
Published: (2024)
by: Lamb, Tom A., et al.
Published: (2024)
Shh, don't say that! Domain Certification in LLMs
by: Emde, Cornelius, et al.
Published: (2025)
by: Emde, Cornelius, et al.
Published: (2025)
Mixture of Experts Made Intrinsically Interpretable
by: Yang, Xingyi, et al.
Published: (2025)
by: Yang, Xingyi, et al.
Published: (2025)
On Pretraining Data Diversity for Self-Supervised Learning
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
Do Multilingual LLMs Think In English?
by: Schut, Lisa, et al.
Published: (2025)
by: Schut, Lisa, et al.
Published: (2025)
Temporal-Difference Variational Continual Learning
by: Melo, Luckeciano C., et al.
Published: (2024)
by: Melo, Luckeciano C., et al.
Published: (2024)
Scaling Up Active Testing to Large Language Models
by: Berrada, Gabrielle, et al.
Published: (2025)
by: Berrada, Gabrielle, et al.
Published: (2025)
Model Merging and Safety Alignment: One Bad Model Spoils the Bunch
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
Label Delay in Online Continual Learning
by: Csaba, Botos, et al.
Published: (2023)
by: Csaba, Botos, et al.
Published: (2023)
Richer Bayesian Last Layers with Subsampled NTK Features
by: Calvo-Ordoñez, Sergio, et al.
Published: (2026)
by: Calvo-Ordoñez, Sergio, et al.
Published: (2026)
Similar Items
-
MIP against Agent: Malicious Image Patches Hijacking Multimodal OS Agents
by: Aichberger, Lukas, et al.
Published: (2025) -
Universal In-Context Approximation By Prompting Fully Recurrent Models
by: Petrov, Aleksandar, et al.
Published: (2024) -
When Do Prompting and Prefix-Tuning Work? A Theory of Capabilities and Limitations
by: Petrov, Aleksandar, et al.
Published: (2023) -
Prompting a Pretrained Transformer Can Be a Universal Approximator
by: Petrov, Aleksandar, et al.
Published: (2024) -
Fine-tuning can cripple your foundation model; preserving features may be the solution
by: Mukhoti, Jishnu, et al.
Published: (2023)