Improving Black-box Robustness with In-Context Rewriting
Fuente:
arXiv
Guardado en:
| Autores principales: | O'Brien, Kyle, Ng, Nathan, Puri, Isha, Mendez, Jorge, Palangi, Hamid, Kim, Yoon, Ghassemi, Marzyeh, Hartvigsen, Thomas |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Reaching Beyond the Mode: RL for Distributional Reasoning in Language Models
por: Puri, Isha, et al.
Publicado: (2026)
por: Puri, Isha, et al.
Publicado: (2026)
Measuring Stochastic Data Complexity with Boltzmann Influence Functions
por: Ng, Nathan, et al.
Publicado: (2024)
por: Ng, Nathan, et al.
Publicado: (2024)
KScope: A Framework for Characterizing the Knowledge Status of Language Models
por: Xiao, Yuxin, et al.
Publicado: (2025)
por: Xiao, Yuxin, et al.
Publicado: (2025)
Can AI Relate: Testing Large Language Model Response for Mental Health Support
por: Gabriel, Saadia, et al.
Publicado: (2024)
por: Gabriel, Saadia, et al.
Publicado: (2024)
BendVLM: Test-Time Debiasing of Vision-Language Embeddings
por: Gerych, Walter, et al.
Publicado: (2024)
por: Gerych, Walter, et al.
Publicado: (2024)
Views Can Be Deceiving: Improved SSL Through Feature Space Augmentation
por: Hamidieh, Kimia, et al.
Publicado: (2024)
por: Hamidieh, Kimia, et al.
Publicado: (2024)
Data Debiasing with Datamodels (D3M): Improving Subgroup Robustness via Data Selection
por: Jain, Saachi, et al.
Publicado: (2024)
por: Jain, Saachi, et al.
Publicado: (2024)
Identifying Implicit Social Biases in Vision-Language Models
por: Hamidieh, Kimia, et al.
Publicado: (2024)
por: Hamidieh, Kimia, et al.
Publicado: (2024)
Robustness Beyond Known Groups with Low-rank Adaptation
por: Gourabathina, Abinitha, et al.
Publicado: (2026)
por: Gourabathina, Abinitha, et al.
Publicado: (2026)
MaskMedPaint: Masked Medical Image Inpainting with Diffusion Models for Mitigation of Spurious Correlations
por: Jin, Qixuan, et al.
Publicado: (2024)
por: Jin, Qixuan, et al.
Publicado: (2024)
ScholarPeer: A Context-Aware Multi-Agent Framework for Automated Peer Review
por: Goyal, Palash, et al.
Publicado: (2026)
por: Goyal, Palash, et al.
Publicado: (2026)
Improving Mutual Information Estimation with Annealed and Energy-Based Bounds
por: Brekelmans, Rob, et al.
Publicado: (2023)
por: Brekelmans, Rob, et al.
Publicado: (2023)
An Investigation of Memorization Risk in Healthcare Foundation Models
por: Tonekaboni, Sana, et al.
Publicado: (2025)
por: Tonekaboni, Sana, et al.
Publicado: (2025)
Aggregation Hides Out-of-Distribution Generalization Failures from Spurious Correlations
por: Salaudeen, Olawale, et al.
Publicado: (2025)
por: Salaudeen, Olawale, et al.
Publicado: (2025)
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
por: Chan, Yik Siu, et al.
Publicado: (2025)
por: Chan, Yik Siu, et al.
Publicado: (2025)
In the Name of Fairness: Assessing the Bias in Clinical Record De-identification
por: Xiao, Yuxin, et al.
Publicado: (2023)
por: Xiao, Yuxin, et al.
Publicado: (2023)
What's in a Query: Polarity-Aware Distribution-Based Fair Ranking
por: Balagopalan, Aparna, et al.
Publicado: (2025)
por: Balagopalan, Aparna, et al.
Publicado: (2025)
Exploring GPT-4 for Robotic Agent Strategy with Real-Time State Feedback and a Reactive Behaviour Framework
por: O'Brien, Thomas, et al.
Publicado: (2025)
por: O'Brien, Thomas, et al.
Publicado: (2025)
SFTMix: Elevating Language Model Instruction Tuning with Mixup Recipe
por: Xiao, Yuxin, et al.
Publicado: (2024)
por: Xiao, Yuxin, et al.
Publicado: (2024)
Sparse identification of nonlinear dynamics in the presence of library and system uncertainty
por: O'Brien, Andrew
Publicado: (2024)
por: O'Brien, Andrew
Publicado: (2024)
Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
por: Damani, Mehul, et al.
Publicado: (2025)
por: Damani, Mehul, et al.
Publicado: (2025)
When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment
por: Xiao, Yuxin, et al.
Publicado: (2025)
por: Xiao, Yuxin, et al.
Publicado: (2025)
OptiGradTrust: Byzantine-Robust Federated Learning with Multi-Feature Gradient Analysis and Reinforcement Learning-Based Trust Weighting
por: Karami, Mohammad, et al.
Publicado: (2025)
por: Karami, Mohammad, et al.
Publicado: (2025)
MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering
por: Hao, Yuexing, et al.
Publicado: (2025)
por: Hao, Yuexing, et al.
Publicado: (2025)
If i die in a combat zone : box me up and ship me home / Tim O'Brien
por: O'Brien, Tim
Publicado: (1989)
por: O'Brien, Tim
Publicado: (1989)
Composable Interventions for Language Models
por: Kolbeinsson, Arinbjorn, et al.
Publicado: (2024)
por: Kolbeinsson, Arinbjorn, et al.
Publicado: (2024)
EvadeDroid: A Practical Evasion Attack on Machine Learning for Black-box Android Malware Detection
por: Bostani, Hamid, et al.
Publicado: (2021)
por: Bostani, Hamid, et al.
Publicado: (2021)
FedMedICL: Towards Holistic Evaluation of Distribution Shifts in Federated Medical Imaging
por: Alhamoud, Kumail, et al.
Publicado: (2024)
por: Alhamoud, Kumail, et al.
Publicado: (2024)
Event-Based Contrastive Learning for Medical Time Series
por: Jeong, Hyewon, et al.
Publicado: (2023)
por: Jeong, Hyewon, et al.
Publicado: (2023)
Improving and Accelerating Offline RL in Large Discrete Action Spaces with Structured Policy Initialization
por: Landers, Matthew, et al.
Publicado: (2026)
por: Landers, Matthew, et al.
Publicado: (2026)
Rollout Roulette: A Probabilistic Inference Approach to Inference-Time Scaling of LLMs using Particle-Based Monte Carlo Methods
por: Puri, Isha, et al.
Publicado: (2025)
por: Puri, Isha, et al.
Publicado: (2025)
A Framework for Improving the Reliability of Black-box Variational Inference
por: Welandawe, Manushi, et al.
Publicado: (2022)
por: Welandawe, Manushi, et al.
Publicado: (2022)
ModelCitizens: Representing Community Voices in Online Safety
por: Suvarna, Ashima, et al.
Publicado: (2025)
por: Suvarna, Ashima, et al.
Publicado: (2025)
Distributed Black-box Attack: Do Not Overestimate Black-box Attacks
por: Wu, Han, et al.
Publicado: (2022)
por: Wu, Han, et al.
Publicado: (2022)
Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment
por: Tice, Cameron, et al.
Publicado: (2026)
por: Tice, Cameron, et al.
Publicado: (2026)
CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs
por: Guo, Han, et al.
Publicado: (2026)
por: Guo, Han, et al.
Publicado: (2026)
Learning to Correct for QA Reasoning with Black-box LLMs
por: Kim, Jaehyung, et al.
Publicado: (2024)
por: Kim, Jaehyung, et al.
Publicado: (2024)
LEMoN: Label Error Detection using Multimodal Neighbors
por: Zhang, Haoran, et al.
Publicado: (2024)
por: Zhang, Haoran, et al.
Publicado: (2024)
Improving Diversity in Black-box Few-shot Knowledge Distillation
por: Vo, Tri-Nhan, et al.
Publicado: (2026)
por: Vo, Tri-Nhan, et al.
Publicado: (2026)
A System for Accurate Tracking and Video Recordings of Rodent Eye Movements using Convolutional Neural Networks for Biomedical Image Segmentation
por: Puri, Isha, et al.
Publicado: (2025)
por: Puri, Isha, et al.
Publicado: (2025)
Ejemplares similares
-
Reaching Beyond the Mode: RL for Distributional Reasoning in Language Models
por: Puri, Isha, et al.
Publicado: (2026) -
Measuring Stochastic Data Complexity with Boltzmann Influence Functions
por: Ng, Nathan, et al.
Publicado: (2024) -
KScope: A Framework for Characterizing the Knowledge Status of Language Models
por: Xiao, Yuxin, et al.
Publicado: (2025) -
Can AI Relate: Testing Large Language Model Response for Mental Health Support
por: Gabriel, Saadia, et al.
Publicado: (2024) -
BendVLM: Test-Time Debiasing of Vision-Language Embeddings
por: Gerych, Walter, et al.
Publicado: (2024)