VITAL: A New Dataset for Benchmarking Pluralistic Alignment in Healthcare
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shetty, Anudeex, Beheshti, Amin, Dras, Mark, Naseem, Usman |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Pluralistic Alignment for Healthcare: A Role-Driven Framework
von: Zhong, Jiayou, et al.
Veröffentlicht: (2025)
von: Zhong, Jiayou, et al.
Veröffentlicht: (2025)
VISPA: Pluralistic Alignment via Automatic Value Selection and Activation
von: Zheng, Shenyan, et al.
Veröffentlicht: (2026)
von: Zheng, Shenyan, et al.
Veröffentlicht: (2026)
In Vino Veritas and Vulnerabilities: Examining LLM Safety via Drunk Language Inducement
von: Shetty, Anudeex, et al.
Veröffentlicht: (2026)
von: Shetty, Anudeex, et al.
Veröffentlicht: (2026)
Steering Towards Fairness: Mitigating Political Bias in LLMs
von: Nadeem, Afrozah, et al.
Veröffentlicht: (2025)
von: Nadeem, Afrozah, et al.
Veröffentlicht: (2025)
Fairness Evaluation and Inference Level Mitigation in LLMs
von: Nadeem, Afrozah, et al.
Veröffentlicht: (2025)
von: Nadeem, Afrozah, et al.
Veröffentlicht: (2025)
Framing Political Bias in Multilingual LLMs Across Pakistani Languages
von: Nadeem, Afrozah, et al.
Veröffentlicht: (2025)
von: Nadeem, Afrozah, et al.
Veröffentlicht: (2025)
Watermarks for Embeddings-as-a-Service Large Language Models
von: Shetty, Anudeex
Veröffentlicht: (2025)
von: Shetty, Anudeex
Veröffentlicht: (2025)
MixDPO: Modeling Preference Strength for Pluralistic Alignment
von: Imai, Saki, et al.
Veröffentlicht: (2026)
von: Imai, Saki, et al.
Veröffentlicht: (2026)
A Persona-Based Evaluation Framework for Pluralistic Alignment in Generative AI
von: Karagoz, Atahan
Veröffentlicht: (2026)
von: Karagoz, Atahan
Veröffentlicht: (2026)
ReflectDiffu:Reflect between Emotion-intent Contagion and Mimicry for Empathetic Response Generation via a RL-Diffusion Framework
von: Yuan, Jiahao, et al.
Veröffentlicht: (2024)
von: Yuan, Jiahao, et al.
Veröffentlicht: (2024)
Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy
von: Afzoon, Saleh, et al.
Veröffentlicht: (2025)
von: Afzoon, Saleh, et al.
Veröffentlicht: (2025)
Do Personality Traits Interfere? Geometric Limitations of Steering in Large Language Models
von: Bhandari, Pranav, et al.
Veröffentlicht: (2026)
von: Bhandari, Pranav, et al.
Veröffentlicht: (2026)
Lying Is Just a Phase: The Hidden Alignment Transition in Language Model Scaling
von: Amin, Adil
Veröffentlicht: (2026)
von: Amin, Adil
Veröffentlicht: (2026)
Pluralistic Alignment Over Time
von: Klassen, Toryn Q., et al.
Veröffentlicht: (2024)
von: Klassen, Toryn Q., et al.
Veröffentlicht: (2024)
VaxGuard: A Multi-Generator, Multi-Type, and Multi-Role Dataset for Detecting LLM-Generated Vaccine Misinformation
von: Ahmad, Syed Talal, et al.
Veröffentlicht: (2025)
von: Ahmad, Syed Talal, et al.
Veröffentlicht: (2025)
Pluralistic Behavior Suite: Stress-Testing Multi-Turn Adherence to Custom Behavioral Policies
von: Varshney, Prasoon, et al.
Veröffentlicht: (2025)
von: Varshney, Prasoon, et al.
Veröffentlicht: (2025)
Myanmar XNLI: Building a Dataset and Exploring Low-resource Approaches to Natural Language Inference with Myanmar
von: Htet, Aung Kyaw, et al.
Veröffentlicht: (2025)
von: Htet, Aung Kyaw, et al.
Veröffentlicht: (2025)
LLMs on a Budget? Say HOLA
von: Siddiqui, Zohaib Hasan, et al.
Veröffentlicht: (2025)
von: Siddiqui, Zohaib Hasan, et al.
Veröffentlicht: (2025)
WET: Overcoming Paraphrasing Vulnerabilities in Embeddings-as-a-Service with Linear Transformation Watermarks
von: Shetty, Anudeex, et al.
Veröffentlicht: (2024)
von: Shetty, Anudeex, et al.
Veröffentlicht: (2024)
MedAidDialog: A Multilingual Multi-Turn Medical Dialogue Dataset for Accessible Healthcare
von: Nigam, Shubham Kumar, et al.
Veröffentlicht: (2026)
von: Nigam, Shubham Kumar, et al.
Veröffentlicht: (2026)
Benchmarking the Medical Understanding and Reasoning of Large Language Models in Arabic Healthcare Tasks
von: AlDahoul, Nouar, et al.
Veröffentlicht: (2025)
von: AlDahoul, Nouar, et al.
Veröffentlicht: (2025)
Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models
von: Rastogi, Charvi, et al.
Veröffentlicht: (2025)
von: Rastogi, Charvi, et al.
Veröffentlicht: (2025)
EngTrace: A Symbolic Benchmark for Verifiable Process Supervision of Engineering Reasoning
von: Gull, Ayesha, et al.
Veröffentlicht: (2025)
von: Gull, Ayesha, et al.
Veröffentlicht: (2025)
AlignBench: Benchmarking Chinese Alignment of Large Language Models
von: Liu, Xiao, et al.
Veröffentlicht: (2023)
von: Liu, Xiao, et al.
Veröffentlicht: (2023)
Latent Principle Discovery for Language Model Self-Improvement
von: Ramji, Keshav, et al.
Veröffentlicht: (2025)
von: Ramji, Keshav, et al.
Veröffentlicht: (2025)
Over-Refusal and Representation Subspaces: A Mechanistic Analysis of Task-Conditioned Refusal in Aligned LLMs
von: Maskey, Utsav, et al.
Veröffentlicht: (2026)
von: Maskey, Utsav, et al.
Veröffentlicht: (2026)
Agentic Moderation: Multi-Agent Design for Safer Vision-Language Models
von: Ren, Juan, et al.
Veröffentlicht: (2025)
von: Ren, Juan, et al.
Veröffentlicht: (2025)
Sequences of Logits Reveal the Low Rank Structure of Language Models
von: Golowich, Noah, et al.
Veröffentlicht: (2025)
von: Golowich, Noah, et al.
Veröffentlicht: (2025)
A Synthetic Dataset for Personal Attribute Inference
von: Yukhymenko, Hanna, et al.
Veröffentlicht: (2024)
von: Yukhymenko, Hanna, et al.
Veröffentlicht: (2024)
Time-IMM: A Dataset and Benchmark for Irregular Multimodal Multivariate Time Series
von: Chang, Ching, et al.
Veröffentlicht: (2025)
von: Chang, Ching, et al.
Veröffentlicht: (2025)
StyleRec: A Benchmark Dataset for Prompt Recovery in Writing Style Transformation
von: Liu, Shenyang, et al.
Veröffentlicht: (2025)
von: Liu, Shenyang, et al.
Veröffentlicht: (2025)
Improving Model Evaluation using SMART Filtering of Benchmark Datasets
von: Gupta, Vipul, et al.
Veröffentlicht: (2024)
von: Gupta, Vipul, et al.
Veröffentlicht: (2024)
Recovered in Translation: Efficient Pipeline for Automated Translation of Benchmarks and Datasets
von: Yukhymenko, Hanna, et al.
Veröffentlicht: (2026)
von: Yukhymenko, Hanna, et al.
Veröffentlicht: (2026)
WARDEN: Multi-Directional Backdoor Watermarks for Embedding-as-a-Service Copyright Protection
von: Shetty, Anudeex, et al.
Veröffentlicht: (2024)
von: Shetty, Anudeex, et al.
Veröffentlicht: (2024)
Benchmarking the Capabilities of Large Language Models in Transportation System Engineering: Accuracy, Consistency, and Reasoning Behaviors
von: Syed, Usman, et al.
Veröffentlicht: (2024)
von: Syed, Usman, et al.
Veröffentlicht: (2024)
SHIELD: Classifier-Guided Prompting for Robust and Safer LVLMs
von: Ren, Juan, et al.
Veröffentlicht: (2025)
von: Ren, Juan, et al.
Veröffentlicht: (2025)
Should LLM Safety Be More Than Refusing Harmful Instructions?
von: Maskey, Utsav, et al.
Veröffentlicht: (2025)
von: Maskey, Utsav, et al.
Veröffentlicht: (2025)
Seeing the Threat: Vulnerabilities in Vision-Language Models to Adversarial Attack
von: Ren, Juan, et al.
Veröffentlicht: (2025)
von: Ren, Juan, et al.
Veröffentlicht: (2025)
Steering Over-refusals Towards Safety in Retrieval Augmented Generation
von: Maskey, Utsav, et al.
Veröffentlicht: (2025)
von: Maskey, Utsav, et al.
Veröffentlicht: (2025)
PersoBench: Benchmarking Personalized Response Generation in Large Language Models
von: Afzoon, Saleh, et al.
Veröffentlicht: (2024)
von: Afzoon, Saleh, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Pluralistic Alignment for Healthcare: A Role-Driven Framework
von: Zhong, Jiayou, et al.
Veröffentlicht: (2025) -
VISPA: Pluralistic Alignment via Automatic Value Selection and Activation
von: Zheng, Shenyan, et al.
Veröffentlicht: (2026) -
In Vino Veritas and Vulnerabilities: Examining LLM Safety via Drunk Language Inducement
von: Shetty, Anudeex, et al.
Veröffentlicht: (2026) -
Steering Towards Fairness: Mitigating Political Bias in LLMs
von: Nadeem, Afrozah, et al.
Veröffentlicht: (2025) -
Fairness Evaluation and Inference Level Mitigation in LLMs
von: Nadeem, Afrozah, et al.
Veröffentlicht: (2025)