Hedging and Non-Affirmation: Quantifying LLM Alignment on Questions of Human Rights
Fuente:
arXiv
Salvato in:
| Autori principali: | Javed, Rafiya, Parent, Cassandra, Kay, Jackie, Yanni, David, Zaini, Abdullah, Sheikh, Anushe, Rauh, Maribeth, Gerych, Walter, Comanescu, Ramona, Gabriel, Iason, Ghassemi, Marzyeh, Weidinger, Laura |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MaskMedPaint: Masked Medical Image Inpainting with Diffusion Models for Mitigation of Spurious Correlations
di: Jin, Qixuan, et al.
Pubblicazione: (2024)
di: Jin, Qixuan, et al.
Pubblicazione: (2024)
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making
di: Gourabathina, Abinitha, et al.
Pubblicazione: (2025)
di: Gourabathina, Abinitha, et al.
Pubblicazione: (2025)
When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment
di: Xiao, Yuxin, et al.
Pubblicazione: (2025)
di: Xiao, Yuxin, et al.
Pubblicazione: (2025)
An Investigation of Memorization Risk in Healthcare Foundation Models
di: Tonekaboni, Sana, et al.
Pubblicazione: (2025)
di: Tonekaboni, Sana, et al.
Pubblicazione: (2025)
Complementing Self-Consistency with Cross-Model Disagreement for Uncertainty Quantification
di: Hamidieh, Kimia, et al.
Pubblicazione: (2026)
di: Hamidieh, Kimia, et al.
Pubblicazione: (2026)
Identifying Implicit Social Biases in Vision-Language Models
di: Hamidieh, Kimia, et al.
Pubblicazione: (2024)
di: Hamidieh, Kimia, et al.
Pubblicazione: (2024)
BendVLM: Test-Time Debiasing of Vision-Language Embeddings
di: Gerych, Walter, et al.
Pubblicazione: (2024)
di: Gerych, Walter, et al.
Pubblicazione: (2024)
Measuring Stochastic Data Complexity with Boltzmann Influence Functions
di: Ng, Nathan, et al.
Pubblicazione: (2024)
di: Ng, Nathan, et al.
Pubblicazione: (2024)
SpaceVLM: Sub-Space Modeling of Negation in Vision-Language Models
di: Ranjbar, Sepehr Kazemi, et al.
Pubblicazione: (2025)
di: Ranjbar, Sepehr Kazemi, et al.
Pubblicazione: (2025)
Views Can Be Deceiving: Improved SSL Through Feature Space Augmentation
di: Hamidieh, Kimia, et al.
Pubblicazione: (2024)
di: Hamidieh, Kimia, et al.
Pubblicazione: (2024)
Linear matrix equations with parameters forming a commuting set of diagonalizable matrices
di: Comănescu, Dan
Pubblicazione: (2024)
di: Comănescu, Dan
Pubblicazione: (2024)
Existence and Stability of 3-Cycles in Quadratic Maps
di: Comănescu, Dan
Pubblicazione: (2026)
di: Comănescu, Dan
Pubblicazione: (2026)
The steady states of antitone electric systems
di: Dan Comănescu
Pubblicazione: (2024)
di: Dan Comănescu
Pubblicazione: (2024)
Suicide Affirmation and the Positive Right to Die
di: Ruth Pearce
Pubblicazione: (2025)
di: Ruth Pearce
Pubblicazione: (2025)
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
di: Chan, Yik Siu, et al.
Pubblicazione: (2025)
di: Chan, Yik Siu, et al.
Pubblicazione: (2025)
In the Name of Fairness: Assessing the Bias in Clinical Record De-identification
di: Xiao, Yuxin, et al.
Pubblicazione: (2023)
di: Xiao, Yuxin, et al.
Pubblicazione: (2023)
Reading and Computers--How Teachers Can Make Them Work Together.
di: Henney, Maribeth
Pubblicazione: (1984)
di: Henney, Maribeth
Pubblicazione: (1984)
Aggregation Hides Out-of-Distribution Generalization Failures from Spurious Correlations
di: Salaudeen, Olawale, et al.
Pubblicazione: (2025)
di: Salaudeen, Olawale, et al.
Pubblicazione: (2025)
Can AI Relate: Testing Large Language Model Response for Mental Health Support
di: Gabriel, Saadia, et al.
Pubblicazione: (2024)
di: Gabriel, Saadia, et al.
Pubblicazione: (2024)
Robustness Beyond Known Groups with Low-rank Adaptation
di: Gourabathina, Abinitha, et al.
Pubblicazione: (2026)
di: Gourabathina, Abinitha, et al.
Pubblicazione: (2026)
What's in a Query: Polarity-Aware Distribution-Based Fair Ranking
di: Balagopalan, Aparna, et al.
Pubblicazione: (2025)
di: Balagopalan, Aparna, et al.
Pubblicazione: (2025)
SFTMix: Elevating Language Model Instruction Tuning with Mixup Recipe
di: Xiao, Yuxin, et al.
Pubblicazione: (2024)
di: Xiao, Yuxin, et al.
Pubblicazione: (2024)
Prediction of ionic liquids toxicity using machine learning models for application to gas hydrate
di: Nurul Hannah Abdullah, et al.
Pubblicazione: (2024)
di: Nurul Hannah Abdullah, et al.
Pubblicazione: (2024)
MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering
di: Hao, Yuexing, et al.
Pubblicazione: (2025)
di: Hao, Yuexing, et al.
Pubblicazione: (2025)
Arming Students against Bad Information
di: Smith, Maribeth D.
Pubblicazione: (2017)
di: Smith, Maribeth D.
Pubblicazione: (2017)
Symptomatic from Clostridioides difficile or symptomatic from inflammatory bowel disease: Highlighting diagnostic challenges
di: Maribeth R. Nicholson
Pubblicazione: (2025)
di: Maribeth R. Nicholson
Pubblicazione: (2025)
Could Micro-Expressions be Quantified? Electromyography Gives Affirmative Evidence
di: Li, Jingting, et al.
Pubblicazione: (2024)
di: Li, Jingting, et al.
Pubblicazione: (2024)
Spinosad‐based insecticide is an effective control method for the potato soil pests Agriotes sp. and Melolontha sp.
di: Gabriel Comanescu, et al.
Pubblicazione: (2026)
di: Gabriel Comanescu, et al.
Pubblicazione: (2026)
Recycling of polyester bottle waste to generate functional acid dye in situ on wool
di: Ankit Singh, et al.
Pubblicazione: (2024)
di: Ankit Singh, et al.
Pubblicazione: (2024)
Nonaffiliated Users in Academic Libraries: Using W.D. Ross's Ethical Pluralism to Make Sense of the Tough Questions
di: Lenker, Mark, et al.
Pubblicazione: (2010)
di: Lenker, Mark, et al.
Pubblicazione: (2010)
Learning the Wrong Lessons: Syntactic-Domain Spurious Correlations in Language Models
di: Shaib, Chantal, et al.
Pubblicazione: (2025)
di: Shaib, Chantal, et al.
Pubblicazione: (2025)
Reaching Beyond the Mode: RL for Distributional Reasoning in Language Models
di: Puri, Isha, et al.
Pubblicazione: (2026)
di: Puri, Isha, et al.
Pubblicazione: (2026)
MisinfoEval: Generative AI in the Era of "Alternative Facts"
di: Gabriel, Saadia, et al.
Pubblicazione: (2024)
di: Gabriel, Saadia, et al.
Pubblicazione: (2024)
KScope: A Framework for Characterizing the Knowledge Status of Language Models
di: Xiao, Yuxin, et al.
Pubblicazione: (2025)
di: Xiao, Yuxin, et al.
Pubblicazione: (2025)
Data Debiasing with Datamodels (D3M): Improving Subgroup Robustness via Data Selection
di: Jain, Saachi, et al.
Pubblicazione: (2024)
di: Jain, Saachi, et al.
Pubblicazione: (2024)
Disparities In Negation Understanding Across Languages In Vision-Language Models
di: Moraitaki, Charikleia, et al.
Pubblicazione: (2026)
di: Moraitaki, Charikleia, et al.
Pubblicazione: (2026)
MOSAIC: Modeling Social AI for Content Dissemination and Regulation in Multi-Agent Simulations
di: Liu, Genglin, et al.
Pubblicazione: (2025)
di: Liu, Genglin, et al.
Pubblicazione: (2025)
Evaluating the Quality of the Quantified Uncertainty for (Re)Calibration of Data-Driven Regression Models
di: Wibbeke, Jelke, et al.
Pubblicazione: (2025)
di: Wibbeke, Jelke, et al.
Pubblicazione: (2025)
Affirmative Action vs. Affirmative Information
di: Reich, Claire Lazar
Pubblicazione: (2021)
di: Reich, Claire Lazar
Pubblicazione: (2021)
Die besondere Atmosphäre
di: Rauh, Andreas
Pubblicazione: (2024)
di: Rauh, Andreas
Pubblicazione: (2024)
Documenti analoghi
-
MaskMedPaint: Masked Medical Image Inpainting with Diffusion Models for Mitigation of Spurious Correlations
di: Jin, Qixuan, et al.
Pubblicazione: (2024) -
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making
di: Gourabathina, Abinitha, et al.
Pubblicazione: (2025) -
When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment
di: Xiao, Yuxin, et al.
Pubblicazione: (2025) -
An Investigation of Memorization Risk in Healthcare Foundation Models
di: Tonekaboni, Sana, et al.
Pubblicazione: (2025) -
Complementing Self-Consistency with Cross-Model Disagreement for Uncertainty Quantification
di: Hamidieh, Kimia, et al.
Pubblicazione: (2026)