Learning the Wrong Lessons: Syntactic-Domain Spurious Correlations in Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shaib, Chantal, Suriyakumar, Vinith M., Sagun, Levent, Wallace, Byron C., Ghassemi, Marzyeh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Detection and Measurement of Syntactic Templates in Generated Text
von: Shaib, Chantal, et al.
Veröffentlicht: (2024)
von: Shaib, Chantal, et al.
Veröffentlicht: (2024)
When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment
von: Xiao, Yuxin, et al.
Veröffentlicht: (2025)
von: Xiao, Yuxin, et al.
Veröffentlicht: (2025)
Who Taught You That? Tracing Teachers in Model Distillation
von: Wadhwa, Somin, et al.
Veröffentlicht: (2025)
von: Wadhwa, Somin, et al.
Veröffentlicht: (2025)
Measuring AI "Slop" in Text
von: Shaib, Chantal, et al.
Veröffentlicht: (2025)
von: Shaib, Chantal, et al.
Veröffentlicht: (2025)
How Much Annotation is Needed to Compare Summarization Models?
von: Shaib, Chantal, et al.
Veröffentlicht: (2024)
von: Shaib, Chantal, et al.
Veröffentlicht: (2024)
UCD: Unlearning in LLMs via Contrastive Decoding
von: Suriyakumar, Vinith M., et al.
Veröffentlicht: (2025)
von: Suriyakumar, Vinith M., et al.
Veröffentlicht: (2025)
MaskMedPaint: Masked Medical Image Inpainting with Diffusion Models for Mitigation of Spurious Correlations
von: Jin, Qixuan, et al.
Veröffentlicht: (2024)
von: Jin, Qixuan, et al.
Veröffentlicht: (2024)
Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence
von: Mo, Kaijie, et al.
Veröffentlicht: (2026)
von: Mo, Kaijie, et al.
Veröffentlicht: (2026)
Standardizing the Measurement of Text Diversity: A Tool and a Comparative Analysis of Scores
von: Shaib, Chantal, et al.
Veröffentlicht: (2024)
von: Shaib, Chantal, et al.
Veröffentlicht: (2024)
Measuring Lexical Diversity of Synthetic Data Generated through Fine-Grained Persona Prompting
von: Kambhatla, Gauri, et al.
Veröffentlicht: (2025)
von: Kambhatla, Gauri, et al.
Veröffentlicht: (2025)
Evaluation without Generation: Non-Generative Assessment of Harmful Model Specialization with Applications to CSAM
von: Suriyakumar, Vinith M., et al.
Veröffentlicht: (2026)
von: Suriyakumar, Vinith M., et al.
Veröffentlicht: (2026)
Can AI Relate: Testing Large Language Model Response for Mental Health Support
von: Gabriel, Saadia, et al.
Veröffentlicht: (2024)
von: Gabriel, Saadia, et al.
Veröffentlicht: (2024)
Reassessing the Validity of Spurious Correlations Benchmarks
von: Bell, Samuel J., et al.
Veröffentlicht: (2024)
von: Bell, Samuel J., et al.
Veröffentlicht: (2024)
Disparities In Negation Understanding Across Languages In Vision-Language Models
von: Moraitaki, Charikleia, et al.
Veröffentlicht: (2026)
von: Moraitaki, Charikleia, et al.
Veröffentlicht: (2026)
Identifying Implicit Social Biases in Vision-Language Models
von: Hamidieh, Kimia, et al.
Veröffentlicht: (2024)
von: Hamidieh, Kimia, et al.
Veröffentlicht: (2024)
SFTMix: Elevating Language Model Instruction Tuning with Mixup Recipe
von: Xiao, Yuxin, et al.
Veröffentlicht: (2024)
von: Xiao, Yuxin, et al.
Veröffentlicht: (2024)
Vision-Language Models Do Not Understand Negation
von: Alhamoud, Kumail, et al.
Veröffentlicht: (2025)
von: Alhamoud, Kumail, et al.
Veröffentlicht: (2025)
Revisiting Relation Extraction in the era of Large Language Models
von: Wadhwa, Somin, et al.
Veröffentlicht: (2023)
von: Wadhwa, Somin, et al.
Veröffentlicht: (2023)
KScope: A Framework for Characterizing the Knowledge Status of Language Models
von: Xiao, Yuxin, et al.
Veröffentlicht: (2025)
von: Xiao, Yuxin, et al.
Veröffentlicht: (2025)
Reaching Beyond the Mode: RL for Distributional Reasoning in Language Models
von: Puri, Isha, et al.
Veröffentlicht: (2026)
von: Puri, Isha, et al.
Veröffentlicht: (2026)
In-context Learning in Presence of Spurious Correlations
von: Harutyunyan, Hrayr, et al.
Veröffentlicht: (2024)
von: Harutyunyan, Hrayr, et al.
Veröffentlicht: (2024)
Assessing Robustness to Spurious Correlations in Post-Training Language Models
von: Shuieh, Julia, et al.
Veröffentlicht: (2025)
von: Shuieh, Julia, et al.
Veröffentlicht: (2025)
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
von: Chan, Yik Siu, et al.
Veröffentlicht: (2025)
von: Chan, Yik Siu, et al.
Veröffentlicht: (2025)
In the Name of Fairness: Assessing the Bias in Clinical Record De-identification
von: Xiao, Yuxin, et al.
Veröffentlicht: (2023)
von: Xiao, Yuxin, et al.
Veröffentlicht: (2023)
Aggregation Hides Out-of-Distribution Generalization Failures from Spurious Correlations
von: Salaudeen, Olawale, et al.
Veröffentlicht: (2025)
von: Salaudeen, Olawale, et al.
Veröffentlicht: (2025)
Explore Spurious Correlations at the Concept Level in Language Models for Text Classification
von: Zhou, Yuhang, et al.
Veröffentlicht: (2023)
von: Zhou, Yuhang, et al.
Veröffentlicht: (2023)
Beg to Differ: Understanding Reasoning-Answer Misalignment Across Languages
von: Ovalle, Anaelia, et al.
Veröffentlicht: (2025)
von: Ovalle, Anaelia, et al.
Veröffentlicht: (2025)
Spurious Correlations and Beyond: Understanding and Mitigating Shortcut Learning in SDOH Extraction with Large Language Models
von: Sakib, Fardin Ahsan, et al.
Veröffentlicht: (2025)
von: Sakib, Fardin Ahsan, et al.
Veröffentlicht: (2025)
MisinfoEval: Generative AI in the Era of "Alternative Facts"
von: Gabriel, Saadia, et al.
Veröffentlicht: (2024)
von: Gabriel, Saadia, et al.
Veröffentlicht: (2024)
Learning from Natural Language Explanations for Generalizable Entity Matching
von: Wadhwa, Somin, et al.
Veröffentlicht: (2024)
von: Wadhwa, Somin, et al.
Veröffentlicht: (2024)
The Root Shapes the Fruit: On the Persistence of Gender-Exclusive Harms in Aligned Language Models
von: Ovalle, Anaelia, et al.
Veröffentlicht: (2024)
von: Ovalle, Anaelia, et al.
Veröffentlicht: (2024)
MOSAIC: Modeling Social AI for Content Dissemination and Regulation in Multi-Agent Simulations
von: Liu, Genglin, et al.
Veröffentlicht: (2025)
von: Liu, Genglin, et al.
Veröffentlicht: (2025)
Measuring Spurious Correlation in Classification: 'Clever Hans' in Translationese
von: Borah, Angana, et al.
Veröffentlicht: (2023)
von: Borah, Angana, et al.
Veröffentlicht: (2023)
SpaceVLM: Sub-Space Modeling of Negation in Vision-Language Models
von: Ranjbar, Sepehr Kazemi, et al.
Veröffentlicht: (2025)
von: Ranjbar, Sepehr Kazemi, et al.
Veröffentlicht: (2025)
LLM Knowledge is Brittle: Truthfulness Representations Rely on Superficial Resemblance
von: Haller, Patrick, et al.
Veröffentlicht: (2025)
von: Haller, Patrick, et al.
Veröffentlicht: (2025)
Can SAEs reveal and mitigate racial biases of LLMs in healthcare?
von: Ahsan, Hiba, et al.
Veröffentlicht: (2025)
von: Ahsan, Hiba, et al.
Veröffentlicht: (2025)
Do Automatic Factuality Metrics Measure Factuality? A Critical Evaluation
von: Ramprasad, Sanjana, et al.
Veröffentlicht: (2024)
von: Ramprasad, Sanjana, et al.
Veröffentlicht: (2024)
Algorithmic Pluralism: A Structural Approach To Equal Opportunity
von: Jain, Shomik, et al.
Veröffentlicht: (2023)
von: Jain, Shomik, et al.
Veröffentlicht: (2023)
Chained Tuning Leads to Biased Forgetting
von: Ung, Megan, et al.
Veröffentlicht: (2024)
von: Ung, Megan, et al.
Veröffentlicht: (2024)
Evaluating the Factuality of Zero-shot Summarizers Across Varied Domains
von: Ramprasad, Sanjana, et al.
Veröffentlicht: (2024)
von: Ramprasad, Sanjana, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Detection and Measurement of Syntactic Templates in Generated Text
von: Shaib, Chantal, et al.
Veröffentlicht: (2024) -
When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment
von: Xiao, Yuxin, et al.
Veröffentlicht: (2025) -
Who Taught You That? Tracing Teachers in Model Distillation
von: Wadhwa, Somin, et al.
Veröffentlicht: (2025) -
Measuring AI "Slop" in Text
von: Shaib, Chantal, et al.
Veröffentlicht: (2025) -
How Much Annotation is Needed to Compare Summarization Models?
von: Shaib, Chantal, et al.
Veröffentlicht: (2024)