Interrogating LLM design under a fair learning doctrine
Fuente:
arXiv
Saved in:
| Main Authors: | Wei, Johnny Tian-Zheng, Wang, Maggie, Godbole, Ameya, Choi, Jonathan H., Jia, Robin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Spiking the training data to correct for test set contamination
by: Wei, Johnny Tian-Zheng, et al.
Published: (2026)
by: Wei, Johnny Tian-Zheng, et al.
Published: (2026)
Verify with Caution: The Pitfalls of Relying on Imperfect Factuality Metrics
by: Godbole, Ameya, et al.
Published: (2025)
by: Godbole, Ameya, et al.
Published: (2025)
SCENE: Self-Labeled Counterfactuals for Extrapolating to Negative Examples
by: Fu, Deqing, et al.
Published: (2023)
by: Fu, Deqing, et al.
Published: (2023)
TokenSmith: Streamlining Data Editing, Search, and Inspection for Large-Scale Language Model Training and Interpretability
by: Khan, Mohammad Aflah, et al.
Published: (2025)
by: Khan, Mohammad Aflah, et al.
Published: (2025)
Hubble: a Model Suite to Advance the Study of LLM Memorization
by: Wei, Johnny Tian-Zheng, et al.
Published: (2025)
by: Wei, Johnny Tian-Zheng, et al.
Published: (2025)
The statistical advantage of automatic NLG metrics at the system level
by: Wei, Johnny Tian-Zheng, et al.
Published: (2021)
by: Wei, Johnny Tian-Zheng, et al.
Published: (2021)
Politics of Questions in News: A Mixed-Methods Study of Interrogative Stances as Markers of Voice and Power
by: Victor, Bros, et al.
Published: (2026)
by: Victor, Bros, et al.
Published: (2026)
Proving membership in LLM pretraining data via data watermarks
by: Wei, Johnny Tian-Zheng, et al.
Published: (2024)
by: Wei, Johnny Tian-Zheng, et al.
Published: (2024)
Operationalizing content moderation "accuracy" in the Digital Services Act
by: Wei, Johnny Tian-Zheng, et al.
Published: (2023)
by: Wei, Johnny Tian-Zheng, et al.
Published: (2023)
Examining Identity Drift in Conversations of LLM Agents
by: Choi, Junhyuk, et al.
Published: (2024)
by: Choi, Junhyuk, et al.
Published: (2024)
Evaluating LLM-Simulated Conversations in Modeling Inconsistent and Uncollaborative Behaviors in Human Social Interaction
by: Kamoi, Ryo, et al.
Published: (2026)
by: Kamoi, Ryo, et al.
Published: (2026)
Defining bias in AI-systems: Biased models are fair models
by: Lindloff, Chiara, et al.
Published: (2025)
by: Lindloff, Chiara, et al.
Published: (2025)
A survey on fairness of large language models in e-commerce: progress, application, and challenge
by: Ren, Qingyang, et al.
Published: (2024)
by: Ren, Qingyang, et al.
Published: (2024)
Handling Students Dropouts in an LLM-driven Interactive Online Course Using Language Models
by: Wang, Yuanchun, et al.
Published: (2025)
by: Wang, Yuanchun, et al.
Published: (2025)
A vibe coding learning design to enhance EFL students' talking to, through, and about AI
by: Woo, David James, et al.
Published: (2025)
by: Woo, David James, et al.
Published: (2025)
The Biased Samaritan: LLM biases in Perceived Kindness
by: Fagan, Jack H, et al.
Published: (2025)
by: Fagan, Jack H, et al.
Published: (2025)
Agree to Disagree? A Meta-Evaluation of LLM Misgendering
by: Subramonian, Arjun, et al.
Published: (2025)
by: Subramonian, Arjun, et al.
Published: (2025)
LLM-MC-Affect: LLM-Based Monte Carlo Modeling of Affective Trajectories and Latent Ambiguity for Interpersonal Dynamic Insight
by: Lin, Yu-Zheng, et al.
Published: (2026)
by: Lin, Yu-Zheng, et al.
Published: (2026)
Multilingual Prompting for Improving LLM Generation Diversity
by: Wang, Qihan, et al.
Published: (2025)
by: Wang, Qihan, et al.
Published: (2025)
An LLM Agent for Automatic Geospatial Data Analysis
by: Chen, Yuxing, et al.
Published: (2024)
by: Chen, Yuxing, et al.
Published: (2024)
Can LLM be a Personalized Judge?
by: Dong, Yijiang River, et al.
Published: (2024)
by: Dong, Yijiang River, et al.
Published: (2024)
Cancer-Myth: Evaluating Large Language Models on Patient Questions with False Presuppositions
by: Zhu, Wang Bill, et al.
Published: (2025)
by: Zhu, Wang Bill, et al.
Published: (2025)
Human or LLM as Standardized Patients? A Comparative Study for Medical Education
by: Zhang, Bingquan, et al.
Published: (2025)
by: Zhang, Bingquan, et al.
Published: (2025)
The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies
by: Zhou, Jiaxu, et al.
Published: (2025)
by: Zhou, Jiaxu, et al.
Published: (2025)
InterrogateLLM: Zero-Resource Hallucination Detection in LLM-Generated Answers
by: Yehuda, Yakir, et al.
Published: (2024)
by: Yehuda, Yakir, et al.
Published: (2024)
SPA: Achieving Consensus in LLM Alignment via Self-Priority Optimization
by: Huang, Yue, et al.
Published: (2025)
by: Huang, Yue, et al.
Published: (2025)
Growth First, Care Second? Tracing the Landscape of LLM Value Preferences in Everyday Dilemmas
by: Chen, Zhiyi, et al.
Published: (2026)
by: Chen, Zhiyi, et al.
Published: (2026)
Robust Data Watermarking in Language Models by Injecting Fictitious Knowledge
by: Cui, Xinyue, et al.
Published: (2025)
by: Cui, Xinyue, et al.
Published: (2025)
Exploring Safety Alignment Evaluation of LLMs in Chinese Mental Health Dialogues via LLM-as-Judge
by: Cai, Yunna, et al.
Published: (2025)
by: Cai, Yunna, et al.
Published: (2025)
The three main doctrines on the future of AI
by: Amadori, Alex, et al.
Published: (2025)
by: Amadori, Alex, et al.
Published: (2025)
Synthetic Data for Robust AI Model Development in Regulated Enterprises
by: Godbole, Aditi
Published: (2025)
by: Godbole, Aditi
Published: (2025)
Answering Students' Questions on Course Forums Using Multiple Chain-of-Thought Reasoning and Finetuning RAG-Enabled LLM
by: Wang, Neo, et al.
Published: (2025)
by: Wang, Neo, et al.
Published: (2025)
The Staircase of Ethics: Probing LLM Value Priorities through Multi-Step Induction to Complex Moral Dilemmas
by: Wu, Ya, et al.
Published: (2025)
by: Wu, Ya, et al.
Published: (2025)
Quantifying the Persona Effect in LLM Simulations
by: Hu, Tiancheng, et al.
Published: (2024)
by: Hu, Tiancheng, et al.
Published: (2024)
Mind the Gap: Assessing Wiktionary's Crowd-Sourced Linguistic Knowledge on Morphological Gaps in Two Related Languages
by: Sakunkoo, Jonathan, et al.
Published: (2025)
by: Sakunkoo, Jonathan, et al.
Published: (2025)
LLM Nepotism in Organizational Governance
by: Mao, Shunqi, et al.
Published: (2026)
by: Mao, Shunqi, et al.
Published: (2026)
TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law
by: Hui, Zheng, et al.
Published: (2025)
by: Hui, Zheng, et al.
Published: (2025)
The Parrot Dilemma: Human-Labeled vs. LLM-augmented Data in Classification Tasks
by: Møller, Anders Giovanni, et al.
Published: (2023)
by: Møller, Anders Giovanni, et al.
Published: (2023)
Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States
by: Xiao, Yang, et al.
Published: (2025)
by: Xiao, Yang, et al.
Published: (2025)
An Annotated Reading of 'The Singer of Tales' in the LLM Era
by: Varshney, Kush R.
Published: (2025)
by: Varshney, Kush R.
Published: (2025)
Similar Items
-
Spiking the training data to correct for test set contamination
by: Wei, Johnny Tian-Zheng, et al.
Published: (2026) -
Verify with Caution: The Pitfalls of Relying on Imperfect Factuality Metrics
by: Godbole, Ameya, et al.
Published: (2025) -
SCENE: Self-Labeled Counterfactuals for Extrapolating to Negative Examples
by: Fu, Deqing, et al.
Published: (2023) -
TokenSmith: Streamlining Data Editing, Search, and Inspection for Large-Scale Language Model Training and Interpretability
by: Khan, Mohammad Aflah, et al.
Published: (2025) -
Hubble: a Model Suite to Advance the Study of LLM Memorization
by: Wei, Johnny Tian-Zheng, et al.
Published: (2025)