Interrogating LLM design under a fair learning doctrine
Fuente:
arXiv
Salvato in:
| Autori principali: | Wei, Johnny Tian-Zheng, Wang, Maggie, Godbole, Ameya, Choi, Jonathan H., Jia, Robin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Spiking the training data to correct for test set contamination
di: Wei, Johnny Tian-Zheng, et al.
Pubblicazione: (2026)
di: Wei, Johnny Tian-Zheng, et al.
Pubblicazione: (2026)
Verify with Caution: The Pitfalls of Relying on Imperfect Factuality Metrics
di: Godbole, Ameya, et al.
Pubblicazione: (2025)
di: Godbole, Ameya, et al.
Pubblicazione: (2025)
SCENE: Self-Labeled Counterfactuals for Extrapolating to Negative Examples
di: Fu, Deqing, et al.
Pubblicazione: (2023)
di: Fu, Deqing, et al.
Pubblicazione: (2023)
TokenSmith: Streamlining Data Editing, Search, and Inspection for Large-Scale Language Model Training and Interpretability
di: Khan, Mohammad Aflah, et al.
Pubblicazione: (2025)
di: Khan, Mohammad Aflah, et al.
Pubblicazione: (2025)
Hubble: a Model Suite to Advance the Study of LLM Memorization
di: Wei, Johnny Tian-Zheng, et al.
Pubblicazione: (2025)
di: Wei, Johnny Tian-Zheng, et al.
Pubblicazione: (2025)
The statistical advantage of automatic NLG metrics at the system level
di: Wei, Johnny Tian-Zheng, et al.
Pubblicazione: (2021)
di: Wei, Johnny Tian-Zheng, et al.
Pubblicazione: (2021)
Politics of Questions in News: A Mixed-Methods Study of Interrogative Stances as Markers of Voice and Power
di: Victor, Bros, et al.
Pubblicazione: (2026)
di: Victor, Bros, et al.
Pubblicazione: (2026)
Proving membership in LLM pretraining data via data watermarks
di: Wei, Johnny Tian-Zheng, et al.
Pubblicazione: (2024)
di: Wei, Johnny Tian-Zheng, et al.
Pubblicazione: (2024)
Operationalizing content moderation "accuracy" in the Digital Services Act
di: Wei, Johnny Tian-Zheng, et al.
Pubblicazione: (2023)
di: Wei, Johnny Tian-Zheng, et al.
Pubblicazione: (2023)
Examining Identity Drift in Conversations of LLM Agents
di: Choi, Junhyuk, et al.
Pubblicazione: (2024)
di: Choi, Junhyuk, et al.
Pubblicazione: (2024)
Evaluating LLM-Simulated Conversations in Modeling Inconsistent and Uncollaborative Behaviors in Human Social Interaction
di: Kamoi, Ryo, et al.
Pubblicazione: (2026)
di: Kamoi, Ryo, et al.
Pubblicazione: (2026)
Defining bias in AI-systems: Biased models are fair models
di: Lindloff, Chiara, et al.
Pubblicazione: (2025)
di: Lindloff, Chiara, et al.
Pubblicazione: (2025)
A survey on fairness of large language models in e-commerce: progress, application, and challenge
di: Ren, Qingyang, et al.
Pubblicazione: (2024)
di: Ren, Qingyang, et al.
Pubblicazione: (2024)
Handling Students Dropouts in an LLM-driven Interactive Online Course Using Language Models
di: Wang, Yuanchun, et al.
Pubblicazione: (2025)
di: Wang, Yuanchun, et al.
Pubblicazione: (2025)
A vibe coding learning design to enhance EFL students' talking to, through, and about AI
di: Woo, David James, et al.
Pubblicazione: (2025)
di: Woo, David James, et al.
Pubblicazione: (2025)
The Biased Samaritan: LLM biases in Perceived Kindness
di: Fagan, Jack H, et al.
Pubblicazione: (2025)
di: Fagan, Jack H, et al.
Pubblicazione: (2025)
Agree to Disagree? A Meta-Evaluation of LLM Misgendering
di: Subramonian, Arjun, et al.
Pubblicazione: (2025)
di: Subramonian, Arjun, et al.
Pubblicazione: (2025)
LLM-MC-Affect: LLM-Based Monte Carlo Modeling of Affective Trajectories and Latent Ambiguity for Interpersonal Dynamic Insight
di: Lin, Yu-Zheng, et al.
Pubblicazione: (2026)
di: Lin, Yu-Zheng, et al.
Pubblicazione: (2026)
Multilingual Prompting for Improving LLM Generation Diversity
di: Wang, Qihan, et al.
Pubblicazione: (2025)
di: Wang, Qihan, et al.
Pubblicazione: (2025)
An LLM Agent for Automatic Geospatial Data Analysis
di: Chen, Yuxing, et al.
Pubblicazione: (2024)
di: Chen, Yuxing, et al.
Pubblicazione: (2024)
Can LLM be a Personalized Judge?
di: Dong, Yijiang River, et al.
Pubblicazione: (2024)
di: Dong, Yijiang River, et al.
Pubblicazione: (2024)
Cancer-Myth: Evaluating Large Language Models on Patient Questions with False Presuppositions
di: Zhu, Wang Bill, et al.
Pubblicazione: (2025)
di: Zhu, Wang Bill, et al.
Pubblicazione: (2025)
Human or LLM as Standardized Patients? A Comparative Study for Medical Education
di: Zhang, Bingquan, et al.
Pubblicazione: (2025)
di: Zhang, Bingquan, et al.
Pubblicazione: (2025)
The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies
di: Zhou, Jiaxu, et al.
Pubblicazione: (2025)
di: Zhou, Jiaxu, et al.
Pubblicazione: (2025)
InterrogateLLM: Zero-Resource Hallucination Detection in LLM-Generated Answers
di: Yehuda, Yakir, et al.
Pubblicazione: (2024)
di: Yehuda, Yakir, et al.
Pubblicazione: (2024)
SPA: Achieving Consensus in LLM Alignment via Self-Priority Optimization
di: Huang, Yue, et al.
Pubblicazione: (2025)
di: Huang, Yue, et al.
Pubblicazione: (2025)
Growth First, Care Second? Tracing the Landscape of LLM Value Preferences in Everyday Dilemmas
di: Chen, Zhiyi, et al.
Pubblicazione: (2026)
di: Chen, Zhiyi, et al.
Pubblicazione: (2026)
Robust Data Watermarking in Language Models by Injecting Fictitious Knowledge
di: Cui, Xinyue, et al.
Pubblicazione: (2025)
di: Cui, Xinyue, et al.
Pubblicazione: (2025)
Exploring Safety Alignment Evaluation of LLMs in Chinese Mental Health Dialogues via LLM-as-Judge
di: Cai, Yunna, et al.
Pubblicazione: (2025)
di: Cai, Yunna, et al.
Pubblicazione: (2025)
The three main doctrines on the future of AI
di: Amadori, Alex, et al.
Pubblicazione: (2025)
di: Amadori, Alex, et al.
Pubblicazione: (2025)
Synthetic Data for Robust AI Model Development in Regulated Enterprises
di: Godbole, Aditi
Pubblicazione: (2025)
di: Godbole, Aditi
Pubblicazione: (2025)
Answering Students' Questions on Course Forums Using Multiple Chain-of-Thought Reasoning and Finetuning RAG-Enabled LLM
di: Wang, Neo, et al.
Pubblicazione: (2025)
di: Wang, Neo, et al.
Pubblicazione: (2025)
The Staircase of Ethics: Probing LLM Value Priorities through Multi-Step Induction to Complex Moral Dilemmas
di: Wu, Ya, et al.
Pubblicazione: (2025)
di: Wu, Ya, et al.
Pubblicazione: (2025)
Quantifying the Persona Effect in LLM Simulations
di: Hu, Tiancheng, et al.
Pubblicazione: (2024)
di: Hu, Tiancheng, et al.
Pubblicazione: (2024)
Mind the Gap: Assessing Wiktionary's Crowd-Sourced Linguistic Knowledge on Morphological Gaps in Two Related Languages
di: Sakunkoo, Jonathan, et al.
Pubblicazione: (2025)
di: Sakunkoo, Jonathan, et al.
Pubblicazione: (2025)
LLM Nepotism in Organizational Governance
di: Mao, Shunqi, et al.
Pubblicazione: (2026)
di: Mao, Shunqi, et al.
Pubblicazione: (2026)
TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law
di: Hui, Zheng, et al.
Pubblicazione: (2025)
di: Hui, Zheng, et al.
Pubblicazione: (2025)
The Parrot Dilemma: Human-Labeled vs. LLM-augmented Data in Classification Tasks
di: Møller, Anders Giovanni, et al.
Pubblicazione: (2023)
di: Møller, Anders Giovanni, et al.
Pubblicazione: (2023)
Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States
di: Xiao, Yang, et al.
Pubblicazione: (2025)
di: Xiao, Yang, et al.
Pubblicazione: (2025)
An Annotated Reading of 'The Singer of Tales' in the LLM Era
di: Varshney, Kush R.
Pubblicazione: (2025)
di: Varshney, Kush R.
Pubblicazione: (2025)
Documenti analoghi
-
Spiking the training data to correct for test set contamination
di: Wei, Johnny Tian-Zheng, et al.
Pubblicazione: (2026) -
Verify with Caution: The Pitfalls of Relying on Imperfect Factuality Metrics
di: Godbole, Ameya, et al.
Pubblicazione: (2025) -
SCENE: Self-Labeled Counterfactuals for Extrapolating to Negative Examples
di: Fu, Deqing, et al.
Pubblicazione: (2023) -
TokenSmith: Streamlining Data Editing, Search, and Inspection for Large-Scale Language Model Training and Interpretability
di: Khan, Mohammad Aflah, et al.
Pubblicazione: (2025) -
Hubble: a Model Suite to Advance the Study of LLM Memorization
di: Wei, Johnny Tian-Zheng, et al.
Pubblicazione: (2025)