ModelCitizens: Representing Community Voices in Online Safety
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Suvarna, Ashima, Chance, Christina, Naranjo, Karolina, Palangi, Hamid, Hao, Sophie, Hartvigsen, Thomas, Gabriel, Saadia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SUPERNOVA: Eliciting General Reasoning in LLMs with Reinforcement Learning on Natural Instructions
von: Suvarna, Ashima, et al.
Veröffentlicht: (2026)
von: Suvarna, Ashima, et al.
Veröffentlicht: (2026)
PhonologyBench: Evaluating Phonological Skills of Large Language Models
von: Suvarna, Ashima, et al.
Veröffentlicht: (2024)
von: Suvarna, Ashima, et al.
Veröffentlicht: (2024)
Disparities in LLM Reasoning Accuracy and Explanations: A Case Study on African American English
von: Zhou, Runtao, et al.
Veröffentlicht: (2025)
von: Zhou, Runtao, et al.
Veröffentlicht: (2025)
Comparing Bad Apples to Good Oranges: Aligning Large Language Models via Joint Preference Optimization
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
When Can LLMs Learn to Reason with Weak Supervision?
von: Rahman, Salman, et al.
Veröffentlicht: (2026)
von: Rahman, Salman, et al.
Veröffentlicht: (2026)
Diversity of Thought Improves Reasoning Abilities of LLMs
von: Naik, Ranjita, et al.
Veröffentlicht: (2023)
von: Naik, Ranjita, et al.
Veröffentlicht: (2023)
Language Models' Factuality Depends on the Language of Inquiry
von: Aggarwal, Tushar, et al.
Veröffentlicht: (2025)
von: Aggarwal, Tushar, et al.
Veröffentlicht: (2025)
Math Neurosurgery: Isolating Language Models' Math Reasoning Abilities Using Only Forward Passes
von: Christ, Bryan R., et al.
Veröffentlicht: (2024)
von: Christ, Bryan R., et al.
Veröffentlicht: (2024)
PLAN-TUNING: Post-Training Language Models to Learn Step-by-Step Planning for Complex Problem Solving
von: Parmar, Mihir, et al.
Veröffentlicht: (2025)
von: Parmar, Mihir, et al.
Veröffentlicht: (2025)
X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents
von: Rahman, Salman, et al.
Veröffentlicht: (2025)
von: Rahman, Salman, et al.
Veröffentlicht: (2025)
Efficient Knowledge Editing via Minimal Precomputation
von: Gupta, Akshat, et al.
Veröffentlicht: (2025)
von: Gupta, Akshat, et al.
Veröffentlicht: (2025)
Multi-Objective Alignment of Language Models for Personalized Psychotherapy
von: Beikzadeh, Mehrab, et al.
Veröffentlicht: (2026)
von: Beikzadeh, Mehrab, et al.
Veröffentlicht: (2026)
Attention Satisfies: A Constraint-Satisfaction Lens on Factual Errors of Language Models
von: Yuksekgonul, Mert, et al.
Veröffentlicht: (2023)
von: Yuksekgonul, Mert, et al.
Veröffentlicht: (2023)
Breaking mBad! Supervised Fine-tuning for Cross-Lingual Detoxification
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2025)
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2025)
MMMT-IF: A Challenging Multimodal Multi-Turn Instruction Following Benchmark
von: Epstein, Elliot L., et al.
Veröffentlicht: (2024)
von: Epstein, Elliot L., et al.
Veröffentlicht: (2024)
Survey of Bias In Text-to-Image Generation: Definition, Evaluation, and Mitigation
von: Wan, Yixin, et al.
Veröffentlicht: (2024)
von: Wan, Yixin, et al.
Veröffentlicht: (2024)
A Glitch in the Matrix? Locating and Detecting Language Model Grounding with Fakepedia
von: Monea, Giovanni, et al.
Veröffentlicht: (2023)
von: Monea, Giovanni, et al.
Veröffentlicht: (2023)
KScope: A Framework for Characterizing the Knowledge Status of Language Models
von: Xiao, Yuxin, et al.
Veröffentlicht: (2025)
von: Xiao, Yuxin, et al.
Veröffentlicht: (2025)
MOSAIC: Modeling Social AI for Content Dissemination and Regulation in Multi-Agent Simulations
von: Liu, Genglin, et al.
Veröffentlicht: (2025)
von: Liu, Genglin, et al.
Veröffentlicht: (2025)
ERAS: Evaluating the Robustness of Chinese NLP Models to Morphological Garden Path Errors
von: Li, Qinchan, et al.
Veröffentlicht: (2024)
von: Li, Qinchan, et al.
Veröffentlicht: (2024)
AI Debate Aids Assessment of Controversial Claims
von: Rahman, Salman, et al.
Veröffentlicht: (2025)
von: Rahman, Salman, et al.
Veröffentlicht: (2025)
Representing the Under-Represented: Cultural and Core Capability Benchmarks for Developing Thai Large Language Models
von: Kim, Dahyun, et al.
Veröffentlicht: (2024)
von: Kim, Dahyun, et al.
Veröffentlicht: (2024)
A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models
von: Kazemi, Hamid, et al.
Veröffentlicht: (2026)
von: Kazemi, Hamid, et al.
Veröffentlicht: (2026)
Language Models Represent Beliefs of Self and Others
von: Zhu, Wentao, et al.
Veröffentlicht: (2024)
von: Zhu, Wentao, et al.
Veröffentlicht: (2024)
Assessing Generalization for Subpopulation Representative Modeling via In-Context Learning
von: Simmons, Gabriel, et al.
Veröffentlicht: (2024)
von: Simmons, Gabriel, et al.
Veröffentlicht: (2024)
Norm Growth and Stability Challenges in Localized Sequential Knowledge Editing
von: Gupta, Akshat, et al.
Veröffentlicht: (2025)
von: Gupta, Akshat, et al.
Veröffentlicht: (2025)
EDUMATH: Generating Standards-aligned Educational Math Word Problems
von: Christ, Bryan R., et al.
Veröffentlicht: (2025)
von: Christ, Bryan R., et al.
Veröffentlicht: (2025)
Learn Globally, Speak Locally: Bridging the Gaps in Multilingual Reasoning
von: Hwang, Jaedong, et al.
Veröffentlicht: (2025)
von: Hwang, Jaedong, et al.
Veröffentlicht: (2025)
LLM for Everyone: Representing the Underrepresented in Large Language Models
von: Cahyawijaya, Samuel
Veröffentlicht: (2024)
von: Cahyawijaya, Samuel
Veröffentlicht: (2024)
Mind the Gesture: Evaluating AI Sensitivity to Culturally Offensive Non-Verbal Gestures
von: Yerukola, Akhila, et al.
Veröffentlicht: (2025)
von: Yerukola, Akhila, et al.
Veröffentlicht: (2025)
CLEAR: A Comprehensive Linguistic Evaluation of Argument Rewriting by Large Language Models
von: Huber, Thomas, et al.
Veröffentlicht: (2025)
von: Huber, Thomas, et al.
Veröffentlicht: (2025)
HALO: An Ontology for Representing and Categorizing Hallucinations in Large Language Models
von: Nananukul, Navapat, et al.
Veröffentlicht: (2023)
von: Nananukul, Navapat, et al.
Veröffentlicht: (2023)
Exploring Group and Symmetry Principles in Large Language Models
von: Imani, Shima, et al.
Veröffentlicht: (2024)
von: Imani, Shima, et al.
Veröffentlicht: (2024)
Evaluating a Multi-Agent Voice-Enabled Smart Speaker for Care Homes: A Safety-Focused Framework
von: Dehghani, Zeinab, et al.
Veröffentlicht: (2026)
von: Dehghani, Zeinab, et al.
Veröffentlicht: (2026)
Identifying Implicit Social Biases in Vision-Language Models
von: Hamidieh, Kimia, et al.
Veröffentlicht: (2024)
von: Hamidieh, Kimia, et al.
Veröffentlicht: (2024)
Sparse Autoencoder Features for Classifications and Transferability
von: Gallifant, Jack, et al.
Veröffentlicht: (2025)
von: Gallifant, Jack, et al.
Veröffentlicht: (2025)
Lifelong Knowledge Editing requires Better Regularization
von: Gupta, Akshat, et al.
Veröffentlicht: (2025)
von: Gupta, Akshat, et al.
Veröffentlicht: (2025)
Voice Communication Analysis in Esports
von: Vinot, Aymeric, et al.
Veröffentlicht: (2024)
von: Vinot, Aymeric, et al.
Veröffentlicht: (2024)
IYKYK (But AI Doesn't): Automated Content Moderation Does Not Capture Communities' Heterogeneous Attitudes Towards Reclaimed Language
von: Chance, Christina, et al.
Veröffentlicht: (2026)
von: Chance, Christina, et al.
Veröffentlicht: (2026)
Language Models Represent Space and Time
von: Gurnee, Wes, et al.
Veröffentlicht: (2023)
von: Gurnee, Wes, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
SUPERNOVA: Eliciting General Reasoning in LLMs with Reinforcement Learning on Natural Instructions
von: Suvarna, Ashima, et al.
Veröffentlicht: (2026) -
PhonologyBench: Evaluating Phonological Skills of Large Language Models
von: Suvarna, Ashima, et al.
Veröffentlicht: (2024) -
Disparities in LLM Reasoning Accuracy and Explanations: A Case Study on African American English
von: Zhou, Runtao, et al.
Veröffentlicht: (2025) -
Comparing Bad Apples to Good Oranges: Aligning Large Language Models via Joint Preference Optimization
von: Bansal, Hritik, et al.
Veröffentlicht: (2024) -
When Can LLMs Learn to Reason with Weak Supervision?
von: Rahman, Salman, et al.
Veröffentlicht: (2026)