Beyond Imperfect Alternatives with Rulemapping: A Neuro-Symbolic Case Study on Online Hate Speech
Fuente:
arXiv
Salvato in:
| Autore principale: | von Cossel, Oskar |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LLM-FACETS: A Privacy-Preserving Framework for Evaluating LLM Transparency and Accountability
di: Lucas, Tom, et al.
Pubblicazione: (2026)
di: Lucas, Tom, et al.
Pubblicazione: (2026)
The Company You Keep: How LLMs Respond to Dark Triad Traits
di: Lu, Zeyi, et al.
Pubblicazione: (2026)
di: Lu, Zeyi, et al.
Pubblicazione: (2026)
Exploring and Mitigating Gender Bias in Encoder-Based Transformer Models
di: Hossain, Ariyan, et al.
Pubblicazione: (2025)
di: Hossain, Ariyan, et al.
Pubblicazione: (2025)
GraphWalk: Enabling Reasoning in Large Language Models through Tool-Based Graph Navigation
di: Ghandi, Taraneh, et al.
Pubblicazione: (2026)
di: Ghandi, Taraneh, et al.
Pubblicazione: (2026)
Aurora: Neuro-Symbolic AI Driven Advising Agent
di: Lugones, Lorena Amanda Quincoso, et al.
Pubblicazione: (2026)
di: Lugones, Lorena Amanda Quincoso, et al.
Pubblicazione: (2026)
Why we need an AI-resilient society
di: Bartz-Beielstein, Thomas
Pubblicazione: (2019)
di: Bartz-Beielstein, Thomas
Pubblicazione: (2019)
Modeling Fairness in Recruitment AI via Information Flow
di: Brännström, Mattias, et al.
Pubblicazione: (2025)
di: Brännström, Mattias, et al.
Pubblicazione: (2025)
Scaling Laws for State Dynamics in Large Language Models
di: Li, Jacob X, et al.
Pubblicazione: (2025)
di: Li, Jacob X, et al.
Pubblicazione: (2025)
Attention-based sequential recommendation system using multimodal data
di: Oh, Hyungtaik, et al.
Pubblicazione: (2024)
di: Oh, Hyungtaik, et al.
Pubblicazione: (2024)
Uncovering Bugs in Formal Explainers: A Case Study with PyXAI
di: Huang, Xuanxiang, et al.
Pubblicazione: (2025)
di: Huang, Xuanxiang, et al.
Pubblicazione: (2025)
Policy Cards: Machine-Readable Runtime Governance for Autonomous AI Agents
di: Mavračić, Juraj
Pubblicazione: (2025)
di: Mavračić, Juraj
Pubblicazione: (2025)
Qwerty AI: Explainable Automated Age Rating and Content Safety Assessment for Russian-Language Screenplays
di: Zmanovskii, Nikita
Pubblicazione: (2025)
di: Zmanovskii, Nikita
Pubblicazione: (2025)
Replicating TEMPEST at Scale: Multi-Turn Adversarial Attacks Against Trillion-Parameter Frontier Models
di: Young, Richard
Pubblicazione: (2025)
di: Young, Richard
Pubblicazione: (2025)
Bye Bye Perspective API: Lessons for Measurement Infrastructure in NLP, CSS and LLM Evaluation
di: Hartmann, David, et al.
Pubblicazione: (2026)
di: Hartmann, David, et al.
Pubblicazione: (2026)
Whose wife is it anyway? Assessing bias against same-gender relationships in machine translation
di: Stewart, Ian, et al.
Pubblicazione: (2024)
di: Stewart, Ian, et al.
Pubblicazione: (2024)
Not Everything That Counts Can Be Counted: A Case for Safe Qualitative AI
di: Beltoft, Stine, et al.
Pubblicazione: (2025)
di: Beltoft, Stine, et al.
Pubblicazione: (2025)
The Epistemic Suite: A Post-Foundational Diagnostic Methodology for Assessing AI Knowledge Claims
di: Kelly, Matthew
Pubblicazione: (2025)
di: Kelly, Matthew
Pubblicazione: (2025)
Gyan: An Explainable Neuro-Symbolic Language Model
di: Srinivasan, Venkat, et al.
Pubblicazione: (2026)
di: Srinivasan, Venkat, et al.
Pubblicazione: (2026)
Demystifying Funding: Reconstructing a Unified Dataset of the UK Funding Lifecycle
di: Thorne, William, et al.
Pubblicazione: (2026)
di: Thorne, William, et al.
Pubblicazione: (2026)
Towards Observation Lakehouses: Living, Interactive Archives of Software Behavior
di: Kessel, Marcus
Pubblicazione: (2025)
di: Kessel, Marcus
Pubblicazione: (2025)
AI and My Values: User Perceptions of LLMs' Ability to Extract, Embody, and Explain Human Values from Casual Conversations
di: Yun, Bhada, et al.
Pubblicazione: (2026)
di: Yun, Bhada, et al.
Pubblicazione: (2026)
Identifying and Mitigating Gender Cues in Academic Recommendation Letters: An Interpretability Case Study
di: Alexander, Charlotte S., et al.
Pubblicazione: (2026)
di: Alexander, Charlotte S., et al.
Pubblicazione: (2026)
The Invisible Coalition Partner: How LLMs Vote When Democracy Gets Concrete
di: Barmettler, Joel
Pubblicazione: (2026)
di: Barmettler, Joel
Pubblicazione: (2026)
Human Values in a Single Sentence: Moral Presence, Hierarchies, and Transformer Ensembles on the Schwartz Continuum
di: Yeste, Víctor, et al.
Pubblicazione: (2026)
di: Yeste, Víctor, et al.
Pubblicazione: (2026)
REMIND: Input Loss Landscapes Reveal Residual Memorization in Post-Unlearning LLMs
di: Cohen, Liran, et al.
Pubblicazione: (2025)
di: Cohen, Liran, et al.
Pubblicazione: (2025)
From Helpfulness to Toxic Proactivity: Diagnosing Behavioral Misalignment in LLM Agents
di: Wang, Xinyue, et al.
Pubblicazione: (2026)
di: Wang, Xinyue, et al.
Pubblicazione: (2026)
Tatemae: Detecting Alignment Faking via Tool Selection in LLMs
di: Leonesi, Matteo, et al.
Pubblicazione: (2026)
di: Leonesi, Matteo, et al.
Pubblicazione: (2026)
AgentLeak: A Full-Stack Benchmark for Privacy Leakage in Multi-Agent LLM Systems
di: Yagoubi, Faouzi El, et al.
Pubblicazione: (2026)
di: Yagoubi, Faouzi El, et al.
Pubblicazione: (2026)
Using a cognitive architecture to consider antiBlackness in design and development of AI systems
di: Dancy, Christopher L.
Pubblicazione: (2022)
di: Dancy, Christopher L.
Pubblicazione: (2022)
CSSDM Ontology to Enable Continuity of Care Data Interoperability
di: Das, Subhashis, et al.
Pubblicazione: (2025)
di: Das, Subhashis, et al.
Pubblicazione: (2025)
Efficacy of a Computer Tutor that Models Expert Human Tutors
di: Olney, Andrew M., et al.
Pubblicazione: (2025)
di: Olney, Andrew M., et al.
Pubblicazione: (2025)
Can LLMs Identify Tax Abuse?
di: Blair-Stanek, Andrew, et al.
Pubblicazione: (2025)
di: Blair-Stanek, Andrew, et al.
Pubblicazione: (2025)
LLMs Simulate Big Five Personality Traits: Further Evidence
di: Sorokovikova, Aleksandra, et al.
Pubblicazione: (2024)
di: Sorokovikova, Aleksandra, et al.
Pubblicazione: (2024)
Scaling In, Not Up? Testing Thick Citation Context Analysis with GPT-5 and Fragile Prompts
di: Simons, Arno
Pubblicazione: (2026)
di: Simons, Arno
Pubblicazione: (2026)
A Graph-based RAG for Energy Efficiency Question Answering
di: Campi, Riccardo, et al.
Pubblicazione: (2025)
di: Campi, Riccardo, et al.
Pubblicazione: (2025)
From Native Memes to Global Moderation: Cross-Cultural Evaluation of Vision-Language Models for Hateful Meme Detection
di: Wang, Mo, et al.
Pubblicazione: (2026)
di: Wang, Mo, et al.
Pubblicazione: (2026)
What Would GPT Click: Practical Effects of Human-AI Behavioral Misalignment and the Cost of Synthetic Participants in User Experience
di: Kuric, Eduard, et al.
Pubblicazione: (2026)
di: Kuric, Eduard, et al.
Pubblicazione: (2026)
LongSumEval: Question-Answering Based Evaluation and Feedback-Driven Refinement for Long Document Summarization
di: Nguyen, Huyen, et al.
Pubblicazione: (2026)
di: Nguyen, Huyen, et al.
Pubblicazione: (2026)
A Field Guide to Decision Making
di: Arthur, Richard B.
Pubblicazione: (2026)
di: Arthur, Richard B.
Pubblicazione: (2026)
A systematic review of relation extraction task since the emergence of Transformers
di: Celian, Ringwald, et al.
Pubblicazione: (2025)
di: Celian, Ringwald, et al.
Pubblicazione: (2025)
Documenti analoghi
-
LLM-FACETS: A Privacy-Preserving Framework for Evaluating LLM Transparency and Accountability
di: Lucas, Tom, et al.
Pubblicazione: (2026) -
The Company You Keep: How LLMs Respond to Dark Triad Traits
di: Lu, Zeyi, et al.
Pubblicazione: (2026) -
Exploring and Mitigating Gender Bias in Encoder-Based Transformer Models
di: Hossain, Ariyan, et al.
Pubblicazione: (2025) -
GraphWalk: Enabling Reasoning in Large Language Models through Tool-Based Graph Navigation
di: Ghandi, Taraneh, et al.
Pubblicazione: (2026) -
Aurora: Neuro-Symbolic AI Driven Advising Agent
di: Lugones, Lorena Amanda Quincoso, et al.
Pubblicazione: (2026)