AmpleHate: Amplifying the Attention for Versatile Implicit Hate Detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Yejin, Hahn, Joonghyuk, Ahn, Hyeseon, Han, Yo-Sub |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RV-HATE: Reinforced Multi-Module Voting for Implicit Hate Speech Detection
von: Lee, Yejin, et al.
Veröffentlicht: (2025)
von: Lee, Yejin, et al.
Veröffentlicht: (2025)
Adaptive Steering and Remasking for Safe Generation in Diffusion Language Models
von: Lee, Yejin, et al.
Veröffentlicht: (2026)
von: Lee, Yejin, et al.
Veröffentlicht: (2026)
TCProF: Time-Complexity Prediction SSL Framework
von: Hahn, Joonghyuk, et al.
Veröffentlicht: (2025)
von: Hahn, Joonghyuk, et al.
Veröffentlicht: (2025)
Repairing Regex Vulnerabilities via Localization-Guided Instructions
von: Sung, Sicheol, et al.
Veröffentlicht: (2025)
von: Sung, Sicheol, et al.
Veröffentlicht: (2025)
Obfuscation Rules for Detecting and Detoxifying Korean Toxicity
von: Lee, Yejin, et al.
Veröffentlicht: (2025)
von: Lee, Yejin, et al.
Veröffentlicht: (2025)
MEC$^3$O: Multi-Expert Consensus for Code Time Complexity Prediction
von: Hahn, Joonghyuk, et al.
Veröffentlicht: (2025)
von: Hahn, Joonghyuk, et al.
Veröffentlicht: (2025)
Bi-Attention HateXplain : Taking into account the sequential aspect of data during explainability in a multi-task context
von: Mondjo, Ghislain Dorian Tchuente
Veröffentlicht: (2026)
von: Mondjo, Ghislain Dorian Tchuente
Veröffentlicht: (2026)
LLM-Based Multi-Task Bangla Hate Speech Detection: Type, Severity, and Target
von: Hasan, Md Arid, et al.
Veröffentlicht: (2025)
von: Hasan, Md Arid, et al.
Veröffentlicht: (2025)
MemeIntel: Explainable Detection of Propagandistic and Hateful Memes
von: Kmainasi, Mohamed Bayan, et al.
Veröffentlicht: (2025)
von: Kmainasi, Mohamed Bayan, et al.
Veröffentlicht: (2025)
Seeing Hate Differently: Hate Subspace Modeling for Culture-Aware Hate Speech Detection
von: Cai, Weibin, et al.
Veröffentlicht: (2025)
von: Cai, Weibin, et al.
Veröffentlicht: (2025)
Train-Attention: Meta-Learning Where to Focus in Continual Knowledge Learning
von: Seo, Yeongbin, et al.
Veröffentlicht: (2024)
von: Seo, Yeongbin, et al.
Veröffentlicht: (2024)
ToxiFrench: Benchmarking and Enhancing Language Models via CoT Fine-Tuning for French Toxicity Detection
von: Delaval, Axel, et al.
Veröffentlicht: (2025)
von: Delaval, Axel, et al.
Veröffentlicht: (2025)
Unpacking Hateful Memes: Presupposed Context and False Claims
von: Cai, Weibin, et al.
Veröffentlicht: (2025)
von: Cai, Weibin, et al.
Veröffentlicht: (2025)
PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review
von: Tu, Songjun, et al.
Veröffentlicht: (2026)
von: Tu, Songjun, et al.
Veröffentlicht: (2026)
WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference
von: Liu, Aiwei, et al.
Veröffentlicht: (2025)
von: Liu, Aiwei, et al.
Veröffentlicht: (2025)
Propaganda to Hate: A Multimodal Analysis of Arabic Memes with Multi-Agent LLMs
von: Alam, Firoj, et al.
Veröffentlicht: (2024)
von: Alam, Firoj, et al.
Veröffentlicht: (2024)
Branching Narratives: Character Decision Points Detection
von: Tikhonov, Alexey
Veröffentlicht: (2024)
von: Tikhonov, Alexey
Veröffentlicht: (2024)
Prompt Compression in Production Task Orchestration: A Pre-Registered Randomized Trial
von: Johnson, Warren, et al.
Veröffentlicht: (2026)
von: Johnson, Warren, et al.
Veröffentlicht: (2026)
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information
von: Park, Seungcheol, et al.
Veröffentlicht: (2025)
von: Park, Seungcheol, et al.
Veröffentlicht: (2025)
WaterSeeker: Pioneering Efficient Detection of Watermarked Segments in Large Documents
von: Pan, Leyi, et al.
Veröffentlicht: (2024)
von: Pan, Leyi, et al.
Veröffentlicht: (2024)
Quantifying Genuine Awareness in Hallucination Prediction Beyond Question-Side Shortcuts
von: Seo, Yeongbin, et al.
Veröffentlicht: (2025)
von: Seo, Yeongbin, et al.
Veröffentlicht: (2025)
SentiCSE: A Sentiment-aware Contrastive Sentence Embedding Framework with Sentiment-guided Textual Similarity
von: Kim, Jaemin, et al.
Veröffentlicht: (2024)
von: Kim, Jaemin, et al.
Veröffentlicht: (2024)
Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models
von: Park, Seungcheol, et al.
Veröffentlicht: (2025)
von: Park, Seungcheol, et al.
Veröffentlicht: (2025)
CritiSense: Critical Digital Literacy and Resilience Against Misinformation
von: Alam, Firoj, et al.
Veröffentlicht: (2026)
von: Alam, Firoj, et al.
Veröffentlicht: (2026)
Quantifying Fairness in LLMs Beyond Tokens: A Semantic and Statistical Perspective
von: Xu, Weijie, et al.
Veröffentlicht: (2025)
von: Xu, Weijie, et al.
Veröffentlicht: (2025)
Change My Frame: Reframing in the Wild in r/ChangeMyView
von: Peguero, Arturo Martínez, et al.
Veröffentlicht: (2024)
von: Peguero, Arturo Martínez, et al.
Veröffentlicht: (2024)
Co-NAML-LSTUR: A Combined Model with Attentive Multi-View Learning and Long- and Short-term User Representations for News Recommendation
von: Nguyen, Minh Hoang, et al.
Veröffentlicht: (2025)
von: Nguyen, Minh Hoang, et al.
Veröffentlicht: (2025)
Align-to-Distill: Trainable Attention Alignment for Knowledge Distillation in Neural Machine Translation
von: Jin, Heegon, et al.
Veröffentlicht: (2024)
von: Jin, Heegon, et al.
Veröffentlicht: (2024)
The Knesset Corpus: An Annotated Corpus of Hebrew Parliamentary Proceedings
von: Goldin, Gili, et al.
Veröffentlicht: (2024)
von: Goldin, Gili, et al.
Veröffentlicht: (2024)
Math Natural Language Inference: this should be easy!
von: de Paiva, Valeria, et al.
Veröffentlicht: (2025)
von: de Paiva, Valeria, et al.
Veröffentlicht: (2025)
Fast Quiet-STaR: Thinking Without Thought Tokens
von: Huang, Wei, et al.
Veröffentlicht: (2025)
von: Huang, Wei, et al.
Veröffentlicht: (2025)
New Skills or Sharper Primitives? A Probabilistic Perspective on the Emergence of Reasoning in RLVR
von: Wang, Zhilin, et al.
Veröffentlicht: (2026)
von: Wang, Zhilin, et al.
Veröffentlicht: (2026)
d-TreeRPO: Towards More Reliable Policy Optimization for Diffusion Language Models
von: Pan, Leyi, et al.
Veröffentlicht: (2025)
von: Pan, Leyi, et al.
Veröffentlicht: (2025)
Towards Effective and Efficient Continual Pre-training of Large Language Models
von: Chen, Jie, et al.
Veröffentlicht: (2024)
von: Chen, Jie, et al.
Veröffentlicht: (2024)
An Unforgeable Publicly Verifiable Watermark for Large Language Models
von: Liu, Aiwei, et al.
Veröffentlicht: (2023)
von: Liu, Aiwei, et al.
Veröffentlicht: (2023)
The Unlikely Duel: Evaluating Creative Writing in LLMs through a Unique Scenario
von: Gómez-Rodríguez, Carlos, et al.
Veröffentlicht: (2024)
von: Gómez-Rodríguez, Carlos, et al.
Veröffentlicht: (2024)
Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt Distillation
von: Liu, Aiwei, et al.
Veröffentlicht: (2024)
von: Liu, Aiwei, et al.
Veröffentlicht: (2024)
Exploiting Pre-trained Encoder-Decoder Transformers for Sequence-to-Sequence Constituent Parsing
von: Fernández-González, Daniel, et al.
Veröffentlicht: (2026)
von: Fernández-González, Daniel, et al.
Veröffentlicht: (2026)
Parametric Social Identity Injection and Diversification in Public Opinion Simulation
von: Wang, Hexi, et al.
Veröffentlicht: (2026)
von: Wang, Hexi, et al.
Veröffentlicht: (2026)
Trusted Uncertainty in Large Language Models: A Unified Framework for Confidence Calibration and Risk-Controlled Refusal
von: Oehri, Markus, et al.
Veröffentlicht: (2025)
von: Oehri, Markus, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
RV-HATE: Reinforced Multi-Module Voting for Implicit Hate Speech Detection
von: Lee, Yejin, et al.
Veröffentlicht: (2025) -
Adaptive Steering and Remasking for Safe Generation in Diffusion Language Models
von: Lee, Yejin, et al.
Veröffentlicht: (2026) -
TCProF: Time-Complexity Prediction SSL Framework
von: Hahn, Joonghyuk, et al.
Veröffentlicht: (2025) -
Repairing Regex Vulnerabilities via Localization-Guided Instructions
von: Sung, Sicheol, et al.
Veröffentlicht: (2025) -
Obfuscation Rules for Detecting and Detoxifying Korean Toxicity
von: Lee, Yejin, et al.
Veröffentlicht: (2025)