Salvato in:
| Autori principali: | Lu, Junyu, Ma, Kai, Wang, Kaichun, Xiao, Kelaiti, Lee, Roy Ka-Wei, Xu, Bo, Yang, Liang, Lin, Hongfei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2502.06207 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Visual Puns from Idioms: An Iterative LLM-T2IM-MLLM Framework
di: Xiao, Kelaiti, et al.
Pubblicazione: (2025)
di: Xiao, Kelaiti, et al.
Pubblicazione: (2025)
Aligning LLM Uncertainty with Human Disagreement in Subjectivity Analysis
di: Lu, Junyu, et al.
Pubblicazione: (2026)
di: Lu, Junyu, et al.
Pubblicazione: (2026)
VisualQuest: A Benchmark for Abstract Visual Reasoning in MLLMs
di: Xiao, Kelaiti, et al.
Pubblicazione: (2025)
di: Xiao, Kelaiti, et al.
Pubblicazione: (2025)
ToxiCloakCN: Evaluating Robustness of Offensive Language Detection in Chinese with Cloaking Perturbations
di: Xiao, Yunze, et al.
Pubblicazione: (2024)
di: Xiao, Yunze, et al.
Pubblicazione: (2024)
Multi-Agent VLMs Guided Self-Training with PNU Loss for Low-Resource Offensive Content Detection
di: Wang, Han, et al.
Pubblicazione: (2025)
di: Wang, Han, et al.
Pubblicazione: (2025)
Harder to Defend: Towards Chinese Toxicity Attacks via Implicit Enhancement and Obfuscation Rewriting
di: Kang, Jingyi, et al.
Pubblicazione: (2026)
di: Kang, Jingyi, et al.
Pubblicazione: (2026)
Towards Patronizing and Condescending Language in Chinese Videos: A Multimodal Dataset and Detector
di: Wang, Hongbo, et al.
Pubblicazione: (2024)
di: Wang, Hongbo, et al.
Pubblicazione: (2024)
PclGPT: A Large Language Model for Patronizing and Condescending Language Detection
di: Wang, Hongbo, et al.
Pubblicazione: (2024)
di: Wang, Hongbo, et al.
Pubblicazione: (2024)
Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution
di: Tian, Zailong, et al.
Pubblicazione: (2025)
di: Tian, Zailong, et al.
Pubblicazione: (2025)
Take its Essence, Discard its Dross! Debiasing for Toxic Language Detection via Counterfactual Causal Effect
di: Lu, Junyu, et al.
Pubblicazione: (2024)
di: Lu, Junyu, et al.
Pubblicazione: (2024)
Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is Offensive
di: Weerasooriya, Tharindu Cyril, et al.
Pubblicazione: (2023)
di: Weerasooriya, Tharindu Cyril, et al.
Pubblicazione: (2023)
D3CODE: Disentangling Disagreements in Data across Cultures on Offensiveness Detection and Evaluation
di: Davani, Aida Mostafazadeh, et al.
Pubblicazione: (2024)
di: Davani, Aida Mostafazadeh, et al.
Pubblicazione: (2024)
Integrating Multi-view Analysis: Multi-view Mixture-of-Expert for Textual Personality Detection
di: Zhu, Haohao, et al.
Pubblicazione: (2024)
di: Zhu, Haohao, et al.
Pubblicazione: (2024)
When Disagreements Elicit Robustness: Investigating Self-Repair Capabilities under LLM Multi-Agent Disagreements
di: Ju, Tianjie, et al.
Pubblicazione: (2025)
di: Ju, Tianjie, et al.
Pubblicazione: (2025)
From Text to Emotion: Unveiling the Emotion Annotation Capabilities of LLMs
di: Niu, Minxue, et al.
Pubblicazione: (2024)
di: Niu, Minxue, et al.
Pubblicazione: (2024)
Guardians of Discourse: Evaluating LLMs on Multilingual Offensive Language Detection
di: He, Jianfei, et al.
Pubblicazione: (2024)
di: He, Jianfei, et al.
Pubblicazione: (2024)
Towards Comprehensive Detection of Chinese Harmful Memes
di: Lu, Junyu, et al.
Pubblicazione: (2024)
di: Lu, Junyu, et al.
Pubblicazione: (2024)
HateClipSeg: A Segment-Level Annotated Dataset for Fine-Grained Hate Video Detection
di: Wang, Han, et al.
Pubblicazione: (2025)
di: Wang, Han, et al.
Pubblicazione: (2025)
Chinese Offensive Language Detection:Current Status and Future Directions
di: Xiao, Yunze, et al.
Pubblicazione: (2024)
di: Xiao, Yunze, et al.
Pubblicazione: (2024)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
di: Shi, Lin, et al.
Pubblicazione: (2024)
di: Shi, Lin, et al.
Pubblicazione: (2024)
With Great Capabilities Come Great Responsibilities: Introducing the Agentic Risk & Capability Framework for Governing Agentic AI Systems
di: Khoo, Shaun, et al.
Pubblicazione: (2025)
di: Khoo, Shaun, et al.
Pubblicazione: (2025)
Language, Culture, and Ideology: Personalizing Offensiveness Detection in Political Tweets with Reasoning LLMs
di: Pihulski, Dzmitry, et al.
Pubblicazione: (2025)
di: Pihulski, Dzmitry, et al.
Pubblicazione: (2025)
OCCULT: Evaluating Large Language Models for Offensive Cyber Operation Capabilities
di: Kouremetis, Michael, et al.
Pubblicazione: (2025)
di: Kouremetis, Michael, et al.
Pubblicazione: (2025)
Evaluating Annotation Consistency in Offensive Language Detection: A Data Analytics Approach on the TweetEval Dataset
di: Fabeela Ali Rawther,Abhinay A K,Anagha Tess B,Alan Joseph,Adham Saheer
Pubblicazione: (2025)
di: Fabeela Ali Rawther,Abhinay A K,Anagha Tess B,Alan Joseph,Adham Saheer
Pubblicazione: (2025)
Leveraging Annotator Disagreement for Text Classification
di: Xu, Jin, et al.
Pubblicazione: (2024)
di: Xu, Jin, et al.
Pubblicazione: (2024)
Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks
di: Merves, Tyler H., et al.
Pubblicazione: (2026)
di: Merves, Tyler H., et al.
Pubblicazione: (2026)
Taming Overconfidence in LLMs: Reward Calibration in RLHF
di: Leng, Jixuan, et al.
Pubblicazione: (2024)
di: Leng, Jixuan, et al.
Pubblicazione: (2024)
Do Language Models Mirror Human Confidence? Exploring Psychological Insights to Address Overconfidence in LLMs
di: Xu, Chenjun, et al.
Pubblicazione: (2025)
di: Xu, Chenjun, et al.
Pubblicazione: (2025)
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
di: Calderon, Nitay, et al.
Pubblicazione: (2025)
di: Calderon, Nitay, et al.
Pubblicazione: (2025)
Heterogeneous Judge-Aware Ranking with Sensitivity, Disagreement, and Confidence
di: Yu, Shibo, et al.
Pubblicazione: (2026)
di: Yu, Shibo, et al.
Pubblicazione: (2026)
Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF Benchmark
di: Shao, Minghao, et al.
Pubblicazione: (2025)
di: Shao, Minghao, et al.
Pubblicazione: (2025)
MemGuard-Alpha: Detecting and Filtering Memorization-Contaminated Signals in LLM-Based Financial Forecasting via Membership Inference and Cross-Model Disagreement
di: Roy, Anisha, et al.
Pubblicazione: (2026)
di: Roy, Anisha, et al.
Pubblicazione: (2026)
Same Verdict, Different Reasons: LLM-as-a-Judge and Clinician Disagreement on Medical Chatbot Completeness
di: DeLucia, Alexandra, et al.
Pubblicazione: (2026)
di: DeLucia, Alexandra, et al.
Pubblicazione: (2026)
Calibrating Probabilistic Object Detectors with Annotator Disagreement
di: Tan, Zhi Qin, et al.
Pubblicazione: (2026)
di: Tan, Zhi Qin, et al.
Pubblicazione: (2026)
Dealing with Annotator Disagreement in Hate Speech Classification
di: Dehghan, Somaiyeh, et al.
Pubblicazione: (2025)
di: Dehghan, Somaiyeh, et al.
Pubblicazione: (2025)
Function-based Labels for Complementary Recommendation: Definition, Annotation, and LLM-as-a-Judge
di: Yamasaki, Chihiro, et al.
Pubblicazione: (2025)
di: Yamasaki, Chihiro, et al.
Pubblicazione: (2025)
Enhancing Textual Personality Detection toward Social Media: Integrating Long-term and Short-term Perspectives
di: Zhu, Haohao, et al.
Pubblicazione: (2024)
di: Zhu, Haohao, et al.
Pubblicazione: (2024)
Distinguishing Right from Wrong in Debates: Attribution Analysis of Chinese Harmful Memes
di: Wang, Weiming, et al.
Pubblicazione: (2026)
di: Wang, Weiming, et al.
Pubblicazione: (2026)
Detecting Offensive Memes with Social Biases in Singapore Context Using Multimodal Large Language Models
di: Yuxuan, Cao, et al.
Pubblicazione: (2025)
di: Yuxuan, Cao, et al.
Pubblicazione: (2025)
Detection and Analysis of Offensive Online Content in Hausa Language
di: Adam, Fatima Muhammad, et al.
Pubblicazione: (2023)
di: Adam, Fatima Muhammad, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Visual Puns from Idioms: An Iterative LLM-T2IM-MLLM Framework
di: Xiao, Kelaiti, et al.
Pubblicazione: (2025) -
Aligning LLM Uncertainty with Human Disagreement in Subjectivity Analysis
di: Lu, Junyu, et al.
Pubblicazione: (2026) -
VisualQuest: A Benchmark for Abstract Visual Reasoning in MLLMs
di: Xiao, Kelaiti, et al.
Pubblicazione: (2025) -
ToxiCloakCN: Evaluating Robustness of Offensive Language Detection in Chinese with Cloaking Perturbations
di: Xiao, Yunze, et al.
Pubblicazione: (2024) -
Multi-Agent VLMs Guided Self-Training with PNU Loss for Low-Resource Offensive Content Detection
di: Wang, Han, et al.
Pubblicazione: (2025)