CSEval: Towards Automated, Multi-Dimensional, and Reference-Free Counterspeech Evaluation using Auto-Calibrated LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Hengle, Amey, Kumar, Aswini, Bandhakavi, Anil, Chakraborty, Tanmoy |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Counterspeech the ultimate shield! Multi-Conditioned Counterspeech Generation through Attributed Prefix Learning
by: Kumar, Aswini, et al.
Published: (2025)
by: Kumar, Aswini, et al.
Published: (2025)
Intent-conditioned and Non-toxic Counterspeech Generation using Multi-Task Instruction Tuning with RLAIF
by: Hengle, Amey, et al.
Published: (2024)
by: Hengle, Amey, et al.
Published: (2024)
Can LLMs reason over extended multilingual contexts? Towards long-context evaluation beyond retrieval and haystacks
by: Hengle, Amey, et al.
Published: (2025)
by: Hengle, Amey, et al.
Published: (2025)
Multilingual Needle in a Haystack: Investigating Long-Context Behavior of Multilingual Large Language Models
by: Hengle, Amey, et al.
Published: (2024)
by: Hengle, Amey, et al.
Published: (2024)
Counterspeech for Mitigating the Influence of Media Bias: Comparing Human and LLM-Generated Responses
by: Lin, Luyang, et al.
Published: (2025)
by: Lin, Luyang, et al.
Published: (2025)
NLP Systems That Can't Tell Use from Mention Censor Counterspeech, but Teaching the Distinction Helps
by: Gligoric, Kristina, et al.
Published: (2024)
by: Gligoric, Kristina, et al.
Published: (2024)
When Do LLMs Generate Realistic Social Networks? A Multi-Dimensional Study of Culture, Language, Scale, and Method
by: Kilaru, Sai Hemanth, et al.
Published: (2026)
by: Kilaru, Sai Hemanth, et al.
Published: (2026)
Decoding Memes: Benchmarking Narrative Role Classification across Multilingual and Multimodal Models
by: Sharma, Shivam, et al.
Published: (2025)
by: Sharma, Shivam, et al.
Published: (2025)
Hate Personified: Investigating the role of LLMs in content moderation
by: Masud, Sarah, et al.
Published: (2024)
by: Masud, Sarah, et al.
Published: (2024)
IndRegBias: A Dataset for Studying Indian Regional Biases in English and Code-Mixed Social Media Comments
by: Panda, Debasmita, et al.
Published: (2026)
by: Panda, Debasmita, et al.
Published: (2026)
Diversity Augmentation of Dynamic User Preference Data for Boosting Personalized Text Summarizers
by: Chatterjee, Parthiv, et al.
Published: (2025)
by: Chatterjee, Parthiv, et al.
Published: (2025)
Why Agents Compromise Safety Under Pressure
by: Jiang, Hengle, et al.
Published: (2026)
by: Jiang, Hengle, et al.
Published: (2026)
SAFE-MEME: Structured Reasoning Framework for Robust Hate Speech Detection in Memes
by: Nandi, Palash, et al.
Published: (2024)
by: Nandi, Palash, et al.
Published: (2024)
Focal Inferential Infusion Coupled with Tractable Density Discrimination for Implicit Hate Detection
by: Masud, Sarah, et al.
Published: (2023)
by: Masud, Sarah, et al.
Published: (2023)
Evaluation of LLMs for Process Model Analysis and Optimization
by: Kumar, Akhil, et al.
Published: (2025)
by: Kumar, Akhil, et al.
Published: (2025)
MemeMQA: Multimodal Question Answering for Memes via Rationale-Based Inferencing
by: Agarwal, Siddhant, et al.
Published: (2024)
by: Agarwal, Siddhant, et al.
Published: (2024)
SUKHSANDESH: An Avatar Therapeutic Question Answering Platform for Sexual Education in Rural India
by: Singh, Salam Michael, et al.
Published: (2024)
by: Singh, Salam Michael, et al.
Published: (2024)
FLAME: Self-Supervised Low-Resource Taxonomy Expansion using Large Language Models
by: Mishra, Sahil, et al.
Published: (2024)
by: Mishra, Sahil, et al.
Published: (2024)
Ancient Wisdom, Modern Tools: Exploring Retrieval-Augmented LLMs for Ancient Indian Philosophy
by: Mandikal, Priyanka
Published: (2024)
by: Mandikal, Priyanka
Published: (2024)
Automating the Information Extraction from Semi-Structured Interview Transcripts
by: Parfenova, Angelina
Published: (2024)
by: Parfenova, Angelina
Published: (2024)
From Facts to Conclusions : Integrating Deductive Reasoning in Retrieval-Augmented LLMs
by: Mishra, Shubham, et al.
Published: (2025)
by: Mishra, Shubham, et al.
Published: (2025)
Credible, Unreliable or Leaked?: Evidence Verification for Enhanced Automated Fact-checking
by: Chrysidis, Zacharias, et al.
Published: (2024)
by: Chrysidis, Zacharias, et al.
Published: (2024)
LLMs Can Infer Political Alignment from Online Conversations
by: Lee, Byunghwee, et al.
Published: (2026)
by: Lee, Byunghwee, et al.
Published: (2026)
You Only Prune Once: Designing Calibration-Free Model Compression With Policy Learning
by: Sengupta, Ayan, et al.
Published: (2025)
by: Sengupta, Ayan, et al.
Published: (2025)
Post-hoc Study of Climate Microtargeting on Social Media Ads with LLMs: Thematic Insights and Fairness Evaluation
by: Islam, Tunazzina, et al.
Published: (2024)
by: Islam, Tunazzina, et al.
Published: (2024)
Evaluating LLMs' Assessment of Mixed-Context Hallucination Through the Lens of Summarization
by: Qi, Siya, et al.
Published: (2025)
by: Qi, Siya, et al.
Published: (2025)
A Guide to Misinformation Detection Data and Evaluation
by: Thibault, Camille, et al.
Published: (2024)
by: Thibault, Camille, et al.
Published: (2024)
Temporal Referential Consistency: Do LLMs Favor Sequences Over Absolute Time References?
by: Bajpai, Ashutosh, et al.
Published: (2025)
by: Bajpai, Ashutosh, et al.
Published: (2025)
Evaluating AI capabilities in detecting conspiracy theories on YouTube
by: La Rocca, Leonardo, et al.
Published: (2025)
by: La Rocca, Leonardo, et al.
Published: (2025)
Automated Assessment of Students' Code Comprehension using LLMs
by: Oli, Priti, et al.
Published: (2023)
by: Oli, Priti, et al.
Published: (2023)
When Can We Trust LLM Graders? Calibrating Confidence for Automated Assessment
by: Ferrer, Robinson, et al.
Published: (2026)
by: Ferrer, Robinson, et al.
Published: (2026)
Leveraging Social Media Data to Identify Factors Influencing Public Attitude Towards Accessibility, Socioeconomic Disparity and Public Transportation
by: Momin, Khondhaker Al, et al.
Published: (2024)
by: Momin, Khondhaker Al, et al.
Published: (2024)
A Multimodal, Multilingual, and Multidimensional Pipeline for Fine-grained Crowdsourcing Earthquake Damage Evaluation
by: Ma, Zihui, et al.
Published: (2025)
by: Ma, Zihui, et al.
Published: (2025)
The Psychology of Falsehood: A Human-Centric Survey of Misinformation Detection
by: Nandi, Arghodeep, et al.
Published: (2025)
by: Nandi, Arghodeep, et al.
Published: (2025)
On Zero-Shot Counterspeech Generation by LLMs
by: Saha, Punyajoy, et al.
Published: (2024)
by: Saha, Punyajoy, et al.
Published: (2024)
CommunityFact: A Dynamic, Multilingual, Multi-domain Benchmark for Misinformation Detection in the Wild
by: Singh, Sahajpreet, et al.
Published: (2026)
by: Singh, Sahajpreet, et al.
Published: (2026)
MURAD: A Large-Scale Multi-Domain Unified Reverse Arabic Dictionary Dataset
by: Sibaee, Serry, et al.
Published: (2026)
by: Sibaee, Serry, et al.
Published: (2026)
Rethinking Hate Speech Detection on Social Media: Can LLMs Replace Traditional Models?
by: Singh, Daman Deep, et al.
Published: (2025)
by: Singh, Daman Deep, et al.
Published: (2025)
Large Language Models' Accuracy in Emulating Human Experts' Evaluation of Public Sentiments about Heated Tobacco Products on Social Media
by: Kim, Kwanho, et al.
Published: (2025)
by: Kim, Kwanho, et al.
Published: (2025)
Generalizing Hate Speech Detection Using Multi-Task Learning: A Case Study of Political Public Figures
by: Yuan, Lanqin, et al.
Published: (2022)
by: Yuan, Lanqin, et al.
Published: (2022)
Similar Items
-
Counterspeech the ultimate shield! Multi-Conditioned Counterspeech Generation through Attributed Prefix Learning
by: Kumar, Aswini, et al.
Published: (2025) -
Intent-conditioned and Non-toxic Counterspeech Generation using Multi-Task Instruction Tuning with RLAIF
by: Hengle, Amey, et al.
Published: (2024) -
Can LLMs reason over extended multilingual contexts? Towards long-context evaluation beyond retrieval and haystacks
by: Hengle, Amey, et al.
Published: (2025) -
Multilingual Needle in a Haystack: Investigating Long-Context Behavior of Multilingual Large Language Models
by: Hengle, Amey, et al.
Published: (2024) -
Counterspeech for Mitigating the Influence of Media Bias: Comparing Human and LLM-Generated Responses
by: Lin, Luyang, et al.
Published: (2025)