A Multi-Aspect Framework for Counter Narrative Evaluation using Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Jones, Jaylen, Mo, Lingbo, Fosler-Lussier, Eric, Sun, Huan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents
di: Jones, Jaylen, et al.
Pubblicazione: (2026)
di: Jones, Jaylen, et al.
Pubblicazione: (2026)
Improving Speech Recognition Error Prediction for Modern and Off-the-shelf Speech Recognizers
di: Serai, Prashant, et al.
Pubblicazione: (2024)
di: Serai, Prashant, et al.
Pubblicazione: (2024)
Beyond Length: Context-Aware Expansion and Independence as Developmentally Sensitive Evaluation in Child Utterances
di: Chun, Jiyun, et al.
Pubblicazione: (2026)
di: Chun, Jiyun, et al.
Pubblicazione: (2026)
Do LLM Decoders Listen Fairly? Benchmarking How Language Model Priors Shape Bias in Speech Recognition
di: Ginjala, Srishti, et al.
Pubblicazione: (2026)
di: Ginjala, Srishti, et al.
Pubblicazione: (2026)
How Trustworthy are Open-Source LLMs? An Assessment under Malicious Demonstrations Shows their Vulnerabilities
di: Mo, Lingbo, et al.
Pubblicazione: (2023)
di: Mo, Lingbo, et al.
Pubblicazione: (2023)
A Trembling House of Cards? Mapping Adversarial Attacks against Language Agents
di: Mo, Lingbo, et al.
Pubblicazione: (2024)
di: Mo, Lingbo, et al.
Pubblicazione: (2024)
RedTeamCUA: Realistic Adversarial Testing of Computer-Use Agents in Hybrid Web-OS Environments
di: Liao, Zeyi, et al.
Pubblicazione: (2025)
di: Liao, Zeyi, et al.
Pubblicazione: (2025)
MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing
di: Zhang, Kai, et al.
Pubblicazione: (2023)
di: Zhang, Kai, et al.
Pubblicazione: (2023)
LETToT: Label-Free Evaluation of Large Language Models On Tourism Using Expert Tree-of-Thought
di: Qi, Ruiyan, et al.
Pubblicazione: (2025)
di: Qi, Ruiyan, et al.
Pubblicazione: (2025)
A Comprehensive Evaluation of Large Language Models on Aspect-Based Sentiment Analysis
di: Zhou, Changzhi, et al.
Pubblicazione: (2024)
di: Zhou, Changzhi, et al.
Pubblicazione: (2024)
A Lightweight Multi Aspect Controlled Text Generation Solution For Large Language Models
di: Zhang, Chenyang, et al.
Pubblicazione: (2024)
di: Zhang, Chenyang, et al.
Pubblicazione: (2024)
What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction
di: Li, Lingbo, et al.
Pubblicazione: (2025)
di: Li, Lingbo, et al.
Pubblicazione: (2025)
Evaluation of Large Language Models for Summarization Tasks in the Medical Domain: A Narrative Review
di: Croxford, Emma, et al.
Pubblicazione: (2024)
di: Croxford, Emma, et al.
Pubblicazione: (2024)
Alternative Speech: Complementary Method to Counter-Narrative for Better Discourse
di: Lee, Seungyoon, et al.
Pubblicazione: (2024)
di: Lee, Seungyoon, et al.
Pubblicazione: (2024)
TRUSTVIS: A Multi-Dimensional Trustworthiness Evaluation Framework for Large Language Models
di: Sun, Ruoyu, et al.
Pubblicazione: (2025)
di: Sun, Ruoyu, et al.
Pubblicazione: (2025)
MTMCS-Bench: Evaluating Contextual Safety of Multimodal Large Language Models in Multi-Turn Dialogues
di: Liu, Zheyuan, et al.
Pubblicazione: (2026)
di: Liu, Zheyuan, et al.
Pubblicazione: (2026)
Dynamic Knowledge Integration for Evidence-Driven Counter-Argument Generation with Large Language Models
di: Yeginbergen, Anar, et al.
Pubblicazione: (2025)
di: Yeginbergen, Anar, et al.
Pubblicazione: (2025)
HierTOD: A Task-Oriented Dialogue System Driven by Hierarchical Goals
di: Mo, Lingbo, et al.
Pubblicazione: (2024)
di: Mo, Lingbo, et al.
Pubblicazione: (2024)
DivLogicEval: A Framework for Benchmarking Logical Reasoning Evaluation in Large Language Models
di: Chung, Tsz Ting, et al.
Pubblicazione: (2025)
di: Chung, Tsz Ting, et al.
Pubblicazione: (2025)
Utilizing Large Language Models for Event Deconstruction to Enhance Multimodal Aspect-Based Sentiment Analysis
di: Huang, Xiaoyong, et al.
Pubblicazione: (2024)
di: Huang, Xiaoyong, et al.
Pubblicazione: (2024)
Narrative-of-Thought: Improving Temporal Reasoning of Large Language Models via Recounted Narratives
di: Zhang, Xinliang Frederick, et al.
Pubblicazione: (2024)
di: Zhang, Xinliang Frederick, et al.
Pubblicazione: (2024)
Evaluating Large Language Models for Abstract Evaluation Tasks: An Empirical Study
di: Liu, Yinuo, et al.
Pubblicazione: (2026)
di: Liu, Yinuo, et al.
Pubblicazione: (2026)
BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali
di: Adib, Shefayat E Shams, et al.
Pubblicazione: (2026)
di: Adib, Shefayat E Shams, et al.
Pubblicazione: (2026)
Improving Neural Diarization through Speaker Attribute Attractors and Local Dependency Modeling
di: Palzer, David, et al.
Pubblicazione: (2025)
di: Palzer, David, et al.
Pubblicazione: (2025)
SemioLLM: Evaluating Large Language Models for Diagnostic Reasoning from Unstructured Clinical Narratives in Epilepsy
di: Dani, Meghal, et al.
Pubblicazione: (2024)
di: Dani, Meghal, et al.
Pubblicazione: (2024)
Can LLMs Narrate Tabular Data? An Evaluation Framework for Natural Language Representations of Text-to-SQL System Outputs
di: Singh, Jyotika, et al.
Pubblicazione: (2025)
di: Singh, Jyotika, et al.
Pubblicazione: (2025)
Catching Chameleons: Detecting Evolving Disinformation Generated using Large Language Models
di: Jiang, Bohan, et al.
Pubblicazione: (2024)
di: Jiang, Bohan, et al.
Pubblicazione: (2024)
Countering Catastrophic Forgetting of Large Language Models for Better Instruction Following via Weight-Space Model Merging
di: Lyu, Mengxian, et al.
Pubblicazione: (2026)
di: Lyu, Mengxian, et al.
Pubblicazione: (2026)
A Novel Self-Evolution Framework for Large Language Models
di: Sun, Haoran, et al.
Pubblicazione: (2025)
di: Sun, Haoran, et al.
Pubblicazione: (2025)
WXImpactBench: A Disruptive Weather Impact Understanding Benchmark for Evaluating Large Language Models
di: Yu, Yongan, et al.
Pubblicazione: (2025)
di: Yu, Yongan, et al.
Pubblicazione: (2025)
A Multi-Domain Red Teaming Framework for Safety, Robustness, and Fairness Evaluation of Medical Large Language Models
di: Feier, Andrei Marian, et al.
Pubblicazione: (2026)
di: Feier, Andrei Marian, et al.
Pubblicazione: (2026)
MulDimIF: A Multi-Dimensional Constraint Framework for Evaluating and Improving Instruction Following in Large Language Models
di: Ye, Junjie, et al.
Pubblicazione: (2025)
di: Ye, Junjie, et al.
Pubblicazione: (2025)
Exploring Narrative Clustering in Large Language Models: A Layerwise Analysis of BERT
di: Banerjee, Awritrojit, et al.
Pubblicazione: (2025)
di: Banerjee, Awritrojit, et al.
Pubblicazione: (2025)
Promptomatix: An Automatic Prompt Optimization Framework for Large Language Models
di: Murthy, Rithesh, et al.
Pubblicazione: (2025)
di: Murthy, Rithesh, et al.
Pubblicazione: (2025)
Language Models Guidance with Multi-Aspect-Cueing: A Case Study for Competitor Analysis
di: Hadifar, Amir, et al.
Pubblicazione: (2025)
di: Hadifar, Amir, et al.
Pubblicazione: (2025)
A Multi-Layered Large Language Model Framework for Disease Prediction
di: Mohamed, Malak, et al.
Pubblicazione: (2025)
di: Mohamed, Malak, et al.
Pubblicazione: (2025)
The Oracle Has Spoken: A Multi-Aspect Evaluation of Dialogue in Pythia
di: Chen, Zixun, et al.
Pubblicazione: (2025)
di: Chen, Zixun, et al.
Pubblicazione: (2025)
CreativityPrism: A Holistic Evaluation Framework for Large Language Model Creativity
di: Hou, Zhaoyi Joey, et al.
Pubblicazione: (2025)
di: Hou, Zhaoyi Joey, et al.
Pubblicazione: (2025)
A Framework for Human Evaluation of Large Language Models in Healthcare Derived from Literature Review
di: Tam, Thomas Yu Chow, et al.
Pubblicazione: (2024)
di: Tam, Thomas Yu Chow, et al.
Pubblicazione: (2024)
Hallucinate Less by Thinking More: Aspect-Based Causal Abstention for Large Language Models
di: Nguyen, Vy, et al.
Pubblicazione: (2025)
di: Nguyen, Vy, et al.
Pubblicazione: (2025)
Documenti analoghi
-
When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents
di: Jones, Jaylen, et al.
Pubblicazione: (2026) -
Improving Speech Recognition Error Prediction for Modern and Off-the-shelf Speech Recognizers
di: Serai, Prashant, et al.
Pubblicazione: (2024) -
Beyond Length: Context-Aware Expansion and Independence as Developmentally Sensitive Evaluation in Child Utterances
di: Chun, Jiyun, et al.
Pubblicazione: (2026) -
Do LLM Decoders Listen Fairly? Benchmarking How Language Model Priors Shape Bias in Speech Recognition
di: Ginjala, Srishti, et al.
Pubblicazione: (2026) -
How Trustworthy are Open-Source LLMs? An Assessment under Malicious Demonstrations Shows their Vulnerabilities
di: Mo, Lingbo, et al.
Pubblicazione: (2023)