DEBATE: Devil's Advocate-Based Assessment and Text Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Alex, Kim, Keonwoo, Yoon, Sangwon |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SLM-Based Agentic AI with P-C-G: Optimized for Korean Tool Use
by: Jeon, Changhyun, et al.
Published: (2025)
by: Jeon, Changhyun, et al.
Published: (2025)
Ontology-Free General-Domain Knowledge Graph-to-Text Generation Dataset Synthesis using Large Language Model
by: Kim, Daehee, et al.
Published: (2024)
by: Kim, Daehee, et al.
Published: (2024)
Multi-Dimensional Optimization for Text Summarization via Reinforcement Learning
by: Ryu, Sangwon, et al.
Published: (2024)
by: Ryu, Sangwon, et al.
Published: (2024)
An Empirical Study of Group Conformity in Multi-Agent Systems
by: Choi, Min, et al.
Published: (2025)
by: Choi, Min, et al.
Published: (2025)
A Multifaceted Analysis of Negative Bias in Large Language Models through the Lens of Parametric Knowledge
by: Song, Jongyoon, et al.
Published: (2025)
by: Song, Jongyoon, et al.
Published: (2025)
Large Language Models are Skeptics: False Negative Problem of Input-conflicting Hallucination
by: Song, Jongyoon, et al.
Published: (2024)
by: Song, Jongyoon, et al.
Published: (2024)
From Static Benchmarks to Dynamic Protocol: Agent-Centric Text Anomaly Detection for Evaluating LLM Reasoning
by: Yoa, Seungdong, et al.
Published: (2026)
by: Yoa, Seungdong, et al.
Published: (2026)
Latent Preference Modeling for Cross-Session Personalized Tool Calling
by: Yoon, Yejin, et al.
Published: (2026)
by: Yoon, Yejin, et al.
Published: (2026)
Exploring Iterative Controllable Summarization with Large Language Models
by: Ryu, Sangwon, et al.
Published: (2024)
by: Ryu, Sangwon, et al.
Published: (2024)
Adapting Text-based Dialogue State Tracker for Spoken Dialogues
by: Yoon, Jaeseok, et al.
Published: (2023)
by: Yoon, Jaeseok, et al.
Published: (2023)
Key-Element-Informed sLLM Tuning for Document Summarization
by: Ryu, Sangwon, et al.
Published: (2024)
by: Ryu, Sangwon, et al.
Published: (2024)
Adaptive Planning for Multi-Attribute Controllable Summarization with Monte Carlo Tree Search
by: Ryu, Sangwon, et al.
Published: (2025)
by: Ryu, Sangwon, et al.
Published: (2025)
Behavior-Aware Item Modeling via Dynamic Procedural Solution Representations for Knowledge Tracing
by: Seo, Jun, et al.
Published: (2026)
by: Seo, Jun, et al.
Published: (2026)
M3-SLU: Evaluating Speaker-Attributed Reasoning in Multimodal Large Language Models
by: Kwon, Yejin, et al.
Published: (2025)
by: Kwon, Yejin, et al.
Published: (2025)
Evaluating LLMs for Police Decision-Making: A Framework Based on Police Action Scenarios
by: Lee, Sangyub, et al.
Published: (2026)
by: Lee, Sangyub, et al.
Published: (2026)
Devil's Advocate: Anticipatory Reflection for LLM Agents
by: Wang, Haoyu, et al.
Published: (2024)
by: Wang, Haoyu, et al.
Published: (2024)
DART: An AIGT Detector using AMR of Rephrased Text
by: Park, Hyeonchu, et al.
Published: (2024)
by: Park, Hyeonchu, et al.
Published: (2024)
Multi-Facet Blending for Faceted Query-by-Example Retrieval
by: Do, Heejin, et al.
Published: (2024)
by: Do, Heejin, et al.
Published: (2024)
RE-RAG: Improving Open-Domain QA Performance and Interpretability with Relevance Estimator in Retrieval-Augmented Generation
by: Kim, Kiseung, et al.
Published: (2024)
by: Kim, Kiseung, et al.
Published: (2024)
Flex-Judge: Text-Only Reasoning Unleashes Zero-Shot Multimodal Evaluators
by: Ko, Jongwoo, et al.
Published: (2025)
by: Ko, Jongwoo, et al.
Published: (2025)
Unsupervised Robust Cross-Lingual Entity Alignment via Neighbor Triple Matching with Entity and Relation Texts
by: Yoon, Soojin, et al.
Published: (2024)
by: Yoon, Soojin, et al.
Published: (2024)
CCQA: Generating Question from Solution Can Improve Inference-Time Reasoning in SLMs
by: Kim, Jin Young, et al.
Published: (2025)
by: Kim, Jin Young, et al.
Published: (2025)
Fine-Grained and Thematic Evaluation of LLMs in Social Deduction Game
by: Kim, Byungjun, et al.
Published: (2024)
by: Kim, Byungjun, et al.
Published: (2024)
The Devil Is in the Details: Tackling Unimodal Spurious Correlations for Generalizable Multimodal Reward Models
by: Li, Zichao, et al.
Published: (2025)
by: Li, Zichao, et al.
Published: (2025)
Evalverse: Unified and Accessible Library for Large Language Model Evaluation
by: Kim, Jihoo, et al.
Published: (2024)
by: Kim, Jihoo, et al.
Published: (2024)
Optimizing Long-Form Clinical Text Generation with Claim-Based Rewards
by: Jhaveri, Samyak, et al.
Published: (2025)
by: Jhaveri, Samyak, et al.
Published: (2025)
Beyond Learning: A Training-Free Alternative to Model Adaptation
by: Yoon, Namkyung, et al.
Published: (2026)
by: Yoon, Namkyung, et al.
Published: (2026)
Playing Devil's Advocate: Unmasking Toxicity and Vulnerabilities in Large Vision-Language Models
by: Erol, Abdulkadir, et al.
Published: (2025)
by: Erol, Abdulkadir, et al.
Published: (2025)
Reasoning Models Better Express Their Confidence
by: Yoon, Dongkeun, et al.
Published: (2025)
by: Yoon, Dongkeun, et al.
Published: (2025)
Understanding LLM Behavior in Multi-Target Cross-Lingual Summarization
by: Ryu, Sangwon, et al.
Published: (2026)
by: Ryu, Sangwon, et al.
Published: (2026)
Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs
by: Kim, Hyeonwoo, et al.
Published: (2024)
by: Kim, Hyeonwoo, et al.
Published: (2024)
Redefining Evaluation Standards: A Unified Framework for Evaluating the Korean Capabilities of Language Models
by: Lee, Hanwool, et al.
Published: (2025)
by: Lee, Hanwool, et al.
Published: (2025)
Designing and Evaluating Multi-Chatbot Interface for Human-AI Communication: Preliminary Findings from a Persuasion Task
by: Yoon, Sion, et al.
Published: (2024)
by: Yoon, Sion, et al.
Published: (2024)
CXR-LLAVA: a multimodal large language model for interpreting chest X-ray images
by: Lee, Seowoo, et al.
Published: (2023)
by: Lee, Seowoo, et al.
Published: (2023)
Towards Verifiable Text Generation with Symbolic References
by: Hennigen, Lucas Torroba, et al.
Published: (2023)
by: Hennigen, Lucas Torroba, et al.
Published: (2023)
ESPERANTO: Evaluating Synthesized Phrases to Enhance Robustness in AI Detection for Text Origination
by: Ayoobi, Navid, et al.
Published: (2024)
by: Ayoobi, Navid, et al.
Published: (2024)
Adverb Is the Key: Simple Text Data Augmentation with Adverb Deletion
by: Choi, Juhwan, et al.
Published: (2024)
by: Choi, Juhwan, et al.
Published: (2024)
Autoregressive Multi-trait Essay Scoring via Reinforcement Learning with Scoring-aware Multiple Rewards
by: Do, Heejin, et al.
Published: (2024)
by: Do, Heejin, et al.
Published: (2024)
Teach-to-Reason with Scoring: Self-Explainable Rationale-Driven Multi-Trait Essay Scoring
by: Do, Heejin, et al.
Published: (2025)
by: Do, Heejin, et al.
Published: (2025)
KatFishNet: Detecting LLM-Generated Korean Text through Linguistic Feature Analysis
by: Park, Shinwoo, et al.
Published: (2025)
by: Park, Shinwoo, et al.
Published: (2025)
Similar Items
-
SLM-Based Agentic AI with P-C-G: Optimized for Korean Tool Use
by: Jeon, Changhyun, et al.
Published: (2025) -
Ontology-Free General-Domain Knowledge Graph-to-Text Generation Dataset Synthesis using Large Language Model
by: Kim, Daehee, et al.
Published: (2024) -
Multi-Dimensional Optimization for Text Summarization via Reinforcement Learning
by: Ryu, Sangwon, et al.
Published: (2024) -
An Empirical Study of Group Conformity in Multi-Agent Systems
by: Choi, Min, et al.
Published: (2025) -
A Multifaceted Analysis of Negative Bias in Large Language Models through the Lens of Parametric Knowledge
by: Song, Jongyoon, et al.
Published: (2025)