RCScore: Quantifying Response Consistency in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jang, Dongjun, Ahn, Youngchae, Shin, Hyopil |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
P-CoT: A Pedagogically-motivated Participatory Chain-of-Thought Prompting for Phonological Reasoning in LLMs
von: Jang, Dongjun, et al.
Veröffentlicht: (2025)
von: Jang, Dongjun, et al.
Veröffentlicht: (2025)
KIT-19: A Comprehensive Korean Instruction Toolkit on 19 Tasks for Fine-Tuning Korean Large Language Models
von: Jang, Dongjun, et al.
Veröffentlicht: (2024)
von: Jang, Dongjun, et al.
Veröffentlicht: (2024)
A Study on How Attention Scores in the BERT Model are Aware of Lexical Categories in Syntactic and Semantic Tasks on the GLUE Benchmark
von: Jang, Dongjun, et al.
Veröffentlicht: (2024)
von: Jang, Dongjun, et al.
Veröffentlicht: (2024)
DAHL: Domain-specific Automated Hallucination Evaluation of Long-Form Text through a Benchmark Dataset in Biomedicine
von: Seo, Jean, et al.
Veröffentlicht: (2024)
von: Seo, Jean, et al.
Veröffentlicht: (2024)
KoBALT: Korean Benchmark For Advanced Linguistic Tasks
von: Shin, Hyopil, et al.
Veröffentlicht: (2025)
von: Shin, Hyopil, et al.
Veröffentlicht: (2025)
CARBD-Ko: A Contextually Annotated Review Benchmark Dataset for Aspect-Level Sentiment Classification in Korean
von: Jang, Dongjun, et al.
Veröffentlicht: (2024)
von: Jang, Dongjun, et al.
Veröffentlicht: (2024)
Korean Bio-Medical Corpus (KBMC) for Medical Named Entity Recognition
von: Byun, Sungjoo, et al.
Veröffentlicht: (2024)
von: Byun, Sungjoo, et al.
Veröffentlicht: (2024)
How does a Language-Specific Tokenizer affect LLMs?
von: Seo, Jean, et al.
Veröffentlicht: (2025)
von: Seo, Jean, et al.
Veröffentlicht: (2025)
MoFE: Mixture of Frozen Experts Architecture
von: Seo, Jean, et al.
Veröffentlicht: (2025)
von: Seo, Jean, et al.
Veröffentlicht: (2025)
Self-Training Large Language Models with Confident Reasoning
von: Jang, Hyosoon, et al.
Veröffentlicht: (2025)
von: Jang, Hyosoon, et al.
Veröffentlicht: (2025)
A General Method for Detecting Information Generated by Large Language Models
von: Mao, Minjia, et al.
Veröffentlicht: (2025)
von: Mao, Minjia, et al.
Veröffentlicht: (2025)
RELIC: Investigating Large Language Model Responses using Self-Consistency
von: Cheng, Furui, et al.
Veröffentlicht: (2023)
von: Cheng, Furui, et al.
Veröffentlicht: (2023)
Watermarking Low-entropy Generation for Large Language Models: An Unbiased and Low-risk Method
von: Mao, Minjia, et al.
Veröffentlicht: (2024)
von: Mao, Minjia, et al.
Veröffentlicht: (2024)
Quantifying Language Disparities in Multilingual Large Language Models
von: Hu, Songbo, et al.
Veröffentlicht: (2025)
von: Hu, Songbo, et al.
Veröffentlicht: (2025)
Quantifying Generalization Complexity for Large Language Models
von: Qi, Zhenting, et al.
Veröffentlicht: (2024)
von: Qi, Zhenting, et al.
Veröffentlicht: (2024)
KITE: A Benchmark for Evaluating Korean Instruction-Following Abilities in Large Language Models
von: Kim, Dongjun, et al.
Veröffentlicht: (2025)
von: Kim, Dongjun, et al.
Veröffentlicht: (2025)
Consistency of Responses and Continuations Generated by Large Language Models on Social Media
von: Xu, Wentao, et al.
Veröffentlicht: (2025)
von: Xu, Wentao, et al.
Veröffentlicht: (2025)
Cross-Lingual Response Consistency in Large Language Models: An ILR-Informed Evaluation of Claude Across Six Languages
von: Baluta, Camelia
Veröffentlicht: (2026)
von: Baluta, Camelia
Veröffentlicht: (2026)
Enhancing Automatic Term Extraction with Large Language Models via Syntactic Retrieval
von: Chun, Yongchan, et al.
Veröffentlicht: (2025)
von: Chun, Yongchan, et al.
Veröffentlicht: (2025)
Calibrating Large Language Models with Sample Consistency
von: Lyu, Qing, et al.
Veröffentlicht: (2024)
von: Lyu, Qing, et al.
Veröffentlicht: (2024)
Recursive Think-Answer Process for LLMs and VLMs
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2026)
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2026)
Zero-Shot Multi-Hop Question Answering via Monte-Carlo Tree Search with Large Language Models
von: Lee, Seongmin, et al.
Veröffentlicht: (2024)
von: Lee, Seongmin, et al.
Veröffentlicht: (2024)
CLLMs: Consistency Large Language Models
von: Kou, Siqi, et al.
Veröffentlicht: (2024)
von: Kou, Siqi, et al.
Veröffentlicht: (2024)
Logical Consistency of Large Language Models in Fact-checking
von: Ghosh, Bishwamittra, et al.
Veröffentlicht: (2024)
von: Ghosh, Bishwamittra, et al.
Veröffentlicht: (2024)
NILE: Internal Consistency Alignment in Large Language Models
von: Hu, Minda, et al.
Veröffentlicht: (2024)
von: Hu, Minda, et al.
Veröffentlicht: (2024)
SCAR: Data Selection via Style Consistency-Aware Response Ranking for Efficient Instruction-Tuning of Large Language Models
von: Li, Zhuang, et al.
Veröffentlicht: (2024)
von: Li, Zhuang, et al.
Veröffentlicht: (2024)
Chain of Empathy: Enhancing Empathetic Response of Large Language Models Based on Psychotherapy Models
von: Lee, Yoon Kyung, et al.
Veröffentlicht: (2023)
von: Lee, Yoon Kyung, et al.
Veröffentlicht: (2023)
Ranked Voting based Self-Consistency of Large Language Models
von: Wang, Weiqin, et al.
Veröffentlicht: (2025)
von: Wang, Weiqin, et al.
Veröffentlicht: (2025)
Improving the Robustness of Large Language Models via Consistency Alignment
von: Zhao, Yukun, et al.
Veröffentlicht: (2024)
von: Zhao, Yukun, et al.
Veröffentlicht: (2024)
Exploring the Factual Consistency in Dialogue Comprehension of Large Language Models
von: She, Shuaijie, et al.
Veröffentlicht: (2023)
von: She, Shuaijie, et al.
Veröffentlicht: (2023)
Large Language Models Lack Understanding of Character Composition of Words
von: Shin, Andrew, et al.
Veröffentlicht: (2024)
von: Shin, Andrew, et al.
Veröffentlicht: (2024)
Unmasking and Quantifying Racial Bias of Large Language Models in Medical Report Generation
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
Quantifying Conversational Reliability of Large Language Models under Multi-Turn Interaction
von: Myung, Jiyoon
Veröffentlicht: (2026)
von: Myung, Jiyoon
Veröffentlicht: (2026)
A Large Collection of Model-generated Contradictory Responses for Consistency-aware Dialogue Systems
von: Sato, Shiki, et al.
Veröffentlicht: (2024)
von: Sato, Shiki, et al.
Veröffentlicht: (2024)
Personality Vector: Modulating Personality of Large Language Models by Model Merging
von: Sun, Seungjong, et al.
Veröffentlicht: (2025)
von: Sun, Seungjong, et al.
Veröffentlicht: (2025)
Semantic Consistency for Assuring Reliability of Large Language Models
von: Raj, Harsh, et al.
Veröffentlicht: (2023)
von: Raj, Harsh, et al.
Veröffentlicht: (2023)
Quantifying Uncertainty in Natural Language Explanations of Large Language Models for Question Answering
von: Li, Yangyi, et al.
Veröffentlicht: (2025)
von: Li, Yangyi, et al.
Veröffentlicht: (2025)
CM-Align: Consistency-based Multilingual Alignment for Large Language Models
von: Zhang, Xue, et al.
Veröffentlicht: (2025)
von: Zhang, Xue, et al.
Veröffentlicht: (2025)
Internal Consistency and Self-Feedback in Large Language Models: A Survey
von: Liang, Xun, et al.
Veröffentlicht: (2024)
von: Liang, Xun, et al.
Veröffentlicht: (2024)
Doc-PP: Document Policy Preservation Benchmark for Large Vision-Language Models
von: Jang, Haeun, et al.
Veröffentlicht: (2026)
von: Jang, Haeun, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
P-CoT: A Pedagogically-motivated Participatory Chain-of-Thought Prompting for Phonological Reasoning in LLMs
von: Jang, Dongjun, et al.
Veröffentlicht: (2025) -
KIT-19: A Comprehensive Korean Instruction Toolkit on 19 Tasks for Fine-Tuning Korean Large Language Models
von: Jang, Dongjun, et al.
Veröffentlicht: (2024) -
A Study on How Attention Scores in the BERT Model are Aware of Lexical Categories in Syntactic and Semantic Tasks on the GLUE Benchmark
von: Jang, Dongjun, et al.
Veröffentlicht: (2024) -
DAHL: Domain-specific Automated Hallucination Evaluation of Long-Form Text through a Benchmark Dataset in Biomedicine
von: Seo, Jean, et al.
Veröffentlicht: (2024) -
KoBALT: Korean Benchmark For Advanced Linguistic Tasks
von: Shin, Hyopil, et al.
Veröffentlicht: (2025)