CLAVE: An Adaptive Framework for Evaluating Values of LLM Generated Responses
Fuente:
arXiv
Saved in:
| Main Authors: | Yao, Jing, Yi, Xiaoyuan, Xie, Xing |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
by: Lee, Jaehyeok, et al.
Published: (2026)
by: Lee, Jaehyeok, et al.
Published: (2026)
AdAEM: An Adaptively and Automated Extensible Measurement of LLMs' Value Difference
by: Yao, Jing, et al.
Published: (2025)
by: Yao, Jing, et al.
Published: (2025)
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
by: Choi, Sooyung, et al.
Published: (2025)
by: Choi, Sooyung, et al.
Published: (2025)
Raising the Bar: Investigating the Values of Large Language Models via Generative Evolving Testing
by: Jiang, Han, et al.
Published: (2024)
by: Jiang, Han, et al.
Published: (2024)
Beyond Human Norms: Unveiling Unique Values of Large Language Models through Interdisciplinary Approaches
by: Biedma, Pablo, et al.
Published: (2024)
by: Biedma, Pablo, et al.
Published: (2024)
Leveraging Implicit Sentiments: Enhancing Reliability and Validity in Psychological Trait Evaluation of LLMs
by: Ma, Huanhuan, et al.
Published: (2025)
by: Ma, Huanhuan, et al.
Published: (2025)
PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization
by: Jiang, Han, et al.
Published: (2025)
by: Jiang, Han, et al.
Published: (2025)
Denevil: Towards Deciphering and Navigating the Ethical Values of Large Language Models via Instruction Learning
by: Duan, Shitong, et al.
Published: (2023)
by: Duan, Shitong, et al.
Published: (2023)
Contextualized Privacy Defense for LLM Agents
by: Wen, Yule, et al.
Published: (2026)
by: Wen, Yule, et al.
Published: (2026)
Elephant in the Room: Unveiling the Impact of Reward Model Quality in Alignment
by: Liu, Yan, et al.
Published: (2024)
by: Liu, Yan, et al.
Published: (2024)
MotiveBench: How Far Are We From Human-Like Motivational Reasoning in Large Language Models?
by: Yong, Xixian, et al.
Published: (2025)
by: Yong, Xixian, et al.
Published: (2025)
The Incomplete Bridge: How AI Research (Mis)Engages with Psychology
by: Jiang, Han, et al.
Published: (2025)
by: Jiang, Han, et al.
Published: (2025)
On the Essence and Prospect: An Investigation of Alignment Approaches for Big Models
by: Wang, Xinpeng, et al.
Published: (2024)
by: Wang, Xinpeng, et al.
Published: (2024)
Integrated Framework for LLM Evaluation with Answer Generation
by: Lee, Sujeong, et al.
Published: (2025)
by: Lee, Sujeong, et al.
Published: (2025)
IROTE: Human-like Traits Elicitation of Large Language Model via In-Context Self-Reflective Optimization
by: Bai, Yuzhuo, et al.
Published: (2025)
by: Bai, Yuzhuo, et al.
Published: (2025)
User Perceptions vs. Proxy LLM Judges: Privacy and Helpfulness in LLM Responses to Privacy-Sensitive Scenarios
by: Wu, Xiaoyuan, et al.
Published: (2025)
by: Wu, Xiaoyuan, et al.
Published: (2025)
ToolNet: Connecting Large Language Models with Massive Tools via Tool Graph
by: Liu, Xukun, et al.
Published: (2024)
by: Liu, Xukun, et al.
Published: (2024)
Evaluating LLM-Generated Versus Human-Authored Responses in Role-Play Dialogues
by: Lu, Dongxu, et al.
Published: (2025)
by: Lu, Dongxu, et al.
Published: (2025)
Domain-Specific Data Generation Framework for RAG Adaptation
by: Tian, Chris Xing, et al.
Published: (2025)
by: Tian, Chris Xing, et al.
Published: (2025)
Negating Negatives: Alignment with Human Negative Samples via Distributional Dispreference Optimization
by: Duan, Shitong, et al.
Published: (2024)
by: Duan, Shitong, et al.
Published: (2024)
STED and Consistency Scoring: A Framework for Evaluating LLM Structured Output Reliability
by: Wang, Guanghui, et al.
Published: (2025)
by: Wang, Guanghui, et al.
Published: (2025)
Citations and Trust in LLM Generated Responses
by: Ding, Yifan, et al.
Published: (2025)
by: Ding, Yifan, et al.
Published: (2025)
A Survey on Evaluation of Large Language Models
by: Chang, Yupeng, et al.
Published: (2023)
by: Chang, Yupeng, et al.
Published: (2023)
DeepSurvey-Bench: Evaluating Academic Value of Automatically Generated Scientific Survey
by: Zhang, Guo-Biao, et al.
Published: (2026)
by: Zhang, Guo-Biao, et al.
Published: (2026)
Variation is the Key: A Variation-Based Framework for LLM-Generated Text Detection
by: Li, Xuecong, et al.
Published: (2026)
by: Li, Xuecong, et al.
Published: (2026)
Developing A Framework to Support Human Evaluation of Bias in Generated Free Response Text
by: Healey, Jennifer, et al.
Published: (2025)
by: Healey, Jennifer, et al.
Published: (2025)
On the Dynamics of Multi-Agent LLM Communities Driven by Value Diversity
by: Huang, Muhua, et al.
Published: (2025)
by: Huang, Muhua, et al.
Published: (2025)
An Investigation into Value Misalignment in LLM-Generated Texts for Cultural Heritage
by: Bu, Fan, et al.
Published: (2025)
by: Bu, Fan, et al.
Published: (2025)
Towards Reliable Detection of LLM-Generated Texts: A Comprehensive Evaluation Framework with CUDRT
by: Tao, Zhen, et al.
Published: (2024)
by: Tao, Zhen, et al.
Published: (2024)
SCORE: A Semantic Evaluation Framework for Generative Document Parsing
by: Li, Renyu, et al.
Published: (2025)
by: Li, Renyu, et al.
Published: (2025)
MVPBench: A Benchmark and Fine-Tuning Framework for Aligning Large Language Models with Diverse Human Values
by: Liang, Yao, et al.
Published: (2025)
by: Liang, Yao, et al.
Published: (2025)
PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
by: Wang, Yidong, et al.
Published: (2023)
by: Wang, Yidong, et al.
Published: (2023)
MMA-ASIA: A Multilingual and Multimodal Alignment Framework for Culturally-Grounded Evaluation
by: Zheng, Weihua, et al.
Published: (2025)
by: Zheng, Weihua, et al.
Published: (2025)
A Graph-Enhanced Defense Framework for Explainable Fake News Detection with LLM
by: Wang, Bo, et al.
Published: (2026)
by: Wang, Bo, et al.
Published: (2026)
LLM-Oriented Token-Adaptive Knowledge Distillation
by: Xie, Xurong, et al.
Published: (2025)
by: Xie, Xurong, et al.
Published: (2025)
Embedding-Informed Adaptive Retrieval-Augmented Generation of Large Language Models
by: Huang, Chengkai, et al.
Published: (2024)
by: Huang, Chengkai, et al.
Published: (2024)
Beyond Pointwise Scores: Decomposed Criteria-Based Evaluation of LLM Responses
by: Yu, Fangyi, et al.
Published: (2025)
by: Yu, Fangyi, et al.
Published: (2025)
Confident Rankings with Fewer Items: Adaptive LLM Evaluation with Continuous Scores
by: Balkır, Esma, et al.
Published: (2026)
by: Balkır, Esma, et al.
Published: (2026)
Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks
by: Li, Peiyu, et al.
Published: (2025)
by: Li, Peiyu, et al.
Published: (2025)
SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text Generation
by: Liu, Xiaoze, et al.
Published: (2024)
by: Liu, Xiaoze, et al.
Published: (2024)
Similar Items
-
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
by: Lee, Jaehyeok, et al.
Published: (2026) -
AdAEM: An Adaptively and Automated Extensible Measurement of LLMs' Value Difference
by: Yao, Jing, et al.
Published: (2025) -
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
by: Choi, Sooyung, et al.
Published: (2025) -
Raising the Bar: Investigating the Values of Large Language Models via Generative Evolving Testing
by: Jiang, Han, et al.
Published: (2024) -
Beyond Human Norms: Unveiling Unique Values of Large Language Models through Interdisciplinary Approaches
by: Biedma, Pablo, et al.
Published: (2024)