DEE: Dual-stage Explainable Evaluation Method for Text Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Shenyu, Li, Yu, Wu, Rui, Huang, Xiutian, Chen, Yongrui, Xu, Wenhao, Qi, Guilin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MATEval: A Multi-Agent Discussion Framework for Advancing Open-Ended Text Evaluation
von: Li, Yu, et al.
Veröffentlicht: (2024)
von: Li, Yu, et al.
Veröffentlicht: (2024)
StressEval: Failure-Driven Dynamic Benchmarking for Knowledge-Intensive Reasoning in Large Language Models
von: Chen, Yongrui, et al.
Veröffentlicht: (2026)
von: Chen, Yongrui, et al.
Veröffentlicht: (2026)
K-DeCore: Facilitating Knowledge Transfer in Continual Structured Knowledge Reasoning via Knowledge Decoupling
von: Chen, Yongrui, et al.
Veröffentlicht: (2025)
von: Chen, Yongrui, et al.
Veröffentlicht: (2025)
Magic Mushroom: A Customizable Benchmark for Fine-grained Analysis of Retrieval Noise Erosion in RAG Systems
von: Zhang, Yuxin, et al.
Veröffentlicht: (2025)
von: Zhang, Yuxin, et al.
Veröffentlicht: (2025)
DoG-Instruct: Towards Premium Instruction-Tuning Data via Text-Grounded Instruction Wrapping
von: Chen, Yongrui, et al.
Veröffentlicht: (2023)
von: Chen, Yongrui, et al.
Veröffentlicht: (2023)
Exploring the Impact of Table-to-Text Methods on Augmenting LLM-based Question Answering with Domain Hybrid Data
von: Min, Dehai, et al.
Veröffentlicht: (2024)
von: Min, Dehai, et al.
Veröffentlicht: (2024)
Pandora: Leveraging Code-driven Knowledge Transfer for Unified Structured Knowledge Reasoning
von: Chen, Yongrui, et al.
Veröffentlicht: (2025)
von: Chen, Yongrui, et al.
Veröffentlicht: (2025)
Pandora: A Code-Driven Large Language Model Agent for Unified Reasoning Across Diverse Structured Knowledge
von: Chen, Yongrui, et al.
Veröffentlicht: (2025)
von: Chen, Yongrui, et al.
Veröffentlicht: (2025)
Can LLMs Evaluate Complex Attribution in QA? Automatic Benchmarking using Knowledge Graphs
von: Hu, Nan, et al.
Veröffentlicht: (2024)
von: Hu, Nan, et al.
Veröffentlicht: (2024)
C$^3$TG: Conflict-aware, Composite, and Collaborative Controlled Text Generation
von: Li, Yu, et al.
Veröffentlicht: (2025)
von: Li, Yu, et al.
Veröffentlicht: (2025)
TIGERScore: Towards Building Explainable Metric for All Text Generation Tasks
von: Jiang, Dongfu, et al.
Veröffentlicht: (2023)
von: Jiang, Dongfu, et al.
Veröffentlicht: (2023)
After Retrieval, Before Generation: Enhancing the Trustworthiness of Large Language Models in Retrieval-Augmented Generation
von: Dai, Xinbang, et al.
Veröffentlicht: (2025)
von: Dai, Xinbang, et al.
Veröffentlicht: (2025)
Table-r1: Self-supervised and Reinforcement Learning for Program-based Table Reasoning in Small Language Models
von: Jin, Rihui, et al.
Veröffentlicht: (2025)
von: Jin, Rihui, et al.
Veröffentlicht: (2025)
AgentBank: Towards Generalized LLM Agents via Fine-Tuning on 50000+ Interaction Trajectories
von: Song, Yifan, et al.
Veröffentlicht: (2024)
von: Song, Yifan, et al.
Veröffentlicht: (2024)
ORCHID: A Chinese Debate Corpus for Target-Independent Stance Detection and Argumentative Dialogue Summarization
von: Zhao, Xiutian, et al.
Veröffentlicht: (2024)
von: Zhao, Xiutian, et al.
Veröffentlicht: (2024)
Measuring the Inconsistency of Large Language Models in Preferential Ranking
von: Zhao, Xiutian, et al.
Veröffentlicht: (2024)
von: Zhao, Xiutian, et al.
Veröffentlicht: (2024)
An Electoral Approach to Diversify LLM-based Multi-Agent Collective Decision-Making
von: Zhao, Xiutian, et al.
Veröffentlicht: (2024)
von: Zhao, Xiutian, et al.
Veröffentlicht: (2024)
Question Answering Over Spatio-Temporal Knowledge Graph
von: Dai, Xinbang, et al.
Veröffentlicht: (2024)
von: Dai, Xinbang, et al.
Veröffentlicht: (2024)
A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations
von: Lan, Tian, et al.
Veröffentlicht: (2025)
von: Lan, Tian, et al.
Veröffentlicht: (2025)
MIKE: A New Benchmark for Fine-grained Multimodal Entity Knowledge Editing
von: Li, Jiaqi, et al.
Veröffentlicht: (2024)
von: Li, Jiaqi, et al.
Veröffentlicht: (2024)
SyntaxShap: Syntax-aware Explainability Method for Text Generation
von: Amara, Kenza, et al.
Veröffentlicht: (2024)
von: Amara, Kenza, et al.
Veröffentlicht: (2024)
LM$^2$otifs : An Explainable Framework for Machine-Generated Texts Detection
von: Zheng, Xu, et al.
Veröffentlicht: (2025)
von: Zheng, Xu, et al.
Veröffentlicht: (2025)
SPOR: A Comprehensive and Practical Evaluation Method for Compositional Generalization in Data-to-Text Generation
von: Xu, Ziyao, et al.
Veröffentlicht: (2024)
von: Xu, Ziyao, et al.
Veröffentlicht: (2024)
HeGTa: Leveraging Heterogeneous Graph-enhanced Large Language Models for Few-shot Complex Table Understanding
von: Jin, Rihui, et al.
Veröffentlicht: (2024)
von: Jin, Rihui, et al.
Veröffentlicht: (2024)
StyleDecipher: Robust and Explainable Detection of LLM-Generated Texts with Stylistic Analysis
von: Li, Siyuan, et al.
Veröffentlicht: (2025)
von: Li, Siyuan, et al.
Veröffentlicht: (2025)
PRIMO: Progressive Induction for Multi-hop Open Rule Generation
von: Liu, Jianyu, et al.
Veröffentlicht: (2024)
von: Liu, Jianyu, et al.
Veröffentlicht: (2024)
Embedding Ontologies via Incorporating Extensional and Intensional Knowledge
von: Wang, Keyu, et al.
Veröffentlicht: (2024)
von: Wang, Keyu, et al.
Veröffentlicht: (2024)
OneEval: Benchmarking LLM Knowledge-intensive Reasoning over Diverse Knowledge Bases
von: Chen, Yongrui, et al.
Veröffentlicht: (2025)
von: Chen, Yongrui, et al.
Veröffentlicht: (2025)
Discovering and Causally Validating Emotion-Sensitive Neurons in Large Audio-Language Models
von: Zhao, Xiutian, et al.
Veröffentlicht: (2026)
von: Zhao, Xiutian, et al.
Veröffentlicht: (2026)
Detecting Machine-Generated Texts: Not Just "AI vs Humans" and Explainability is Complicated
von: Ji, Jiazhou, et al.
Veröffentlicht: (2024)
von: Ji, Jiazhou, et al.
Veröffentlicht: (2024)
Decision-Oriented Text Evaluation
von: Huang, Yu-Shiang, et al.
Veröffentlicht: (2025)
von: Huang, Yu-Shiang, et al.
Veröffentlicht: (2025)
Challenges and Opportunities in Text Generation Explainability
von: Amara, Kenza, et al.
Veröffentlicht: (2024)
von: Amara, Kenza, et al.
Veröffentlicht: (2024)
Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement
von: Xiong, Weimin, et al.
Veröffentlicht: (2024)
von: Xiong, Weimin, et al.
Veröffentlicht: (2024)
UniHGKR: Unified Instruction-aware Heterogeneous Knowledge Retrievers
von: Min, Dehai, et al.
Veröffentlicht: (2024)
von: Min, Dehai, et al.
Veröffentlicht: (2024)
SQLStructEval: Structural Evaluation of LLM Text-to-SQL Generation
von: Zhou, Yixi, et al.
Veröffentlicht: (2026)
von: Zhou, Yixi, et al.
Veröffentlicht: (2026)
Neuron-Level Emotion Control in Speech-Generative Large Audio-Language Models
von: Zhao, Xiutian, et al.
Veröffentlicht: (2026)
von: Zhao, Xiutian, et al.
Veröffentlicht: (2026)
CoTKR: Chain-of-Thought Enhanced Knowledge Rewriting for Complex Knowledge Graph Question Answering
von: Wu, Yike, et al.
Veröffentlicht: (2024)
von: Wu, Yike, et al.
Veröffentlicht: (2024)
Holistic Evaluation for Interleaved Text-and-Image Generation
von: Liu, Minqian, et al.
Veröffentlicht: (2024)
von: Liu, Minqian, et al.
Veröffentlicht: (2024)
EMMM, Explain Me My Model! Explainable Machine Generated Text Detection in Dialogues
von: Yuan, Angela Yifei, et al.
Veröffentlicht: (2025)
von: Yuan, Angela Yifei, et al.
Veröffentlicht: (2025)
ExPerT: Effective and Explainable Evaluation of Personalized Long-Form Text Generation
von: Salemi, Alireza, et al.
Veröffentlicht: (2025)
von: Salemi, Alireza, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MATEval: A Multi-Agent Discussion Framework for Advancing Open-Ended Text Evaluation
von: Li, Yu, et al.
Veröffentlicht: (2024) -
StressEval: Failure-Driven Dynamic Benchmarking for Knowledge-Intensive Reasoning in Large Language Models
von: Chen, Yongrui, et al.
Veröffentlicht: (2026) -
K-DeCore: Facilitating Knowledge Transfer in Continual Structured Knowledge Reasoning via Knowledge Decoupling
von: Chen, Yongrui, et al.
Veröffentlicht: (2025) -
Magic Mushroom: A Customizable Benchmark for Fine-grained Analysis of Retrieval Noise Erosion in RAG Systems
von: Zhang, Yuxin, et al.
Veröffentlicht: (2025) -
DoG-Instruct: Towards Premium Instruction-Tuning Data via Text-Grounded Instruction Wrapping
von: Chen, Yongrui, et al.
Veröffentlicht: (2023)