SEOE: A Scalable and Reliable Semantic Evaluation Framework for Open Domain Event Detection
Fuente:
arXiv
Guardado en:
| Autores principales: | Lu, Yi-Fan, Mao, Xian-Ling, Lan, Tian, Zhang, Tong, Zhu, Yu-Shi, Huang, Heyan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Beyond Exact Match: Semantically Reassessing Event Extraction by Large Language Models
por: Lu, Yi-Fan, et al.
Publicado: (2024)
por: Lu, Yi-Fan, et al.
Publicado: (2024)
EXCEEDS: Extracting Complex Events via Nugget-based Grid Modeling in Scientific Domain
por: Lu, Yi-Fan, et al.
Publicado: (2024)
por: Lu, Yi-Fan, et al.
Publicado: (2024)
Multi-modal Retrieval Augmented Multi-modal Generation: Datasets, Evaluation Metrics and Strong Baselines
por: Ma, Zi-Ao, et al.
Publicado: (2024)
por: Ma, Zi-Ao, et al.
Publicado: (2024)
Automatic Evaluation for Text-to-image Generation: Task-decomposed Framework, Distilled Training, and Meta-evaluation Benchmark
por: Tu, Rong-Cheng, et al.
Publicado: (2024)
por: Tu, Rong-Cheng, et al.
Publicado: (2024)
DeepSurvey-Bench: Evaluating Academic Value of Automatically Generated Scientific Survey
por: Zhang, Guo-Biao, et al.
Publicado: (2026)
por: Zhang, Guo-Biao, et al.
Publicado: (2026)
Mix-Initiative Response Generation with Dynamic Prefix Tuning
por: Nie, Yuxiang, et al.
Publicado: (2024)
por: Nie, Yuxiang, et al.
Publicado: (2024)
CriticEval: Evaluating Large Language Model as Critic
por: Lan, Tian, et al.
Publicado: (2024)
por: Lan, Tian, et al.
Publicado: (2024)
T2I-Eval-R1: Reinforcement Learning-Driven Reasoning for Interpretable Text-to-Image Evaluation
por: Ma, Zi-Ao, et al.
Publicado: (2025)
por: Ma, Zi-Ao, et al.
Publicado: (2025)
Training Language Models to Critique With Multi-agent Feedback
por: Lan, Tian, et al.
Publicado: (2024)
por: Lan, Tian, et al.
Publicado: (2024)
MMWOZ: Building Multimodal Agent for Task-oriented Dialogue
por: Yang, Pu-Hai, et al.
Publicado: (2025)
por: Yang, Pu-Hai, et al.
Publicado: (2025)
Word Matters: What Influences Domain Adaptation in Summarization?
por: Li, Yinghao, et al.
Publicado: (2024)
por: Li, Yinghao, et al.
Publicado: (2024)
Building Knowledge-Grounded Dialogue Systems with Graph-Based Semantic Modeling
por: Yang, Yizhe, et al.
Publicado: (2022)
por: Yang, Yizhe, et al.
Publicado: (2022)
A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations
por: Lan, Tian, et al.
Publicado: (2025)
por: Lan, Tian, et al.
Publicado: (2025)
REGen: A Reliable Evaluation Framework for Generative Event Argument Extraction
por: Sharif, Omar, et al.
Publicado: (2025)
por: Sharif, Omar, et al.
Publicado: (2025)
Training-free Truthfulness Detection via Value Vectors in LLMs
por: Liu, Runheng, et al.
Publicado: (2025)
por: Liu, Runheng, et al.
Publicado: (2025)
MaP: A Unified Framework for Reliable Evaluation of Pre-training Dynamics
por: Wang, Jiapeng, et al.
Publicado: (2025)
por: Wang, Jiapeng, et al.
Publicado: (2025)
LLMs Judge Themselves: A Game-Theoretic Framework for Human-Aligned Evaluation
por: Yang, Gao, et al.
Publicado: (2025)
por: Yang, Gao, et al.
Publicado: (2025)
Beyond Literal Mapping: Benchmarking and Improving Non-Literal Translation Evaluation
por: Tian, Yanzhi, et al.
Publicado: (2026)
por: Tian, Yanzhi, et al.
Publicado: (2026)
ProtLLM: An Interleaved Protein-Language LLM with Protein-as-Word Pre-Training
por: Zhuo, Le, et al.
Publicado: (2024)
por: Zhuo, Le, et al.
Publicado: (2024)
Zero-Shot Detection of LLM-Generated Text via Implicit Reward Model
por: Liu, Runheng, et al.
Publicado: (2026)
por: Liu, Runheng, et al.
Publicado: (2026)
Utilizing and Calibrating Hindsight Process Rewards via Reinforcement with Mutual Information Self-Evaluation
por: Yao, Jiashu, et al.
Publicado: (2026)
por: Yao, Jiashu, et al.
Publicado: (2026)
A Distributed Collaborative Retrieval Framework Excelling in All Queries and Corpora based on Zero-shot Rank-Oriented Automatic Evaluation
por: Che, Tian-Yi, et al.
Publicado: (2024)
por: Che, Tian-Yi, et al.
Publicado: (2024)
Facilitating NSFW Text Detection in Open-Domain Dialogue Systems via Knowledge Distillation
por: Qiu, Huachuan, et al.
Publicado: (2023)
por: Qiu, Huachuan, et al.
Publicado: (2023)
Leveraging Open Information Extraction for More Robust Domain Transfer of Event Trigger Detection
por: Dukić, David, et al.
Publicado: (2023)
por: Dukić, David, et al.
Publicado: (2023)
Open-Domain Text Evaluation via Contrastive Distribution Methods
por: Lu, Sidi, et al.
Publicado: (2023)
por: Lu, Sidi, et al.
Publicado: (2023)
CEO: Corpus-based Open-Domain Event Ontology Induction
por: Xu, Nan, et al.
Publicado: (2023)
por: Xu, Nan, et al.
Publicado: (2023)
MEDAL: A Framework for Benchmarking LLMs as Multilingual Open-Domain Dialogue Evaluators
por: Mendonça, John, et al.
Publicado: (2025)
por: Mendonça, John, et al.
Publicado: (2025)
Debate, Reflect, and Distill: Multi-Agent Feedback with Tree-Structured Preference Optimization for Efficient Language Model Enhancement
por: Zhou, Xiaofeng, et al.
Publicado: (2025)
por: Zhou, Xiaofeng, et al.
Publicado: (2025)
Assessing LLM Reliability on Temporally Recent Open-Domain Questions
por: Krishnappa, Pushwitha, et al.
Publicado: (2026)
por: Krishnappa, Pushwitha, et al.
Publicado: (2026)
SA-MDKIF: A Scalable and Adaptable Medical Domain Knowledge Injection Framework for Large Language Models
por: Xu, Tianhan, et al.
Publicado: (2024)
por: Xu, Tianhan, et al.
Publicado: (2024)
Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models
por: Zhang, Fan, et al.
Publicado: (2024)
por: Zhang, Fan, et al.
Publicado: (2024)
CMNEE: A Large-Scale Document-Level Event Extraction Dataset based on Open-Source Chinese Military News
por: Zhu, Mengna, et al.
Publicado: (2024)
por: Zhu, Mengna, et al.
Publicado: (2024)
Facilitating Pornographic Text Detection for Open-Domain Dialogue Systems via Knowledge Distillation of Large Language Models
por: Qiu, Huachuan, et al.
Publicado: (2024)
por: Qiu, Huachuan, et al.
Publicado: (2024)
How Far Are We? Systematic Evaluation of LLMs vs. Human Experts in Mathematical Contest in Modeling
por: Liu, Yuhang, et al.
Publicado: (2026)
por: Liu, Yuhang, et al.
Publicado: (2026)
Forgetting Curve: A Reliable Method for Evaluating Memorization Capability for Long-context Models
por: Liu, Xinyu, et al.
Publicado: (2024)
por: Liu, Xinyu, et al.
Publicado: (2024)
Controllable and Diverse Data Augmentation with Large Language Model for Low-Resource Open-Domain Dialogue Generation
por: Liu, Zhenhua, et al.
Publicado: (2024)
por: Liu, Zhenhua, et al.
Publicado: (2024)
Optimizing Chain-of-Thought Reasoning: Tackling Arranging Bottleneck via Plan Augmentation
por: Qiu, Yuli, et al.
Publicado: (2024)
por: Qiu, Yuli, et al.
Publicado: (2024)
Deterministic Reversible Data Augmentation for Neural Machine Translation
por: Yao, Jiashu, et al.
Publicado: (2024)
por: Yao, Jiashu, et al.
Publicado: (2024)
MASS-RAG: Multi-Agent Synthesis Retrieval-Augmented Generation
por: Xiao, Xingchen, et al.
Publicado: (2026)
por: Xiao, Xingchen, et al.
Publicado: (2026)
On the Benchmarking of LLMs for Open-Domain Dialogue Evaluation
por: Mendonça, John, et al.
Publicado: (2024)
por: Mendonça, John, et al.
Publicado: (2024)
Ejemplares similares
-
Beyond Exact Match: Semantically Reassessing Event Extraction by Large Language Models
por: Lu, Yi-Fan, et al.
Publicado: (2024) -
EXCEEDS: Extracting Complex Events via Nugget-based Grid Modeling in Scientific Domain
por: Lu, Yi-Fan, et al.
Publicado: (2024) -
Multi-modal Retrieval Augmented Multi-modal Generation: Datasets, Evaluation Metrics and Strong Baselines
por: Ma, Zi-Ao, et al.
Publicado: (2024) -
Automatic Evaluation for Text-to-image Generation: Task-decomposed Framework, Distilled Training, and Meta-evaluation Benchmark
por: Tu, Rong-Cheng, et al.
Publicado: (2024) -
DeepSurvey-Bench: Evaluating Academic Value of Automatically Generated Scientific Survey
por: Zhang, Guo-Biao, et al.
Publicado: (2026)