Rethinking Meeting Effectiveness: A Benchmark and Framework for Temporal Fine-grained Automatic Meeting Effectiveness Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Yihang, Chu, Chenhui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EMS: Efficient and Effective Massively Multilingual Sentence Embedding Learning
von: Mao, Zhuoyuan, et al.
Veröffentlicht: (2022)
von: Mao, Zhuoyuan, et al.
Veröffentlicht: (2022)
When Large Language Models Meet Speech: A Survey on Integration Approaches
von: Yang, Zhengdong, et al.
Veröffentlicht: (2025)
von: Yang, Zhengdong, et al.
Veröffentlicht: (2025)
FActBench: A Benchmark for Fine-grained Automatic Evaluation of LLM-Generated Text in the Medical Domain
von: Afzal, Anum, et al.
Veröffentlicht: (2025)
von: Afzal, Anum, et al.
Veröffentlicht: (2025)
CFSafety: Comprehensive Fine-grained Safety Assessment for LLMs
von: Liu, Zhihao, et al.
Veröffentlicht: (2024)
von: Liu, Zhihao, et al.
Veröffentlicht: (2024)
Evaluating the Effectiveness of Black-Box Prompt Optimization as the Scale of LLMs Continues to Grow
von: Zhou, Ziyu, et al.
Veröffentlicht: (2025)
von: Zhou, Ziyu, et al.
Veröffentlicht: (2025)
CREAM: Comparison-Based Reference-Free ELO-Ranked Automatic Evaluation for Meeting Summarization
von: Gong, Ziwei, et al.
Veröffentlicht: (2024)
von: Gong, Ziwei, et al.
Veröffentlicht: (2024)
A Comprehensive Analysis of the Effectiveness of Large Language Models as Automatic Dialogue Evaluators
von: Zhang, Chen, et al.
Veröffentlicht: (2023)
von: Zhang, Chen, et al.
Veröffentlicht: (2023)
LoRA-FA: Efficient and Effective Low Rank Representation Fine-tuning
von: Zhang, Longteng, et al.
Veröffentlicht: (2023)
von: Zhang, Longteng, et al.
Veröffentlicht: (2023)
MathOPEval: A Fine-grained Evaluation Benchmark for Visual Operations of MLLMs in Mathematical Reasoning
von: Li, Xiaoyuan, et al.
Veröffentlicht: (2025)
von: Li, Xiaoyuan, et al.
Veröffentlicht: (2025)
Generalist Scanner Meets Specialist Locator: A Synergistic Coarse-to-Fine Framework for Robust GUI Grounding
von: Li, Zhecheng, et al.
Veröffentlicht: (2025)
von: Li, Zhecheng, et al.
Veröffentlicht: (2025)
When Scale Meets Diversity: Evaluating Language Models on Fine-Grained Multilingual Claim Verification
von: Shcharbakova, Hanna, et al.
Veröffentlicht: (2025)
von: Shcharbakova, Hanna, et al.
Veröffentlicht: (2025)
Factcheck-Bench: Fine-Grained Evaluation Benchmark for Automatic Fact-checkers
von: Wang, Yuxia, et al.
Veröffentlicht: (2023)
von: Wang, Yuxia, et al.
Veröffentlicht: (2023)
Texts or Images? A Fine-grained Analysis on the Effectiveness of Input Representations and Models for Table Question Answering
von: Zhou, Wei, et al.
Veröffentlicht: (2025)
von: Zhou, Wei, et al.
Veröffentlicht: (2025)
TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
von: Cai, Mu, et al.
Veröffentlicht: (2024)
von: Cai, Mu, et al.
Veröffentlicht: (2024)
What's under the hood: Investigating Automatic Metrics on Meeting Summarization
von: Kirstein, Frederic, et al.
Veröffentlicht: (2024)
von: Kirstein, Frederic, et al.
Veröffentlicht: (2024)
M$^3$FinMeeting: A Multilingual, Multi-Sector, and Multi-Task Financial Meeting Understanding Evaluation Dataset
von: Zhu, Jie, et al.
Veröffentlicht: (2025)
von: Zhu, Jie, et al.
Veröffentlicht: (2025)
Rethinking the Unsolvable: When In-Context Search Meets Test-Time Scaling
von: Xia, Fanzeng, et al.
Veröffentlicht: (2025)
von: Xia, Fanzeng, et al.
Veröffentlicht: (2025)
A Dual-Perspective NLG Meta-Evaluation Framework with Automatic Benchmark and Better Interpretability
von: Hu, Xinyu, et al.
Veröffentlicht: (2025)
von: Hu, Xinyu, et al.
Veröffentlicht: (2025)
Looking for the Bottleneck in Fine-grained Temporal Relation Classification
von: Sousa, Hugo, et al.
Veröffentlicht: (2026)
von: Sousa, Hugo, et al.
Veröffentlicht: (2026)
When LLMs Meet Cunning Texts: A Fallacy Understanding Benchmark for Large Language Models
von: Li, Yinghui, et al.
Veröffentlicht: (2024)
von: Li, Yinghui, et al.
Veröffentlicht: (2024)
PersLitEval: Fine-grained Benchmark and Evaluation of LLMs on Persian Literature Questions
von: Niazi, Ruhallah, et al.
Veröffentlicht: (2026)
von: Niazi, Ruhallah, et al.
Veröffentlicht: (2026)
Text Meets Topology: Rethinking Out-of-distribution Detection in Text-Rich Networks
von: Wang, Danny, et al.
Veröffentlicht: (2025)
von: Wang, Danny, et al.
Veröffentlicht: (2025)
AraHalluEval: A Fine-grained Hallucination Evaluation Framework for Arabic LLMs
von: Alansari, Aisha, et al.
Veröffentlicht: (2025)
von: Alansari, Aisha, et al.
Veröffentlicht: (2025)
Fine-grained Stateful Knowledge Exploration: Effective and Efficient Graph Retrieval with Large Language Models
von: Tao, Dehao, et al.
Veröffentlicht: (2024)
von: Tao, Dehao, et al.
Veröffentlicht: (2024)
The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models
von: Kim, Seungone, et al.
Veröffentlicht: (2024)
von: Kim, Seungone, et al.
Veröffentlicht: (2024)
Large Language Models Meet Symbolic Provers for Logical Reasoning Evaluation
von: Qi, Chengwen, et al.
Veröffentlicht: (2025)
von: Qi, Chengwen, et al.
Veröffentlicht: (2025)
MELD-ST: An Emotion-aware Speech Translation Dataset
von: Chen, Sirou, et al.
Veröffentlicht: (2024)
von: Chen, Sirou, et al.
Veröffentlicht: (2024)
Model Tells Itself Where to Attend: Faithfulness Meets Automatic Attention Steering
von: Zhang, Qingru, et al.
Veröffentlicht: (2024)
von: Zhang, Qingru, et al.
Veröffentlicht: (2024)
Speculative Decoding Meets Quantization: Compatibility Evaluation and Hierarchical Framework Design
von: Zhang, Yudi, et al.
Veröffentlicht: (2025)
von: Zhang, Yudi, et al.
Veröffentlicht: (2025)
Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context
von: Alizadeh, Keivan, et al.
Veröffentlicht: (2026)
von: Alizadeh, Keivan, et al.
Veröffentlicht: (2026)
Rethinking Composed Image Retrieval Evaluation: A Fine-Grained Benchmark from Image Editing
von: Song, Tingyu, et al.
Veröffentlicht: (2026)
von: Song, Tingyu, et al.
Veröffentlicht: (2026)
Understanding the Prompt Sensitivity
von: Liu, Yang, et al.
Veröffentlicht: (2026)
von: Liu, Yang, et al.
Veröffentlicht: (2026)
Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs
von: Liu, Yang, et al.
Veröffentlicht: (2025)
von: Liu, Yang, et al.
Veröffentlicht: (2025)
Meet Dynamic Individual Preferences: Resolving Conflicting Human Value with Paired Fine-Tuning
von: Wang, Shanyong, et al.
Veröffentlicht: (2026)
von: Wang, Shanyong, et al.
Veröffentlicht: (2026)
MEETING DELEGATE: Benchmarking LLMs on Attending Meetings on Our Behalf
von: Hu, Lingxiang, et al.
Veröffentlicht: (2025)
von: Hu, Lingxiang, et al.
Veröffentlicht: (2025)
A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains
von: Wang, Shirui, et al.
Veröffentlicht: (2025)
von: Wang, Shirui, et al.
Veröffentlicht: (2025)
Policies and Evaluation for Online Meeting Summarization
von: Schneider, Felix, et al.
Veröffentlicht: (2025)
von: Schneider, Felix, et al.
Veröffentlicht: (2025)
RPC-Bench: A Fine-grained Benchmark for Research Paper Comprehension
von: Chen, Yelin, et al.
Veröffentlicht: (2026)
von: Chen, Yelin, et al.
Veröffentlicht: (2026)
MedRCube: A Multidimensional Framework for Fine-Grained and In-Depth Evaluation of MLLMs in Medical Imaging
von: Bao, Zhijie, et al.
Veröffentlicht: (2026)
von: Bao, Zhijie, et al.
Veröffentlicht: (2026)
Evaluating the Effectiveness of Cost-Efficient Large Language Models in Benchmark Biomedical Tasks
von: Jahan, Israt, et al.
Veröffentlicht: (2025)
von: Jahan, Israt, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
EMS: Efficient and Effective Massively Multilingual Sentence Embedding Learning
von: Mao, Zhuoyuan, et al.
Veröffentlicht: (2022) -
When Large Language Models Meet Speech: A Survey on Integration Approaches
von: Yang, Zhengdong, et al.
Veröffentlicht: (2025) -
FActBench: A Benchmark for Fine-grained Automatic Evaluation of LLM-Generated Text in the Medical Domain
von: Afzal, Anum, et al.
Veröffentlicht: (2025) -
CFSafety: Comprehensive Fine-grained Safety Assessment for LLMs
von: Liu, Zhihao, et al.
Veröffentlicht: (2024) -
Evaluating the Effectiveness of Black-Box Prompt Optimization as the Scale of LLMs Continues to Grow
von: Zhou, Ziyu, et al.
Veröffentlicht: (2025)