When F1 Fails: Granularity-Aware Evaluation for Dialogue Topic Segmentation
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Coen, Michael H. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
When Does Data Augmentation Help? Evaluating LLM and Back-Translation Methods for Hausa and Fongbe NLP
von: Adjovi, Mahounan Pericles, et al.
Veröffentlicht: (2026)
von: Adjovi, Mahounan Pericles, et al.
Veröffentlicht: (2026)
OpenGloss: A Synthetic Encyclopedic Dictionary and Semantic Knowledge Graph
von: Bommarito II, Michael J.
Veröffentlicht: (2025)
von: Bommarito II, Michael J.
Veröffentlicht: (2025)
Mining Large Language Models for Low-Resource Language Data: Comparing Elicitation Strategies for Hausa and Fongbe
von: Adjovi, Mahounan Pericles, et al.
Veröffentlicht: (2026)
von: Adjovi, Mahounan Pericles, et al.
Veröffentlicht: (2026)
CIG: Measuring Conversational Information Gain in Deliberative Dialogues with Semantic Memory Dynamics
von: Chen, Ming-Bin, et al.
Veröffentlicht: (2026)
von: Chen, Ming-Bin, et al.
Veröffentlicht: (2026)
LegalBench-BR: A Benchmark for Evaluating Large Language Models on Brazilian Legal Decision Classification
von: Neto, Pedro Barbosa de Carvalho
Veröffentlicht: (2026)
von: Neto, Pedro Barbosa de Carvalho
Veröffentlicht: (2026)
TMT: A Simple Way to Translate Topic Models Using Dictionaries
von: Engl, Felix, et al.
Veröffentlicht: (2025)
von: Engl, Felix, et al.
Veröffentlicht: (2025)
Data Processing for the OpenGPT-X Model Family
von: Brandizzi, Nicolo', et al.
Veröffentlicht: (2024)
von: Brandizzi, Nicolo', et al.
Veröffentlicht: (2024)
Leveraging LLMs to Create Content Corpora for Niche Domains
von: Zhang, Franklin, et al.
Veröffentlicht: (2025)
von: Zhang, Franklin, et al.
Veröffentlicht: (2025)
LLM-based Extraction of Contradictions from Patents
von: Trapp, Stefan, et al.
Veröffentlicht: (2024)
von: Trapp, Stefan, et al.
Veröffentlicht: (2024)
Supervised Semantic Differential for Cross-Cultural Concept Analysis: A Case Study of Human Affect
von: Sikora, Jan, et al.
Veröffentlicht: (2026)
von: Sikora, Jan, et al.
Veröffentlicht: (2026)
Using Instruction-Tuned Large Language Models to Identify Indicators of Vulnerability in Police Incident Narratives
von: Relins, Sam, et al.
Veröffentlicht: (2024)
von: Relins, Sam, et al.
Veröffentlicht: (2024)
How Real Are Synthetic Therapy Conversations? Evaluating Fidelity in Prolonged Exposure Dialogues
von: BN, Suhas, et al.
Veröffentlicht: (2025)
von: BN, Suhas, et al.
Veröffentlicht: (2025)
Reading Between the Waves: Robust Topic Segmentation Using Inter-Sentence Audio Features
von: Freisinger, Steffen, et al.
Veröffentlicht: (2026)
von: Freisinger, Steffen, et al.
Veröffentlicht: (2026)
MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
von: Li, Yingyun, et al.
Veröffentlicht: (2026)
von: Li, Yingyun, et al.
Veröffentlicht: (2026)
A Survey of Text and Speech Resources for Hausa and Fongbe: Availability, Quality, and Gaps for NLP Development
von: Adjovi, Mahounan Pericles, et al.
Veröffentlicht: (2026)
von: Adjovi, Mahounan Pericles, et al.
Veröffentlicht: (2026)
LLMs as Architects and Critics for Multi-Source Opinion Summarization
von: Attri, Anuj, et al.
Veröffentlicht: (2025)
von: Attri, Anuj, et al.
Veröffentlicht: (2025)
Why We Feel What We Feel: Joint Detection of Emotions and Their Opinion Triggers in E-commerce
von: Attri, Arnav, et al.
Veröffentlicht: (2025)
von: Attri, Arnav, et al.
Veröffentlicht: (2025)
Enhancing Scientific Literature Chatbots with Retrieval-Augmented Generation: A Performance Evaluation of Vector and Graph-Based Systems
von: Ghanadian, Hamideh, et al.
Veröffentlicht: (2026)
von: Ghanadian, Hamideh, et al.
Veröffentlicht: (2026)
DEUCE: Dual-diversity Enhancement and Uncertainty-awareness for Cold-start Active Learning
von: Guo, Jiaxin, et al.
Veröffentlicht: (2025)
von: Guo, Jiaxin, et al.
Veröffentlicht: (2025)
AI-Friendly LaTeX: Using LaTeX Code as a Knowledge Source for Retrieval-Augmented Generation
von: Verhoeff, Tom
Veröffentlicht: (2026)
von: Verhoeff, Tom
Veröffentlicht: (2026)
IndiaFinBench: An Evaluation Benchmark for Large Language Model Performance on Indian Financial Regulatory Text
von: Pall, Rajveer Singh
Veröffentlicht: (2026)
von: Pall, Rajveer Singh
Veröffentlicht: (2026)
Low-Resource English-Tigrinya MT: Leveraging Multilingual Models, Custom Tokenizers, and Clean Evaluation Benchmarks
von: Teklehaymanot, Hailay Kidu, et al.
Veröffentlicht: (2025)
von: Teklehaymanot, Hailay Kidu, et al.
Veröffentlicht: (2025)
Generative AI-Based Virtual Assistant using Retrieval-Augmented Generation: An evaluation study for bachelor projects
von: Verşebeniuc, Dumitru, et al.
Veröffentlicht: (2026)
von: Verşebeniuc, Dumitru, et al.
Veröffentlicht: (2026)
Detection of Personal Data in Structured Datasets Using a Large Language Model
von: Ntwali, Albert Agisha, et al.
Veröffentlicht: (2025)
von: Ntwali, Albert Agisha, et al.
Veröffentlicht: (2025)
Neuromem: A Granular Decomposition of the Streaming Lifecycle in External Memory for LLMs
von: Zhang, Ruicheng, et al.
Veröffentlicht: (2026)
von: Zhang, Ruicheng, et al.
Veröffentlicht: (2026)
Beyond End-to-End Video Models: An LLM-Based Multi-Agent System for Educational Video Generation
von: Yan, Lingyong, et al.
Veröffentlicht: (2026)
von: Yan, Lingyong, et al.
Veröffentlicht: (2026)
IMLJD: A Computational Dataset for Indian Matrimonial Litigation Analysis
von: Bose, Joy
Veröffentlicht: (2026)
von: Bose, Joy
Veröffentlicht: (2026)
Argument Quality Annotation and Gender Bias Detection in Financial Communication through Large Language Models
von: Alhamzeh, Alaa, et al.
Veröffentlicht: (2025)
von: Alhamzeh, Alaa, et al.
Veröffentlicht: (2025)
Advancing Uto-Aztecan Language Technologies: A Case Study on the Endangered Comanche Language
von: C, Jesus Alvarez, et al.
Veröffentlicht: (2025)
von: C, Jesus Alvarez, et al.
Veröffentlicht: (2025)
Do Small Language Models Know When They're Wrong? Confidence-Based Cascade Scoring for Educational Assessment
von: Burleigh, Tyler
Veröffentlicht: (2026)
von: Burleigh, Tyler
Veröffentlicht: (2026)
"This Suits You the Best": Query Focused Comparative Explainable Summarization
von: Attri, Arnav, et al.
Veröffentlicht: (2025)
von: Attri, Arnav, et al.
Veröffentlicht: (2025)
Subjective Question Generation and Answer Evaluation using NLP
von: Islam, G. M. Refatul, et al.
Veröffentlicht: (2025)
von: Islam, G. M. Refatul, et al.
Veröffentlicht: (2025)
Assessing Crime Disclosure Patterns in a Large-Scale Cybercrime Forum
von: Hoheisel, Raphael, et al.
Veröffentlicht: (2026)
von: Hoheisel, Raphael, et al.
Veröffentlicht: (2026)
The Syntactic Acceptability Dataset (Preview): A Resource for Machine Learning and Linguistic Analysis of English
von: Juzek, Tom S
Veröffentlicht: (2025)
von: Juzek, Tom S
Veröffentlicht: (2025)
Tokens with Meaning: A Hybrid Tokenization Approach for Turkish
von: Bayram, M. Ali, et al.
Veröffentlicht: (2025)
von: Bayram, M. Ali, et al.
Veröffentlicht: (2025)
HiPS: Hierarchical PDF Segmentation of Textbooks
von: Wehnert, Sabine, et al.
Veröffentlicht: (2025)
von: Wehnert, Sabine, et al.
Veröffentlicht: (2025)
EPIC-EuroParl-UdS: Information-Theoretic Perspectives on Translation and Interpreting
von: Kunilovskaya, Maria, et al.
Veröffentlicht: (2026)
von: Kunilovskaya, Maria, et al.
Veröffentlicht: (2026)
Translationese as a Rational Response to Translation Task Difficulty
von: Kunilovskaya, Maria
Veröffentlicht: (2026)
von: Kunilovskaya, Maria
Veröffentlicht: (2026)
Leveraging OpenFlamingo for Multimodal Embedding Analysis of C2C Car Parts Data
von: Rashid, Maisha Binte, et al.
Veröffentlicht: (2025)
von: Rashid, Maisha Binte, et al.
Veröffentlicht: (2025)
A ripple in time: a discontinuity in American history
von: Kolpakov, Alexander, et al.
Veröffentlicht: (2023)
von: Kolpakov, Alexander, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
When Does Data Augmentation Help? Evaluating LLM and Back-Translation Methods for Hausa and Fongbe NLP
von: Adjovi, Mahounan Pericles, et al.
Veröffentlicht: (2026) -
OpenGloss: A Synthetic Encyclopedic Dictionary and Semantic Knowledge Graph
von: Bommarito II, Michael J.
Veröffentlicht: (2025) -
Mining Large Language Models for Low-Resource Language Data: Comparing Elicitation Strategies for Hausa and Fongbe
von: Adjovi, Mahounan Pericles, et al.
Veröffentlicht: (2026) -
CIG: Measuring Conversational Information Gain in Deliberative Dialogues with Semantic Memory Dynamics
von: Chen, Ming-Bin, et al.
Veröffentlicht: (2026) -
LegalBench-BR: A Benchmark for Evaluating Large Language Models on Brazilian Legal Decision Classification
von: Neto, Pedro Barbosa de Carvalho
Veröffentlicht: (2026)