Depth $F_1$: Improving Evaluation of Cross-Domain Text Classification by Measuring Semantic Generalizability
Fuente:
arXiv
Saved in:
| Main Authors: | Seegmiller, Parker, Gatto, Joseph, Preum, Sarah Masud |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
In-Context Learning for Preserving Patient Privacy: A Framework for Synthesizing Realistic Patient Portal Messages
by: Gatto, Joseph, et al.
Published: (2024)
by: Gatto, Joseph, et al.
Published: (2024)
Measuring Distribution Shift in User Prompts and Its Effects on LLM Performance
by: Seegmiller, Parker, et al.
Published: (2026)
by: Seegmiller, Parker, et al.
Published: (2026)
How Much Would a Clinician Edit This Draft? Evaluating LLM Alignment for Patient Message Response Drafting
by: Seegmiller, Parker, et al.
Published: (2026)
by: Seegmiller, Parker, et al.
Published: (2026)
Follow-up Question Generation For Enhanced Patient-Provider Conversations
by: Gatto, Joseph, et al.
Published: (2025)
by: Gatto, Joseph, et al.
Published: (2025)
Do LLMs Find Human Answers To Fact-Driven Questions Perplexing? A Case Study on Reddit
by: Seegmiller, Parker, et al.
Published: (2024)
by: Seegmiller, Parker, et al.
Published: (2024)
Large Language Models for Document-Level Event-Argument Data Augmentation for Challenging Role Types
by: Gatto, Joseph, et al.
Published: (2024)
by: Gatto, Joseph, et al.
Published: (2024)
Medical Triage as Pairwise Ranking: A Benchmark for Urgency in Patient Portal Messages
by: Gatto, Joseph, et al.
Published: (2026)
by: Gatto, Joseph, et al.
Published: (2026)
REGen: A Reliable Evaluation Framework for Generative Event Argument Extraction
by: Sharif, Omar, et al.
Published: (2025)
by: Sharif, Omar, et al.
Published: (2025)
A Two-Stage Framework with Self-Supervised Distillation For Cross-Domain Text Classification
by: Feng, Yunlong, et al.
Published: (2023)
by: Feng, Yunlong, et al.
Published: (2023)
FLAMES: Improving LLM Math Reasoning via a Fine-Grained Analysis of the Data Synthesis Pipeline
by: Seegmiller, Parker, et al.
Published: (2025)
by: Seegmiller, Parker, et al.
Published: (2025)
Explicit, Implicit, and Scattered: Revisiting Event Extraction to Capture Complex Arguments
by: Sharif, Omar, et al.
Published: (2024)
by: Sharif, Omar, et al.
Published: (2024)
Cross-lingual Text Classification Transfer: The Case of Ukrainian
by: Dementieva, Daryna, et al.
Published: (2024)
by: Dementieva, Daryna, et al.
Published: (2024)
Exploiting Text Semantics for Few and Zero Shot Node Classification on Text-attributed Graph
by: Wang, Yuxiang, et al.
Published: (2025)
by: Wang, Yuxiang, et al.
Published: (2025)
Token Masking Improves Transformer-Based Text Classification
by: Xu, Xianglong, et al.
Published: (2025)
by: Xu, Xianglong, et al.
Published: (2025)
DetectAnyLLM: Towards Generalizable and Robust Detection of Machine-Generated Text Across Domains and Models
by: Fu, Jiachen, et al.
Published: (2025)
by: Fu, Jiachen, et al.
Published: (2025)
A Benchmark for Cross-Domain Argumentative Stance Classification on Social Media
by: Yuan, Jiaqing, et al.
Published: (2024)
by: Yuan, Jiaqing, et al.
Published: (2024)
Your Next State-of-the-Art Could Come from Another Domain: A Cross-Domain Analysis of Hierarchical Text Classification
by: Li, Nan, et al.
Published: (2024)
by: Li, Nan, et al.
Published: (2024)
Controlling Out-of-Domain Gaps in LLMs for Genre Classification and Generated Text Detection
by: Roussinov, Dmitri, et al.
Published: (2024)
by: Roussinov, Dmitri, et al.
Published: (2024)
DepthCharge: A Domain-Agnostic Framework for Measuring Depth-Dependent Knowledge in Large Language Models
by: Sheppert, Alexander
Published: (2026)
by: Sheppert, Alexander
Published: (2026)
Bengali Text Classification: An Evaluation of Large Language Model Approaches
by: Hoque, Md Mahmudul, et al.
Published: (2026)
by: Hoque, Md Mahmudul, et al.
Published: (2026)
Discovering Multi-Scale Semantic Structure in Text Corpora Using Density-Based Trees and LLM Embeddings
by: Haschka, Thomas, et al.
Published: (2025)
by: Haschka, Thomas, et al.
Published: (2025)
Rethinking Word Similarity: Semantic Similarity through Classification Confusion
by: Zhou, Kaitlyn, et al.
Published: (2025)
by: Zhou, Kaitlyn, et al.
Published: (2025)
Reasoning before Comparison: LLM-Enhanced Semantic Similarity Metrics for Domain Specialized Text Analysis
by: Xu, Shaochen, et al.
Published: (2024)
by: Xu, Shaochen, et al.
Published: (2024)
OptBA: Optimizing Hyperparameters with the Bees Algorithm for Improved Medical Text Classification
by: Shaaban, Mai A., et al.
Published: (2023)
by: Shaaban, Mai A., et al.
Published: (2023)
Selective Attention Federated Learning: Improving Privacy and Efficiency for Clinical Text Classification
by: Li, Yue, et al.
Published: (2025)
by: Li, Yue, et al.
Published: (2025)
A New HOPE: Domain-agnostic Automatic Evaluation of Text Chunking
by: Brådland, Henrik, et al.
Published: (2025)
by: Brådland, Henrik, et al.
Published: (2025)
Evaluating AI Meeting Summaries with a Reusable Cross-Domain Pipeline
by: Zhong, Philip, et al.
Published: (2026)
by: Zhong, Philip, et al.
Published: (2026)
Breaking the Silence: A Dataset and Benchmark for Bangla Text-to-Gloss Translation
by: Abdullah, Sharif Mohammad, et al.
Published: (2025)
by: Abdullah, Sharif Mohammad, et al.
Published: (2025)
Attention-Guided Feature Fusion (AGFF) Model for Integrating Statistical and Semantic Features in News Text Classification
by: Zare, Mohammad
Published: (2025)
by: Zare, Mohammad
Published: (2025)
Detecting AI-Generated Texts in Cross-Domains
by: Zhou, You, et al.
Published: (2024)
by: Zhou, You, et al.
Published: (2024)
Towards Compositionally Generalizable Semantic Parsing in Large Language Models: A Survey
by: Mannekote, Amogh
Published: (2024)
by: Mannekote, Amogh
Published: (2024)
A Necessary Step toward Faithfulness: Measuring and Improving Consistency in Free-Text Explanations
by: Zhao, Lingjun, et al.
Published: (2025)
by: Zhao, Lingjun, et al.
Published: (2025)
Attribute Structuring Improves LLM-Based Evaluation of Clinical Text Summaries
by: Gero, Zelalem, et al.
Published: (2024)
by: Gero, Zelalem, et al.
Published: (2024)
AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models
by: Jackson, Declan, et al.
Published: (2025)
by: Jackson, Declan, et al.
Published: (2025)
Text2Zinc: A Cross-Domain Dataset for Modeling Optimization and Satisfaction Problems in MiniZinc
by: Singirikonda, Akash, et al.
Published: (2025)
by: Singirikonda, Akash, et al.
Published: (2025)
TexIm FAST: Text-to-Image Representation for Semantic Similarity Evaluation using Transformers
by: Ansar, Wazib, et al.
Published: (2024)
by: Ansar, Wazib, et al.
Published: (2024)
X-Eval: Generalizable Multi-aspect Text Evaluation via Augmented Instruction Tuning with Auxiliary Evaluation Aspects
by: Liu, Minqian, et al.
Published: (2023)
by: Liu, Minqian, et al.
Published: (2023)
LLM-Guided Semantic Bootstrapping for Interpretable Text Classification with Tsetlin Machines
by: Gao, Jiechao, et al.
Published: (2026)
by: Gao, Jiechao, et al.
Published: (2026)
Hallucination-Resistant, Domain-Specific Research Assistant with Self-Evaluation and Vector-Grounded Retrieval
by: Bhavsar, Vivek, et al.
Published: (2025)
by: Bhavsar, Vivek, et al.
Published: (2025)
KGPA: Robustness Evaluation for Large Language Models via Cross-Domain Knowledge Graphs
by: Pei, Aihua, et al.
Published: (2024)
by: Pei, Aihua, et al.
Published: (2024)
Similar Items
-
In-Context Learning for Preserving Patient Privacy: A Framework for Synthesizing Realistic Patient Portal Messages
by: Gatto, Joseph, et al.
Published: (2024) -
Measuring Distribution Shift in User Prompts and Its Effects on LLM Performance
by: Seegmiller, Parker, et al.
Published: (2026) -
How Much Would a Clinician Edit This Draft? Evaluating LLM Alignment for Patient Message Response Drafting
by: Seegmiller, Parker, et al.
Published: (2026) -
Follow-up Question Generation For Enhanced Patient-Provider Conversations
by: Gatto, Joseph, et al.
Published: (2025) -
Do LLMs Find Human Answers To Fact-Driven Questions Perplexing? A Case Study on Reddit
by: Seegmiller, Parker, et al.
Published: (2024)