Saved in:
| Main Authors: | Zhong, Philip, Wang, Don, Zhang, Jason |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.21345 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unitxt: Flexible, Shareable and Reusable Data Preparation and Evaluation for Generative AI
by: Bandel, Elron, et al.
Published: (2024)
by: Bandel, Elron, et al.
Published: (2024)
Is my Meeting Summary Good? Estimating Quality with a Multi-LLM Evaluator
by: Kirstein, Frederic, et al.
Published: (2024)
by: Kirstein, Frederic, et al.
Published: (2024)
Evaluating Embedding Models and Pipeline Optimization for AI Search Quality
by: Zhong, Philip, et al.
Published: (2025)
by: Zhong, Philip, et al.
Published: (2025)
When Thoughts Meet Facts: Reusable Reasoning for Long-Context LMs
by: Jeong, Soyeong, et al.
Published: (2025)
by: Jeong, Soyeong, et al.
Published: (2025)
What's Wrong? Refining Meeting Summaries with LLM Feedback
by: Kirstein, Frederic, et al.
Published: (2024)
by: Kirstein, Frederic, et al.
Published: (2024)
Toward Reusability of AI Models Using Dynamic Updates of AI Documentation
by: Bajcsy, Peter, et al.
Published: (2026)
by: Bajcsy, Peter, et al.
Published: (2026)
Ethical and Explainable AI in Reusable MLOps Pipelines
by: Hossain, Rakib, et al.
Published: (2026)
by: Hossain, Rakib, et al.
Published: (2026)
Evaluating Chain-of-Thought Reasoning through Reusability and Verifiability
by: Aggarwal, Shashank, et al.
Published: (2026)
by: Aggarwal, Shashank, et al.
Published: (2026)
Detecting AI-Generated Texts in Cross-Domains
by: Zhou, You, et al.
Published: (2024)
by: Zhou, You, et al.
Published: (2024)
Leveraging Multi-AI Agents for Cross-Domain Knowledge Discovery
by: Aryal, Shiva, et al.
Published: (2024)
by: Aryal, Shiva, et al.
Published: (2024)
Evaluating Novelty in AI-Generated Research Plans Using Multi-Workflow LLM Pipelines
by: Saraogi, Devesh, et al.
Published: (2025)
by: Saraogi, Devesh, et al.
Published: (2025)
KGPA: Robustness Evaluation for Large Language Models via Cross-Domain Knowledge Graphs
by: Pei, Aihua, et al.
Published: (2024)
by: Pei, Aihua, et al.
Published: (2024)
Attribute Structuring Improves LLM-Based Evaluation of Clinical Text Summaries
by: Gero, Zelalem, et al.
Published: (2024)
by: Gero, Zelalem, et al.
Published: (2024)
Evaluating Text Summaries Generated by Large Language Models Using OpenAI's GPT
by: Shakil, Hassan, et al.
Published: (2024)
by: Shakil, Hassan, et al.
Published: (2024)
AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models
by: Jackson, Declan, et al.
Published: (2025)
by: Jackson, Declan, et al.
Published: (2025)
SCURank: Ranking Multiple Candidate Summaries with Summary Content Units for Enhanced Summarization
by: Wang, Bo-Jyun, et al.
Published: (2026)
by: Wang, Bo-Jyun, et al.
Published: (2026)
LangGPT: Rethinking Structured Reusable Prompt Design Framework for LLMs from the Programming Language
by: Wang, Ming, et al.
Published: (2024)
by: Wang, Ming, et al.
Published: (2024)
LCDS: A Logic-Controlled Discharge Summary Generation System Supporting Source Attribution and Expert Review
by: Yuan, Cheng, et al.
Published: (2025)
by: Yuan, Cheng, et al.
Published: (2025)
DIAL-SUMMER: A Structured Evaluation Framework of Hierarchical Errors in Dialogue Summaries
by: Ramnath, Sahana, et al.
Published: (2026)
by: Ramnath, Sahana, et al.
Published: (2026)
WorkRB: A Community-Driven Evaluation Framework for AI in the Work Domain
by: De Lange, Matthias, et al.
Published: (2026)
by: De Lange, Matthias, et al.
Published: (2026)
Ontology-Constrained Generation of Domain-Specific Clinical Summaries
by: Mehenni, Gaya, et al.
Published: (2024)
by: Mehenni, Gaya, et al.
Published: (2024)
Evaluation Ethics of LLMs in Legal Domain
by: Zhang, Ruizhe, et al.
Published: (2024)
by: Zhang, Ruizhe, et al.
Published: (2024)
Towards Multi-dimensional Evaluation of LLM Summarization across Domains and Languages
by: Min, Hyangsuk, et al.
Published: (2025)
by: Min, Hyangsuk, et al.
Published: (2025)
BEADs: Bias Evaluation Across Domains
by: Raza, Shaina, et al.
Published: (2024)
by: Raza, Shaina, et al.
Published: (2024)
Depth $F_1$: Improving Evaluation of Cross-Domain Text Classification by Measuring Semantic Generalizability
by: Seegmiller, Parker, et al.
Published: (2024)
by: Seegmiller, Parker, et al.
Published: (2024)
Agent Primitives: Reusable Latent Building Blocks for Multi-Agent Systems
by: Jin, Haibo, et al.
Published: (2026)
by: Jin, Haibo, et al.
Published: (2026)
When Compression Meets Model Compression: Memory-Efficient Double Compression for Large Language Models
by: Wang, Weilan, et al.
Published: (2025)
by: Wang, Weilan, et al.
Published: (2025)
Log-Augmented Generation: Scaling Test-Time Reasoning with Reusable Computation
by: Chen, Peter Baile, et al.
Published: (2025)
by: Chen, Peter Baile, et al.
Published: (2025)
Cross-Domain Content Generation with Domain-Specific Small Language Models
by: Maloo, Ankit, et al.
Published: (2024)
by: Maloo, Ankit, et al.
Published: (2024)
Pushing on Text Readability Assessment: A Transformer Meets Handcrafted Linguistic Features
by: Lee, Bruce W., et al.
Published: (2021)
by: Lee, Bruce W., et al.
Published: (2021)
Benchmarking Complex Multimodal Document Processing Pipelines: A Unified Evaluation Framework for Enterprise AI
by: Singh, Saurabh K., et al.
Published: (2026)
by: Singh, Saurabh K., et al.
Published: (2026)
GenKnowSub: Improving Modularity and Reusability of LLMs through General Knowledge Subtraction
by: Bagherifard, Mohammadtaha, et al.
Published: (2025)
by: Bagherifard, Mohammadtaha, et al.
Published: (2025)
Decoding Time Series with LLMs: A Multi-Agent Framework for Cross-Domain Annotation
by: Lin, Minhua, et al.
Published: (2024)
by: Lin, Minhua, et al.
Published: (2024)
The Moral Consistency Pipeline: Continuous Ethical Evaluation for Large Language Models
by: Jamshidi, Saeid, et al.
Published: (2025)
by: Jamshidi, Saeid, et al.
Published: (2025)
Reheat Nachos for Dinner? Evaluating AI Support for Cross-Cultural Communication of Neologisms
by: Ki, Dayeon, et al.
Published: (2026)
by: Ki, Dayeon, et al.
Published: (2026)
Agent KB: Leveraging Cross-Domain Experience for Agentic Problem Solving
by: Tang, Xiangru, et al.
Published: (2025)
by: Tang, Xiangru, et al.
Published: (2025)
Large Language Models Meet NLP: A Survey
by: Qin, Libo, et al.
Published: (2024)
by: Qin, Libo, et al.
Published: (2024)
TestAgent: Automatic Benchmarking and Exploratory Interaction for Evaluating LLMs in Vertical Domains
by: Wang, Wanying, et al.
Published: (2024)
by: Wang, Wanying, et al.
Published: (2024)
Deep Research with Open-Domain Evaluation and Multi-Stage Guardrails for Safety
by: Huang, Wei-Chieh, et al.
Published: (2025)
by: Huang, Wei-Chieh, et al.
Published: (2025)
A Functionality-Grounded Benchmark for Evaluating Web Agents in E-commerce Domains
by: Zhang, Xianren, et al.
Published: (2025)
by: Zhang, Xianren, et al.
Published: (2025)
Similar Items
-
Unitxt: Flexible, Shareable and Reusable Data Preparation and Evaluation for Generative AI
by: Bandel, Elron, et al.
Published: (2024) -
Is my Meeting Summary Good? Estimating Quality with a Multi-LLM Evaluator
by: Kirstein, Frederic, et al.
Published: (2024) -
Evaluating Embedding Models and Pipeline Optimization for AI Search Quality
by: Zhong, Philip, et al.
Published: (2025) -
When Thoughts Meet Facts: Reusable Reasoning for Long-Context LMs
by: Jeong, Soyeong, et al.
Published: (2025) -
What's Wrong? Refining Meeting Summaries with LLM Feedback
by: Kirstein, Frederic, et al.
Published: (2024)