AI Agents-as-Judge: Automated Assessment of Accuracy, Consistency, Completeness and Clarity for Enterprise Documents
Fuente:
arXiv
Saved in:
| Main Authors: | Dasgupta, Sudip, Shankar, Himanshu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics
by: Peyronnet, Antoine, et al.
Published: (2026)
by: Peyronnet, Antoine, et al.
Published: (2026)
BMAM: Brain-inspired Multi-Agent Memory Framework
by: Li, Yang, et al.
Published: (2026)
by: Li, Yang, et al.
Published: (2026)
Thinking Machines: Mathematical Reasoning in the Age of LLMs
by: Asperti, Andrea, et al.
Published: (2025)
by: Asperti, Andrea, et al.
Published: (2025)
PerkwE_COQA: Enhanced Persian Conversational Question Answering by combining contextual keyword extraction with Large Language Models
by: Moradbeiki, Pardis, et al.
Published: (2024)
by: Moradbeiki, Pardis, et al.
Published: (2024)
Challenges and Applications of Large Language Models: A Comparison of GPT and DeepSeek family of models
by: Sharma, Shubham, et al.
Published: (2025)
by: Sharma, Shubham, et al.
Published: (2025)
Mitigating LLM Hallucinations through Domain-Grounded Tiered Retrieval
by: Haque, Md. Asraful, et al.
Published: (2026)
by: Haque, Md. Asraful, et al.
Published: (2026)
Advancing Explainability in Neural Machine Translation: Analytical Metrics for Attention and Alignment Consistency
by: Mishra, Anurag
Published: (2024)
by: Mishra, Anurag
Published: (2024)
PolyTruth: Multilingual Disinformation Detection using Transformer-Based Language Models
by: Gouliev, Zaur, et al.
Published: (2025)
by: Gouliev, Zaur, et al.
Published: (2025)
Beyond Long Context: When Semantics Matter More than Tokens
by: Chawdhury, Tarun Kumar, et al.
Published: (2025)
by: Chawdhury, Tarun Kumar, et al.
Published: (2025)
Mubeen AI: A Specialized Arabic Language Model for Heritage Preservation and User Intent Understanding
by: Aljafari, Mohammed, et al.
Published: (2025)
by: Aljafari, Mohammed, et al.
Published: (2025)
DPDisc: From Factoid Questions to Data Product Requests for Open-World Data Product Discovery over Tables and Text
by: Zhang, Liangliang, et al.
Published: (2025)
by: Zhang, Liangliang, et al.
Published: (2025)
GE-Chat: A Graph Enhanced RAG Framework for Evidential Response Generation of LLMs
by: Da, Longchao, et al.
Published: (2025)
by: Da, Longchao, et al.
Published: (2025)
Towards Robust Retrieval-Augmented Generation Based on Knowledge Graph: A Comparative Analysis
by: Amamou, Hazem, et al.
Published: (2026)
by: Amamou, Hazem, et al.
Published: (2026)
MedMemoryBench: Benchmarking Agent Memory in Personalized Healthcare
by: Wang, Yihao, et al.
Published: (2026)
by: Wang, Yihao, et al.
Published: (2026)
A Multi-granularity Concept Sparse Activation and Hierarchical Knowledge Graph Fusion Framework for Rare Disease Diagnosis
by: Zhang, Mingda, et al.
Published: (2025)
by: Zhang, Mingda, et al.
Published: (2025)
Fact Grounded Attention: Eliminating Hallucination in Large Language Models Through Attention Level Knowledge Integration
by: Gupta, Aayush
Published: (2025)
by: Gupta, Aayush
Published: (2025)
Optimizing Retrieval-Augmented Generation (RAG) for Colloquial Cantonese: A LoRA-Based Systematic Review
by: Calonge, David Santandreu, et al.
Published: (2025)
by: Calonge, David Santandreu, et al.
Published: (2025)
EdgeJury: Cross-Reviewed Small-Model Ensembles for Truthful Question Answering on Serverless Edge Inference
by: Kumar, Aayush
Published: (2025)
by: Kumar, Aayush
Published: (2025)
Representing LLMs in Prompt Semantic Task Space
by: Kashani, Idan, et al.
Published: (2025)
by: Kashani, Idan, et al.
Published: (2025)
Adaptive Multi-Stage Patent Claim Generation with Unified Quality Assessment
by: Liang, Chen-Wei, et al.
Published: (2026)
by: Liang, Chen-Wei, et al.
Published: (2026)
From Unstructured Recall to Schema-Grounded Memory: Reliable AI Memory via Iterative, Schema-Aware Extraction
by: Petrov, Alex, et al.
Published: (2026)
by: Petrov, Alex, et al.
Published: (2026)
SQuARE: Structured Query & Adaptive Retrieval Engine For Tabular Formats
by: Gondhalekar, Chinmay, et al.
Published: (2025)
by: Gondhalekar, Chinmay, et al.
Published: (2025)
Maat: The Agentic Legal Research Assistant for Competition Protection
by: Mounir, Basant, et al.
Published: (2026)
by: Mounir, Basant, et al.
Published: (2026)
ByteRover: Agent-Native Memory Through LLM-Curated Hierarchical Context
by: Nguyen, Andy, et al.
Published: (2026)
by: Nguyen, Andy, et al.
Published: (2026)
AgenticAI-DialogGen: Topic-Guided Conversation Generation for Fine-Tuning and Evaluating Short- and Long-Term Memories of LLMs
by: Perera, Manoj Madushanka, et al.
Published: (2026)
by: Perera, Manoj Madushanka, et al.
Published: (2026)
Optimizing Large Language Models for OpenAPI Code Completion
by: Petryshyn, Bohdan, et al.
Published: (2024)
by: Petryshyn, Bohdan, et al.
Published: (2024)
Reducing Labeling Costs in Sentiment Analysis via Semi-Supervised Learning
by: Jafarlou, Minoo, et al.
Published: (2024)
by: Jafarlou, Minoo, et al.
Published: (2024)
WebMap -- Large Language Model-assisted Semantic Link Induction in the Web
by: Pokharel, Shiraj, et al.
Published: (2025)
by: Pokharel, Shiraj, et al.
Published: (2025)
On Self-improving Token Embeddings
by: Kubek, Mario M., et al.
Published: (2025)
by: Kubek, Mario M., et al.
Published: (2025)
Intersymbolic AI: Interlinking Symbolic AI and Subsymbolic AI
by: Platzer, André
Published: (2024)
by: Platzer, André
Published: (2024)
ARISE: Agentic Rubric-Guided Iterative Survey Engine for Automated Scholarly Paper Generation
by: Wang, Zi, et al.
Published: (2025)
by: Wang, Zi, et al.
Published: (2025)
Combating data scarcity in recommendation services: Integrating cognitive types of VARK and neural network technologies (LLM)
by: Zmanovskii, Nikita
Published: (2026)
by: Zmanovskii, Nikita
Published: (2026)
HumanMCP: A Human-Like Query Dataset for Evaluating MCP Tool Retrieval Performance
by: Laddha, Shubh, et al.
Published: (2025)
by: Laddha, Shubh, et al.
Published: (2025)
Heterogeneous LLM Methods for Ontology Learning (Few-Shot Prompting, Ensemble Typing, and Attention-Based Taxonomies)
by: Beliaeva, Aleksandra, et al.
Published: (2025)
by: Beliaeva, Aleksandra, et al.
Published: (2025)
Language processing in humans and computers
by: Pavlovic, Dusko
Published: (2024)
by: Pavlovic, Dusko
Published: (2024)
A Case Study of Balanced Query Recommendation on Wikipedia
by: Mishra, Harshit, et al.
Published: (2025)
by: Mishra, Harshit, et al.
Published: (2025)
MeVer at CheckThat! 2026: Cluster-Aware Hard-Negative Mining for Multilingual Scientific-Source Retrieval
by: Bakagianni, Juli, et al.
Published: (2026)
by: Bakagianni, Juli, et al.
Published: (2026)
IndoBERT-Sentiment: Context-Conditioned Sentiment Classification for Indonesian Text
by: Saputra, Muhammad Apriandito Arya, et al.
Published: (2026)
by: Saputra, Muhammad Apriandito Arya, et al.
Published: (2026)
IndoBERT-Relevancy: A Context-Conditioned Relevancy Classifier for Indonesian Text
by: Saputra, Muhammad Apriandito Arya, et al.
Published: (2026)
by: Saputra, Muhammad Apriandito Arya, et al.
Published: (2026)
An Ensemble Embedding Approach for Improving Semantic Caching Performance in LLM-based Systems
by: Ghaffari, Shervin, et al.
Published: (2025)
by: Ghaffari, Shervin, et al.
Published: (2025)
Similar Items
-
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics
by: Peyronnet, Antoine, et al.
Published: (2026) -
BMAM: Brain-inspired Multi-Agent Memory Framework
by: Li, Yang, et al.
Published: (2026) -
Thinking Machines: Mathematical Reasoning in the Age of LLMs
by: Asperti, Andrea, et al.
Published: (2025) -
PerkwE_COQA: Enhanced Persian Conversational Question Answering by combining contextual keyword extraction with Large Language Models
by: Moradbeiki, Pardis, et al.
Published: (2024) -
Challenges and Applications of Large Language Models: A Comparison of GPT and DeepSeek family of models
by: Sharma, Shubham, et al.
Published: (2025)