ExtractBench: A Benchmark and Evaluation Methodology for Complex Structured Extraction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ferguson, Nick, Pennington, Josh, Beghian, Narek, Mohan, Aravind, Kiela, Douwe, Agrawal, Sheshansh, Nguyen, Thien Hang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BlitzRank: Principled Zero-shot Ranking Agents with Tournament Graphs
von: Agrawal, Sheshansh, et al.
Veröffentlicht: (2026)
von: Agrawal, Sheshansh, et al.
Veröffentlicht: (2026)
Classification is a RAG problem: A case study on hate speech detection
von: Willats, Richard, et al.
Veröffentlicht: (2025)
von: Willats, Richard, et al.
Veröffentlicht: (2025)
Lynx: An Open Source Hallucination Evaluation Model
von: Ravi, Selvan Sunitha, et al.
Veröffentlicht: (2024)
von: Ravi, Selvan Sunitha, et al.
Veröffentlicht: (2024)
Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets
von: Tran, Dat, et al.
Veröffentlicht: (2026)
von: Tran, Dat, et al.
Veröffentlicht: (2026)
KTO: Model Alignment as Prospect Theoretic Optimization
von: Ethayarajh, Kawin, et al.
Veröffentlicht: (2024)
von: Ethayarajh, Kawin, et al.
Veröffentlicht: (2024)
I am a Strange Dataset: Metalinguistic Tests for Language Models
von: Thrush, Tristan, et al.
Veröffentlicht: (2024)
von: Thrush, Tristan, et al.
Veröffentlicht: (2024)
Leveraging Diffusion Perturbations for Measuring Fairness in Computer Vision
von: Lui, Nicholas, et al.
Veröffentlicht: (2023)
von: Lui, Nicholas, et al.
Veröffentlicht: (2023)
Anchor Points: Benchmarking Models with Much Fewer Examples
von: Vivek, Rajan, et al.
Veröffentlicht: (2023)
von: Vivek, Rajan, et al.
Veröffentlicht: (2023)
Reflective Context Learning: Studying the Optimization Primitives of Context Space
von: Vassilyev, Nikita, et al.
Veröffentlicht: (2026)
von: Vassilyev, Nikita, et al.
Veröffentlicht: (2026)
Nearest Neighbor Normalization Improves Multimodal Retrieval
von: Chowdhury, Neil, et al.
Veröffentlicht: (2024)
von: Chowdhury, Neil, et al.
Veröffentlicht: (2024)
LMUnit: Fine-grained Evaluation with Natural Language Unit Tests
von: Saad-Falcon, Jon, et al.
Veröffentlicht: (2024)
von: Saad-Falcon, Jon, et al.
Veröffentlicht: (2024)
Few-shot Continual Relation Extraction via Open Information Extraction
von: Nguyen, Thiem, et al.
Veröffentlicht: (2025)
von: Nguyen, Thiem, et al.
Veröffentlicht: (2025)
Document Optimization for Black-Box Retrieval via Reinforcement Learning
von: Uzan, Omri, et al.
Veröffentlicht: (2026)
von: Uzan, Omri, et al.
Veröffentlicht: (2026)
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing
von: Fein, Daniel, et al.
Veröffentlicht: (2025)
von: Fein, Daniel, et al.
Veröffentlicht: (2025)
Generative Representational Instruction Tuning
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2024)
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2024)
Zero-shot Cross-lingual Transfer Learning with Multiple Source and Target Languages for Information Extraction: Language Selection and Adversarial Training
von: Ngo, Nghia Trung, et al.
Veröffentlicht: (2024)
von: Ngo, Nghia Trung, et al.
Veröffentlicht: (2024)
Anchored Preference Optimization and Contrastive Revisions: Addressing Underspecification in Alignment
von: D'Oosterlinck, Karel, et al.
Veröffentlicht: (2024)
von: D'Oosterlinck, Karel, et al.
Veröffentlicht: (2024)
Breaking the Protocol: Security Analysis of the Model Context Protocol Specification and Prompt Injection Vulnerabilities in Tool-Integrated LLM Agents
von: Maloyan, Narek, et al.
Veröffentlicht: (2026)
von: Maloyan, Narek, et al.
Veröffentlicht: (2026)
Realistic Evaluation of Toxicity in Large Language Models
von: Luong, Tinh Son, et al.
Veröffentlicht: (2024)
von: Luong, Tinh Son, et al.
Veröffentlicht: (2024)
ColorBench: Benchmarking Mobile Agents with Graph-Structured Framework for Complex Long-Horizon Tasks
von: Song, Yuanyi, et al.
Veröffentlicht: (2025)
von: Song, Yuanyi, et al.
Veröffentlicht: (2025)
FoldA: Computing Partial-Order Alignments Using Directed Net Unfoldings
von: Geurtjens, Douwe, et al.
Veröffentlicht: (2025)
von: Geurtjens, Douwe, et al.
Veröffentlicht: (2025)
Evaluating the Meta- and Object-Level Reasoning of Large Language Models for Question Answering
von: Ferguson, Nick, et al.
Veröffentlicht: (2025)
von: Ferguson, Nick, et al.
Veröffentlicht: (2025)
mSCoRe: a $M$ultilingual and Scalable Benchmark for $S$kill-based $Co$mmonsense $Re$asoning
von: Ngo, Nghia Trung, et al.
Veröffentlicht: (2025)
von: Ngo, Nghia Trung, et al.
Veröffentlicht: (2025)
Great Models Think Alike and this Undermines AI Oversight
von: Goel, Shashwat, et al.
Veröffentlicht: (2025)
von: Goel, Shashwat, et al.
Veröffentlicht: (2025)
SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs
von: Liu, Zhiqiang, et al.
Veröffentlicht: (2025)
von: Liu, Zhiqiang, et al.
Veröffentlicht: (2025)
Preserving Generalization of Language models in Few-shot Continual Relation Extraction
von: Tran, Quyen, et al.
Veröffentlicht: (2024)
von: Tran, Quyen, et al.
Veröffentlicht: (2024)
FeatureBench: Benchmarking Agentic Coding for Complex Feature Development
von: Zhou, Qixing, et al.
Veröffentlicht: (2026)
von: Zhou, Qixing, et al.
Veröffentlicht: (2026)
Decompose, Enrich, and Extract! Schema-aware Event Extraction using LLMs
von: Shiri, Fatemeh, et al.
Veröffentlicht: (2024)
von: Shiri, Fatemeh, et al.
Veröffentlicht: (2024)
CXReasonBench: A Benchmark for Evaluating Structured Diagnostic Reasoning in Chest X-rays
von: Lee, Hyungyung, et al.
Veröffentlicht: (2025)
von: Lee, Hyungyung, et al.
Veröffentlicht: (2025)
CCR-Bench: A Comprehensive Benchmark for Evaluating LLMs on Complex Constraints, Control Flows, and Real-World Cases
von: Xue, Xiaona, et al.
Veröffentlicht: (2026)
von: Xue, Xiaona, et al.
Veröffentlicht: (2026)
NC-Bench: An LLM Benchmark for Evaluating Conversational Competence
von: Moore, Robert J., et al.
Veröffentlicht: (2026)
von: Moore, Robert J., et al.
Veröffentlicht: (2026)
AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery
von: Xiong, Lei, et al.
Veröffentlicht: (2026)
von: Xiong, Lei, et al.
Veröffentlicht: (2026)
SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents
von: Nandi, Subhrangshu, et al.
Veröffentlicht: (2025)
von: Nandi, Subhrangshu, et al.
Veröffentlicht: (2025)
LegalCiteBench: Evaluating Citation Reliability in Legal Language Models
von: Chen, Sijia, et al.
Veröffentlicht: (2026)
von: Chen, Sijia, et al.
Veröffentlicht: (2026)
CaseReportBench: An LLM Benchmark Dataset for Dense Information Extraction in Clinical Case Reports
von: Zhang, Xiao Yu Cindy, et al.
Veröffentlicht: (2025)
von: Zhang, Xiao Yu Cindy, et al.
Veröffentlicht: (2025)
SproutBench: A Benchmark for Safe and Ethical Large Language Models for Youth
von: Xing, Wenpeng, et al.
Veröffentlicht: (2025)
von: Xing, Wenpeng, et al.
Veröffentlicht: (2025)
Precision Guided Approach to Mitigate Data Poisoning Attacks in Federated Learning
von: Kumar, K Naveen, et al.
Veröffentlicht: (2024)
von: Kumar, K Naveen, et al.
Veröffentlicht: (2024)
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)
T2S-Bench & Structure-of-Thought: Benchmarking and Prompting Comprehensive Text-to-Structure Reasoning
von: Wang, Qinsi, et al.
Veröffentlicht: (2026)
von: Wang, Qinsi, et al.
Veröffentlicht: (2026)
Comprehensive and Practical Evaluation of Retrieval-Augmented Generation Systems for Medical Question Answering
von: Ngo, Nghia Trung, et al.
Veröffentlicht: (2024)
von: Ngo, Nghia Trung, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
BlitzRank: Principled Zero-shot Ranking Agents with Tournament Graphs
von: Agrawal, Sheshansh, et al.
Veröffentlicht: (2026) -
Classification is a RAG problem: A case study on hate speech detection
von: Willats, Richard, et al.
Veröffentlicht: (2025) -
Lynx: An Open Source Hallucination Evaluation Model
von: Ravi, Selvan Sunitha, et al.
Veröffentlicht: (2024) -
Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets
von: Tran, Dat, et al.
Veröffentlicht: (2026) -
KTO: Model Alignment as Prospect Theoretic Optimization
von: Ethayarajh, Kawin, et al.
Veröffentlicht: (2024)