LLM-Based Section Identifiers Excel on Open Source but Stumble in Real World Applications
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Krishnamoorthy, Saranya, Singh, Ayush, Tafreshi, Shabnam |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Can GPT Improve the State of Prior Authorization via Guideline Based Automated Question Answering?
von: Vatsal, Shubham, et al.
Veröffentlicht: (2024)
von: Vatsal, Shubham, et al.
Veröffentlicht: (2024)
Emotion Classification in Low and Moderate Resource Languages
von: Tafreshi, Shabnam, et al.
Veröffentlicht: (2024)
von: Tafreshi, Shabnam, et al.
Veröffentlicht: (2024)
Cornerstones or Stumbling Blocks? Deciphering the Rock Tokens in On-Policy Distillation
von: Jiang, Yuxuan, et al.
Veröffentlicht: (2026)
von: Jiang, Yuxuan, et al.
Veröffentlicht: (2026)
CollabStory: Multi-LLM Collaborative Story Generation and Authorship Analysis
von: Venkatraman, Saranya, et al.
Veröffentlicht: (2024)
von: Venkatraman, Saranya, et al.
Veröffentlicht: (2024)
MiniLingua: A Small Open-Source LLM for European Languages
von: Aksenova, Anna, et al.
Veröffentlicht: (2025)
von: Aksenova, Anna, et al.
Veröffentlicht: (2025)
Wide-Horizon Thinking and Simulation-Based Evaluation for Real-World LLM Planning with Multifaceted Constraints
von: Yang, Dongjie, et al.
Veröffentlicht: (2025)
von: Yang, Dongjie, et al.
Veröffentlicht: (2025)
Medmarks: A Comprehensive Open-Source LLM Benchmark Suite for Medical Tasks
von: Warner, Benjamin, et al.
Veröffentlicht: (2026)
von: Warner, Benjamin, et al.
Veröffentlicht: (2026)
Tracing Thought: Using Chain-of-Thought Reasoning to Identify the LLM Behind AI-Generated Text
von: Agrahari, Shifali, et al.
Veröffentlicht: (2025)
von: Agrahari, Shifali, et al.
Veröffentlicht: (2025)
OpenSeal: Good, Fast, and Cheap Construction of an Open-Source Southeast Asian LLM via Parallel Data
von: Nguyen, Tan Sang, et al.
Veröffentlicht: (2026)
von: Nguyen, Tan Sang, et al.
Veröffentlicht: (2026)
BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
von: Zhang, Zehua, et al.
Veröffentlicht: (2025)
von: Zhang, Zehua, et al.
Veröffentlicht: (2025)
Stay Tuned: An Empirical Study of the Impact of Hyperparameters on LLM Tuning in Real-World Applications
von: Halfon, Alon, et al.
Veröffentlicht: (2024)
von: Halfon, Alon, et al.
Veröffentlicht: (2024)
Steel-LLM:From Scratch to Open Source -- A Personal Journey in Building a Chinese-Centric LLM
von: Gu, Qingshui, et al.
Veröffentlicht: (2025)
von: Gu, Qingshui, et al.
Veröffentlicht: (2025)
Why Synthetic Isn't Real Yet: A Diagnostic Framework for Contact Center Dialogue Generation
von: Devanathan, Rishikesh, et al.
Veröffentlicht: (2025)
von: Devanathan, Rishikesh, et al.
Veröffentlicht: (2025)
Aqulia-Med LLM: Pioneering Full-Process Open-Source Medical Language Models
von: Zhao, Lulu, et al.
Veröffentlicht: (2024)
von: Zhao, Lulu, et al.
Veröffentlicht: (2024)
From Long to Short: LLMs Excel at Trimming Own Reasoning Chains
von: Han, Wei, et al.
Veröffentlicht: (2025)
von: Han, Wei, et al.
Veröffentlicht: (2025)
Can GPT Redefine Medical Understanding? Evaluating GPT on Biomedical Machine Reading Comprehension
von: Vatsal, Shubham, et al.
Veröffentlicht: (2024)
von: Vatsal, Shubham, et al.
Veröffentlicht: (2024)
LiteWebAgent: The Open-Source Suite for VLM-Based Web-Agent Applications
von: Zhang, Danqing, et al.
Veröffentlicht: (2025)
von: Zhang, Danqing, et al.
Veröffentlicht: (2025)
Spot the BlindSpots: Systematic Identification and Quantification of Fine-Grained LLM Biases in Contact Center Summaries
von: Mayilvaghanan, Kawin, et al.
Veröffentlicht: (2025)
von: Mayilvaghanan, Kawin, et al.
Veröffentlicht: (2025)
DetectRL: Benchmarking LLM-Generated Text Detection in Real-World Scenarios
von: Wu, Junchao, et al.
Veröffentlicht: (2024)
von: Wu, Junchao, et al.
Veröffentlicht: (2024)
Do Large Language Models Excel in Complex Logical Reasoning with Formal Language?
von: Jiang, Jin, et al.
Veröffentlicht: (2025)
von: Jiang, Jin, et al.
Veröffentlicht: (2025)
Optimus-1: Hybrid Multimodal Memory Empowered Agents Excel in Long-Horizon Tasks
von: Li, Zaijing, et al.
Veröffentlicht: (2024)
von: Li, Zaijing, et al.
Veröffentlicht: (2024)
F2LLM Technical Report: Matching SOTA Embedding Performance with 6 Million Open-Source Data
von: Zhang, Ziyin, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyin, et al.
Veröffentlicht: (2025)
LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset
von: Zheng, Lianmin, et al.
Veröffentlicht: (2023)
von: Zheng, Lianmin, et al.
Veröffentlicht: (2023)
Can LLM Agents Generate Real-World Evidence? Evaluating Observational Studies in Medical Databases
von: Li, Dubai, et al.
Veröffentlicht: (2026)
von: Li, Dubai, et al.
Veröffentlicht: (2026)
DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
von: DeepSeek-AI, et al.
Veröffentlicht: (2024)
von: DeepSeek-AI, et al.
Veröffentlicht: (2024)
VITA: Towards Open-Source Interactive Omni Multimodal LLM
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
OpenWebVoyager: Building Multimodal Web Agents via Iterative Real-World Exploration, Feedback and Optimization
von: He, Hongliang, et al.
Veröffentlicht: (2024)
von: He, Hongliang, et al.
Veröffentlicht: (2024)
DeepSeek in Healthcare: A Survey of Capabilities, Risks, and Clinical Applications of Open-Source Large Language Models
von: Ye, Jiancheng, et al.
Veröffentlicht: (2025)
von: Ye, Jiancheng, et al.
Veröffentlicht: (2025)
LLM-Match: An Open-Sourced Patient Matching Model Based on Large Language Models and Retrieval-Augmented Generation
von: Li, Xiaodi, et al.
Veröffentlicht: (2025)
von: Li, Xiaodi, et al.
Veröffentlicht: (2025)
VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
von: He, Wei, et al.
Veröffentlicht: (2025)
von: He, Wei, et al.
Veröffentlicht: (2025)
Qabas: An Open-Source Arabic Lexicographic Database
von: Jarrar, Mustafa, et al.
Veröffentlicht: (2024)
von: Jarrar, Mustafa, et al.
Veröffentlicht: (2024)
Lynx: An Open Source Hallucination Evaluation Model
von: Ravi, Selvan Sunitha, et al.
Veröffentlicht: (2024)
von: Ravi, Selvan Sunitha, et al.
Veröffentlicht: (2024)
Orchard: An Open-Source Agentic Modeling Framework
von: Peng, Baolin, et al.
Veröffentlicht: (2026)
von: Peng, Baolin, et al.
Veröffentlicht: (2026)
The Challenges of Evaluating LLM Applications: An Analysis of Automated, Human, and LLM-Based Approaches
von: Abeysinghe, Bhashithe, et al.
Veröffentlicht: (2024)
von: Abeysinghe, Bhashithe, et al.
Veröffentlicht: (2024)
Is Open-Source There Yet? A Comparative Study on Commercial and Open-Source LLMs in Their Ability to Label Chest X-Ray Reports
von: Dorfner, Felix J., et al.
Veröffentlicht: (2024)
von: Dorfner, Felix J., et al.
Veröffentlicht: (2024)
Beyond Perfect APIs: A Comprehensive Evaluation of LLM Agents Under Real-World API Complexity
von: Kim, Doyoung, et al.
Veröffentlicht: (2026)
von: Kim, Doyoung, et al.
Veröffentlicht: (2026)
OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models
von: Wang, Jun, et al.
Veröffentlicht: (2024)
von: Wang, Jun, et al.
Veröffentlicht: (2024)
OpenFActScore: Open-Source Atomic Evaluation of Factuality in Text Generation
von: Lage, Lucas Fonseca, et al.
Veröffentlicht: (2025)
von: Lage, Lucas Fonseca, et al.
Veröffentlicht: (2025)
Leveraging LLM-GNN Integration for Open-World Question Answering over Knowledge Graphs
von: Abdallah, Hussein, et al.
Veröffentlicht: (2026)
von: Abdallah, Hussein, et al.
Veröffentlicht: (2026)
Primus: A Pioneering Collection of Open-Source Datasets for Cybersecurity LLM Training
von: Yu, Yao-Ching, et al.
Veröffentlicht: (2025)
von: Yu, Yao-Ching, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Can GPT Improve the State of Prior Authorization via Guideline Based Automated Question Answering?
von: Vatsal, Shubham, et al.
Veröffentlicht: (2024) -
Emotion Classification in Low and Moderate Resource Languages
von: Tafreshi, Shabnam, et al.
Veröffentlicht: (2024) -
Cornerstones or Stumbling Blocks? Deciphering the Rock Tokens in On-Policy Distillation
von: Jiang, Yuxuan, et al.
Veröffentlicht: (2026) -
CollabStory: Multi-LLM Collaborative Story Generation and Authorship Analysis
von: Venkatraman, Saranya, et al.
Veröffentlicht: (2024) -
MiniLingua: A Small Open-Source LLM for European Languages
von: Aksenova, Anna, et al.
Veröffentlicht: (2025)