Lynx: An Open Source Hallucination Evaluation Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ravi, Selvan Sunitha, Mielczarek, Bartosz, Kannappan, Anand, Kiela, Douwe, Qian, Rebecca |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GLIDER: Grading LLM Interactions and Decisions using Explainable Ranking
von: Deshpande, Darshan, et al.
Veröffentlicht: (2024)
von: Deshpande, Darshan, et al.
Veröffentlicht: (2024)
I am a Strange Dataset: Metalinguistic Tests for Language Models
von: Thrush, Tristan, et al.
Veröffentlicht: (2024)
von: Thrush, Tristan, et al.
Veröffentlicht: (2024)
MEMTRACK: Evaluating Long-Term Memory and State Tracking in Multi-Platform Dynamic Agent Environments
von: Deshpande, Darshan, et al.
Veröffentlicht: (2025)
von: Deshpande, Darshan, et al.
Veröffentlicht: (2025)
TRAIL: Trace Reasoning and Agentic Issue Localization
von: Deshpande, Darshan, et al.
Veröffentlicht: (2025)
von: Deshpande, Darshan, et al.
Veröffentlicht: (2025)
Integrating Supertag Features into Neural Discontinuous Constituent Parsing
von: Mielczarek, Lukas
Veröffentlicht: (2024)
von: Mielczarek, Lukas
Veröffentlicht: (2024)
LMUnit: Fine-grained Evaluation with Natural Language Unit Tests
von: Saad-Falcon, Jon, et al.
Veröffentlicht: (2024)
von: Saad-Falcon, Jon, et al.
Veröffentlicht: (2024)
Browsing Lost Unformed Recollections: A Benchmark for Tip-of-the-Tongue Search and Reasoning
von: CH-Wang, Sky, et al.
Veröffentlicht: (2025)
von: CH-Wang, Sky, et al.
Veröffentlicht: (2025)
Nearest Neighbor Normalization Improves Multimodal Retrieval
von: Chowdhury, Neil, et al.
Veröffentlicht: (2024)
von: Chowdhury, Neil, et al.
Veröffentlicht: (2024)
Defenses & Enablers For Skill Injection Attacks on Terminal Based Agents
von: Fujinuma, Yoshinari, et al.
Veröffentlicht: (2026)
von: Fujinuma, Yoshinari, et al.
Veröffentlicht: (2026)
Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets
von: Tran, Dat, et al.
Veröffentlicht: (2026)
von: Tran, Dat, et al.
Veröffentlicht: (2026)
Generative Representational Instruction Tuning
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2024)
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2024)
Great Models Think Alike and this Undermines AI Oversight
von: Goel, Shashwat, et al.
Veröffentlicht: (2025)
von: Goel, Shashwat, et al.
Veröffentlicht: (2025)
Anchored Preference Optimization and Contrastive Revisions: Addressing Underspecification in Alignment
von: D'Oosterlinck, Karel, et al.
Veröffentlicht: (2024)
von: D'Oosterlinck, Karel, et al.
Veröffentlicht: (2024)
Toolink: Linking Toolkit Creation and Using through Chain-of-Solving on Open-Source Model
von: Qian, Cheng, et al.
Veröffentlicht: (2023)
von: Qian, Cheng, et al.
Veröffentlicht: (2023)
OpenFActScore: Open-Source Atomic Evaluation of Factuality in Text Generation
von: Lage, Lucas Fonseca, et al.
Veröffentlicht: (2025)
von: Lage, Lucas Fonseca, et al.
Veröffentlicht: (2025)
Hallucinations and Key Information Extraction in Medical Texts: A Comprehensive Assessment of Open-Source Large Language Models
von: Das, Anindya Bijoy, et al.
Veröffentlicht: (2025)
von: Das, Anindya Bijoy, et al.
Veröffentlicht: (2025)
HalluciNot: Hallucination Detection Through Context and Common Knowledge Verification
von: Paudel, Bibek, et al.
Veröffentlicht: (2025)
von: Paudel, Bibek, et al.
Veröffentlicht: (2025)
Multilingual Mathematical Reasoning: Advancing Open-Source LLMs in Hindi and English
von: Anand, Avinash, et al.
Veröffentlicht: (2024)
von: Anand, Avinash, et al.
Veröffentlicht: (2024)
Anchor Points: Benchmarking Models with Much Fewer Examples
von: Vivek, Rajan, et al.
Veröffentlicht: (2023)
von: Vivek, Rajan, et al.
Veröffentlicht: (2023)
Benchmarking Reward Hack Detection in Code Environments via Contrastive Analysis
von: Deshpande, Darshan, et al.
Veröffentlicht: (2026)
von: Deshpande, Darshan, et al.
Veröffentlicht: (2026)
Hallucination is Inevitable for LLMs with the Open World Assumption
von: Xu, Bowen
Veröffentlicht: (2025)
von: Xu, Bowen
Veröffentlicht: (2025)
OLMoE: Open Mixture-of-Experts Language Models
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2024)
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2024)
Performance Evaluation of Open-Source Large Language Models for Assisting Pathology Report Writing in Japanese
von: Kawai, Masataka, et al.
Veröffentlicht: (2026)
von: Kawai, Masataka, et al.
Veröffentlicht: (2026)
Orchard: An Open-Source Agentic Modeling Framework
von: Peng, Baolin, et al.
Veröffentlicht: (2026)
von: Peng, Baolin, et al.
Veröffentlicht: (2026)
PRISM: Probing Reasoning, Instruction, and Source Memory in LLM Hallucinations
von: Wu, Yuhe, et al.
Veröffentlicht: (2026)
von: Wu, Yuhe, et al.
Veröffentlicht: (2026)
Beyond Facts: Evaluating Intent Hallucination in Large Language Models
von: Hao, Yijie, et al.
Veröffentlicht: (2025)
von: Hao, Yijie, et al.
Veröffentlicht: (2025)
Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights
von: Abbas, Alexandra, et al.
Veröffentlicht: (2025)
von: Abbas, Alexandra, et al.
Veröffentlicht: (2025)
Leveraging Diffusion Perturbations for Measuring Fairness in Computer Vision
von: Lui, Nicholas, et al.
Veröffentlicht: (2023)
von: Lui, Nicholas, et al.
Veröffentlicht: (2023)
OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models
von: Wang, Jun, et al.
Veröffentlicht: (2024)
von: Wang, Jun, et al.
Veröffentlicht: (2024)
TinyLlama: An Open-Source Small Language Model
von: Zhang, Peiyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Peiyuan, et al.
Veröffentlicht: (2024)
Benchmarking Open-Source Safety Guard Models: A Comprehensive Evaluation
von: Harsh, Reetu Raj, et al.
Veröffentlicht: (2026)
von: Harsh, Reetu Raj, et al.
Veröffentlicht: (2026)
HypeLoRA: Hyper-Network-Generated LoRA Adapters for Calibrated Language Model Fine-Tuning
von: Trojan, Bartosz, et al.
Veröffentlicht: (2026)
von: Trojan, Bartosz, et al.
Veröffentlicht: (2026)
DETOUR: An Interactive Benchmark for Dual-Agent Search and Reasoning
von: Siyan, Li, et al.
Veröffentlicht: (2026)
von: Siyan, Li, et al.
Veröffentlicht: (2026)
Alleviating Hallucinations of Large Language Models through Induced Hallucinations
von: Zhang, Yue, et al.
Veröffentlicht: (2023)
von: Zhang, Yue, et al.
Veröffentlicht: (2023)
An Open Source Data Contamination Report for Large Language Models
von: Li, Yucheng, et al.
Veröffentlicht: (2023)
von: Li, Yucheng, et al.
Veröffentlicht: (2023)
Luna: An Evaluation Foundation Model to Catch Language Model Hallucinations with High Accuracy and Low Cost
von: Belyi, Masha, et al.
Veröffentlicht: (2024)
von: Belyi, Masha, et al.
Veröffentlicht: (2024)
The System Hallucination Scale (SHS): A Minimal yet Effective Human-Centered Instrument for Evaluating Hallucination-Related Behavior in Large Language Models
von: Müller, Heimo, et al.
Veröffentlicht: (2026)
von: Müller, Heimo, et al.
Veröffentlicht: (2026)
Evaluating Open-Weight Large Language Models for Structured Data Extraction from Narrative Medical Reports Across Multiple Use Cases and Languages
von: Spaanderman, Douwe J., et al.
Veröffentlicht: (2025)
von: Spaanderman, Douwe J., et al.
Veröffentlicht: (2025)
DiaHalu: A Dialogue-level Hallucination Evaluation Benchmark for Large Language Models
von: Chen, Kedi, et al.
Veröffentlicht: (2024)
von: Chen, Kedi, et al.
Veröffentlicht: (2024)
Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models
von: Xiong, Guangzhi, et al.
Veröffentlicht: (2025)
von: Xiong, Guangzhi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
GLIDER: Grading LLM Interactions and Decisions using Explainable Ranking
von: Deshpande, Darshan, et al.
Veröffentlicht: (2024) -
I am a Strange Dataset: Metalinguistic Tests for Language Models
von: Thrush, Tristan, et al.
Veröffentlicht: (2024) -
MEMTRACK: Evaluating Long-Term Memory and State Tracking in Multi-Platform Dynamic Agent Environments
von: Deshpande, Darshan, et al.
Veröffentlicht: (2025) -
TRAIL: Trace Reasoning and Agentic Issue Localization
von: Deshpande, Darshan, et al.
Veröffentlicht: (2025) -
Integrating Supertag Features into Neural Discontinuous Constituent Parsing
von: Mielczarek, Lukas
Veröffentlicht: (2024)