Don't Fight Hallucinations, Use Them: Estimating Image Realism using NLI over Atomic Facts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rykov, Elisei, Petrushina, Kseniia, Titova, Kseniia, Panchenko, Alexander, Konovalov, Vasily |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Through the Looking Glass: Common Sense Consistency Evaluation of Weird Images
von: Rykov, Elisei, et al.
Veröffentlicht: (2025)
von: Rykov, Elisei, et al.
Veröffentlicht: (2025)
SmurfCat at SemEval-2024 Task 6: Leveraging Synthetic Data for Hallucination Detection
von: Rykov, Elisei, et al.
Veröffentlicht: (2024)
von: Rykov, Elisei, et al.
Veröffentlicht: (2024)
When Models Lie, We Learn: Multilingual Span-Level Hallucination Detection with PsiloQA
von: Rykov, Elisei, et al.
Veröffentlicht: (2025)
von: Rykov, Elisei, et al.
Veröffentlicht: (2025)
S3: A Simple Strong Sample-effective Multimodal Dialog System
von: Rykov, Elisei, et al.
Veröffentlicht: (2024)
von: Rykov, Elisei, et al.
Veröffentlicht: (2024)
Beyond Detection: Rethinking Education in the Age of AI-writing
von: Marina, Maria, et al.
Veröffentlicht: (2026)
von: Marina, Maria, et al.
Veröffentlicht: (2026)
Atomic Inference for NLI with Generated Facts as Atoms
von: Stacey, Joe, et al.
Veröffentlicht: (2023)
von: Stacey, Joe, et al.
Veröffentlicht: (2023)
SmurfCat at PAN 2024 TextDetox: Alignment of Multilingual Transformers for Text Detoxification
von: Rykov, Elisei, et al.
Veröffentlicht: (2024)
von: Rykov, Elisei, et al.
Veröffentlicht: (2024)
Leveraging LLM Parametric Knowledge for Fact Checking without Retrieval
von: Vazhentsev, Artem, et al.
Veröffentlicht: (2026)
von: Vazhentsev, Artem, et al.
Veröffentlicht: (2026)
Intent Matters: Enhancing AI Tutoring with Fine-Grained Pedagogical Intent Annotation
von: Petukhova, Kseniia, et al.
Veröffentlicht: (2025)
von: Petukhova, Kseniia, et al.
Veröffentlicht: (2025)
A Fully Automated Pipeline for Conversational Discourse Annotation: Tree Scheme Generation and Labeling with Large Language Models
von: Petukhova, Kseniia, et al.
Veröffentlicht: (2025)
von: Petukhova, Kseniia, et al.
Veröffentlicht: (2025)
Towards Reward Modeling for AI Tutors in Math Mistake Remediation
von: Petukhova, Kseniia, et al.
Veröffentlicht: (2026)
von: Petukhova, Kseniia, et al.
Veröffentlicht: (2026)
How Much Knowledge Can You Pack into a LoRA Adapter without Harming LLM?
von: Pletenev, Sergey, et al.
Veröffentlicht: (2025)
von: Pletenev, Sergey, et al.
Veröffentlicht: (2025)
Multimodal Evaluation of Russian-language Architectures
von: Chervyakov, Artem, et al.
Veröffentlicht: (2025)
von: Chervyakov, Artem, et al.
Veröffentlicht: (2025)
The Chronicles of RiDiC: Generating Datasets with Controlled Popularity Distribution for Long-form Factuality Evaluation
von: Braslavski, Pavel, et al.
Veröffentlicht: (2026)
von: Braslavski, Pavel, et al.
Veröffentlicht: (2026)
DeepPavlov at SemEval-2024 Task 8: Leveraging Transfer Learning for Detecting Boundaries of Machine-Generated Texts
von: Voznyuk, Anastasia, et al.
Veröffentlicht: (2024)
von: Voznyuk, Anastasia, et al.
Veröffentlicht: (2024)
Don't Trust Generative Agents to Mimic Communication on Social Networks Unless You Benchmarked their Empirical Realism
von: Münker, Simon, et al.
Veröffentlicht: (2025)
von: Münker, Simon, et al.
Veröffentlicht: (2025)
Anatomy of Unlearning: The Dual Impact of Fact Salience and Model Fine-Tuning
von: Borisiuk, Anna, et al.
Veröffentlicht: (2026)
von: Borisiuk, Anna, et al.
Veröffentlicht: (2026)
Will It Still Be True Tomorrow? Multilingual Evergreen Question Classification to Improve Trustworthy QA
von: Pletenev, Sergey, et al.
Veröffentlicht: (2025)
von: Pletenev, Sergey, et al.
Veröffentlicht: (2025)
If You Don't Understand It, Don't Use It: Eliminating Trojans with Filters Between Layers
von: Hernandez, Adriano
Veröffentlicht: (2024)
von: Hernandez, Adriano
Veröffentlicht: (2024)
Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration
von: Feng, Shangbin, et al.
Veröffentlicht: (2024)
von: Feng, Shangbin, et al.
Veröffentlicht: (2024)
Learning Selective LLM Autonomy from Copilot Feedback in Enterprise Customer Support Workflows
von: Borovkov, Nikita, et al.
Veröffentlicht: (2026)
von: Borovkov, Nikita, et al.
Veröffentlicht: (2026)
NLI under the Microscope: What Atomic Hypothesis Decomposition Reveals
von: Srikanth, Neha, et al.
Veröffentlicht: (2025)
von: Srikanth, Neha, et al.
Veröffentlicht: (2025)
LLM-Independent Adaptive RAG: Let the Question Speak for Itself
von: Marina, Maria, et al.
Veröffentlicht: (2025)
von: Marina, Maria, et al.
Veröffentlicht: (2025)
Don't Let It Hallucinate: Premise Verification via Retrieval-Augmented Logical Reasoning
von: Qin, Yuehan, et al.
Veröffentlicht: (2025)
von: Qin, Yuehan, et al.
Veröffentlicht: (2025)
Your Students Don't Use LLMs Like You Wish They Did
von: Kobler, Sebastian, et al.
Veröffentlicht: (2026)
von: Kobler, Sebastian, et al.
Veröffentlicht: (2026)
Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Literal Extraction, Logical Inference, and Hallucination Risks in Long-Context LLMs
von: Ebrahimzadeh, Amirali, et al.
Veröffentlicht: (2026)
von: Ebrahimzadeh, Amirali, et al.
Veröffentlicht: (2026)
Don't Touch My Diacritics
von: Gorman, Kyle, et al.
Veröffentlicht: (2024)
von: Gorman, Kyle, et al.
Veröffentlicht: (2024)
AITutor-EvalKit: Exploring the Capabilities of AI Tutors
von: Naeem, Numaan, et al.
Veröffentlicht: (2025)
von: Naeem, Numaan, et al.
Veröffentlicht: (2025)
sDPO: Don't Use Your Data All at Once
von: Kim, Dahyun, et al.
Veröffentlicht: (2024)
von: Kim, Dahyun, et al.
Veröffentlicht: (2024)
PetKaz at SemEval-2024 Task 3: Advancing Emotion Classification with an LLM for Emotion-Cause Pair Extraction in Conversations
von: Kazakov, Roman, et al.
Veröffentlicht: (2024)
von: Kazakov, Roman, et al.
Veröffentlicht: (2024)
PetKaz at SemEval-2024 Task 8: Can Linguistics Capture the Specifics of LLM-generated Text?
von: Petukhova, Kseniia, et al.
Veröffentlicht: (2024)
von: Petukhova, Kseniia, et al.
Veröffentlicht: (2024)
Adaptive Retrieval Without Self-Knowledge? Bringing Uncertainty Back Home
von: Moskvoretskii, Viktor, et al.
Veröffentlicht: (2025)
von: Moskvoretskii, Viktor, et al.
Veröffentlicht: (2025)
Don't Pay Attention
von: Hammoud, Mohammad, et al.
Veröffentlicht: (2025)
von: Hammoud, Mohammad, et al.
Veröffentlicht: (2025)
ELTEX: A Framework for Domain-Driven Synthetic Data Generation
von: Razmyslovich, Arina, et al.
Veröffentlicht: (2025)
von: Razmyslovich, Arina, et al.
Veröffentlicht: (2025)
Unifying AI Tutor Evaluation: An Evaluation Taxonomy for Pedagogical Ability Assessment of LLM-Powered AI Tutors
von: Maurya, Kaushal Kumar, et al.
Veröffentlicht: (2024)
von: Maurya, Kaushal Kumar, et al.
Veröffentlicht: (2024)
Taming Object Hallucinations with Verified Atomic Confidence Estimation
von: Liu, Jiarui, et al.
Veröffentlicht: (2025)
von: Liu, Jiarui, et al.
Veröffentlicht: (2025)
Calibrate, Don't Curate: Label-Efficient Estimation from Noisy LLM Judges
von: Li, Yanran
Veröffentlicht: (2026)
von: Li, Yanran
Veröffentlicht: (2026)
Don't Throw Away Your Pretrained Model
von: Feng, Shangbin, et al.
Veröffentlicht: (2025)
von: Feng, Shangbin, et al.
Veröffentlicht: (2025)
Don't Say No: Jailbreaking LLM by Suppressing Refusal
von: Zhou, Yukai, et al.
Veröffentlicht: (2024)
von: Zhou, Yukai, et al.
Veröffentlicht: (2024)
Show, Don't Tell: Uncovering Implicit Character Portrayal using LLMs
von: Jaipersaud, Brandon, et al.
Veröffentlicht: (2024)
von: Jaipersaud, Brandon, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Through the Looking Glass: Common Sense Consistency Evaluation of Weird Images
von: Rykov, Elisei, et al.
Veröffentlicht: (2025) -
SmurfCat at SemEval-2024 Task 6: Leveraging Synthetic Data for Hallucination Detection
von: Rykov, Elisei, et al.
Veröffentlicht: (2024) -
When Models Lie, We Learn: Multilingual Span-Level Hallucination Detection with PsiloQA
von: Rykov, Elisei, et al.
Veröffentlicht: (2025) -
S3: A Simple Strong Sample-effective Multimodal Dialog System
von: Rykov, Elisei, et al.
Veröffentlicht: (2024) -
Beyond Detection: Rethinking Education in the Age of AI-writing
von: Marina, Maria, et al.
Veröffentlicht: (2026)