Salvato in:
| Autori principali: | Ren, Yujie, Gruhlke, Niklas, Lauscher, Anne |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2510.10539 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
How Much Do LLMs Hallucinate across Languages? On Realistic Multilingual Estimation of LLM Hallucination
di: Islam, Saad Obaid ul, et al.
Pubblicazione: (2025)
di: Islam, Saad Obaid ul, et al.
Pubblicazione: (2025)
GRUFF: LLM Pronoun Fidelity, Reasoning, and Biases in German
di: Mewes, Fabian, et al.
Pubblicazione: (2026)
di: Mewes, Fabian, et al.
Pubblicazione: (2026)
Around the World in 24 Hours: Probing LLM Knowledge of Time and Place
di: Holtermann, Carolin, et al.
Pubblicazione: (2025)
di: Holtermann, Carolin, et al.
Pubblicazione: (2025)
Decoding Multilingual Moral Preferences: Unveiling LLM's Biases Through the Moral Machine Experiment
di: Vida, Karina, et al.
Pubblicazione: (2024)
di: Vida, Karina, et al.
Pubblicazione: (2024)
Large Language Models for Human-Machine Collaborative Particle Accelerator Tuning through Natural Language
di: Kaiser, Jan, et al.
Pubblicazione: (2024)
di: Kaiser, Jan, et al.
Pubblicazione: (2024)
Reviewing the Reviewer: Elevating Peer Review Quality through LLM-Guided Feedback
di: Purkayastha, Sukannya, et al.
Pubblicazione: (2026)
di: Purkayastha, Sukannya, et al.
Pubblicazione: (2026)
Building Bridges: A Dataset for Evaluating Gender-Fair Machine Translation into German
di: Lardelli, Manuel, et al.
Pubblicazione: (2024)
di: Lardelli, Manuel, et al.
Pubblicazione: (2024)
The Echoes of Multilinguality: Tracing Cultural Value Shifts during LM Fine-tuning
di: Choenni, Rochelle, et al.
Pubblicazione: (2024)
di: Choenni, Rochelle, et al.
Pubblicazione: (2024)
Multi3Hate: Multimodal, Multilingual, and Multicultural Hate Speech Detection with Vision-Language Models
di: Bui, Minh Duc, et al.
Pubblicazione: (2024)
di: Bui, Minh Duc, et al.
Pubblicazione: (2024)
AutomaTikZ: Text-Guided Synthesis of Scientific Vector Graphics with TikZ
di: Belouadi, Jonas, et al.
Pubblicazione: (2023)
di: Belouadi, Jonas, et al.
Pubblicazione: (2023)
SoS: Analysis of Surface over Semantics in Multilingual Text-To-Image Generation
di: Holtermann, Carolin, et al.
Pubblicazione: (2026)
di: Holtermann, Carolin, et al.
Pubblicazione: (2026)
TempViz: On the Evaluation of Temporal Knowledge in Text-to-Image Models
di: Holtermann, Carolin, et al.
Pubblicazione: (2026)
di: Holtermann, Carolin, et al.
Pubblicazione: (2026)
Local Contrastive Editing of Gender Stereotypes
di: Lutz, Marlene, et al.
Pubblicazione: (2024)
di: Lutz, Marlene, et al.
Pubblicazione: (2024)
GIMMICK -- Globally Inclusive Multimodal Multitask Cultural Knowledge Benchmarking
di: Schneider, Florian, et al.
Pubblicazione: (2025)
di: Schneider, Florian, et al.
Pubblicazione: (2025)
Decision-Making with Deliberation: Meta-reviewing as a Document-grounded Dialogue
di: Purkayastha, Sukannya, et al.
Pubblicazione: (2025)
di: Purkayastha, Sukannya, et al.
Pubblicazione: (2025)
The Lou Dataset -- Exploring the Impact of Gender-Fair Language in German Text Classification
di: Waldis, Andreas, et al.
Pubblicazione: (2024)
di: Waldis, Andreas, et al.
Pubblicazione: (2024)
Stop! In the Name of Flaws: Disentangling Personal Names and Sociodemographic Attributes in NLP
di: Gautam, Vagrant, et al.
Pubblicazione: (2024)
di: Gautam, Vagrant, et al.
Pubblicazione: (2024)
LLM Hallucination Detection: HSAD
di: Li, JinXin, et al.
Pubblicazione: (2025)
di: Li, JinXin, et al.
Pubblicazione: (2025)
The Curious Case of Factual (Mis)Alignment between LLMs' Short- and Long-Form Answers
di: Islam, Saad Obaid ul, et al.
Pubblicazione: (2025)
di: Islam, Saad Obaid ul, et al.
Pubblicazione: (2025)
What the Weight?! A Unified Framework for Zero-Shot Knowledge Composition
di: Holtermann, Carolin, et al.
Pubblicazione: (2024)
di: Holtermann, Carolin, et al.
Pubblicazione: (2024)
Evaluating the Elementary Multilingual Capabilities of Large Language Models with MultiQ
di: Holtermann, Carolin, et al.
Pubblicazione: (2024)
di: Holtermann, Carolin, et al.
Pubblicazione: (2024)
Sensitivity, Performance, Robustness: Deconstructing the Effect of Sociodemographic Prompting
di: Beck, Tilman, et al.
Pubblicazione: (2023)
di: Beck, Tilman, et al.
Pubblicazione: (2023)
Robust Pronoun Fidelity with English LLMs: Are they Reasoning, Repeating, or Just Biased?
di: Gautam, Vagrant, et al.
Pubblicazione: (2024)
di: Gautam, Vagrant, et al.
Pubblicazione: (2024)
Aligned Probing: Relating Toxic Behavior and Model Internals
di: Waldis, Andreas, et al.
Pubblicazione: (2025)
di: Waldis, Andreas, et al.
Pubblicazione: (2025)
LazyReview A Dataset for Uncovering Lazy Thinking in NLP Peer Reviews
di: Purkayastha, Sukannya, et al.
Pubblicazione: (2025)
di: Purkayastha, Sukannya, et al.
Pubblicazione: (2025)
Towards Ethical Multi-Agent Systems of Large Language Models: A Mechanistic Interpretability Perspective
di: Lee, Jae Hee, et al.
Pubblicazione: (2025)
di: Lee, Jae Hee, et al.
Pubblicazione: (2025)
Cultural Authenticity: Comparing LLM Cultural Representations to Native Human Expectations
di: van Liemt, Erin MacMurray, et al.
Pubblicazione: (2026)
di: van Liemt, Erin MacMurray, et al.
Pubblicazione: (2026)
Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment
di: Atasoy, I. F., et al.
Pubblicazione: (2026)
di: Atasoy, I. F., et al.
Pubblicazione: (2026)
Blending Human and LLM Expertise to Detect Hallucinations and Omissions in Mental Health Chatbot Responses
di: Hussain, Khizar, et al.
Pubblicazione: (2026)
di: Hussain, Khizar, et al.
Pubblicazione: (2026)
Span-Level Hallucination Detection for LLM-Generated Answers
di: Elchafei, Passant, et al.
Pubblicazione: (2025)
di: Elchafei, Passant, et al.
Pubblicazione: (2025)
WinoPron: Revisiting English Winogender Schemas for Consistency, Coverage, and Grammatical Case
di: Gautam, Vagrant, et al.
Pubblicazione: (2024)
di: Gautam, Vagrant, et al.
Pubblicazione: (2024)
Steer LLM Latents for Hallucination Detection
di: Park, Seongheon, et al.
Pubblicazione: (2025)
di: Park, Seongheon, et al.
Pubblicazione: (2025)
PRISM: Probing Reasoning, Instruction, and Source Memory in LLM Hallucinations
di: Wu, Yuhe, et al.
Pubblicazione: (2026)
di: Wu, Yuhe, et al.
Pubblicazione: (2026)
ScaLearn: Simple and Highly Parameter-Efficient Task Transfer by Learning to Scale
di: Frohmann, Markus, et al.
Pubblicazione: (2023)
di: Frohmann, Markus, et al.
Pubblicazione: (2023)
HalluCounter: Reference-free LLM Hallucination Detection in the Wild!
di: Urlana, Ashok, et al.
Pubblicazione: (2025)
di: Urlana, Ashok, et al.
Pubblicazione: (2025)
Large Language Models Discriminate Against Speakers of German Dialects
di: Bui, Minh Duc, et al.
Pubblicazione: (2025)
di: Bui, Minh Duc, et al.
Pubblicazione: (2025)
InterrogateLLM: Zero-Resource Hallucination Detection in LLM-Generated Answers
di: Yehuda, Yakir, et al.
Pubblicazione: (2024)
di: Yehuda, Yakir, et al.
Pubblicazione: (2024)
Ethical Concern Identification in NLP: A Corpus of ACL Anthology Ethics Statements
di: Karamolegkou, Antonia, et al.
Pubblicazione: (2024)
di: Karamolegkou, Antonia, et al.
Pubblicazione: (2024)
HaloScope: Harnessing Unlabeled LLM Generations for Hallucination Detection
di: Du, Xuefeng, et al.
Pubblicazione: (2024)
di: Du, Xuefeng, et al.
Pubblicazione: (2024)
Hallucination Detection and Hallucination Mitigation: An Investigation
di: Luo, Junliang, et al.
Pubblicazione: (2024)
di: Luo, Junliang, et al.
Pubblicazione: (2024)
Documenti analoghi
-
How Much Do LLMs Hallucinate across Languages? On Realistic Multilingual Estimation of LLM Hallucination
di: Islam, Saad Obaid ul, et al.
Pubblicazione: (2025) -
GRUFF: LLM Pronoun Fidelity, Reasoning, and Biases in German
di: Mewes, Fabian, et al.
Pubblicazione: (2026) -
Around the World in 24 Hours: Probing LLM Knowledge of Time and Place
di: Holtermann, Carolin, et al.
Pubblicazione: (2025) -
Decoding Multilingual Moral Preferences: Unveiling LLM's Biases Through the Moral Machine Experiment
di: Vida, Karina, et al.
Pubblicazione: (2024) -
Large Language Models for Human-Machine Collaborative Particle Accelerator Tuning through Natural Language
di: Kaiser, Jan, et al.
Pubblicazione: (2024)