LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
Fuente:
arXiv
Salvato in:
| Autori principali: | Orgad, Hadas, Toker, Michael, Gekhman, Zorik, Reichart, Roi, Szpektor, Idan, Kotek, Hadas, Belinkov, Yonatan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ReFACT: Updating Text-to-Image Models by Editing the Text Encoder
di: Arad, Dana, et al.
Pubblicazione: (2023)
di: Arad, Dana, et al.
Pubblicazione: (2023)
Position-aware Automatic Circuit Discovery
di: Haklay, Tal, et al.
Pubblicazione: (2025)
di: Haklay, Tal, et al.
Pubblicazione: (2025)
HACK: Hallucinations Along Certainty and Knowledge Axes
di: Simhi, Adi, et al.
Pubblicazione: (2025)
di: Simhi, Adi, et al.
Pubblicazione: (2025)
Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines
di: Toker, Michael, et al.
Pubblicazione: (2024)
di: Toker, Michael, et al.
Pubblicazione: (2024)
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
di: Simhi, Adi, et al.
Pubblicazione: (2024)
di: Simhi, Adi, et al.
Pubblicazione: (2024)
Inside-Out: Hidden Factual Knowledge in LLMs
di: Gekhman, Zorik, et al.
Pubblicazione: (2025)
di: Gekhman, Zorik, et al.
Pubblicazione: (2025)
Distinguishing Ignorance from Error in LLM Hallucinations
di: Simhi, Adi, et al.
Pubblicazione: (2024)
di: Simhi, Adi, et al.
Pubblicazione: (2024)
Pitfalls in Evaluating Interpretability Agents
di: Haklay, Tal, et al.
Pubblicazione: (2026)
di: Haklay, Tal, et al.
Pubblicazione: (2026)
Tokenization Is More Than Compression
di: Schmidt, Craig W., et al.
Pubblicazione: (2024)
di: Schmidt, Craig W., et al.
Pubblicazione: (2024)
Mitigating LLM Hallucinations through Domain-Grounded Tiered Retrieval
di: Haque, Md. Asraful, et al.
Pubblicazione: (2026)
di: Haque, Md. Asraful, et al.
Pubblicazione: (2026)
Reasoning Models Know What's Important, and Encode It in Their Activations
di: Nikankin, Yaniv, et al.
Pubblicazione: (2026)
di: Nikankin, Yaniv, et al.
Pubblicazione: (2026)
Quantifying Genuine Awareness in Hallucination Prediction Beyond Question-Side Shortcuts
di: Seo, Yeongbin, et al.
Pubblicazione: (2025)
di: Seo, Yeongbin, et al.
Pubblicazione: (2025)
MetaCheckGPT -- A Multi-task Hallucination Detector Using LLM Uncertainty and Meta-models
di: Mehta, Rahul, et al.
Pubblicazione: (2024)
di: Mehta, Rahul, et al.
Pubblicazione: (2024)
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
di: Simhi, Adi, et al.
Pubblicazione: (2025)
di: Simhi, Adi, et al.
Pubblicazione: (2025)
d-TreeRPO: Towards More Reliable Policy Optimization for Diffusion Language Models
di: Pan, Leyi, et al.
Pubblicazione: (2025)
di: Pan, Leyi, et al.
Pubblicazione: (2025)
Can LLMs Compute with Reasons?
di: Sandilya, Harshit, et al.
Pubblicazione: (2024)
di: Sandilya, Harshit, et al.
Pubblicazione: (2024)
Reducing Hallucinations in Summarization via Reinforcement Learning with Entity Hallucination Index
di: Katwe, Praveenkumar, et al.
Pubblicazione: (2025)
di: Katwe, Praveenkumar, et al.
Pubblicazione: (2025)
LLMs for Legal Subsumption in German Employment Contracts
di: Wardas, Oliver, et al.
Pubblicazione: (2025)
di: Wardas, Oliver, et al.
Pubblicazione: (2025)
Retrieval-Based Multi-Label Legal Annotation: Extensible, Data-Efficient and Hallucination-Free
di: Zhang, Li, et al.
Pubblicazione: (2026)
di: Zhang, Li, et al.
Pubblicazione: (2026)
Multiplication in Multimodal LLMs: Computation with Text, Image, and Audio Inputs
di: Balter, Samuel G., et al.
Pubblicazione: (2026)
di: Balter, Samuel G., et al.
Pubblicazione: (2026)
When Retrieval Succeeds and Fails: Rethinking Retrieval-Augmented Generation for LLMs
di: Wang, Yongjie, et al.
Pubblicazione: (2025)
di: Wang, Yongjie, et al.
Pubblicazione: (2025)
Co-NAML-LSTUR: A Combined Model with Attentive Multi-View Learning and Long- and Short-term User Representations for News Recommendation
di: Nguyen, Minh Hoang, et al.
Pubblicazione: (2025)
di: Nguyen, Minh Hoang, et al.
Pubblicazione: (2025)
The Unlikely Duel: Evaluating Creative Writing in LLMs through a Unique Scenario
di: Gómez-Rodríguez, Carlos, et al.
Pubblicazione: (2024)
di: Gómez-Rodríguez, Carlos, et al.
Pubblicazione: (2024)
Large Language Models(LLMs) on Tabular Data: Prediction, Generation, and Understanding -- A Survey
di: Fang, Xi, et al.
Pubblicazione: (2024)
di: Fang, Xi, et al.
Pubblicazione: (2024)
Representing LLMs in Prompt Semantic Task Space
di: Kashani, Idan, et al.
Pubblicazione: (2025)
di: Kashani, Idan, et al.
Pubblicazione: (2025)
Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation?
di: Pan, Leyi, et al.
Pubblicazione: (2025)
di: Pan, Leyi, et al.
Pubblicazione: (2025)
Contrasting Linguistic Patterns in Human and LLM-Generated News Text
di: Muñoz-Ortiz, Alberto, et al.
Pubblicazione: (2023)
di: Muñoz-Ortiz, Alberto, et al.
Pubblicazione: (2023)
How much do LLMs learn from negative examples?
di: Hamdan, Shadi, et al.
Pubblicazione: (2025)
di: Hamdan, Shadi, et al.
Pubblicazione: (2025)
Communicative Agents for Slideshow Storytelling Video Generation based on LLMs
di: Fan, Jingxing, et al.
Pubblicazione: (2025)
di: Fan, Jingxing, et al.
Pubblicazione: (2025)
Compression Method Matters: Benchmark-Dependent Output Dynamics in LLM Prompt Compression
di: Johnson, Warren
Pubblicazione: (2026)
di: Johnson, Warren
Pubblicazione: (2026)
RoleRAG: Enhancing LLM Role-Playing via Graph Guided Retrieval
di: Wang, Yongjie, et al.
Pubblicazione: (2025)
di: Wang, Yongjie, et al.
Pubblicazione: (2025)
Semantic Convergence: Investigating Shared Representations Across Scaled LLMs
di: Son, Daniel, et al.
Pubblicazione: (2025)
di: Son, Daniel, et al.
Pubblicazione: (2025)
From Noise to Diversity: Random Embedding Injection in LLM Reasoning
di: Kim, Heejun, et al.
Pubblicazione: (2026)
di: Kim, Heejun, et al.
Pubblicazione: (2026)
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
di: Basu, Abhinaba
Pubblicazione: (2026)
di: Basu, Abhinaba
Pubblicazione: (2026)
Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
di: Simhi, Adi, et al.
Pubblicazione: (2025)
di: Simhi, Adi, et al.
Pubblicazione: (2025)
D-SMART: Enhancing LLM Dialogue Consistency via Dynamic Structured Memory And Reasoning Tree
di: Lei, Xiang, et al.
Pubblicazione: (2025)
di: Lei, Xiang, et al.
Pubblicazione: (2025)
An NLP-Driven Framework for Curriculum-Labor Market Alignment: Schema-Constrained LLM Extraction, ESCO-Anchored Semantic Matching, and Multi-Dimensional Gap Quantification
di: Turaev, Sherzod, et al.
Pubblicazione: (2026)
di: Turaev, Sherzod, et al.
Pubblicazione: (2026)
Large Language Models Generate Harmful Content Using a Distinct, Unified Mechanism
di: Orgad, Hadas, et al.
Pubblicazione: (2026)
di: Orgad, Hadas, et al.
Pubblicazione: (2026)
Tethered Reasoning: Decoupling Entropy from Hallucination in Quantized LLMs via Manifold Steering
di: Atkinson, Craig
Pubblicazione: (2026)
di: Atkinson, Craig
Pubblicazione: (2026)
Prompt Engineering and the Effectiveness of Large Language Models in Enhancing Human Productivity
di: Anam, Rizal Khoirul
Pubblicazione: (2025)
di: Anam, Rizal Khoirul
Pubblicazione: (2025)
Documenti analoghi
-
ReFACT: Updating Text-to-Image Models by Editing the Text Encoder
di: Arad, Dana, et al.
Pubblicazione: (2023) -
Position-aware Automatic Circuit Discovery
di: Haklay, Tal, et al.
Pubblicazione: (2025) -
HACK: Hallucinations Along Certainty and Knowledge Axes
di: Simhi, Adi, et al.
Pubblicazione: (2025) -
Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines
di: Toker, Michael, et al.
Pubblicazione: (2024) -
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
di: Simhi, Adi, et al.
Pubblicazione: (2024)