Embedding Trust: Semantic Isotropy Predicts Nonfactuality in Long-Form Text Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Bhardwaj, Dhrupad, Kempe, Julia, Rudner, Tim G. J. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Attacking Bayes: On the Adversarial Robustness of Bayesian Neural Networks
by: Feng, Yunzhen, et al.
Published: (2024)
by: Feng, Yunzhen, et al.
Published: (2024)
Mind the GAP: Improving Robustness to Subpopulation Shifts with Group-Aware Priors
by: Rudner, Tim G. J., et al.
Published: (2024)
by: Rudner, Tim G. J., et al.
Published: (2024)
Text Rationalization for Robust Causal Effect Estimation
by: Zhang, Lijinghua, et al.
Published: (2025)
by: Zhang, Lijinghua, et al.
Published: (2025)
LETS-C: Leveraging Text Embedding for Time Series Classification
by: Kaur, Rachneet, et al.
Published: (2024)
by: Kaur, Rachneet, et al.
Published: (2024)
Adaptive Uncertainty Quantification for Generative AI
by: Kim, Jungeum, et al.
Published: (2024)
by: Kim, Jungeum, et al.
Published: (2024)
Language Models as Causal Effect Generators
by: Bynum, Lucius E. J., et al.
Published: (2024)
by: Bynum, Lucius E. J., et al.
Published: (2024)
A General Framework for Producing Interpretable Semantic Text Embeddings
by: Sun, Yiqun, et al.
Published: (2024)
by: Sun, Yiqun, et al.
Published: (2024)
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers
by: Su, Jingtong, et al.
Published: (2025)
by: Su, Jingtong, et al.
Published: (2025)
Mission Impossible: A Statistical Perspective on Jailbreaking LLMs
by: Su, Jingtong, et al.
Published: (2024)
by: Su, Jingtong, et al.
Published: (2024)
A Causal Lens for Evaluating Faithfulness Metrics
by: Zaman, Kerem, et al.
Published: (2025)
by: Zaman, Kerem, et al.
Published: (2025)
RCT Rejection Sampling for Causal Estimation Evaluation
by: Keith, Katherine A., et al.
Published: (2023)
by: Keith, Katherine A., et al.
Published: (2023)
From Ground Truth to Measurement: A Statistical Framework for Human Labeling
by: Chew, Robert, et al.
Published: (2026)
by: Chew, Robert, et al.
Published: (2026)
Evaluating Interventional Reasoning Capabilities of Large Language Models
by: Kasetty, Tejas, et al.
Published: (2024)
by: Kasetty, Tejas, et al.
Published: (2024)
The Leaderboard Illusion
by: Singh, Shivalika, et al.
Published: (2025)
by: Singh, Shivalika, et al.
Published: (2025)
Majority of the Bests: Improving Best-of-N via Bootstrapping
by: Rakhsha, Amin, et al.
Published: (2025)
by: Rakhsha, Amin, et al.
Published: (2025)
AutoEval Done Right: Using Synthetic Data for Model Evaluation
by: Boyeau, Pierre, et al.
Published: (2024)
by: Boyeau, Pierre, et al.
Published: (2024)
Efficient Exploration for LLMs
by: Dwaracherla, Vikranth, et al.
Published: (2024)
by: Dwaracherla, Vikranth, et al.
Published: (2024)
ALCM: Autonomous LLM-Augmented Causal Discovery Framework
by: Khatibi, Elahe, et al.
Published: (2024)
by: Khatibi, Elahe, et al.
Published: (2024)
Industrial-Grade Smart Troubleshooting through Causal Technical Language Processing: a Proof of Concept
by: Trilla, Alexandre, et al.
Published: (2024)
by: Trilla, Alexandre, et al.
Published: (2024)
(Mis)Fitting: A Survey of Scaling Laws
by: Li, Margaret, et al.
Published: (2025)
by: Li, Margaret, et al.
Published: (2025)
Propagation and Pitfalls: Reasoning-based Assessment of Knowledge Editing through Counterfactual Tasks
by: Hua, Wenyue, et al.
Published: (2024)
by: Hua, Wenyue, et al.
Published: (2024)
Enhancing Causal Reasoning in Large Language Models: A Causal Attribution Model for Precision Fine-Tuning
by: Cai, Hengrui, et al.
Published: (2023)
by: Cai, Hengrui, et al.
Published: (2023)
CLEAR: Can Language Models Really Understand Causal Graphs?
by: Chen, Sirui, et al.
Published: (2024)
by: Chen, Sirui, et al.
Published: (2024)
When Should LLMs Be Less Specific? Selective Abstraction for Reliable Long-Form Text Generation
by: Goren, Shani, et al.
Published: (2026)
by: Goren, Shani, et al.
Published: (2026)
Linguistic Calibration of Long-Form Generations
by: Band, Neil, et al.
Published: (2024)
by: Band, Neil, et al.
Published: (2024)
Conformal Prediction Sets with Improved Conditional Coverage using Trust Scores
by: Kaur, Jivat Neet, et al.
Published: (2025)
by: Kaur, Jivat Neet, et al.
Published: (2025)
Removing Spurious Correlation from Neural Network Interpretations
by: Fotouhi, Milad, et al.
Published: (2024)
by: Fotouhi, Milad, et al.
Published: (2024)
Soft Tokens, Hard Truths
by: Butt, Natasha, et al.
Published: (2025)
by: Butt, Natasha, et al.
Published: (2025)
A Tale of Tails: Model Collapse as a Change of Scaling Laws
by: Dohmatob, Elvis, et al.
Published: (2024)
by: Dohmatob, Elvis, et al.
Published: (2024)
Real-Time Detection of Hallucinated Entities in Long-Form Generation
by: Obeso, Oscar, et al.
Published: (2025)
by: Obeso, Oscar, et al.
Published: (2025)
Empowering Diffusion Models on the Embedding Space for Text Generation
by: Gao, Zhujin, et al.
Published: (2022)
by: Gao, Zhujin, et al.
Published: (2022)
Pre-trained Text-to-Image Diffusion Models Are Versatile Representation Learners for Control
by: Gupta, Gunshi, et al.
Published: (2024)
by: Gupta, Gunshi, et al.
Published: (2024)
Black Box Causal Inference: Effect Estimation via Meta Prediction
by: Bynum, Lucius E. J., et al.
Published: (2025)
by: Bynum, Lucius E. J., et al.
Published: (2025)
Iteration Head: A Mechanistic Study of Chain-of-Thought
by: Cabannes, Vivien, et al.
Published: (2024)
by: Cabannes, Vivien, et al.
Published: (2024)
A Comparative Study of DSPy Teleprompter Algorithms for Aligning Large Language Models Evaluation Metrics to Human Evaluation
by: Sarmah, Bhaskarjit, et al.
Published: (2024)
by: Sarmah, Bhaskarjit, et al.
Published: (2024)
IUQ: Interrogative Uncertainty Quantification for Long-Form Large Language Model Generation
by: Fan, Haozhi, et al.
Published: (2026)
by: Fan, Haozhi, et al.
Published: (2026)
LongRecall: A Structured Approach for Robust Recall Evaluation in Long-Form Text
by: Ardestani, MohamamdJavad, et al.
Published: (2025)
by: Ardestani, MohamamdJavad, et al.
Published: (2025)
Text Detoxification: Data Efficiency, Semantic Preservation and Model Generalization
by: Yu, Jing, et al.
Published: (2025)
by: Yu, Jing, et al.
Published: (2025)
Mind Your Tone: Investigating How Prompt Politeness Affects LLM Accuracy (short paper)
by: Dobariya, Om, et al.
Published: (2025)
by: Dobariya, Om, et al.
Published: (2025)
LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning
by: Wu, Yuhao, et al.
Published: (2025)
by: Wu, Yuhao, et al.
Published: (2025)
Similar Items
-
Attacking Bayes: On the Adversarial Robustness of Bayesian Neural Networks
by: Feng, Yunzhen, et al.
Published: (2024) -
Mind the GAP: Improving Robustness to Subpopulation Shifts with Group-Aware Priors
by: Rudner, Tim G. J., et al.
Published: (2024) -
Text Rationalization for Robust Causal Effect Estimation
by: Zhang, Lijinghua, et al.
Published: (2025) -
LETS-C: Leveraging Text Embedding for Time Series Classification
by: Kaur, Rachneet, et al.
Published: (2024) -
Adaptive Uncertainty Quantification for Generative AI
by: Kim, Jungeum, et al.
Published: (2024)