Atomic Calibration of LLMs in Long-Form Generations
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Caiqi, Yang, Ruihan, Zhang, Zhisong, Huang, Xinting, Yang, Sen, Yu, Dong, Collier, Nigel |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LoGU: Long-form Generation with Uncertainty Expressions
by: Yang, Ruihan, et al.
Published: (2024)
by: Yang, Ruihan, et al.
Published: (2024)
UNCLE: Benchmarking Uncertainty Expressions in Long-Form Generation
by: Yang, Ruihan, et al.
Published: (2025)
by: Yang, Ruihan, et al.
Published: (2025)
LoVeC: Reinforcement Learning for Better Verbalized Confidence in Long-Form Generations
by: Zhang, Caiqi, et al.
Published: (2025)
by: Zhang, Caiqi, et al.
Published: (2025)
All Roads Lead to Rome: Graph-Based Confidence Estimation for Large Language Model Reasoning
by: Zhang, Caiqi, et al.
Published: (2025)
by: Zhang, Caiqi, et al.
Published: (2025)
Conformity in Large Language Models
by: Zhu, Xiaochen, et al.
Published: (2024)
by: Zhu, Xiaochen, et al.
Published: (2024)
LUQ: Long-text Uncertainty Quantification for LLMs
by: Zhang, Caiqi, et al.
Published: (2024)
by: Zhang, Caiqi, et al.
Published: (2024)
Confidence Estimation for LLMs in Multi-turn Interactions
by: Zhang, Caiqi, et al.
Published: (2026)
by: Zhang, Caiqi, et al.
Published: (2026)
Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging
by: Hu, Tiancheng, et al.
Published: (2025)
by: Hu, Tiancheng, et al.
Published: (2025)
Linguistic Calibration of Long-Form Generations
by: Band, Neil, et al.
Published: (2024)
by: Band, Neil, et al.
Published: (2024)
Multi-Trigger Poisoning Amplifies Backdoor Vulnerabilities in LLMs
by: Sivapiromrat, Sanhanat, et al.
Published: (2025)
by: Sivapiromrat, Sanhanat, et al.
Published: (2025)
DYCP: Dynamic Context Pruning for Long-Form Dialogue with LLMs
by: Choi, Nayoung, et al.
Published: (2026)
by: Choi, Nayoung, et al.
Published: (2026)
MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue
by: Zhang, Naifan, et al.
Published: (2026)
by: Zhang, Naifan, et al.
Published: (2026)
When Personalization Meets Reality: A Multi-Faceted Analysis of Personalized Preference Learning
by: Dong, Yijiang River, et al.
Published: (2025)
by: Dong, Yijiang River, et al.
Published: (2025)
Tug-of-war between idioms' figurative and literal interpretations in LLMs
by: Oh, Soyoung, et al.
Published: (2025)
by: Oh, Soyoung, et al.
Published: (2025)
ReasonGraph: Visualisation of Reasoning Paths
by: Li, Zongqian, et al.
Published: (2025)
by: Li, Zongqian, et al.
Published: (2025)
WebRollback: Enhancing Web Agents with Explicit Rollback Mechanisms
by: Zhang, Zhisong, et al.
Published: (2025)
by: Zhang, Zhisong, et al.
Published: (2025)
A Decomposition Perspective to Long-context Reasoning for LLMs
by: Xiao, Yanling, et al.
Published: (2026)
by: Xiao, Yanling, et al.
Published: (2026)
Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation
by: Yang, Zixuan, et al.
Published: (2026)
by: Yang, Zixuan, et al.
Published: (2026)
Empathy-R1: A Chain-of-Empathy and Reinforcement Learning Framework for Long-Form Mental Health Support
by: Yao, Xianrong, et al.
Published: (2025)
by: Yao, Xianrong, et al.
Published: (2025)
Automated Detection of Pre-training Text in Black-box LLMs
by: Hu, Ruihan, et al.
Published: (2025)
by: Hu, Ruihan, et al.
Published: (2025)
LongWeave: A Long-Form Generation Benchmark Bridging Real-World Relevance and Verifiability
by: Xiao, Zikai, et al.
Published: (2025)
by: Xiao, Zikai, et al.
Published: (2025)
Long-Form Information Alignment Evaluation Beyond Atomic Facts
by: Zheng, Danna, et al.
Published: (2025)
by: Zheng, Danna, et al.
Published: (2025)
Argument Collapse: LLMs Flatten Long-Form Public Debate
by: Kim, Yekyung, et al.
Published: (2026)
by: Kim, Yekyung, et al.
Published: (2026)
Integrating Planning into Single-Turn Long-Form Text Generation
by: Liang, Yi, et al.
Published: (2024)
by: Liang, Yi, et al.
Published: (2024)
Towards Generalization of Block Attention via Automatic Segmentation and Block Distillation
by: Li, Shuaiyi, et al.
Published: (2026)
by: Li, Shuaiyi, et al.
Published: (2026)
SoLoPO: Unlocking Long-Context Capabilities in LLMs via Short-to-Long Preference Optimization
by: Sun, Huashan, et al.
Published: (2025)
by: Sun, Huashan, et al.
Published: (2025)
Decomposing Representation Space into Interpretable Subspaces with Unsupervised Learning
by: Huang, Xinting, et al.
Published: (2025)
by: Huang, Xinting, et al.
Published: (2025)
Beyond the Final Layer: Intermediate Representations for Better Multilingual Calibration in Large Language Models
by: Zhou, Ej, et al.
Published: (2025)
by: Zhou, Ej, et al.
Published: (2025)
Confident Rankings with Fewer Items: Adaptive LLM Evaluation with Continuous Scores
by: Balkır, Esma, et al.
Published: (2026)
by: Balkır, Esma, et al.
Published: (2026)
Deep-Reporter: Deep Research for Grounded Multimodal Long-Form Generation
by: Ye, Fangda, et al.
Published: (2026)
by: Ye, Fangda, et al.
Published: (2026)
Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA
by: Wang, Minzheng, et al.
Published: (2024)
by: Wang, Minzheng, et al.
Published: (2024)
Use of Retrieval-Augmented Large Language Model Agent for Long-Form COVID-19 Fact-Checking
by: Huang, Jingyi, et al.
Published: (2025)
by: Huang, Jingyi, et al.
Published: (2025)
The Curious Case of Factual (Mis)Alignment between LLMs' Short- and Long-Form Answers
by: Islam, Saad Obaid ul, et al.
Published: (2025)
by: Islam, Saad Obaid ul, et al.
Published: (2025)
Long-Context Long-Form Question Answering for Legal Domain
by: Kulkarni, Anagha, et al.
Published: (2026)
by: Kulkarni, Anagha, et al.
Published: (2026)
Draft Model Knows When to Stop: Self-Verification Speculative Decoding for Long-Form Generation
by: Zhang, Ziyin, et al.
Published: (2024)
by: Zhang, Ziyin, et al.
Published: (2024)
On the Transformations across Reward Model, Parameter Update, and In-Context Prompt
by: Cai, Deng, et al.
Published: (2024)
by: Cai, Deng, et al.
Published: (2024)
BalanceRAG: Joint Risk Calibration for Cascaded Retrieval-Augmented Generation
by: Jia, Zijun, et al.
Published: (2026)
by: Jia, Zijun, et al.
Published: (2026)
InversionView: A General-Purpose Method for Reading Information from Neural Activations
by: Huang, Xinting, et al.
Published: (2024)
by: Huang, Xinting, et al.
Published: (2024)
LongSafety: Enhance Safety for Long-Context LLMs
by: Huang, Mianqiu, et al.
Published: (2024)
by: Huang, Mianqiu, et al.
Published: (2024)
Span-level Emotion-Cause-Category Triplet Extraction with Instruction Tuning LLMs and Data Augmentation
by: Li, Xiangju, et al.
Published: (2025)
by: Li, Xiangju, et al.
Published: (2025)
Similar Items
-
LoGU: Long-form Generation with Uncertainty Expressions
by: Yang, Ruihan, et al.
Published: (2024) -
UNCLE: Benchmarking Uncertainty Expressions in Long-Form Generation
by: Yang, Ruihan, et al.
Published: (2025) -
LoVeC: Reinforcement Learning for Better Verbalized Confidence in Long-Form Generations
by: Zhang, Caiqi, et al.
Published: (2025) -
All Roads Lead to Rome: Graph-Based Confidence Estimation for Large Language Model Reasoning
by: Zhang, Caiqi, et al.
Published: (2025) -
Conformity in Large Language Models
by: Zhu, Xiaochen, et al.
Published: (2024)