Surprise Calibration for Better In-Context Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Tan, Zhihang, Hou, Jingrui, Wang, Ping, Hu, Qibiao, Zhu, Peng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Surprisingly Fragile: Assessing and Addressing Prompt Instability in Multimodal Foundation Models
by: Stewart, Ian, et al.
Published: (2024)
by: Stewart, Ian, et al.
Published: (2024)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
by: Ashuach, Tomer, et al.
Published: (2025)
by: Ashuach, Tomer, et al.
Published: (2025)
Retrieve-Refine-Calibrate: A Framework for Complex Claim Fact-Checking
by: Sun, Mingwei, et al.
Published: (2026)
by: Sun, Mingwei, et al.
Published: (2026)
Are Generative Models Underconfident? Better Quality Estimation with Boosted Model Probability
by: Dinh, Tu Anh, et al.
Published: (2025)
by: Dinh, Tu Anh, et al.
Published: (2025)
Hidden Failures in Robustness: Why Supervised Uncertainty Quantification Needs Better Evaluation
by: Stacey, Joe, et al.
Published: (2026)
by: Stacey, Joe, et al.
Published: (2026)
Calibrated Confidence Estimation for Tabular Question Answering
by: Voss, Lukas
Published: (2026)
by: Voss, Lukas
Published: (2026)
LCFO: Long Context and Long Form Output Dataset and Benchmarking
by: Costa-jussà, Marta R., et al.
Published: (2024)
by: Costa-jussà, Marta R., et al.
Published: (2024)
Scaling Synthetic Logical Reasoning Datasets with Context-Sensitive Declarative Grammars
by: Sileo, Damien
Published: (2024)
by: Sileo, Damien
Published: (2024)
Comonadic Morphophonology: A Compositional Framework for Context-Dependent Morphological Rules in Finnish
by: Jang, Yongseok
Published: (2026)
by: Jang, Yongseok
Published: (2026)
Is Less More? Quality, Quantity and Context in Idiom Processing with Natural Language Models
by: Knietaite, Agne, et al.
Published: (2024)
by: Knietaite, Agne, et al.
Published: (2024)
Homogeneous Keys, Heterogeneous Values: Exploiting Local KV Cache Asymmetry for Long-Context LLMs
by: Cui, Wanyun, et al.
Published: (2025)
by: Cui, Wanyun, et al.
Published: (2025)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
by: Saji, Alan, et al.
Published: (2025)
by: Saji, Alan, et al.
Published: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025)
by: Peters, Sydney, et al.
Published: (2025)
SITA: Learning Speaker-Invariant and Tone-Aware Speech Representations for Low-Resource Tonal Languages
by: Xu, Tianyi, et al.
Published: (2026)
by: Xu, Tianyi, et al.
Published: (2026)
AsyncTLS: Efficient Generative LLM Inference with Asynchronous Two-level Sparse Attention
by: Hu, Yuxuan, et al.
Published: (2026)
by: Hu, Yuxuan, et al.
Published: (2026)
Fin-ExBERT: User Intent based Text Extraction in Financial Context using Graph-Augmented BERT and trainable Plugin
by: Sarker, Soumick, et al.
Published: (2025)
by: Sarker, Soumick, et al.
Published: (2025)
Can LLM Graph Reasoning Generalize beyond Pattern Memorization?
by: Zhang, Yizhuo, et al.
Published: (2024)
by: Zhang, Yizhuo, et al.
Published: (2024)
STLM Engineering Report: Dropout
by: Hillier, Dylan, et al.
Published: (2024)
by: Hillier, Dylan, et al.
Published: (2024)
BabelDOC: Better Layout-Preserving PDF Translation via Intermediate Representation
by: Yang, Qi, et al.
Published: (2026)
by: Yang, Qi, et al.
Published: (2026)
LLM-GLOBE: A Benchmark Evaluating the Cultural Values Embedded in LLM Output
by: Karinshak, Elise, et al.
Published: (2024)
by: Karinshak, Elise, et al.
Published: (2024)
Overcoming Long-Context Limitations of State-Space Models via Context-Dependent Sparse Attention
by: Zhan, Zhihao, et al.
Published: (2025)
by: Zhan, Zhihao, et al.
Published: (2025)
UM_FHS at the CLEF 2025 SimpleText Track: Comparing No-Context and Fine-Tune Approaches for GPT-4.1 Models in Sentence and Document-Level Text Simplification
by: Kocbek, Primoz, et al.
Published: (2025)
by: Kocbek, Primoz, et al.
Published: (2025)
HACHIMI: Scalable and Controllable Student Persona Generation via Orchestrated Agents
by: Jiang, Yilin, et al.
Published: (2026)
by: Jiang, Yilin, et al.
Published: (2026)
Evaluating an evidence-guided reinforcement learning framework in aligning light-parameter large language models with decision-making cognition in psychiatric clinical reasoning
by: Lin, Xinxin, et al.
Published: (2026)
by: Lin, Xinxin, et al.
Published: (2026)
Learning to Generate Structured Output with Schema Reinforcement Learning
by: Lu, Yaxi, et al.
Published: (2025)
by: Lu, Yaxi, et al.
Published: (2025)
Learning From Failure: Integrating Negative Examples when Fine-tuning Large Language Models as Agents
by: Wang, Renxi, et al.
Published: (2024)
by: Wang, Renxi, et al.
Published: (2024)
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
by: Collado-Montañez, Jaime, et al.
Published: (2025)
by: Collado-Montañez, Jaime, et al.
Published: (2025)
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)
RAG-Optimized Tibetan Tourism LLMs: Enhancing Accuracy and Personalization
by: Qi, Jinhu, et al.
Published: (2024)
by: Qi, Jinhu, et al.
Published: (2024)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
by: Oketunji, Abiodun Finbarrs
Published: (2023)
by: Oketunji, Abiodun Finbarrs
Published: (2023)
After Retrieval, Before Generation: Enhancing the Trustworthiness of Large Language Models in Retrieval-Augmented Generation
by: Dai, Xinbang, et al.
Published: (2025)
by: Dai, Xinbang, et al.
Published: (2025)
Advancing Expert Specialization for Better MoE
by: Guo, Hongcan, et al.
Published: (2025)
by: Guo, Hongcan, et al.
Published: (2025)
Technical Report of TeleChat2, TeleChat2.5 and T1
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
Better and Worse with Scale: How Contextual Entrainment Diverges with Model Size
by: Kukreja, Dikshant, et al.
Published: (2026)
by: Kukreja, Dikshant, et al.
Published: (2026)
Towards Massive Multilingual Holistic Bias
by: Tan, Xiaoqing Ellen, et al.
Published: (2024)
by: Tan, Xiaoqing Ellen, et al.
Published: (2024)
LinkNER: Linking Local Named Entity Recognition Models to Large Language Models using Uncertainty
by: Zhang, Zhen, et al.
Published: (2024)
by: Zhang, Zhen, et al.
Published: (2024)
Learning Translations via Matrix Completion
by: Wijaya, Derry, et al.
Published: (2024)
by: Wijaya, Derry, et al.
Published: (2024)
AI Can Learn Scientific Taste
by: Tong, Jingqi, et al.
Published: (2026)
by: Tong, Jingqi, et al.
Published: (2026)
Detecting Data Contamination in LLMs via In-Context Learning
by: Zawalski, Michał, et al.
Published: (2025)
by: Zawalski, Michał, et al.
Published: (2025)
Dual Debiasing for Noisy In-Context Learning for Text Generation
by: Liang, Siqi, et al.
Published: (2025)
by: Liang, Siqi, et al.
Published: (2025)
Similar Items
-
Surprisingly Fragile: Assessing and Addressing Prompt Instability in Multimodal Foundation Models
by: Stewart, Ian, et al.
Published: (2024) -
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
by: Ashuach, Tomer, et al.
Published: (2025) -
Retrieve-Refine-Calibrate: A Framework for Complex Claim Fact-Checking
by: Sun, Mingwei, et al.
Published: (2026) -
Are Generative Models Underconfident? Better Quality Estimation with Boosted Model Probability
by: Dinh, Tu Anh, et al.
Published: (2025) -
Hidden Failures in Robustness: Why Supervised Uncertainty Quantification Needs Better Evaluation
by: Stacey, Joe, et al.
Published: (2026)