Evolutionary Search for Automated Design of Uncertainty Quantification Methods
Fuente:
arXiv
Guardado en:
| Autores principales: | Seleznyov, Mikhail, Korbut, Daniil, Moskvoretskii, Viktor, Somov, Oleg, Panchenko, Alexander, Tutubalina, Elena |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs
por: Seleznyov, Mikhail, et al.
Publicado: (2025)
por: Seleznyov, Mikhail, et al.
Publicado: (2025)
Breaking the Chain: A Causal Analysis of LLM Faithfulness to Intermediate Structures
por: Somov, Oleg, et al.
Publicado: (2026)
por: Somov, Oleg, et al.
Publicado: (2026)
Leveraging LLM Parametric Knowledge for Fact Checking without Retrieval
por: Vazhentsev, Artem, et al.
Publicado: (2026)
por: Vazhentsev, Artem, et al.
Publicado: (2026)
Confidence Estimation for Error Detection in Text-to-SQL Systems
por: Somov, Oleg, et al.
Publicado: (2025)
por: Somov, Oleg, et al.
Publicado: (2025)
The benefits of query-based KGQA systems for complex and temporal questions in LLM era
por: Alekseev, Artem, et al.
Publicado: (2025)
por: Alekseev, Artem, et al.
Publicado: (2025)
<think> So let's replace this phrase with insult... </think> Lessons learned from generation of toxic texts with LLMs
por: Pletenev, Sergey, et al.
Publicado: (2025)
por: Pletenev, Sergey, et al.
Publicado: (2025)
Harnessing non-adversarial robustness in large language models
por: Zhou, Qinghua, et al.
Publicado: (2026)
por: Zhou, Qinghua, et al.
Publicado: (2026)
The Chronicles of RiDiC: Generating Datasets with Controlled Popularity Distribution for Long-form Factuality Evaluation
por: Braslavski, Pavel, et al.
Publicado: (2026)
por: Braslavski, Pavel, et al.
Publicado: (2026)
xCOMET-lite: Bridging the Gap Between Efficiency and Quality in Learned MT Evaluation Metrics
por: Larionov, Daniil, et al.
Publicado: (2024)
por: Larionov, Daniil, et al.
Publicado: (2024)
SynthDetoxM: Modern LLMs are Few-Shot Parallel Detoxification Data Annotators
por: Moskovskiy, Daniil, et al.
Publicado: (2025)
por: Moskovskiy, Daniil, et al.
Publicado: (2025)
Tracing Persona Vectors Through LLM Pretraining
por: Moskvoretskii, Viktor, et al.
Publicado: (2026)
por: Moskvoretskii, Viktor, et al.
Publicado: (2026)
Fact-Checking the Output of Large Language Models via Token-Level Uncertainty Quantification
por: Fadeeva, Ekaterina, et al.
Publicado: (2024)
por: Fadeeva, Ekaterina, et al.
Publicado: (2024)
Facilitating large language model Russian adaptation with Learned Embedding Propagation
por: Tikhomirov, Mikhail, et al.
Publicado: (2024)
por: Tikhomirov, Mikhail, et al.
Publicado: (2024)
Evaluating Uncertainty Quantification Methods in Argumentative Large Language Models
por: Zhou, Kevin, et al.
Publicado: (2025)
por: Zhou, Kevin, et al.
Publicado: (2025)
SparseGrad: A Selective Method for Efficient Fine-tuning of MLP Layers
por: Chekalina, Viktoriia, et al.
Publicado: (2024)
por: Chekalina, Viktoriia, et al.
Publicado: (2024)
MultiParaDetox: Extending Text Detoxification with Parallel Data to New Languages
por: Dementieva, Daryna, et al.
Publicado: (2024)
por: Dementieva, Daryna, et al.
Publicado: (2024)
S3: A Simple Strong Sample-effective Multimodal Dialog System
por: Rykov, Elisei, et al.
Publicado: (2024)
por: Rykov, Elisei, et al.
Publicado: (2024)
TaxoLLaMA: WordNet-based Model for Solving Multiple Lexical Semantic Tasks
por: Moskvoretskii, Viktor, et al.
Publicado: (2024)
por: Moskvoretskii, Viktor, et al.
Publicado: (2024)
Agentic Uncertainty Quantification
por: Zhang, Jiaxin, et al.
Publicado: (2026)
por: Zhang, Jiaxin, et al.
Publicado: (2026)
RuCCoD: Towards Automated ICD Coding in Russian
por: Nesterov, Aleksandr, et al.
Publicado: (2025)
por: Nesterov, Aleksandr, et al.
Publicado: (2025)
Anatomy of Unlearning: The Dual Impact of Fact Salience and Model Fine-Tuning
por: Borisiuk, Anna, et al.
Publicado: (2026)
por: Borisiuk, Anna, et al.
Publicado: (2026)
Emergent Misalignment via In-Context Learning: Narrow in-context examples can produce broadly misaligned LLMs
por: Afonin, Nikita, et al.
Publicado: (2025)
por: Afonin, Nikita, et al.
Publicado: (2025)
Combining Confidence Elicitation and Sample-based Methods for Uncertainty Quantification in Misinformation Mitigation
por: Rivera, Mauricio, et al.
Publicado: (2024)
por: Rivera, Mauricio, et al.
Publicado: (2024)
MAQA: Evaluating Uncertainty Quantification in LLMs Regarding Data Uncertainty
por: Yang, Yongjin, et al.
Publicado: (2024)
por: Yang, Yongjin, et al.
Publicado: (2024)
BALI: Enhancing Biomedical Language Representations through Knowledge Graph and Language Model Alignment
por: Sakhovskiy, Andrey, et al.
Publicado: (2025)
por: Sakhovskiy, Andrey, et al.
Publicado: (2025)
CSS: Contrastive Semantic Similarity for Uncertainty Quantification of LLMs
por: Ao, Shuang, et al.
Publicado: (2024)
por: Ao, Shuang, et al.
Publicado: (2024)
Adaptive Retrieval Without Self-Knowledge? Bringing Uncertainty Back Home
por: Moskvoretskii, Viktor, et al.
Publicado: (2025)
por: Moskvoretskii, Viktor, et al.
Publicado: (2025)
SPUQ: Perturbation-Based Uncertainty Quantification for Large Language Models
por: Gao, Xiang, et al.
Publicado: (2024)
por: Gao, Xiang, et al.
Publicado: (2024)
Uncertainty-Based Methods for Automated Process Reward Data Construction and Output Aggregation in Mathematical Reasoning
por: Han, Jiuzhou, et al.
Publicado: (2025)
por: Han, Jiuzhou, et al.
Publicado: (2025)
Semantic Consistency-Based Uncertainty Quantification for Factuality in Radiology Report Generation
por: Wang, Chenyu, et al.
Publicado: (2024)
por: Wang, Chenyu, et al.
Publicado: (2024)
Uncertainty Quantification of Large Language Models through Multi-Dimensional Responses
por: Chen, Tiejin, et al.
Publicado: (2025)
por: Chen, Tiejin, et al.
Publicado: (2025)
Uncertainty Quantification in Large Language Models Through Convex Hull Analysis
por: Catak, Ferhat Ozgur, et al.
Publicado: (2024)
por: Catak, Ferhat Ozgur, et al.
Publicado: (2024)
Do I look like a `cat.n.01` to you? A Taxonomy Image Generation Benchmark
por: Moskvoretskii, Viktor, et al.
Publicado: (2025)
por: Moskvoretskii, Viktor, et al.
Publicado: (2025)
SmurfCat at SemEval-2024 Task 6: Leveraging Synthetic Data for Hallucination Detection
por: Rykov, Elisei, et al.
Publicado: (2024)
por: Rykov, Elisei, et al.
Publicado: (2024)
Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic Similarities
por: Nikitin, Alexander, et al.
Publicado: (2024)
por: Nikitin, Alexander, et al.
Publicado: (2024)
Mind the Ambiguity: Aleatoric Uncertainty Quantification in LLMs for Safe Medical Question Answering
por: Liu, Yaokun, et al.
Publicado: (2026)
por: Liu, Yaokun, et al.
Publicado: (2026)
SIMBA UQ: Similarity-Based Aggregation for Uncertainty Quantification in Large Language Models
por: Bhattacharjya, Debarun, et al.
Publicado: (2025)
por: Bhattacharjya, Debarun, et al.
Publicado: (2025)
Optimizing Instruction Synthesis: Effective Exploration of Evolutionary Space with Tree Search
por: Li, Chenglin, et al.
Publicado: (2024)
por: Li, Chenglin, et al.
Publicado: (2024)
Search Wisely: Mitigating Sub-optimal Agentic Searches By Reducing Uncertainty
por: Wu, Peilin, et al.
Publicado: (2025)
por: Wu, Peilin, et al.
Publicado: (2025)
Calibrating Uncertainty Quantification of Multi-Modal LLMs using Grounding
por: Padhi, Trilok, et al.
Publicado: (2025)
por: Padhi, Trilok, et al.
Publicado: (2025)
Ejemplares similares
-
When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs
por: Seleznyov, Mikhail, et al.
Publicado: (2025) -
Breaking the Chain: A Causal Analysis of LLM Faithfulness to Intermediate Structures
por: Somov, Oleg, et al.
Publicado: (2026) -
Leveraging LLM Parametric Knowledge for Fact Checking without Retrieval
por: Vazhentsev, Artem, et al.
Publicado: (2026) -
Confidence Estimation for Error Detection in Text-to-SQL Systems
por: Somov, Oleg, et al.
Publicado: (2025) -
The benefits of query-based KGQA systems for complex and temporal questions in LLM era
por: Alekseev, Artem, et al.
Publicado: (2025)