Guardado en:
| Autores principales: | Saadat, Mohammadreza, Nemzer, Steve |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2603.03330 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning
por: Lu, Jinghui, et al.
Publicado: (2025)
por: Lu, Jinghui, et al.
Publicado: (2025)
Fact-Checking with Large Language Models via Probabilistic Certainty and Consistency
por: Wang, Haoran, et al.
Publicado: (2026)
por: Wang, Haoran, et al.
Publicado: (2026)
Sequence-Level Certainty Reduces Hallucination In Knowledge-Grounded Dialogue Generation
por: Wan, Yixin, et al.
Publicado: (2023)
por: Wan, Yixin, et al.
Publicado: (2023)
DayDreamer at CQs-Gen 2025: Generating Critical Questions through Argument Scheme Completion
por: Zhou, Wendi, et al.
Publicado: (2025)
por: Zhou, Wendi, et al.
Publicado: (2025)
Prompt-prompted Adaptive Structured Pruning for Efficient LLM Generation
por: Dong, Harry, et al.
Publicado: (2024)
por: Dong, Harry, et al.
Publicado: (2024)
Certainty-Guided Reasoning in Large Language Models: A Dynamic Thinking Budget Approach
por: Nogueira, João Paulo, et al.
Publicado: (2025)
por: Nogueira, João Paulo, et al.
Publicado: (2025)
Soft-prompt Tuning for Large Language Models to Evaluate Bias
por: Tian, Jacob-Junqi, et al.
Publicado: (2023)
por: Tian, Jacob-Junqi, et al.
Publicado: (2023)
Auto prompting without training labels: An LLM cascade for product quality assessment in e-commerce catalogs
por: Satyadharma, Soham, et al.
Publicado: (2025)
por: Satyadharma, Soham, et al.
Publicado: (2025)
Meaningless is better: hashing bias-inducing words in LLM prompts improves performance in logical reasoning and statistical learning
por: Chadimová, Milena, et al.
Publicado: (2024)
por: Chadimová, Milena, et al.
Publicado: (2024)
Knowledge prompt chaining for semantic modeling
por: Ding, Ning Pei, et al.
Publicado: (2025)
por: Ding, Ning Pei, et al.
Publicado: (2025)
StateAct: Enhancing LLM Base Agents via Self-prompting and State-tracking
por: Rozanov, Nikolai, et al.
Publicado: (2024)
por: Rozanov, Nikolai, et al.
Publicado: (2024)
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations
por: Chaudhary, Manav, et al.
Publicado: (2024)
por: Chaudhary, Manav, et al.
Publicado: (2024)
Scalable Best-of-N Selection for Large Language Models via Self-Certainty
por: Kang, Zhewei, et al.
Publicado: (2025)
por: Kang, Zhewei, et al.
Publicado: (2025)
Fooling LLM graders into giving better grades through neural activity guided adversarial prompting
por: Yamamura, Atsushi, et al.
Publicado: (2024)
por: Yamamura, Atsushi, et al.
Publicado: (2024)
MOSLIM:Align with diverse preferences in prompts through reward classification
por: Zhang, Yu, et al.
Publicado: (2025)
por: Zhang, Yu, et al.
Publicado: (2025)
Efficient multi-prompt evaluation of LLMs
por: Polo, Felipe Maia, et al.
Publicado: (2024)
por: Polo, Felipe Maia, et al.
Publicado: (2024)
Efficient Reasoning for Large Reasoning Language Models via Certainty-Guided Reflection Suppression
por: Huang, Jiameng, et al.
Publicado: (2025)
por: Huang, Jiameng, et al.
Publicado: (2025)
Mis-prompt: Benchmarking Large Language Models for Proactive Error Handling
por: Zeng, Jiayi, et al.
Publicado: (2025)
por: Zeng, Jiayi, et al.
Publicado: (2025)
A Comprehensive Evaluation of LLM Unlearning Robustness under Multi-Turn Interaction
por: Pan, Ruihao, et al.
Publicado: (2026)
por: Pan, Ruihao, et al.
Publicado: (2026)
Improving Complex Reasoning with Dynamic Prompt Corruption: A soft prompt Optimization Approach
por: Fan, Sinan, et al.
Publicado: (2025)
por: Fan, Sinan, et al.
Publicado: (2025)
Trusting CHATGPT: how minor tweaks in the prompts lead to major differences in sentiment classification
por: Cuellar, Jaime E., et al.
Publicado: (2025)
por: Cuellar, Jaime E., et al.
Publicado: (2025)
Language hooks: a modular framework for augmenting LLM reasoning that decouples tool usage from the model and its prompt
por: de Mijolla, Damien, et al.
Publicado: (2024)
por: de Mijolla, Damien, et al.
Publicado: (2024)
Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement
por: Han, Steve, et al.
Publicado: (2025)
por: Han, Steve, et al.
Publicado: (2025)
Do LLMs Align with My Task? Evaluating Text-to-SQL via Dataset Alignment
por: Rafiei, Davood, et al.
Publicado: (2025)
por: Rafiei, Davood, et al.
Publicado: (2025)
RTLC -- Research, Teach-to-Learn, Critique: A three-stage prompting paradigm inspired by the Feynman Learning Technique that lifts LLM-as-judge accuracy on JudgeBench with no fine-tuning
por: Morandi, Andrea
Publicado: (2026)
por: Morandi, Andrea
Publicado: (2026)
Zero-shot prompt-based classification: topic labeling in times of foundation models in German Tweets
por: Münker, Simon, et al.
Publicado: (2024)
por: Münker, Simon, et al.
Publicado: (2024)
CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks
por: Yu, Ping, et al.
Publicado: (2025)
por: Yu, Ping, et al.
Publicado: (2025)
Evil twins are not that evil: Qualitative insights into machine-generated prompts
por: Rakotonirina, Nathanaël Carraz, et al.
Publicado: (2024)
por: Rakotonirina, Nathanaël Carraz, et al.
Publicado: (2024)
RETUYT-INCO at BEA 2026 Shared Task 2: Meta-prompting in Rubric-based Scoring for German
por: Sastre, Ignacio, et al.
Publicado: (2026)
por: Sastre, Ignacio, et al.
Publicado: (2026)
Metaphor identification using large language models: A comparison of RAG, prompt engineering, and fine-tuning
por: Fuoli, Matteo, et al.
Publicado: (2025)
por: Fuoli, Matteo, et al.
Publicado: (2025)
PII-Compass: Guiding LLM training data extraction prompts towards the target PII via grounding
por: Nakka, Krishna Kanth, et al.
Publicado: (2024)
por: Nakka, Krishna Kanth, et al.
Publicado: (2024)
Unleashing the potential of prompt engineering for large language models
por: Chen, Banghao, et al.
Publicado: (2023)
por: Chen, Banghao, et al.
Publicado: (2023)
Evaluate Summarization in Fine-Granularity: Auto Evaluation with LLM
por: Yuan, Dong, et al.
Publicado: (2024)
por: Yuan, Dong, et al.
Publicado: (2024)
GreenTEA: Gradient Descent with Topic-modeling and Evolutionary Auto-prompting
por: Dong, Zheng, et al.
Publicado: (2025)
por: Dong, Zheng, et al.
Publicado: (2025)
SQL-Exchange: Transforming SQL Queries Across Domains
por: Daviran, Mohammadreza, et al.
Publicado: (2025)
por: Daviran, Mohammadreza, et al.
Publicado: (2025)
DaVinci at SemEval-2024 Task 9: Few-shot prompting GPT-3.5 for Unconventional Reasoning
por: Mathur, Suyash Vardhan, et al.
Publicado: (2024)
por: Mathur, Suyash Vardhan, et al.
Publicado: (2024)
Emergent misalignment as prompt sensitivity: A research note
por: Wyse, Tim, et al.
Publicado: (2025)
por: Wyse, Tim, et al.
Publicado: (2025)
LLM Prompt Evaluation for Educational Applications
por: Holmes, Langdon, et al.
Publicado: (2026)
por: Holmes, Langdon, et al.
Publicado: (2026)
Evaluating Metrics for Safety with LLM-as-Judges
por: Clegg, Kester, et al.
Publicado: (2025)
por: Clegg, Kester, et al.
Publicado: (2025)
Three Models of RLHF Annotation: Extension, Evidence, and Authority
por: Coyne, Steve
Publicado: (2026)
por: Coyne, Steve
Publicado: (2026)
Ejemplares similares
-
Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning
por: Lu, Jinghui, et al.
Publicado: (2025) -
Fact-Checking with Large Language Models via Probabilistic Certainty and Consistency
por: Wang, Haoran, et al.
Publicado: (2026) -
Sequence-Level Certainty Reduces Hallucination In Knowledge-Grounded Dialogue Generation
por: Wan, Yixin, et al.
Publicado: (2023) -
DayDreamer at CQs-Gen 2025: Generating Critical Questions through Argument Scheme Completion
por: Zhou, Wendi, et al.
Publicado: (2025) -
Prompt-prompted Adaptive Structured Pruning for Efficient LLM Generation
por: Dong, Harry, et al.
Publicado: (2024)