QUIET: A Multi-Blank Cascaded Story Cloze Benchmark for LLM Creative Generation Capability
Fuente:
arXiv
Guardado en:
| Autores principales: | Zou, Bo, Xu, Chao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Calibrated Surprise: An Information-Theoretic Account of Creative Quality
por: Zou, Bo, et al.
Publicado: (2026)
por: Zou, Bo, et al.
Publicado: (2026)
Fill in the Blank: Exploring and Enhancing LLM Capabilities for Backward Reasoning in Math Word Problems
por: Deb, Aniruddha, et al.
Publicado: (2023)
por: Deb, Aniruddha, et al.
Publicado: (2023)
CDGP: Automatic Cloze Distractor Generation based on Pre-trained Language Model
por: Chiang, Shang-Hsuan, et al.
Publicado: (2024)
por: Chiang, Shang-Hsuan, et al.
Publicado: (2024)
BC Protocol: Structured Dual-Expert Dialogue for Eliciting High-Quality Chain-of-Thought Post-Training Data
por: Zou, Bo, et al.
Publicado: (2026)
por: Zou, Bo, et al.
Publicado: (2026)
Unveiling Intrinsic Dimension of Texts: from Academic Abstract to Creative Story
por: Pedashenko, Vladislav, et al.
Publicado: (2025)
por: Pedashenko, Vladislav, et al.
Publicado: (2025)
Cooking Up Creativity: Enhancing LLM Creativity through Structured Recombination
por: Mizrahi, Moran, et al.
Publicado: (2025)
por: Mizrahi, Moran, et al.
Publicado: (2025)
MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM Tutors
por: Macina, Jakub, et al.
Publicado: (2025)
por: Macina, Jakub, et al.
Publicado: (2025)
ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
por: Lu, Jiarui, et al.
Publicado: (2024)
por: Lu, Jiarui, et al.
Publicado: (2024)
CodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesis
por: Wei, Anjiang, et al.
Publicado: (2025)
por: Wei, Anjiang, et al.
Publicado: (2025)
ModelGPT: Unleashing LLM's Capabilities for Tailored Model Generation
por: Tang, Zihao, et al.
Publicado: (2024)
por: Tang, Zihao, et al.
Publicado: (2024)
PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier
por: Jiang, Yuhua, et al.
Publicado: (2025)
por: Jiang, Yuhua, et al.
Publicado: (2025)
ECG-Reasoning-Benchmark: A Benchmark for Evaluating Clinical Reasoning Capabilities in ECG Interpretation
por: Oh, Jungwoo, et al.
Publicado: (2026)
por: Oh, Jungwoo, et al.
Publicado: (2026)
Is Escalation Worth It? A Decision-Theoretic Characterization of LLM Cascades
por: Bouchard, Dylan
Publicado: (2026)
por: Bouchard, Dylan
Publicado: (2026)
ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities
por: Karger, Ezra, et al.
Publicado: (2024)
por: Karger, Ezra, et al.
Publicado: (2024)
Elias in the Lighthouse, Again? Diagnosing Low Diversity in LLM Stories
por: Hamilton, Sil, et al.
Publicado: (2026)
por: Hamilton, Sil, et al.
Publicado: (2026)
Exploring and Benchmarking the Planning Capabilities of Large Language Models
por: Bohnet, Bernd, et al.
Publicado: (2024)
por: Bohnet, Bernd, et al.
Publicado: (2024)
Nemotron-Cascade: Scaling Cascaded Reinforcement Learning for General-Purpose Reasoning Models
por: Wang, Boxin, et al.
Publicado: (2025)
por: Wang, Boxin, et al.
Publicado: (2025)
Attention Flows: Tracing LLM Conceptual Engagement via Story Summaries
por: Hicke, Rebecca M. M., et al.
Publicado: (2026)
por: Hicke, Rebecca M. M., et al.
Publicado: (2026)
Affective and Dynamic Beam Search for Story Generation
por: Huang, Tenghao, et al.
Publicado: (2023)
por: Huang, Tenghao, et al.
Publicado: (2023)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
por: Andriushchenko, Maksym, et al.
Publicado: (2024)
por: Andriushchenko, Maksym, et al.
Publicado: (2024)
Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation
por: Yang, Zhuolin, et al.
Publicado: (2026)
por: Yang, Zhuolin, et al.
Publicado: (2026)
HindSight: Evaluating LLM-Generated Research Ideas via Future Impact
por: Jiang, Bo
Publicado: (2026)
por: Jiang, Bo
Publicado: (2026)
Capability Instruction Tuning: A New Paradigm for Dynamic LLM Routing
por: Zhang, Yi-Kai, et al.
Publicado: (2025)
por: Zhang, Yi-Kai, et al.
Publicado: (2025)
Creative Beam Search: LLM-as-a-Judge For Improving Response Generation
por: Franceschelli, Giorgio, et al.
Publicado: (2024)
por: Franceschelli, Giorgio, et al.
Publicado: (2024)
Efficient Contextual LLM Cascades through Budget-Constrained Policy Learning
por: Zhang, Xuechen, et al.
Publicado: (2024)
por: Zhang, Xuechen, et al.
Publicado: (2024)
CaresAI at BioCreative IX Track 1 -- LLM for Biomedical QA
por: Abdel-Salam, Reem, et al.
Publicado: (2025)
por: Abdel-Salam, Reem, et al.
Publicado: (2025)
CTBench: A Comprehensive Benchmark for Evaluating Language Model Capabilities in Clinical Trial Design
por: Neehal, Nafis, et al.
Publicado: (2024)
por: Neehal, Nafis, et al.
Publicado: (2024)
MMLU-SR: A Benchmark for Stress-Testing Reasoning Capability of Large Language Models
por: Wang, Wentian, et al.
Publicado: (2024)
por: Wang, Wentian, et al.
Publicado: (2024)
Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability
por: Lin, Zicheng, et al.
Publicado: (2024)
por: Lin, Zicheng, et al.
Publicado: (2024)
Enhancing the General Agent Capabilities of Low-Parameter LLMs through Tuning and Multi-Branch Reasoning
por: Zhou, Qinhao, et al.
Publicado: (2024)
por: Zhou, Qinhao, et al.
Publicado: (2024)
Breaking the Capability Ceiling of LLM Post-Training by Reintroducing Markov States
por: Yuan, Yurun, et al.
Publicado: (2026)
por: Yuan, Yurun, et al.
Publicado: (2026)
ALHD: A Large-Scale and Multigenre Benchmark Dataset for Arabic LLM-Generated Text Detection
por: Khairallah, Ali, et al.
Publicado: (2025)
por: Khairallah, Ali, et al.
Publicado: (2025)
Weaver: Foundation Models for Creative Writing
por: Wang, Tiannan, et al.
Publicado: (2024)
por: Wang, Tiannan, et al.
Publicado: (2024)
AraLingBench A Human-Annotated Benchmark for Evaluating Arabic Linguistic Capabilities of Large Language Models
por: Zbeeb, Mohammad, et al.
Publicado: (2025)
por: Zbeeb, Mohammad, et al.
Publicado: (2025)
Benchmarking the Capabilities of Large Language Models in Transportation System Engineering: Accuracy, Consistency, and Reasoning Behaviors
por: Syed, Usman, et al.
Publicado: (2024)
por: Syed, Usman, et al.
Publicado: (2024)
Top-H Decoding: Adapting the Creativity and Coherence with Bounded Entropy in Text Generation
por: Potraghloo, Erfan Baghaei, et al.
Publicado: (2025)
por: Potraghloo, Erfan Baghaei, et al.
Publicado: (2025)
CreativityBench: Evaluating Agent Creative Reasoning via Affordance-Based Tool Repurposing
por: Qian, Cheng, et al.
Publicado: (2026)
por: Qian, Cheng, et al.
Publicado: (2026)
Confident or Seek Stronger: Exploring Uncertainty-Based On-device LLM Routing From Benchmarking to Generalization
por: Chuang, Yu-Neng, et al.
Publicado: (2025)
por: Chuang, Yu-Neng, et al.
Publicado: (2025)
Mutual Reinforcement of LLM Dialogue Synthesis and Summarization Capabilities for Few-Shot Dialogue Summarization
por: Lu, Yen-Ju, et al.
Publicado: (2025)
por: Lu, Yen-Ju, et al.
Publicado: (2025)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
por: Tan, Sijun, et al.
Publicado: (2024)
por: Tan, Sijun, et al.
Publicado: (2024)
Ejemplares similares
-
Calibrated Surprise: An Information-Theoretic Account of Creative Quality
por: Zou, Bo, et al.
Publicado: (2026) -
Fill in the Blank: Exploring and Enhancing LLM Capabilities for Backward Reasoning in Math Word Problems
por: Deb, Aniruddha, et al.
Publicado: (2023) -
CDGP: Automatic Cloze Distractor Generation based on Pre-trained Language Model
por: Chiang, Shang-Hsuan, et al.
Publicado: (2024) -
BC Protocol: Structured Dual-Expert Dialogue for Eliciting High-Quality Chain-of-Thought Post-Training Data
por: Zou, Bo, et al.
Publicado: (2026) -
Unveiling Intrinsic Dimension of Texts: from Academic Abstract to Creative Story
por: Pedashenko, Vladislav, et al.
Publicado: (2025)