Choice-75: A Dataset on Decision Branching in Script Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Hou, Zhaoyi Joey, Zhang, Li, Callison-Burch, Chris |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Autorubric: Unifying Rubric-based LLM Evaluation
by: Rao, Delip, et al.
Published: (2026)
by: Rao, Delip, et al.
Published: (2026)
Overhearing LLM Agents: A Survey, Taxonomy, and Roadmap
by: Zhu, Andrew, et al.
Published: (2025)
by: Zhu, Andrew, et al.
Published: (2025)
ThinknCheck: Grounded Claim Verification with Compact, Reasoning-Driven, and Interpretable Models
by: Rao, Delip, et al.
Published: (2026)
by: Rao, Delip, et al.
Published: (2026)
You Have Thirteen Hours in Which to Solve the Labyrinth: Enhancing AI Game Masters with Function Calling
by: Song, Jaewoo, et al.
Published: (2024)
by: Song, Jaewoo, et al.
Published: (2024)
First Steps Towards Overhearing LLM Agents: A Case Study With Dungeons & Dragons Gameplay
by: Zhu, Andrew, et al.
Published: (2025)
by: Zhu, Andrew, et al.
Published: (2025)
Evaluating Vision-Language Models on Bistable Images
by: Panagopoulou, Artemis, et al.
Published: (2024)
by: Panagopoulou, Artemis, et al.
Published: (2024)
FanOutQA: A Multi-Hop, Multi-Document Question Answering Benchmark for Large Language Models
by: Zhu, Andrew, et al.
Published: (2024)
by: Zhu, Andrew, et al.
Published: (2024)
When Verification Fails: How Compositionally Infeasible Claims Escape Rejection
by: Liu, Muxin, et al.
Published: (2026)
by: Liu, Muxin, et al.
Published: (2026)
Improve LLM-based Automatic Essay Scoring with Linguistic Features
by: Hou, Zhaoyi Joey, et al.
Published: (2025)
by: Hou, Zhaoyi Joey, et al.
Published: (2025)
Leveraging Large Models to Evaluate Novel Content: A Case Study on Advertisement Creativity
by: Hou, Zhaoyi Joey, et al.
Published: (2025)
by: Hou, Zhaoyi Joey, et al.
Published: (2025)
ParaGuide: Guided Diffusion Paraphrasers for Plug-and-Play Textual Style Transfer
by: Horvitz, Zachary, et al.
Published: (2023)
by: Horvitz, Zachary, et al.
Published: (2023)
What Do Claim Verification Datasets Actually Test? A Reasoning Trace Analysis
by: Rao, Delip, et al.
Published: (2026)
by: Rao, Delip, et al.
Published: (2026)
Large Language Models Can Self-Improve At Web Agent Tasks
by: Patel, Ajay, et al.
Published: (2024)
by: Patel, Ajay, et al.
Published: (2024)
Contra4: Evaluating Contrastive Cross-Modal Reasoning in Audio, Video, Image, and 3D
by: Panagopoulou, Artemis, et al.
Published: (2025)
by: Panagopoulou, Artemis, et al.
Published: (2025)
Concept Lancet: Image Editing with Compositional Representation Transplant
by: Luo, Jinqi, et al.
Published: (2025)
by: Luo, Jinqi, et al.
Published: (2025)
Domain Gating Ensemble Networks for AI-Generated Text Detection
by: Tripathi, Arihant, et al.
Published: (2025)
by: Tripathi, Arihant, et al.
Published: (2025)
LLM4Branch: Large Language Model for Discovering Efficient Branching Policies of Integer Programs
by: Hou, Zhinan, et al.
Published: (2026)
by: Hou, Zhinan, et al.
Published: (2026)
Multimedia Generative Script Learning for Task Planning
by: Wang, Qingyun, et al.
Published: (2022)
by: Wang, Qingyun, et al.
Published: (2022)
PaCE: Parsimonious Concept Engineering for Large Language Models
by: Luo, Jinqi, et al.
Published: (2024)
by: Luo, Jinqi, et al.
Published: (2024)
Open Problems in Differentiable Social Choice: Learning Mechanisms, Decisions, and Alignment
by: An, Zhiyu, et al.
Published: (2026)
by: An, Zhiyu, et al.
Published: (2026)
WHAT-IF: Exploring Branching Narratives by Meta-Prompting Large Language Models
by: Huang, Runsheng "Anson", et al.
Published: (2024)
by: Huang, Runsheng "Anson", et al.
Published: (2024)
Beyond the Birkhoff Polytope: Spectral-Sphere-Constrained Hyper-Connections
by: Liu, Zhaoyi, et al.
Published: (2026)
by: Liu, Zhaoyi, et al.
Published: (2026)
Unlocking Non-Block-Structured Decisions: Inductive Mining with Choice Graphs
by: Kourani, Humam, et al.
Published: (2025)
by: Kourani, Humam, et al.
Published: (2025)
IB-Net: Initial Branch Network for Variable Decision in Boolean Satisfiability
by: Chan, Tsz Ho, et al.
Published: (2024)
by: Chan, Tsz Ho, et al.
Published: (2024)
Dark Distillation: Backdooring Distilled Datasets without Accessing Raw Data
by: Yang, Ziyuan, et al.
Published: (2025)
by: Yang, Ziyuan, et al.
Published: (2025)
HealthBranches: Synthesizing Clinically-Grounded Question Answering Datasets via Decision Pathways
by: Cosentino, Cristian, et al.
Published: (2025)
by: Cosentino, Cristian, et al.
Published: (2025)
Classical AI vs. LLMs for Decision-Maker Alignment in Health Insurance Choices
by: Mainali, Mallika, et al.
Published: (2025)
by: Mainali, Mallika, et al.
Published: (2025)
Dataset Distillation via Committee Voting
by: Cui, Jiacheng, et al.
Published: (2025)
by: Cui, Jiacheng, et al.
Published: (2025)
SD-VSum: A Method and Dataset for Script-Driven Video Summarization
by: Mylonas, Manolis, et al.
Published: (2025)
by: Mylonas, Manolis, et al.
Published: (2025)
Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why
by: Rao, Delip, et al.
Published: (2026)
by: Rao, Delip, et al.
Published: (2026)
BibTeX Citation Hallucinations in Scientific Publishing Agents: Evaluation and Mitigation
by: Rao, Delip, et al.
Published: (2026)
by: Rao, Delip, et al.
Published: (2026)
A Vietnamese Dataset for Text Segmentation and Multiple Choices Reading Comprehension
by: Hai, Toan Nguyen, et al.
Published: (2025)
by: Hai, Toan Nguyen, et al.
Published: (2025)
Learning to Be A Doctor: Searching for Effective Medical Agent Architectures
by: Zhuang, Yangyang, et al.
Published: (2025)
by: Zhuang, Yangyang, et al.
Published: (2025)
XChoice: Explainable Evaluation of AI-Human Alignment in LLM-based Constrained Choice Decision Making
by: Qi, Weihong, et al.
Published: (2026)
by: Qi, Weihong, et al.
Published: (2026)
CreativityPrism: A Holistic Evaluation Framework for Large Language Model Creativity
by: Hou, Zhaoyi Joey, et al.
Published: (2025)
by: Hou, Zhaoyi Joey, et al.
Published: (2025)
Exploring Student Choice and the Use of Multimodal Generative AI in Programming Learning
by: Hou, Xinying, et al.
Published: (2025)
by: Hou, Xinying, et al.
Published: (2025)
Refine Large Language Model Fine-tuning via Instruction Vector
by: Jiang, Gangwei, et al.
Published: (2024)
by: Jiang, Gangwei, et al.
Published: (2024)
WithdrarXiv: A Large-Scale Dataset for Retraction Study
by: Rao, Delip, et al.
Published: (2024)
by: Rao, Delip, et al.
Published: (2024)
Script Sensitivity: Benchmarking Language Models on Unicode, Romanized and Mixed-Script Sinhala
by: Rajapakse, Minuri, et al.
Published: (2026)
by: Rajapakse, Minuri, et al.
Published: (2026)
Efficiently Generating Expressive Quadruped Behaviors via Language-Guided Preference Learning
by: Clark, Jaden, et al.
Published: (2025)
by: Clark, Jaden, et al.
Published: (2025)
Similar Items
-
Autorubric: Unifying Rubric-based LLM Evaluation
by: Rao, Delip, et al.
Published: (2026) -
Overhearing LLM Agents: A Survey, Taxonomy, and Roadmap
by: Zhu, Andrew, et al.
Published: (2025) -
ThinknCheck: Grounded Claim Verification with Compact, Reasoning-Driven, and Interpretable Models
by: Rao, Delip, et al.
Published: (2026) -
You Have Thirteen Hours in Which to Solve the Labyrinth: Enhancing AI Game Masters with Function Calling
by: Song, Jaewoo, et al.
Published: (2024) -
First Steps Towards Overhearing LLM Agents: A Case Study With Dungeons & Dragons Gameplay
by: Zhu, Andrew, et al.
Published: (2025)