CREATE: Testing LLMs for Associative Creativity
Fuente:
arXiv
Saved in:
| Main Authors: | Wadhwa, Manya, Roy, Tiasa Singha, Lederman, Harvey, Li, Junyi Jessy, Durrett, Greg |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Using Natural Language Explanations to Rescale Human Judgments
by: Wadhwa, Manya, et al.
Published: (2023)
by: Wadhwa, Manya, et al.
Published: (2023)
Learning to Refine with Fine-Grained Natural Language Feedback
by: Wadhwa, Manya, et al.
Published: (2024)
by: Wadhwa, Manya, et al.
Published: (2024)
EvalAgent: Discovering Implicit Evaluation Criteria from the Web
by: Wadhwa, Manya, et al.
Published: (2025)
by: Wadhwa, Manya, et al.
Published: (2025)
QUDsim: Quantifying Discourse Similarities in LLM-Generated Text
by: Namuduri, Ramya, et al.
Published: (2025)
by: Namuduri, Ramya, et al.
Published: (2025)
Interpreting and Mitigating Unwanted Uncertainty in LLMs
by: Roy, Tiasa Singha, et al.
Published: (2025)
by: Roy, Tiasa Singha, et al.
Published: (2025)
SkillFactory: Self-Distillation For Learning Cognitive Behaviors
by: Sprague, Zayne, et al.
Published: (2025)
by: Sprague, Zayne, et al.
Published: (2025)
Which questions should I answer? Salience Prediction of Inquisitive Questions
by: Wu, Yating, et al.
Published: (2024)
by: Wu, Yating, et al.
Published: (2024)
Can LLMs $\textit{understand}$ Math? -- Exploring the Pitfalls in Mathematical Reasoning
by: Roy, Tiasa Singha, et al.
Published: (2025)
by: Roy, Tiasa Singha, et al.
Published: (2025)
Are Language Models More Like Libraries or Like Librarians? Bibliotechnism, the Novel Reference Problem, and the Attitudes of LLMs
by: Lederman, Harvey, et al.
Published: (2024)
by: Lederman, Harvey, et al.
Published: (2024)
MiniCheck: Efficient Fact-Checking of LLMs on Grounding Documents
by: Tang, Liyan, et al.
Published: (2024)
by: Tang, Liyan, et al.
Published: (2024)
Pairwise or Pointwise? Evaluating Feedback Protocols for Bias in LLM-Based Evaluation
by: Tripathi, Tuhina, et al.
Published: (2025)
by: Tripathi, Tuhina, et al.
Published: (2025)
Molecular Facts: Desiderata for Decontextualization in LLM Fact Verification
by: Gunjal, Anisha, et al.
Published: (2024)
by: Gunjal, Anisha, et al.
Published: (2024)
PropMEND: Hypernetworks for Knowledge Propagation in LLMs
by: Liu, Zeyu Leo, et al.
Published: (2025)
by: Liu, Zeyu Leo, et al.
Published: (2025)
To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
by: Sprague, Zayne, et al.
Published: (2024)
by: Sprague, Zayne, et al.
Published: (2024)
SynthesizRR: Generating Diverse Datasets with Retrieval Augmentation
by: Divekar, Abhishek, et al.
Published: (2024)
by: Divekar, Abhishek, et al.
Published: (2024)
Emergent Introspection in AI is Content-Agnostic
by: Lederman, Harvey, et al.
Published: (2026)
by: Lederman, Harvey, et al.
Published: (2026)
MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning
by: Sprague, Zayne, et al.
Published: (2023)
by: Sprague, Zayne, et al.
Published: (2023)
Understanding Synthetic Context Extension via Retrieval Heads
by: Zhao, Xinyu, et al.
Published: (2024)
by: Zhao, Xinyu, et al.
Published: (2024)
LoFiT: Localized Fine-tuning on LLM Representations
by: Yin, Fangcong, et al.
Published: (2024)
by: Yin, Fangcong, et al.
Published: (2024)
ChartMuseum: Testing Visual Reasoning Capabilities of Large Vision-Language Models
by: Tang, Liyan, et al.
Published: (2025)
by: Tang, Liyan, et al.
Published: (2025)
Robo-Instruct: Simulator-Augmented Instruction Alignment For Finetuning Code LLMs
by: Hu, Zichao, et al.
Published: (2024)
by: Hu, Zichao, et al.
Published: (2024)
LLMs Lean on Priors, Not Programming Language Semantics
by: Thimmaiah, Aditya, et al.
Published: (2025)
by: Thimmaiah, Aditya, et al.
Published: (2025)
X-PARADE: Cross-Lingual Textual Entailment and Information Divergence across Paragraphs
by: Rodriguez, Juan Diego, et al.
Published: (2023)
by: Rodriguez, Juan Diego, et al.
Published: (2023)
Calibrate-Then-Act: Cost-Aware Exploration in LLM Agents
by: Ding, Wenxuan, et al.
Published: (2026)
by: Ding, Wenxuan, et al.
Published: (2026)
Which course? Discourse! Teaching Discourse and Generation in the Era of LLMs
by: Li, Junyi Jessy, et al.
Published: (2026)
by: Li, Junyi Jessy, et al.
Published: (2026)
From Distributional to Overton Pluralism: Investigating Large Language Model Alignment
by: Lake, Thom, et al.
Published: (2024)
by: Lake, Thom, et al.
Published: (2024)
Language Models (Mostly) Do Not Consider Emotion Triggers When Predicting Emotion
by: Singh, Smriti, et al.
Published: (2023)
by: Singh, Smriti, et al.
Published: (2023)
Improving the Distributional Alignment of LLMs using Supervision
by: Kambhatla, Gauri, et al.
Published: (2025)
by: Kambhatla, Gauri, et al.
Published: (2025)
VeriSoftBench: Repository-Scale Formal Verification Benchmarks for Lean
by: Xin, Yutong, et al.
Published: (2026)
by: Xin, Yutong, et al.
Published: (2026)
RankAlign: A Ranking View of the Generator-Validator Gap in Large Language Models
by: Rodriguez, Juan Diego, et al.
Published: (2025)
by: Rodriguez, Juan Diego, et al.
Published: (2025)
Is It JUST Semantics? A Case Study of Discourse Particle Understanding in LLMs
by: Sheffield, William, et al.
Published: (2025)
by: Sheffield, William, et al.
Published: (2025)
Behavioral Analysis of Information Salience in Large Language Models
by: Trienes, Jan, et al.
Published: (2025)
by: Trienes, Jan, et al.
Published: (2025)
WUGNECTIVES: Novel Entity Inferences of Language Models from Discourse Connectives
by: Brubaker, Daniel, et al.
Published: (2025)
by: Brubaker, Daniel, et al.
Published: (2025)
Strategic Dialogue Assessment: The Crooked Path to Innocence
by: Zheng, Anshun Asher, et al.
Published: (2025)
by: Zheng, Anshun Asher, et al.
Published: (2025)
A Long Way to Go: Investigating Length Correlations in RLHF
by: Singhal, Prasann, et al.
Published: (2023)
by: Singhal, Prasann, et al.
Published: (2023)
SPARTA ALIGNMENT: Collectively Aligning Multiple Language Models through Combat
by: Jiang, Yuru, et al.
Published: (2025)
by: Jiang, Yuru, et al.
Published: (2025)
Complex Claim Verification with Evidence Retrieved in the Wild
by: Chen, Jifan, et al.
Published: (2023)
by: Chen, Jifan, et al.
Published: (2023)
D2PO: Discriminator-Guided DPO with Response Evaluation Models
by: Singhal, Prasann, et al.
Published: (2024)
by: Singhal, Prasann, et al.
Published: (2024)
Assessing the Creativity of LLMs in Proposing Novel Solutions to Mathematical Problems
by: Ye, Junyi, et al.
Published: (2024)
by: Ye, Junyi, et al.
Published: (2024)
Detection and Measurement of Syntactic Templates in Generated Text
by: Shaib, Chantal, et al.
Published: (2024)
by: Shaib, Chantal, et al.
Published: (2024)
Similar Items
-
Using Natural Language Explanations to Rescale Human Judgments
by: Wadhwa, Manya, et al.
Published: (2023) -
Learning to Refine with Fine-Grained Natural Language Feedback
by: Wadhwa, Manya, et al.
Published: (2024) -
EvalAgent: Discovering Implicit Evaluation Criteria from the Web
by: Wadhwa, Manya, et al.
Published: (2025) -
QUDsim: Quantifying Discourse Similarities in LLM-Generated Text
by: Namuduri, Ramya, et al.
Published: (2025) -
Interpreting and Mitigating Unwanted Uncertainty in LLMs
by: Roy, Tiasa Singha, et al.
Published: (2025)