MacGyver: Are Large Language Models Creative Problem Solvers?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tian, Yufei, Ravichander, Abhilasha, Qin, Lianhui, Bras, Ronan Le, Marjieh, Raja, Peng, Nanyun, Choi, Yejin, Griffiths, Thomas L., Brahman, Faeze |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RESTOR: Knowledge Recovery in Machine Unlearning
von: Rezaei, Keivan, et al.
Veröffentlicht: (2024)
von: Rezaei, Keivan, et al.
Veröffentlicht: (2024)
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2024)
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2024)
Agent Lumos: Unified and Modular Training for Open-Source Language Agents
von: Yin, Da, et al.
Veröffentlicht: (2023)
von: Yin, Da, et al.
Veröffentlicht: (2023)
Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
von: Jung, Jaehun, et al.
Veröffentlicht: (2024)
von: Jung, Jaehun, et al.
Veröffentlicht: (2024)
HALoGEN: Fantastic LLM Hallucinations and Where to Find Them
von: Ravichander, Abhilasha, et al.
Veröffentlicht: (2025)
von: Ravichander, Abhilasha, et al.
Veröffentlicht: (2025)
Multi-Attribute Constraint Satisfaction via Language Model Rewriting
von: Baheti, Ashutosh, et al.
Veröffentlicht: (2024)
von: Baheti, Ashutosh, et al.
Veröffentlicht: (2024)
Leftover Lunch: Advantage-based Offline Reinforcement Learning for Language Models
von: Baheti, Ashutosh, et al.
Veröffentlicht: (2023)
von: Baheti, Ashutosh, et al.
Veröffentlicht: (2023)
Creativity Support in the Age of Large Language Models: An Empirical Study Involving Emerging Writers
von: Chakrabarty, Tuhin, et al.
Veröffentlicht: (2023)
von: Chakrabarty, Tuhin, et al.
Veröffentlicht: (2023)
From Dogwhistles to Bullhorns: Unveiling Coded Rhetoric with Language Models
von: Mendelsohn, Julia, et al.
Veröffentlicht: (2023)
von: Mendelsohn, Julia, et al.
Veröffentlicht: (2023)
WildHallucinations: Evaluating Long-form Factuality in LLMs with Real-World Entity Queries
von: Zhao, Wenting, et al.
Veröffentlicht: (2024)
von: Zhao, Wenting, et al.
Veröffentlicht: (2024)
Why and How LLMs Hallucinate: Connecting the Dots with Subsequence Associations
von: Sun, Yiyou, et al.
Veröffentlicht: (2025)
von: Sun, Yiyou, et al.
Veröffentlicht: (2025)
What is a Number, That a Large Language Model May Know It?
von: Marjieh, Raja, et al.
Veröffentlicht: (2025)
von: Marjieh, Raja, et al.
Veröffentlicht: (2025)
HAICOSYSTEM: An Ecosystem for Sandboxing Safety Risks in Human-AI Interactions
von: Zhou, Xuhui, et al.
Veröffentlicht: (2024)
von: Zhou, Xuhui, et al.
Veröffentlicht: (2024)
The Art of Saying No: Contextual Noncompliance in Language Models
von: Brahman, Faeze, et al.
Veröffentlicht: (2024)
von: Brahman, Faeze, et al.
Veröffentlicht: (2024)
What Has Been Lost with Synthetic Evaluation?
von: Gill, Alexander, et al.
Veröffentlicht: (2025)
von: Gill, Alexander, et al.
Veröffentlicht: (2025)
Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
von: Balepur, Nishant, et al.
Veröffentlicht: (2024)
von: Balepur, Nishant, et al.
Veröffentlicht: (2024)
The Dynamics of Collective Creativity in Human-AI Hybrid Societies
von: Shiiku, Shota, et al.
Veröffentlicht: (2025)
von: Shiiku, Shota, et al.
Veröffentlicht: (2025)
Information-Guided Identification of Training Data Imprint in (Proprietary) Large Language Models
von: Ravichander, Abhilasha, et al.
Veröffentlicht: (2025)
von: Ravichander, Abhilasha, et al.
Veröffentlicht: (2025)
Rethinking Creativity Evaluation: A Critical Analysis of Existing Creativity Evaluations
von: Lu, Li-Chun, et al.
Veröffentlicht: (2025)
von: Lu, Li-Chun, et al.
Veröffentlicht: (2025)
Information-Theoretic Distillation for Reference-less Summarization
von: Jung, Jaehun, et al.
Veröffentlicht: (2024)
von: Jung, Jaehun, et al.
Veröffentlicht: (2024)
Characterizing the Interaction of Cultural Evolution Mechanisms in Experimental Social Networks
von: Marjieh, Raja, et al.
Veröffentlicht: (2025)
von: Marjieh, Raja, et al.
Veröffentlicht: (2025)
The Curious Case of Factuality Finetuning: Models' Internal Beliefs Can Improve Factuality
von: Newman, Benjamin, et al.
Veröffentlicht: (2025)
von: Newman, Benjamin, et al.
Veröffentlicht: (2025)
Revisiting the Past: Data Unlearning with Model State History
von: Rezaei, Keivan, et al.
Veröffentlicht: (2025)
von: Rezaei, Keivan, et al.
Veröffentlicht: (2025)
A Call for Clarity in Beam Search: How It Works and When It Stops
von: Kasai, Jungo, et al.
Veröffentlicht: (2022)
von: Kasai, Jungo, et al.
Veröffentlicht: (2022)
Detecting Machine-Generated Long-Form Content with Latent-Space Variables
von: Tian, Yufei, et al.
Veröffentlicht: (2024)
von: Tian, Yufei, et al.
Veröffentlicht: (2024)
RoboWits: Unexpected Challenges for Robotic Creative Problem Solving
von: Lin, Chunru, et al.
Veröffentlicht: (2026)
von: Lin, Chunru, et al.
Veröffentlicht: (2026)
What Makes it Ok to Set a Fire? Iterative Self-distillation of Contexts and Rationales for Disambiguating Defeasible Social and Moral Situations
von: Rao, Kavel, et al.
Veröffentlicht: (2023)
von: Rao, Kavel, et al.
Veröffentlicht: (2023)
Impossible Distillation: from Low-Quality Model to High-Quality Dataset & Model for Summarization and Paraphrasing
von: Jung, Jaehun, et al.
Veröffentlicht: (2023)
von: Jung, Jaehun, et al.
Veröffentlicht: (2023)
How to Train Your Fact Verifier: Knowledge Transfer with Multimodal Open Models
von: Lee, Jaeyoung, et al.
Veröffentlicht: (2024)
von: Lee, Jaeyoung, et al.
Veröffentlicht: (2024)
Characterizing the Large‐Scale Structure of Multimodal Semantic Networks
von: Raja Marjieh, et al.
Veröffentlicht: (2025)
von: Raja Marjieh, et al.
Veröffentlicht: (2025)
SimpleToM: Exposing the Gap between Explicit ToM Inference and Implicit ToM Application in LLMs
von: Gu, Yuling, et al.
Veröffentlicht: (2024)
von: Gu, Yuling, et al.
Veröffentlicht: (2024)
Structured Chemistry Reasoning with Large Language Models
von: Ouyang, Siru, et al.
Veröffentlicht: (2023)
von: Ouyang, Siru, et al.
Veröffentlicht: (2023)
Large-Scale Data Selection for Instruction Tuning
von: Ivison, Hamish, et al.
Veröffentlicht: (2025)
von: Ivison, Hamish, et al.
Veröffentlicht: (2025)
ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2025)
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2025)
CreativityPrism: A Holistic Evaluation Framework for Large Language Model Creativity
von: Hou, Zhaoyi Joey, et al.
Veröffentlicht: (2025)
von: Hou, Zhaoyi Joey, et al.
Veröffentlicht: (2025)
Human-AI Synergy Supports Collective Creative Search
von: Li, Chenyi, et al.
Veröffentlicht: (2026)
von: Li, Chenyi, et al.
Veröffentlicht: (2026)
Fractional Rotation, Full Potential? Investigating Performance and Convergence of Partial RoPE
von: Khan, Mohammad Aflah, et al.
Veröffentlicht: (2026)
von: Khan, Mohammad Aflah, et al.
Veröffentlicht: (2026)
REFFLY: Melody-Constrained Lyrics Editing Model
von: Zhao, Songyan, et al.
Veröffentlicht: (2024)
von: Zhao, Songyan, et al.
Veröffentlicht: (2024)
SkillVerse : Assessing and Enhancing LLMs with Tree Evaluation
von: Tian, Yufei, et al.
Veröffentlicht: (2025)
von: Tian, Yufei, et al.
Veröffentlicht: (2025)
The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage
von: Hallinan, Skyler, et al.
Veröffentlicht: (2025)
von: Hallinan, Skyler, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
RESTOR: Knowledge Recovery in Machine Unlearning
von: Rezaei, Keivan, et al.
Veröffentlicht: (2024) -
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2024) -
Agent Lumos: Unified and Modular Training for Open-Source Language Agents
von: Yin, Da, et al.
Veröffentlicht: (2023) -
Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
von: Jung, Jaehun, et al.
Veröffentlicht: (2024) -
HALoGEN: Fantastic LLM Hallucinations and Where to Find Them
von: Ravichander, Abhilasha, et al.
Veröffentlicht: (2025)