IDEAFix: Evaluation Framework for Creative Defixation Prompting in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Carichon, F., Sharma, S., Girard, M., Rampa, R., Farnadi, G. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Crossing Boundaries: Leveraging Semantic Divergences to Explore Cultural Novelty in Cooking Recipes
by: Carichon, Florian, et al.
Published: (2025)
by: Carichon, Florian, et al.
Published: (2025)
Can LLMs Cook Jamaican Couscous? A Study of Cultural Novelty in Recipe Generation
by: Carichon, F., et al.
Published: (2026)
by: Carichon, F., et al.
Published: (2026)
Understanding Intrinsic Socioeconomic Biases in Large Language Models
by: Arzaghi, Mina, et al.
Published: (2024)
by: Arzaghi, Mina, et al.
Published: (2024)
Mechanics of Bias and Reasoning: Interpreting the Impact of Chain-of-Thought Prompting on Gender Bias in LLMs
by: Pearman, Edie, et al.
Published: (2026)
by: Pearman, Edie, et al.
Published: (2026)
Intrinsic Meets Extrinsic Fairness: Assessing the Downstream Impact of Bias Mitigation in Large Language Models
by: Arzaghi', 'Mina, et al.
Published: (2025)
by: Arzaghi', 'Mina, et al.
Published: (2025)
The Coming Crisis of Multi-Agent Misalignment: AI Alignment Must Be a Dynamic and Social Process
by: Carichon, Florian, et al.
Published: (2025)
by: Carichon, Florian, et al.
Published: (2025)
Embedding Cultural Diversity in Prototype-based Recommender Systems
by: Moradi, Armin, et al.
Published: (2024)
by: Moradi, Armin, et al.
Published: (2024)
Say It Another Way: Auditing LLMs with a User-Grounded Automated Paraphrasing Framework
by: Chataigner, Cléa, et al.
Published: (2025)
by: Chataigner, Cléa, et al.
Published: (2025)
Evaluating the Creativity of LLMs in Persian Literary Text Generation
by: Tourajmehr, Armin, et al.
Published: (2025)
by: Tourajmehr, Armin, et al.
Published: (2025)
Designing and Evaluating Dialogue LLMs for Co-Creative Improvised Theatre
by: Branch, Boyd, et al.
Published: (2024)
by: Branch, Boyd, et al.
Published: (2024)
Do LLMs Agree on the Creativity Evaluation of Alternative Uses?
by: Rabeyah, Abdullah Al, et al.
Published: (2024)
by: Rabeyah, Abdullah Al, et al.
Published: (2024)
CreativityPrism: A Holistic Evaluation Framework for Large Language Model Creativity
by: Hou, Zhaoyi Joey, et al.
Published: (2025)
by: Hou, Zhaoyi Joey, et al.
Published: (2025)
Multilingual Hallucination Gaps in Large Language Models
by: Chataigner, Cléa, et al.
Published: (2024)
by: Chataigner, Cléa, et al.
Published: (2024)
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges
by: Chehbouni, Khaoula, et al.
Published: (2025)
by: Chehbouni, Khaoula, et al.
Published: (2025)
Rethinking Creativity Evaluation: A Critical Analysis of Existing Creativity Evaluations
by: Lu, Li-Chun, et al.
Published: (2025)
by: Lu, Li-Chun, et al.
Published: (2025)
CREATE: Testing LLMs for Associative Creativity
by: Wadhwa, Manya, et al.
Published: (2026)
by: Wadhwa, Manya, et al.
Published: (2026)
ProMoral-Bench: Evaluating Prompting Strategies for Moral Reasoning and Safety in LLMs
by: Thomas, Rohan Subramanian, et al.
Published: (2026)
by: Thomas, Rohan Subramanian, et al.
Published: (2026)
LoRA Provides Differential Privacy by Design via Random Sketching
by: Malekmohammadi, Saber, et al.
Published: (2024)
by: Malekmohammadi, Saber, et al.
Published: (2024)
Multilingual Amnesia: On the Transferability of Unlearning in Multilingual LLMs
by: Farashah, Alireza Dehghanpour, et al.
Published: (2026)
by: Farashah, Alireza Dehghanpour, et al.
Published: (2026)
Adaptive Originality Filtering: Rejection Based Prompting and RiddleScore for Culturally Grounded Multilingual Riddle Generation
by: Le, Duy, et al.
Published: (2025)
by: Le, Duy, et al.
Published: (2025)
Evaluation Framework for AI Creativity: A Case Study Based on Story Generation
by: Sathya, Pharath, et al.
Published: (2026)
by: Sathya, Pharath, et al.
Published: (2026)
Structured Prompt Language: Declarative Context Management for LLMs
by: Gong, Wen G.
Published: (2026)
by: Gong, Wen G.
Published: (2026)
GRP: Goal-Reversed Prompting for Zero-Shot Evaluation with LLMs
by: Song, Mingyang, et al.
Published: (2025)
by: Song, Mingyang, et al.
Published: (2025)
The power of Prompts: Evaluating and Mitigating Gender Bias in MT with LLMs
by: Sant, Aleix, et al.
Published: (2024)
by: Sant, Aleix, et al.
Published: (2024)
Creativity Bias: How Machine Evaluation Struggles with Creativity in Literary Translations
by: Gerrits, Kyo, et al.
Published: (2026)
by: Gerrits, Kyo, et al.
Published: (2026)
CreativEval: Evaluating Creativity of LLM-Based Hardware Code Generation
by: DeLorenzo, Matthew, et al.
Published: (2024)
by: DeLorenzo, Matthew, et al.
Published: (2024)
Role-playing Prompt Framework: Generation and Evaluation
by: Liu, Xun, et al.
Published: (2024)
by: Liu, Xun, et al.
Published: (2024)
Can Prompts Rewind Time for LLMs? Evaluating the Effectiveness of Prompted Knowledge Cutoffs
by: Gao, Xin, et al.
Published: (2025)
by: Gao, Xin, et al.
Published: (2025)
One Prompt To Rule Them All: LLMs for Opinion Summary Evaluation
by: Siledar, Tejpalsingh, et al.
Published: (2024)
by: Siledar, Tejpalsingh, et al.
Published: (2024)
Palisade -- Prompt Injection Detection Framework
by: Kokkula, Sahasra, et al.
Published: (2024)
by: Kokkula, Sahasra, et al.
Published: (2024)
Beyond Divergent Creativity: A Human-Based Evaluation of Creativity in Large Language Models
by: Nakajima, Kumiko, et al.
Published: (2026)
by: Nakajima, Kumiko, et al.
Published: (2026)
Shakespearean Sparks: The Dance of Hallucination and Creativity in LLMs' Decoding Layers
by: He, Zicong, et al.
Published: (2025)
by: He, Zicong, et al.
Published: (2025)
Confidence, Not Perplexity: A Better Metric for the Creative Era of LLMs
by: Parupudi, V. S. Raghu
Published: (2025)
by: Parupudi, V. S. Raghu
Published: (2025)
Evaluating the Effectiveness of Black-Box Prompt Optimization as the Scale of LLMs Continues to Grow
by: Zhou, Ziyu, et al.
Published: (2025)
by: Zhou, Ziyu, et al.
Published: (2025)
IQ Test for LLMs: An Evaluation Framework for Uncovering Core Skills in LLMs
by: Maimon, Aviya, et al.
Published: (2025)
by: Maimon, Aviya, et al.
Published: (2025)
Reviving Your MNEME: Predicting The Side Effects of LLM Unlearning and Fine-Tuning via Sparse Model Diffing
by: Kassem, Aly M., et al.
Published: (2025)
by: Kassem, Aly M., et al.
Published: (2025)
Towards More Realistic Extraction Attacks: An Adversarial Perspective
by: More, Yash, et al.
Published: (2024)
by: More, Yash, et al.
Published: (2024)
Modular Prompt Optimization: Optimizing Structured Prompts with Section-Local Textual Gradients
by: Sharma, Prith, et al.
Published: (2026)
by: Sharma, Prith, et al.
Published: (2026)
Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs
by: Naderi, Nariman, et al.
Published: (2025)
by: Naderi, Nariman, et al.
Published: (2025)
LLAMADRS: Evaluating Open-Source LLMs on Real Clinical Interviews--To Reason or Not to Reason?
by: Kebe, Gaoussou Youssouf, et al.
Published: (2025)
by: Kebe, Gaoussou Youssouf, et al.
Published: (2025)
Similar Items
-
Crossing Boundaries: Leveraging Semantic Divergences to Explore Cultural Novelty in Cooking Recipes
by: Carichon, Florian, et al.
Published: (2025) -
Can LLMs Cook Jamaican Couscous? A Study of Cultural Novelty in Recipe Generation
by: Carichon, F., et al.
Published: (2026) -
Understanding Intrinsic Socioeconomic Biases in Large Language Models
by: Arzaghi, Mina, et al.
Published: (2024) -
Mechanics of Bias and Reasoning: Interpreting the Impact of Chain-of-Thought Prompting on Gender Bias in LLMs
by: Pearman, Edie, et al.
Published: (2026) -
Intrinsic Meets Extrinsic Fairness: Assessing the Downstream Impact of Bias Mitigation in Large Language Models
by: Arzaghi', 'Mina, et al.
Published: (2025)