Saved in:
| Main Authors: | Haase, Jennifer, Klessascheck, Finn, Mendling, Jan, Pokutta, Sebastian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2504.13217 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Has the Creativity of Large-Language Models peaked? An analysis of inter- and intra-LLM variability
by: Haase, Jennifer, et al.
Published: (2025)
by: Haase, Jennifer, et al.
Published: (2025)
Unlocking Sustainability Compliance: Characterizing the EU Taxonomy for Business Process Management
by: Klessascheck, Finn, et al.
Published: (2024)
by: Klessascheck, Finn, et al.
Published: (2024)
Within-Model vs Between-Prompt Variability in Large Language Models for Creative Tasks
by: Haase, Jennifer, et al.
Published: (2026)
by: Haase, Jennifer, et al.
Published: (2026)
Agentic MIP Research: Accelerated Constraint Handler Generation
by: Xu, Liding, et al.
Published: (2026)
by: Xu, Liding, et al.
Published: (2026)
S-DAT: A Multilingual, GenAI-Driven Framework for Automated Divergent Thinking Assessment
by: Haase, Jennifer, et al.
Published: (2025)
by: Haase, Jennifer, et al.
Published: (2025)
From Associations to Activations: Comparing Behavioral and Hidden-State Semantic Geometry in LLMs
by: Schiekiera, Louis, et al.
Published: (2026)
by: Schiekiera, Louis, et al.
Published: (2026)
Teaching People LLM's Errors and Getting it Right
by: Stringham, Nathan, et al.
Published: (2025)
by: Stringham, Nathan, et al.
Published: (2025)
Effect of Gender Fair Job Description on Generative AI Images
by: Böckling, Finn, et al.
Published: (2025)
by: Böckling, Finn, et al.
Published: (2025)
Are We on the Right Way to Assessing LLM-as-a-Judge?
by: Feng, Yuanning, et al.
Published: (2025)
by: Feng, Yuanning, et al.
Published: (2025)
Counting Clues: A Lightweight Probabilistic Baseline Can Match an LLM
by: Jia, Furong, et al.
Published: (2025)
by: Jia, Furong, et al.
Published: (2025)
This Treatment Works, Right? Evaluating LLM Sensitivity to Patient Question Framing in Medical QA
by: Yun, Hye Sun, et al.
Published: (2026)
by: Yun, Hye Sun, et al.
Published: (2026)
Human-AI Co-Creativity: Exploring Synergies Across Levels of Creative Collaboration
by: Haase, Jennifer, et al.
Published: (2024)
by: Haase, Jennifer, et al.
Published: (2024)
Fairness in Healthcare Processes: A Quantitative Analysis of Decision Making in Triage
by: Andreswari, Rachmadita, et al.
Published: (2026)
by: Andreswari, Rachmadita, et al.
Published: (2026)
Comparing LLM Text Annotation Skills: A Study on Human Rights Violations in Social Media Data
by: Nemkova, Poli Apollinaire, et al.
Published: (2025)
by: Nemkova, Poli Apollinaire, et al.
Published: (2025)
The Knowledge-Behaviour Disconnect in LLM-based Chatbots
by: Broersen, Jan
Published: (2025)
by: Broersen, Jan
Published: (2025)
Towards Nudging in BPM: A Human-Centric Approach for Sustainable Business Processes
by: Moyano, Cielo Gonzalez, et al.
Published: (2024)
by: Moyano, Cielo Gonzalez, et al.
Published: (2024)
Plausibility Vaccine: Injecting LLM Knowledge for Event Plausibility
by: Chmura, Jacob, et al.
Published: (2025)
by: Chmura, Jacob, et al.
Published: (2025)
Asking the Right Question at the Right Time: Human and Model Uncertainty Guidance to Ask Clarification Questions
by: Testoni, Alberto, et al.
Published: (2024)
by: Testoni, Alberto, et al.
Published: (2024)
YESciEval: Robust LLM-as-a-Judge for Scientific Question Answering
by: D'Souza, Jennifer, et al.
Published: (2025)
by: D'Souza, Jennifer, et al.
Published: (2025)
Insight Agents: An LLM-Based Multi-Agent System for Data Insights
by: Bai, Jincheng, et al.
Published: (2026)
by: Bai, Jincheng, et al.
Published: (2026)
Batch Speculative Decoding Done Right
by: Zhang, Ranran Haoran, et al.
Published: (2025)
by: Zhang, Ranran Haoran, et al.
Published: (2025)
Efficient LLM-Jailbreaking via Multimodal-LLM Jailbreak
by: Ji, Haoxuan, et al.
Published: (2024)
by: Ji, Haoxuan, et al.
Published: (2024)
Set-LLM: A Permutation-Invariant LLM
by: Egressy, Beni, et al.
Published: (2025)
by: Egressy, Beni, et al.
Published: (2025)
ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks
by: Schmidt, Jan-Philipp
Published: (2026)
by: Schmidt, Jan-Philipp
Published: (2026)
Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic
by: Collot, Stephane, et al.
Published: (2025)
by: Collot, Stephane, et al.
Published: (2025)
Epistemic Context Learning: Building Trust the Right Way in LLM-Based Multi-Agent Systems
by: Zhou, Ruiwen, et al.
Published: (2026)
by: Zhou, Ruiwen, et al.
Published: (2026)
Sustainable Digitalization of Business with Multi-Agent RAG and LLM
by: Arslan, Muhammad, et al.
Published: (2025)
by: Arslan, Muhammad, et al.
Published: (2025)
Evaluating the Impact of Advanced LLM Techniques on AI-Lecture Tutors for a Robotics Course
by: Kahl, Sebastian, et al.
Published: (2024)
by: Kahl, Sebastian, et al.
Published: (2024)
ToxiGAN: Toxic Data Augmentation via LLM-Guided Directional Adversarial Generation
by: Li, Peiran, et al.
Published: (2026)
by: Li, Peiran, et al.
Published: (2026)
Talk Less, Call Right: Enhancing Role-Play LLM Agents with Automatic Prompt Optimization and Role Prompting
by: Ruangtanusak, Saksorn, et al.
Published: (2025)
by: Ruangtanusak, Saksorn, et al.
Published: (2025)
LIME: Making LLM Data More Efficient with Linguistic Metadata Embeddings
by: Sztwiertnia, Sebastian, et al.
Published: (2025)
by: Sztwiertnia, Sebastian, et al.
Published: (2025)
DLER: Doing Length pEnalty Right - Incentivizing More Intelligence per Token via Reinforcement Learning
by: Liu, Shih-Yang, et al.
Published: (2025)
by: Liu, Shih-Yang, et al.
Published: (2025)
Backtranslation and paraphrasing in the LLM era? Comparing data augmentation methods for emotion classification
by: Radliński, Łukasz, et al.
Published: (2025)
by: Radliński, Łukasz, et al.
Published: (2025)
LLM Olympiad: Why Model Evaluation Needs a Sealed Exam
by: Cruz, Jan Christian Blaise, et al.
Published: (2026)
by: Cruz, Jan Christian Blaise, et al.
Published: (2026)
LLMSQL: Upgrading WikiSQL for the LLM Era of Text-to-SQL
by: Pihulski, Dzmitry, et al.
Published: (2025)
by: Pihulski, Dzmitry, et al.
Published: (2025)
Diagnosing Structural Failures in LLM-Based Evidence Extraction for Meta-Analysis
by: Tan, Zhiyin, et al.
Published: (2026)
by: Tan, Zhiyin, et al.
Published: (2026)
Two-dimensional early exit optimisation of LLM inference
by: Hůla, Jan, et al.
Published: (2026)
by: Hůla, Jan, et al.
Published: (2026)
The LLM Already Knows: Estimating LLM-Perceived Question Difficulty via Hidden Representations
by: Zhu, Yubo, et al.
Published: (2025)
by: Zhu, Yubo, et al.
Published: (2025)
Catastrophic Failure of LLM Unlearning via Quantization
by: Zhang, Zhiwei, et al.
Published: (2024)
by: Zhang, Zhiwei, et al.
Published: (2024)
Beyond Static Responses: Multi-Agent LLM Systems as a New Paradigm for Social Science Research
by: Haase, Jennifer, et al.
Published: (2025)
by: Haase, Jennifer, et al.
Published: (2025)
Similar Items
-
Has the Creativity of Large-Language Models peaked? An analysis of inter- and intra-LLM variability
by: Haase, Jennifer, et al.
Published: (2025) -
Unlocking Sustainability Compliance: Characterizing the EU Taxonomy for Business Process Management
by: Klessascheck, Finn, et al.
Published: (2024) -
Within-Model vs Between-Prompt Variability in Large Language Models for Creative Tasks
by: Haase, Jennifer, et al.
Published: (2026) -
Agentic MIP Research: Accelerated Constraint Handler Generation
by: Xu, Liding, et al.
Published: (2026) -
S-DAT: A Multilingual, GenAI-Driven Framework for Automated Divergent Thinking Assessment
by: Haase, Jennifer, et al.
Published: (2025)