Evaluating Novelty in AI-Generated Research Plans Using Multi-Workflow LLM Pipelines
Fuente:
arXiv
Salvato in:
| Autori principali: | Saraogi, Devesh, Singhee, Rohit, Kumar, Dhruv |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Infinite Problem Generator: Verifiably Scaling Physics Reasoning Data with Agentic Workflows
di: Sharan, Aditya, et al.
Pubblicazione: (2026)
di: Sharan, Aditya, et al.
Pubblicazione: (2026)
A Hybrid Supervised-LLM Pipeline for Actionable Suggestion Mining in Unstructured Customer Reviews
di: Trivedi, Aakash, et al.
Pubblicazione: (2026)
di: Trivedi, Aakash, et al.
Pubblicazione: (2026)
Nova: An Iterative Planning and Search Approach to Enhance Novelty and Diversity of LLM Generated Ideas
di: Hu, Xiang, et al.
Pubblicazione: (2024)
di: Hu, Xiang, et al.
Pubblicazione: (2024)
Novelty-based Tree-of-Thought Search for LLM Reasoning and Planning
di: Hamm, Leon, et al.
Pubblicazione: (2026)
di: Hamm, Leon, et al.
Pubblicazione: (2026)
A Hybrid DeBERTa and Gated Broad Learning System for Cyberbullying Detection in English Text
di: Kumar, Devesh
Pubblicazione: (2025)
di: Kumar, Devesh
Pubblicazione: (2025)
Not Just Novelty: A Longitudinal Study on Utility and Customization of an AI Workflow
di: Long, Tao, et al.
Pubblicazione: (2024)
di: Long, Tao, et al.
Pubblicazione: (2024)
Beyond Sentiment: A Multi-Agent Pipeline for Actionable Business Advice from Reviews
di: Bhandari, Kartikey Singh, et al.
Pubblicazione: (2026)
di: Bhandari, Kartikey Singh, et al.
Pubblicazione: (2026)
ReviewEval: An Evaluation Framework for AI-Generated Reviews
di: Garg, Madhav Krishan, et al.
Pubblicazione: (2025)
di: Garg, Madhav Krishan, et al.
Pubblicazione: (2025)
ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks
di: Schmidt, Jan-Philipp
Pubblicazione: (2026)
di: Schmidt, Jan-Philipp
Pubblicazione: (2026)
Evaluating $n$-Gram Novelty of Language Models Using Rusty-DAWG
di: Merrill, William, et al.
Pubblicazione: (2024)
di: Merrill, William, et al.
Pubblicazione: (2024)
PromptScreen: Efficient Jailbreak Mitigation Using Semantic Linear Classification in a Multi-Staged Pipeline
di: Rao, Akshaj Prashanth, et al.
Pubblicazione: (2025)
di: Rao, Akshaj Prashanth, et al.
Pubblicazione: (2025)
The LLM Fallacy: Misattribution in AI-Assisted Cognitive Workflows
di: Kim, Hyunwoo, et al.
Pubblicazione: (2026)
di: Kim, Hyunwoo, et al.
Pubblicazione: (2026)
OpenNovelty: An LLM-powered Agentic System for Verifiable Scholarly Novelty Assessment
di: Zhang, Ming, et al.
Pubblicazione: (2026)
di: Zhang, Ming, et al.
Pubblicazione: (2026)
Beyond Static Pipelines: Learning Dynamic Workflows for Text-to-SQL
di: Wang, Yihan, et al.
Pubblicazione: (2026)
di: Wang, Yihan, et al.
Pubblicazione: (2026)
Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations
di: Gupta, Manan, et al.
Pubblicazione: (2026)
di: Gupta, Manan, et al.
Pubblicazione: (2026)
LLM-as-a-Judge for Time Series Explanations
di: Sivalingam, Preetham, et al.
Pubblicazione: (2026)
di: Sivalingam, Preetham, et al.
Pubblicazione: (2026)
Detecting Errors through Ensembling Prompts (DEEP): An End-to-End LLM Framework for Detecting Factual Errors
di: Chandler, Alex, et al.
Pubblicazione: (2024)
di: Chandler, Alex, et al.
Pubblicazione: (2024)
Magellan: Guided MCTS for Latent Space Exploration and Novelty Generation
di: Chang, Lufan
Pubblicazione: (2025)
di: Chang, Lufan
Pubblicazione: (2025)
Evaluating LLMs for Zeolite Synthesis Event Extraction (ZSEE): A Systematic Analysis of Prompting Strategies
di: Rathore, Charan Prakash, et al.
Pubblicazione: (2025)
di: Rathore, Charan Prakash, et al.
Pubblicazione: (2025)
SwissNYF: Tool Grounded LLM Agents for Black Box Setting
di: Kumar, Somnath Sendhil, et al.
Pubblicazione: (2024)
di: Kumar, Somnath Sendhil, et al.
Pubblicazione: (2024)
Psittacines of Innovation? Assessing the True Novelty of AI Creations
di: Mukherjee, Anirban
Pubblicazione: (2024)
di: Mukherjee, Anirban
Pubblicazione: (2024)
Evaluating AI Meeting Summaries with a Reusable Cross-Domain Pipeline
di: Zhong, Philip, et al.
Pubblicazione: (2026)
di: Zhong, Philip, et al.
Pubblicazione: (2026)
How Trustworthy Are LLM-as-Judge Ratings for Interpretive Responses? Implications for Qualitative Research Workflows
di: Han, Songhee, et al.
Pubblicazione: (2026)
di: Han, Songhee, et al.
Pubblicazione: (2026)
LUDOBENCH: Evaluating LLM Behavioural Decision-Making Through Spot-Based Board Game Scenarios in Ludo
di: Jain, Ojas, et al.
Pubblicazione: (2026)
di: Jain, Ojas, et al.
Pubblicazione: (2026)
CodeRefine: A Pipeline for Enhancing LLM-Generated Code Implementations of Research Papers
di: Trofimova, Ekaterina, et al.
Pubblicazione: (2024)
di: Trofimova, Ekaterina, et al.
Pubblicazione: (2024)
FinRpt: Dataset, Evaluation System and LLM-based Multi-agent Framework for Equity Research Report Generation
di: Jin, Song, et al.
Pubblicazione: (2025)
di: Jin, Song, et al.
Pubblicazione: (2025)
AI Planning Framework for LLM-Based Web Agents
di: Shahnovsky, Orit, et al.
Pubblicazione: (2026)
di: Shahnovsky, Orit, et al.
Pubblicazione: (2026)
ResearchCodeAgent: An LLM Multi-Agent System for Automated Codification of Research Methodologies
di: Gandhi, Shubham, et al.
Pubblicazione: (2025)
di: Gandhi, Shubham, et al.
Pubblicazione: (2025)
The Fellowship of the LLMs: Multi-Model Workflows for Synthetic Preference Optimization Dataset Generation
di: Arif, Samee, et al.
Pubblicazione: (2024)
di: Arif, Samee, et al.
Pubblicazione: (2024)
An LLM + ASP Workflow for Joint Entity-Relation Extraction
di: Tran, Trang, et al.
Pubblicazione: (2025)
di: Tran, Trang, et al.
Pubblicazione: (2025)
Conversational Education at Scale: A Multi-LLM Agent Workflow for Procedural Learning and Pedagogic Quality Assessment
di: Pei, Jiahuan, et al.
Pubblicazione: (2025)
di: Pei, Jiahuan, et al.
Pubblicazione: (2025)
STACK: Adversarial Attacks on LLM Safeguard Pipelines
di: McKenzie, Ian R., et al.
Pubblicazione: (2025)
di: McKenzie, Ian R., et al.
Pubblicazione: (2025)
PatentScore: Multi-dimensional Evaluation of LLM-Generated Patent Claims
di: Yoo, Yongmin, et al.
Pubblicazione: (2025)
di: Yoo, Yongmin, et al.
Pubblicazione: (2025)
TextMineX: Data, Evaluation Framework and Ontology-guided LLM Pipeline for Humanitarian Mine Action
di: Zhou, Chenyue, et al.
Pubblicazione: (2025)
di: Zhou, Chenyue, et al.
Pubblicazione: (2025)
AI-Assisted Systematization for Evaluating GenAI Systems
di: Agarwal, Dhruv, et al.
Pubblicazione: (2026)
di: Agarwal, Dhruv, et al.
Pubblicazione: (2026)
Generation Z's Ability to Discriminate Between AI-generated and Human-Authored Text on Discord
di: Ramu, Dhruv, et al.
Pubblicazione: (2023)
di: Ramu, Dhruv, et al.
Pubblicazione: (2023)
Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
di: Saha, Swarnadeep, et al.
Pubblicazione: (2025)
di: Saha, Swarnadeep, et al.
Pubblicazione: (2025)
LLM-BABYBENCH: Understanding and Evaluating Grounded Planning and Reasoning in LLMs
di: Choukrani, Omar, et al.
Pubblicazione: (2025)
di: Choukrani, Omar, et al.
Pubblicazione: (2025)
CricBench: A Multilingual Benchmark for Evaluating LLMs in Cricket Analytics
di: Agarwal, Parth, et al.
Pubblicazione: (2025)
di: Agarwal, Parth, et al.
Pubblicazione: (2025)
Multi-Step Dialogue Workflow Action Prediction
di: Ramakrishnan, Ramya, et al.
Pubblicazione: (2023)
di: Ramakrishnan, Ramya, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Infinite Problem Generator: Verifiably Scaling Physics Reasoning Data with Agentic Workflows
di: Sharan, Aditya, et al.
Pubblicazione: (2026) -
A Hybrid Supervised-LLM Pipeline for Actionable Suggestion Mining in Unstructured Customer Reviews
di: Trivedi, Aakash, et al.
Pubblicazione: (2026) -
Nova: An Iterative Planning and Search Approach to Enhance Novelty and Diversity of LLM Generated Ideas
di: Hu, Xiang, et al.
Pubblicazione: (2024) -
Novelty-based Tree-of-Thought Search for LLM Reasoning and Planning
di: Hamm, Leon, et al.
Pubblicazione: (2026) -
A Hybrid DeBERTa and Gated Broad Learning System for Cyberbullying Detection in English Text
di: Kumar, Devesh
Pubblicazione: (2025)