Reinforcement learning for question answering in programming domain using public community scoring as a human feedback
Fuente:
arXiv
Salvato in:
| Autori principali: | Gorbatovski, Alexey, Kovalchuk, Sergey |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
How to build trust in answers given by Generative AI for specific, and vague, financial questions
di: Zarifis, Alex, et al.
Pubblicazione: (2024)
di: Zarifis, Alex, et al.
Pubblicazione: (2024)
Automated stereotactic radiosurgery planning using a human-in-the-loop reasoning large language model agent
di: Nusrat, Humza, et al.
Pubblicazione: (2025)
di: Nusrat, Humza, et al.
Pubblicazione: (2025)
Performance of leading large language models in May 2025 in Membership of the Royal College of General Practitioners-style examination questions: a cross-sectional analysis
di: Armitage, Richard
Pubblicazione: (2025)
di: Armitage, Richard
Pubblicazione: (2025)
Large language models provide unsafe answers to patient-posed medical questions
di: Draelos, Rachel L., et al.
Pubblicazione: (2025)
di: Draelos, Rachel L., et al.
Pubblicazione: (2025)
Everyone prefers human writers, including AI
di: Haverals, Wouter, et al.
Pubblicazione: (2025)
di: Haverals, Wouter, et al.
Pubblicazione: (2025)
Can AI grade your essays? A comparative analysis of large language models and teacher ratings in multidimensional essay scoring
di: Seßler, Kathrin, et al.
Pubblicazione: (2024)
di: Seßler, Kathrin, et al.
Pubblicazione: (2024)
Clinical knowledge in LLMs does not translate to human interactions
di: Bean, Andrew M., et al.
Pubblicazione: (2025)
di: Bean, Andrew M., et al.
Pubblicazione: (2025)
Maximizing the efficiency of human feedback in AI alignment: a comparative analysis
di: Chouliaras, Andreas, et al.
Pubblicazione: (2025)
di: Chouliaras, Andreas, et al.
Pubblicazione: (2025)
TouchAI: Exploring human-AI perceptual alignment in touch through language model representations
di: Zhong, Shu, et al.
Pubblicazione: (2024)
di: Zhong, Shu, et al.
Pubblicazione: (2024)
Playing 20 Question Game with Policy-Based Reinforcement Learning
di: Hu, Huang, et al.
Pubblicazione: (2018)
di: Hu, Huang, et al.
Pubblicazione: (2018)
AI-Enabled grading with near-domain data for scaling feedback with human-level accuracy
di: Agarwal, Shyam, et al.
Pubblicazione: (2025)
di: Agarwal, Shyam, et al.
Pubblicazione: (2025)
Multi-agent KTO: Reinforcing Strategic Interactions of Large Language Model in Language Game
di: Ye, Rong, et al.
Pubblicazione: (2025)
di: Ye, Rong, et al.
Pubblicazione: (2025)
WatChat: Explaining perplexing programs by debugging mental models
di: Chandra, Kartik, et al.
Pubblicazione: (2024)
di: Chandra, Kartik, et al.
Pubblicazione: (2024)
BRAIn: Bayesian Reward-conditioned Amortized Inference for natural language generation from feedback
di: Pandey, Gaurav, et al.
Pubblicazione: (2024)
di: Pandey, Gaurav, et al.
Pubblicazione: (2024)
Towards a copilot in BIM authoring tool using a large language model-based agent for intelligent human-machine interaction
di: Du, Changyu, et al.
Pubblicazione: (2024)
di: Du, Changyu, et al.
Pubblicazione: (2024)
From tools to thieves: Measuring and understanding public perceptions of AI through crowdsourced metaphors
di: Cheng, Myra, et al.
Pubblicazione: (2025)
di: Cheng, Myra, et al.
Pubblicazione: (2025)
"Like a Nesting Doll": Analyzing Recursion Analogies Generated by CS Students using Large Language Models
di: Bernstein, Seth, et al.
Pubblicazione: (2024)
di: Bernstein, Seth, et al.
Pubblicazione: (2024)
Are the confidence scores of reviewers consistent with the review content? Evidence from top conference proceedings in AI
di: Wu, Wenqing, et al.
Pubblicazione: (2025)
di: Wu, Wenqing, et al.
Pubblicazione: (2025)
Empirical evidence of Large Language Model's influence on human spoken communication
di: Yakura, Hiromu, et al.
Pubblicazione: (2024)
di: Yakura, Hiromu, et al.
Pubblicazione: (2024)
Twin-2K-500: A dataset for building digital twins of over 2,000 people based on their answers to over 500 questions
di: Toubia, Olivier, et al.
Pubblicazione: (2025)
di: Toubia, Olivier, et al.
Pubblicazione: (2025)
Logic-Scaffolding: Personalized Aspect-Instructed Recommendation Explanation Generation using LLMs
di: Rahdari, Behnam, et al.
Pubblicazione: (2023)
di: Rahdari, Behnam, et al.
Pubblicazione: (2023)
Step Guided Reasoning: Improving Mathematical Reasoning using Guidance Generation and Step Reasoning
di: Cao, Lang, et al.
Pubblicazione: (2024)
di: Cao, Lang, et al.
Pubblicazione: (2024)
Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
di: Thakkar, Nitya, et al.
Pubblicazione: (2025)
di: Thakkar, Nitya, et al.
Pubblicazione: (2025)
RealitySummary: Exploring On-Demand Mixed Reality Text Summarization and Question Answering using Large Language Models
di: Gunturu, Aditya, et al.
Pubblicazione: (2024)
di: Gunturu, Aditya, et al.
Pubblicazione: (2024)
From Novice to Expert: LLM Agent Policy Optimization via Step-wise Reinforcement Learning
di: Deng, Zhirui, et al.
Pubblicazione: (2024)
di: Deng, Zhirui, et al.
Pubblicazione: (2024)
Adult learners recall and recognition performance and affective feedback when learning from an AI-generated synthetic video
di: Li, Zoe Ruo-Yu, et al.
Pubblicazione: (2024)
di: Li, Zoe Ruo-Yu, et al.
Pubblicazione: (2024)
Reinforcement Learning from Human Feedback: Whose Culture, Whose Values, Whose Perspectives?
di: Barman, Kristian González, et al.
Pubblicazione: (2024)
di: Barman, Kristian González, et al.
Pubblicazione: (2024)
Digital assistant in a point of sales
di: Lesiak, Emilia, et al.
Pubblicazione: (2024)
di: Lesiak, Emilia, et al.
Pubblicazione: (2024)
Prompt Engineering a Schizophrenia Chatbot: Utilizing a Multi-Agent Approach for Enhanced Compliance with Prompt Instructions
di: Waaler, Per Niklas, et al.
Pubblicazione: (2024)
di: Waaler, Per Niklas, et al.
Pubblicazione: (2024)
Writing as a testbed for open ended agents
di: Gooding, Sian, et al.
Pubblicazione: (2025)
di: Gooding, Sian, et al.
Pubblicazione: (2025)
Reinforcement Learning for Personalized Dialogue Management
di: Hengst, Floris den, et al.
Pubblicazione: (2019)
di: Hengst, Floris den, et al.
Pubblicazione: (2019)
Designing a Dashboard for Transparency and Control of Conversational AI
di: Chen, Yida, et al.
Pubblicazione: (2024)
di: Chen, Yida, et al.
Pubblicazione: (2024)
ChatVis: Automating Scientific Visualization with a Large Language Model
di: Mallick, Tanwi, et al.
Pubblicazione: (2024)
di: Mallick, Tanwi, et al.
Pubblicazione: (2024)
Performance Gains of LLMs With Humans in a World of LLMs Versus Humans
di: McCullum, Lucas, et al.
Pubblicazione: (2025)
di: McCullum, Lucas, et al.
Pubblicazione: (2025)
GPT Models in Construction Industry: Opportunities, Limitations, and a Use Case Validation
di: Saka, Abdullahi, et al.
Pubblicazione: (2023)
di: Saka, Abdullahi, et al.
Pubblicazione: (2023)
From Control to Foresight: Simulation as a New Paradigm for Human-Agent Collaboration
di: He, Gaole, et al.
Pubblicazione: (2026)
di: He, Gaole, et al.
Pubblicazione: (2026)
A Generalized LLM-Augmented BIM Framework: Application to a Speech-to-BIM system
di: Lee, Ghang, et al.
Pubblicazione: (2024)
di: Lee, Ghang, et al.
Pubblicazione: (2024)
DreamGarden: A Designer Assistant for Growing Games from a Single Prompt
di: Earle, Sam, et al.
Pubblicazione: (2024)
di: Earle, Sam, et al.
Pubblicazione: (2024)
Death of the Novel(ty): Beyond n-Gram Novelty as a Metric for Textual Creativity
di: Saakyan, Arkadiy, et al.
Pubblicazione: (2025)
di: Saakyan, Arkadiy, et al.
Pubblicazione: (2025)
GPT-4 Emulates Average-Human Emotional Cognition from a Third-Person Perspective
di: Tak, Ala N., et al.
Pubblicazione: (2024)
di: Tak, Ala N., et al.
Pubblicazione: (2024)
Documenti analoghi
-
How to build trust in answers given by Generative AI for specific, and vague, financial questions
di: Zarifis, Alex, et al.
Pubblicazione: (2024) -
Automated stereotactic radiosurgery planning using a human-in-the-loop reasoning large language model agent
di: Nusrat, Humza, et al.
Pubblicazione: (2025) -
Performance of leading large language models in May 2025 in Membership of the Royal College of General Practitioners-style examination questions: a cross-sectional analysis
di: Armitage, Richard
Pubblicazione: (2025) -
Large language models provide unsafe answers to patient-posed medical questions
di: Draelos, Rachel L., et al.
Pubblicazione: (2025) -
Everyone prefers human writers, including AI
di: Haverals, Wouter, et al.
Pubblicazione: (2025)