Applying RLAIF for Code Generation with API-usage in Lightweight LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dutta, Sujan, Mahinder, Sayantan, Anantha, Raviteja, Bandyopadhyay, Bortik |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Retrieval Augmented Correction of Named Entity Speech Recognition Errors
von: Pusateri, Ernest, et al.
Veröffentlicht: (2024)
von: Pusateri, Ernest, et al.
Veröffentlicht: (2024)
CAMPHOR: Collaborative Agents for Multi-input Planning and High-Order Reasoning On Device
von: Fu, Yicheng, et al.
Veröffentlicht: (2024)
von: Fu, Yicheng, et al.
Veröffentlicht: (2024)
UI-JEPA: Towards Active Perception of User Intent through Onscreen User Activity
von: Fu, Yicheng, et al.
Veröffentlicht: (2024)
von: Fu, Yicheng, et al.
Veröffentlicht: (2024)
Down the Toxicity Rabbit Hole: A Novel Framework to Bias Audit Large Language Models
von: Dutta, Arka, et al.
Veröffentlicht: (2023)
von: Dutta, Arka, et al.
Veröffentlicht: (2023)
Intent-conditioned and Non-toxic Counterspeech Generation using Multi-Task Instruction Tuning with RLAIF
von: Hengle, Amey, et al.
Veröffentlicht: (2024)
von: Hengle, Amey, et al.
Veröffentlicht: (2024)
What About the Scene with the Hitler Reference? HAUNT: A Framework to Probe LLMs' Self-consistency Via Adversarial Nudge
von: Dutta, Arka, et al.
Veröffentlicht: (2025)
von: Dutta, Arka, et al.
Veröffentlicht: (2025)
Sparse Semantic Dimension as a Generalization Certificate for LLMs
von: Bandyopadhyay, Dibyanayan, et al.
Veröffentlicht: (2026)
von: Bandyopadhyay, Dibyanayan, et al.
Veröffentlicht: (2026)
Semantic Anchors in In-Context Learning: Why Small LLMs Cannot Flip Their Labels
von: Kumar, Anantha Padmanaban Krishna
Veröffentlicht: (2025)
von: Kumar, Anantha Padmanaban Krishna
Veröffentlicht: (2025)
Reverse Constitutional AI: A Framework for Controllable Toxic Data Generation via Probability-Clamped RLAIF
von: Fang, Yuan, et al.
Veröffentlicht: (2026)
von: Fang, Yuan, et al.
Veröffentlicht: (2026)
API-BLEND: A Comprehensive Corpora for Training and Benchmarking API LLMs
von: Basu, Kinjal, et al.
Veröffentlicht: (2024)
von: Basu, Kinjal, et al.
Veröffentlicht: (2024)
Harnessing LLMs for API Interactions: A Framework for Classification and Synthetic Data Generation
von: Tao, Chunliang, et al.
Veröffentlicht: (2024)
von: Tao, Chunliang, et al.
Veröffentlicht: (2024)
RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness
von: Yu, Tianyu, et al.
Veröffentlicht: (2024)
von: Yu, Tianyu, et al.
Veröffentlicht: (2024)
RLAIF-SPA: Structured AI Feedback for Semantic-Prosodic Alignment in Speech Synthesis
von: Yang, Qing, et al.
Veröffentlicht: (2025)
von: Yang, Qing, et al.
Veröffentlicht: (2025)
Compositional API Recommendation for Library-Oriented Code Generation
von: Ma, Zexiong, et al.
Veröffentlicht: (2024)
von: Ma, Zexiong, et al.
Veröffentlicht: (2024)
API-Assisted Code Generation for Question Answering on Varied Table Structures
von: Cao, Yihan, et al.
Veröffentlicht: (2023)
von: Cao, Yihan, et al.
Veröffentlicht: (2023)
Embedding Retrofitting: Data Engineering for better RAG
von: Sharma, Anantha
Veröffentlicht: (2026)
von: Sharma, Anantha
Veröffentlicht: (2026)
Generation-Time vs. Post-hoc Citation: A Holistic Evaluation of LLM Attribution
von: Saxena, Yash, et al.
Veröffentlicht: (2025)
von: Saxena, Yash, et al.
Veröffentlicht: (2025)
Gender Representation and Bias in Indian Civil Service Mock Interviews
von: Banerjee, Somonnoy, et al.
Veröffentlicht: (2024)
von: Banerjee, Somonnoy, et al.
Veröffentlicht: (2024)
ARGUS: Adaptive Rotation-Invariant Geometric Unsupervised System
von: Sharma, Anantha
Veröffentlicht: (2026)
von: Sharma, Anantha
Veröffentlicht: (2026)
Private Federated Learning In Real World Application -- A Case Study
von: Ji, An, et al.
Veröffentlicht: (2025)
von: Ji, An, et al.
Veröffentlicht: (2025)
Evaluating LLMs on Sequential API Call Through Automated Test Generation
von: Huang, Yuheng, et al.
Veröffentlicht: (2025)
von: Huang, Yuheng, et al.
Veröffentlicht: (2025)
RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
von: Lee, Harrison, et al.
Veröffentlicht: (2023)
von: Lee, Harrison, et al.
Veröffentlicht: (2023)
SpeCrawler: Generating OpenAPI Specifications from API Documentation Using Large Language Models
von: Lazar, Koren, et al.
Veröffentlicht: (2024)
von: Lazar, Koren, et al.
Veröffentlicht: (2024)
JU-NLP at Touché: Covert Advertisement in Conversational AI-Generation and Detection Strategies
von: Dutta, Arka, et al.
Veröffentlicht: (2025)
von: Dutta, Arka, et al.
Veröffentlicht: (2025)
On Mitigating Code LLM Hallucinations with API Documentation
von: Jain, Nihal, et al.
Veröffentlicht: (2024)
von: Jain, Nihal, et al.
Veröffentlicht: (2024)
Applying Large Language Models API to Issue Classification Problem
von: Aracena, Gabriel, et al.
Veröffentlicht: (2024)
von: Aracena, Gabriel, et al.
Veröffentlicht: (2024)
Leveraging Self-Attention for Input-Dependent Soft Prompting in LLMs
von: Muppidi, Ananth, et al.
Veröffentlicht: (2025)
von: Muppidi, Ananth, et al.
Veröffentlicht: (2025)
CodeUpdateArena: Benchmarking Knowledge Editing on API Updates
von: Liu, Zeyu Leo, et al.
Veröffentlicht: (2024)
von: Liu, Zeyu Leo, et al.
Veröffentlicht: (2024)
LiFi: Lightweight Controlled Text Generation with Fine-Grained Control Codes
von: Shi, Chufan, et al.
Veröffentlicht: (2024)
von: Shi, Chufan, et al.
Veröffentlicht: (2024)
Enhancing Presentation Slide Generation by LLMs with a Multi-Staged End-to-End Approach
von: Bandyopadhyay, Sambaran, et al.
Veröffentlicht: (2024)
von: Bandyopadhyay, Sambaran, et al.
Veröffentlicht: (2024)
Evaluation of Code LLMs on Geospatial Code Generation
von: Gramacki, Piotr, et al.
Veröffentlicht: (2024)
von: Gramacki, Piotr, et al.
Veröffentlicht: (2024)
Enhancing Project-Specific Code Completion by Inferring Internal API Information
von: Deng, Le, et al.
Veröffentlicht: (2025)
von: Deng, Le, et al.
Veröffentlicht: (2025)
NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls
von: Basu, Kinjal, et al.
Veröffentlicht: (2024)
von: Basu, Kinjal, et al.
Veröffentlicht: (2024)
The Anatomy of Evidence: An Investigation Into Explainable ICD Coding
von: Beckh, Katharina, et al.
Veröffentlicht: (2025)
von: Beckh, Katharina, et al.
Veröffentlicht: (2025)
AlienLM: Alienization of Language for API-Boundary Privacy in Black-Box LLMs
von: Kim, Jaehee, et al.
Veröffentlicht: (2026)
von: Kim, Jaehee, et al.
Veröffentlicht: (2026)
API Pack: A Massive Multi-Programming Language Dataset for API Call Generation
von: Guo, Zhen, et al.
Veröffentlicht: (2024)
von: Guo, Zhen, et al.
Veröffentlicht: (2024)
ReCode: Updating Code API Knowledge with Reinforcement Learning
von: Wu, Haoze, et al.
Veröffentlicht: (2025)
von: Wu, Haoze, et al.
Veröffentlicht: (2025)
On Code-Induced Reasoning in LLMs
von: Waheed, Abdul, et al.
Veröffentlicht: (2025)
von: Waheed, Abdul, et al.
Veröffentlicht: (2025)
Identity Lock: Locking API Fine-tuned LLMs With Identity-based Wake Words
von: Su, Hongyu, et al.
Veröffentlicht: (2025)
von: Su, Hongyu, et al.
Veröffentlicht: (2025)
TreeDiff: AST-Guided Code Generation with Diffusion LLMs
von: Zeng, Yiming, et al.
Veröffentlicht: (2025)
von: Zeng, Yiming, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Retrieval Augmented Correction of Named Entity Speech Recognition Errors
von: Pusateri, Ernest, et al.
Veröffentlicht: (2024) -
CAMPHOR: Collaborative Agents for Multi-input Planning and High-Order Reasoning On Device
von: Fu, Yicheng, et al.
Veröffentlicht: (2024) -
UI-JEPA: Towards Active Perception of User Intent through Onscreen User Activity
von: Fu, Yicheng, et al.
Veröffentlicht: (2024) -
Down the Toxicity Rabbit Hole: A Novel Framework to Bias Audit Large Language Models
von: Dutta, Arka, et al.
Veröffentlicht: (2023) -
Intent-conditioned and Non-toxic Counterspeech Generation using Multi-Task Instruction Tuning with RLAIF
von: Hengle, Amey, et al.
Veröffentlicht: (2024)