Balancing Exploration and Exploitation in LLM using Soft RLLF for Enhanced Negation Understanding
Fuente:
arXiv
Guardado en:
| Autores principales: | Nguyen, Ha-Thanh, Satoh, Ken |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
KRAG Framework for Enhancing LLMs in the Legal Domain
por: Thanh, Nguyen Ha, et al.
Publicado: (2024)
por: Thanh, Nguyen Ha, et al.
Publicado: (2024)
GPTs and Language Barrier: A Cross-Lingual Legal QA Examination
por: Nguyen, Ha-Thanh, et al.
Publicado: (2024)
por: Nguyen, Ha-Thanh, et al.
Publicado: (2024)
Layer-of-Thoughts Prompting (LoT): Leveraging LLM-Based Retrieval with Constraint Hierarchies
por: Fungwacharakorn, Wachara, et al.
Publicado: (2024)
por: Fungwacharakorn, Wachara, et al.
Publicado: (2024)
Enhancing Legal Document Retrieval: A Multi-Phase Approach with Large Language Models
por: Nguyen, Hai-Long, et al.
Publicado: (2024)
por: Nguyen, Hai-Long, et al.
Publicado: (2024)
Legal2LogicICL: Improving Generalization in Transforming Legal Cases to Logical Formulas via Diverse Few-Shot Learning
por: Xue, Jieying, et al.
Publicado: (2026)
por: Xue, Jieying, et al.
Publicado: (2026)
PYTHEN: A Flexible Framework for Legal Reasoning in Python
por: Nguyen, Ha-Thanh, et al.
Publicado: (2026)
por: Nguyen, Ha-Thanh, et al.
Publicado: (2026)
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners
por: Zeng, Weihao, et al.
Publicado: (2024)
por: Zeng, Weihao, et al.
Publicado: (2024)
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
por: Chen, Zhipeng, et al.
Publicado: (2025)
por: Chen, Zhipeng, et al.
Publicado: (2025)
$ϕ$-Decoding: Adaptive Foresight Sampling for Balanced Inference-Time Exploration and Exploitation
por: Xu, Fangzhi, et al.
Publicado: (2025)
por: Xu, Fangzhi, et al.
Publicado: (2025)
HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control
por: Yao, Xincheng, et al.
Publicado: (2026)
por: Yao, Xincheng, et al.
Publicado: (2026)
Multi-Agent Legal Verifier Systems for Data Transfer Planning
por: Nguyen, Ha-Thanh, et al.
Publicado: (2025)
por: Nguyen, Ha-Thanh, et al.
Publicado: (2025)
Enhancing Semantics in Multimodal Chain of Thought via Soft Negative Sampling
por: Zheng, Guangmin, et al.
Publicado: (2024)
por: Zheng, Guangmin, et al.
Publicado: (2024)
Code Repair with LLMs gives an Exploration-Exploitation Tradeoff
por: Tang, Hao, et al.
Publicado: (2024)
por: Tang, Hao, et al.
Publicado: (2024)
Disentangling Exploration of Large Language Models by Optimal Exploitation
por: Grams, Tim, et al.
Publicado: (2025)
por: Grams, Tim, et al.
Publicado: (2025)
A Scalable Multi-LLM Collaboration System with Retrieval-based Selection and Exploration-Exploitation-Driven Enhancement
por: Tang, Shengji, et al.
Publicado: (2025)
por: Tang, Shengji, et al.
Publicado: (2025)
Enhancing Retrieval Augmented Generation with Hierarchical Text Segmentation Chunking
por: Nguyen, Hai Toan, et al.
Publicado: (2025)
por: Nguyen, Hai Toan, et al.
Publicado: (2025)
ClaimPKG: Enhancing Claim Verification via Pseudo-Subgraph Generation with Lightweight Specialized LLM
por: Pham, Hoang, et al.
Publicado: (2025)
por: Pham, Hoang, et al.
Publicado: (2025)
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
por: Chen, Peter, et al.
Publicado: (2025)
por: Chen, Peter, et al.
Publicado: (2025)
Exploiting LLMs' Reasoning Capability to Infer Implicit Concepts in Legal Information Retrieval
por: Nguyen, Hai-Long, et al.
Publicado: (2024)
por: Nguyen, Hai-Long, et al.
Publicado: (2024)
Detection of Illicit Content on Online Marketplaces using Large Language Models
por: Tran, Quoc Khoa, et al.
Publicado: (2026)
por: Tran, Quoc Khoa, et al.
Publicado: (2026)
Enhancing LLM Medical Coding with Structured External Knowledge
por: Gan, Yidong, et al.
Publicado: (2026)
por: Gan, Yidong, et al.
Publicado: (2026)
Misinformation Detection using Large Language Models with Explainability
por: Patel, Jainee, et al.
Publicado: (2025)
por: Patel, Jainee, et al.
Publicado: (2025)
VLQA: The First Comprehensive, Large, and High-Quality Vietnamese Dataset for Legal Question Answering
por: Nguyen, Tan-Minh, et al.
Publicado: (2025)
por: Nguyen, Tan-Minh, et al.
Publicado: (2025)
Balancing Faithfulness and Performance in Reasoning via Multi-Listener Soft Execution
por: Sivakumaran, Nithin, et al.
Publicado: (2026)
por: Sivakumaran, Nithin, et al.
Publicado: (2026)
MEDSAGE: Enhancing Robustness of Medical Dialogue Summarization to ASR Errors with LLM-generated Synthetic Dialogues
por: Binici, Kuluhan, et al.
Publicado: (2024)
por: Binici, Kuluhan, et al.
Publicado: (2024)
Graph Counselor: Adaptive Graph Exploration via Multi-Agent Synergy to Enhance LLM Reasoning
por: Gao, Junqi, et al.
Publicado: (2025)
por: Gao, Junqi, et al.
Publicado: (2025)
NOWJ@COLIEE 2025: A Multi-stage Framework Integrating Embedding Models and Large Language Models for Legal Retrieval and Entailment
por: Nguyen, Hoang-Trung, et al.
Publicado: (2025)
por: Nguyen, Hoang-Trung, et al.
Publicado: (2025)
AdaSwitch: Balancing Exploration and Guidance in Knowledge Distillation via Adaptive Switching
por: Peng, Jingyu, et al.
Publicado: (2025)
por: Peng, Jingyu, et al.
Publicado: (2025)
Enhancing Adversarial Transferability by Balancing Exploration and Exploitation with Gradient-Guided Sampling
por: Niu, Zenghao, et al.
Publicado: (2025)
por: Niu, Zenghao, et al.
Publicado: (2025)
Data Augmented Pipeline for Legal Information Extraction and Reasoning
por: Phuong, Nguyen Minh, et al.
Publicado: (2026)
por: Phuong, Nguyen Minh, et al.
Publicado: (2026)
When Greedy Wins: Emergent Exploitation Bias in Meta-Bandit LLM Training
por: Chen, Sanxing, et al.
Publicado: (2025)
por: Chen, Sanxing, et al.
Publicado: (2025)
A Weakly Supervised Data Labeling Framework for Machine Lexical Normalization in Vietnamese Social Media
por: Nguyen, Dung Ha, et al.
Publicado: (2024)
por: Nguyen, Dung Ha, et al.
Publicado: (2024)
ViSoLex: An Open-Source Repository for Vietnamese Social Media Lexical Normalization
por: Nguyen, Anh Thi-Hoang, et al.
Publicado: (2025)
por: Nguyen, Anh Thi-Hoang, et al.
Publicado: (2025)
BIS Reasoning 1.0: The First Large-Scale Japanese Benchmark for Belief-Inconsistent Syllogistic Reasoning
por: Nguyen, Ha-Thanh, et al.
Publicado: (2025)
por: Nguyen, Ha-Thanh, et al.
Publicado: (2025)
Calibrate-Then-Act: Cost-Aware Exploration in LLM Agents
por: Ding, Wenxuan, et al.
Publicado: (2026)
por: Ding, Wenxuan, et al.
Publicado: (2026)
From No to Know: Taxonomy, Challenges, and Opportunities for Negation Understanding in Multimodal Foundation Models
por: Vatsa, Mayank, et al.
Publicado: (2025)
por: Vatsa, Mayank, et al.
Publicado: (2025)
Dual-Space Smoothness for Robust and Balanced LLM Unlearning
por: Yan, Han, et al.
Publicado: (2025)
por: Yan, Han, et al.
Publicado: (2025)
Optimus: Accelerating Large-Scale Multi-Modal LLM Training by Bubble Exploitation
por: Feng, Weiqi, et al.
Publicado: (2024)
por: Feng, Weiqi, et al.
Publicado: (2024)
ARUQULA -- An LLM based Text2SPARQL Approach using ReAct and Knowledge Graph Exploration Utilities
por: Brei, Felix, et al.
Publicado: (2025)
por: Brei, Felix, et al.
Publicado: (2025)
Integrating Symbolic Natural Language Understanding and Language Models for Word Sense Disambiguation
por: Zhao, Kexin, et al.
Publicado: (2025)
por: Zhao, Kexin, et al.
Publicado: (2025)
Ejemplares similares
-
KRAG Framework for Enhancing LLMs in the Legal Domain
por: Thanh, Nguyen Ha, et al.
Publicado: (2024) -
GPTs and Language Barrier: A Cross-Lingual Legal QA Examination
por: Nguyen, Ha-Thanh, et al.
Publicado: (2024) -
Layer-of-Thoughts Prompting (LoT): Leveraging LLM-Based Retrieval with Constraint Hierarchies
por: Fungwacharakorn, Wachara, et al.
Publicado: (2024) -
Enhancing Legal Document Retrieval: A Multi-Phase Approach with Large Language Models
por: Nguyen, Hai-Long, et al.
Publicado: (2024) -
Legal2LogicICL: Improving Generalization in Transforming Legal Cases to Logical Formulas via Diverse Few-Shot Learning
por: Xue, Jieying, et al.
Publicado: (2026)