TRE: Encouraging Exploration in the Trust Region
Fuente:
arXiv
Salvato in:
| Autori principali: | Huang, Chao, Lu, Yujing, Li, Quangang, Wang, Shenghe, Wang, Yan, Zhang, Yueyang, Xia, Long, Zhao, Jiashu, Sun, Zhiyuan, Shi, Daiting, Liu, Tingwen |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training
di: Liang, Yu, et al.
Pubblicazione: (2026)
di: Liang, Yu, et al.
Pubblicazione: (2026)
ReflectRM: Boosting Generative Reward Models via Self-Reflection within a Unified Judgment Framework
di: Qin, Kai, et al.
Pubblicazione: (2026)
di: Qin, Kai, et al.
Pubblicazione: (2026)
Advancing General-Purpose Reasoning Models with Modular Gradient Surgery
di: Cai, Min, et al.
Pubblicazione: (2026)
di: Cai, Min, et al.
Pubblicazione: (2026)
Safety Alignment Should Be Made More Than Just A Few Attention Heads
di: Huang, Chao, et al.
Pubblicazione: (2025)
di: Huang, Chao, et al.
Pubblicazione: (2025)
GenCRF: Generative Clustering and Reformulation Framework for Enhanced Intent-Driven Information Retrieval
di: Seo, Wonduk, et al.
Pubblicazione: (2024)
di: Seo, Wonduk, et al.
Pubblicazione: (2024)
Enhancing Multimodal Entity and Relation Extraction with Variational Information Bottleneck
di: Cui, Shiyao, et al.
Pubblicazione: (2023)
di: Cui, Shiyao, et al.
Pubblicazione: (2023)
When Less is More: The LLM Scaling Paradox in Context Compression
di: Guo, Ruishan, et al.
Pubblicazione: (2026)
di: Guo, Ruishan, et al.
Pubblicazione: (2026)
Reinforced Efficient Reasoning via Semantically Diverse Exploration
di: Zhao, Ziqi, et al.
Pubblicazione: (2026)
di: Zhao, Ziqi, et al.
Pubblicazione: (2026)
Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate
di: Liang, Tian, et al.
Pubblicazione: (2023)
di: Liang, Tian, et al.
Pubblicazione: (2023)
DARL: Encouraging Diverse Answers for General Reasoning without Verifiers
di: Huang, Chongxuan, et al.
Pubblicazione: (2026)
di: Huang, Chongxuan, et al.
Pubblicazione: (2026)
Trust Region On-Policy Distillation
di: Xing, Xingrun, et al.
Pubblicazione: (2026)
di: Xing, Xingrun, et al.
Pubblicazione: (2026)
ClueAnchor: Clue-Anchored Knowledge Reasoning Exploration and Optimization for Retrieval-Augmented Generation
di: Chen, Hao, et al.
Pubblicazione: (2025)
di: Chen, Hao, et al.
Pubblicazione: (2025)
HomeSafeBench: A Benchmark for Embodied Vision-Language Models in Free-Exploration Home Safety Inspection
di: Gao, Siyuan, et al.
Pubblicazione: (2025)
di: Gao, Siyuan, et al.
Pubblicazione: (2025)
Encourage or Inhibit Monosemanticity? Revisit Monosemanticity from a Feature Decorrelation Perspective
di: Yan, Hanqi, et al.
Pubblicazione: (2024)
di: Yan, Hanqi, et al.
Pubblicazione: (2024)
SynPlanResearch-R1: Encouraging Tool Exploration for Deep Research with Synthetic Plans
di: Zeng, Hansi, et al.
Pubblicazione: (2026)
di: Zeng, Hansi, et al.
Pubblicazione: (2026)
Policy Split: Incentivizing Dual-Mode Exploration in LLM Reinforcement with Dual-Mode Entropy Regularization
di: Yao, Jiashu, et al.
Pubblicazione: (2026)
di: Yao, Jiashu, et al.
Pubblicazione: (2026)
Chain-of-Procedure: Hierarchical Visual-Language Reasoning for Procedural QA
di: Chen, Guanhua, et al.
Pubblicazione: (2026)
di: Chen, Guanhua, et al.
Pubblicazione: (2026)
Adaptive Data Augmentation for Aspect Sentiment Quad Prediction
di: Zhang, Wenyuan, et al.
Pubblicazione: (2024)
di: Zhang, Wenyuan, et al.
Pubblicazione: (2024)
SED-SFT: Selectively Encouraging Diversity in Supervised Fine-Tuning
di: Chen, Yijie, et al.
Pubblicazione: (2026)
di: Chen, Yijie, et al.
Pubblicazione: (2026)
Thinking as Compression: Your Reasoning Model is Secretly a Context Compressor
di: Ma, Guoxin, et al.
Pubblicazione: (2026)
di: Ma, Guoxin, et al.
Pubblicazione: (2026)
Deterministic Reversible Data Augmentation for Neural Machine Translation
di: Yao, Jiashu, et al.
Pubblicazione: (2024)
di: Yao, Jiashu, et al.
Pubblicazione: (2024)
Agentic-R: Learning to Retrieve for Agentic Search
di: Liu, Wenhan, et al.
Pubblicazione: (2026)
di: Liu, Wenhan, et al.
Pubblicazione: (2026)
Towards AI Search Paradigm
di: Li, Yuchen, et al.
Pubblicazione: (2025)
di: Li, Yuchen, et al.
Pubblicazione: (2025)
Unleashing the Power of LLMs in Dense Retrieval with Query Likelihood Modeling
di: Zhang, Hengran, et al.
Pubblicazione: (2025)
di: Zhang, Hengran, et al.
Pubblicazione: (2025)
Graph-GRPO: Stabilizing Multi-Agent Topology Learning via Group Relative Policy Optimization
di: Cang, Yueyang, et al.
Pubblicazione: (2026)
di: Cang, Yueyang, et al.
Pubblicazione: (2026)
EvoSpec: Evolving Speculative Decoding via Real-Time Vocabulary and Parameter Adaptation
di: Zhang, Shuyu, et al.
Pubblicazione: (2026)
di: Zhang, Shuyu, et al.
Pubblicazione: (2026)
TrustUQA: A Trustful Framework for Unified Structured Data Question Answering
di: Zhang, Wen, et al.
Pubblicazione: (2024)
di: Zhang, Wen, et al.
Pubblicazione: (2024)
Improving Reasoning Capabilities in Small Models through Mixture-of-Layers Distillation with Stepwise Attention on Key Information
di: Chen, Yao, et al.
Pubblicazione: (2026)
di: Chen, Yao, et al.
Pubblicazione: (2026)
DINT Transformer
di: Cang, Yueyang, et al.
Pubblicazione: (2025)
di: Cang, Yueyang, et al.
Pubblicazione: (2025)
Large Language Model-Powered Query-Driven Event Timeline Summarization in Industrial Search
di: Wang, Mingyue, et al.
Pubblicazione: (2026)
di: Wang, Mingyue, et al.
Pubblicazione: (2026)
TourRank: Utilizing Large Language Models for Documents Ranking with a Tournament-Inspired Strategy
di: Chen, Yiqun, et al.
Pubblicazione: (2024)
di: Chen, Yiqun, et al.
Pubblicazione: (2024)
Hit-RAG: Learning to Reason with Long Contexts via Preference Alignment
di: Liu, Junming, et al.
Pubblicazione: (2026)
di: Liu, Junming, et al.
Pubblicazione: (2026)
Utilizing and Calibrating Hindsight Process Rewards via Reinforcement with Mutual Information Self-Evaluation
di: Yao, Jiashu, et al.
Pubblicazione: (2026)
di: Yao, Jiashu, et al.
Pubblicazione: (2026)
Entropy-Tree: Tree-Based Decoding with Entropy-Guided Exploration
di: Wei, Longxuan, et al.
Pubblicazione: (2026)
di: Wei, Longxuan, et al.
Pubblicazione: (2026)
HomeBench: Evaluating LLMs in Smart Homes with Valid and Invalid Instructions Across Single and Multiple Devices
di: Li, Silin, et al.
Pubblicazione: (2025)
di: Li, Silin, et al.
Pubblicazione: (2025)
Text-Video Retrieval via Variational Multi-Modal Hypergraph Networks
di: Li, Qian, et al.
Pubblicazione: (2024)
di: Li, Qian, et al.
Pubblicazione: (2024)
Mixture of Hidden-Dimensions Transformer
di: Chen, Yilong, et al.
Pubblicazione: (2024)
di: Chen, Yilong, et al.
Pubblicazione: (2024)
GRADE: Probing Knowledge Gaps in LLMs through Gradient Subspace Dynamics
di: Wang, Yujing, et al.
Pubblicazione: (2026)
di: Wang, Yujing, et al.
Pubblicazione: (2026)
IBMEA: Exploring Variational Information Bottleneck for Multi-modal Entity Alignment
di: Su, Taoyu, et al.
Pubblicazione: (2024)
di: Su, Taoyu, et al.
Pubblicazione: (2024)
DeepNote: Note-Centric Deep Retrieval-Augmented Generation
di: Wang, Ruobing, et al.
Pubblicazione: (2024)
di: Wang, Ruobing, et al.
Pubblicazione: (2024)
Documenti analoghi
-
ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training
di: Liang, Yu, et al.
Pubblicazione: (2026) -
ReflectRM: Boosting Generative Reward Models via Self-Reflection within a Unified Judgment Framework
di: Qin, Kai, et al.
Pubblicazione: (2026) -
Advancing General-Purpose Reasoning Models with Modular Gradient Surgery
di: Cai, Min, et al.
Pubblicazione: (2026) -
Safety Alignment Should Be Made More Than Just A Few Attention Heads
di: Huang, Chao, et al.
Pubblicazione: (2025) -
GenCRF: Generative Clustering and Reformulation Framework for Enhanced Intent-Driven Information Retrieval
di: Seo, Wonduk, et al.
Pubblicazione: (2024)