First Return, Entropy-Eliciting Explore
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Tianyu, Xing, Tianshun, Gu, Qingshui, Liang, Taoran, Qu, Xingwei, Zhou, Xin, Li, Yizhi, Wen, Zhoufutu, Lin, Chenghua, Huang, Wenhao, Liu, Qian, Zhang, Ge, Ma, Zejun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
by: Li, Yizhi, et al.
Published: (2025)
by: Li, Yizhi, et al.
Published: (2025)
Overview of the NLPCC 2024 Shared Task on Chinese Metaphor Generation
by: Qu, Xingwei, et al.
Published: (2024)
by: Qu, Xingwei, et al.
Published: (2024)
Steel-LLM:From Scratch to Open Source -- A Personal Journey in Building a Chinese-Centric LLM
by: Gu, Qingshui, et al.
Published: (2025)
by: Gu, Qingshui, et al.
Published: (2025)
Kun: Answer Polishment for Chinese Self-Alignment with Instruction Back-Translation
by: Zheng, Tianyu, et al.
Published: (2024)
by: Zheng, Tianyu, et al.
Published: (2024)
I-SHEEP: Self-Alignment of LLM from Scratch through an Iterative Self-Enhancement Paradigm
by: Liang, Yiming, et al.
Published: (2024)
by: Liang, Yiming, et al.
Published: (2024)
KOR-Bench: Benchmarking Language Models on Knowledge-Orthogonal Reasoning Tasks
by: Ma, Kaijing, et al.
Published: (2024)
by: Ma, Kaijing, et al.
Published: (2024)
Dynamic Large Concept Models: Latent Reasoning in an Adaptive Semantic Space
by: Qu, Xingwei, et al.
Published: (2025)
by: Qu, Xingwei, et al.
Published: (2025)
MMRA: A Benchmark for Evaluating Multi-Granularity and Multi-Image Relational Association Capabilities in Large Visual Language Models
by: Wu, Siwei, et al.
Published: (2024)
by: Wu, Siwei, et al.
Published: (2024)
MORE-3S:Multimodal-based Offline Reinforcement Learning with Shared Semantic Spaces
by: Zheng, Tianyu, et al.
Published: (2024)
by: Zheng, Tianyu, et al.
Published: (2024)
TEGEE: Task dEfinition Guided Expert Ensembling for Generalizable and Few-shot Learning
by: Qu, Xingwei, et al.
Published: (2024)
by: Qu, Xingwei, et al.
Published: (2024)
Overview of the NLPCC 2025 Shared Task: Gender Bias Mitigation Challenge
by: Li, Yizhi, et al.
Published: (2025)
by: Li, Yizhi, et al.
Published: (2025)
Observing Micromotives and Macrobehavior of Large Language Models
by: Cheng, Yuyang, et al.
Published: (2024)
by: Cheng, Yuyang, et al.
Published: (2024)
COIG-Writer: A High-Quality Dataset for Chinese Creative Writing with Thought Processes
by: Li, Yunwen, et al.
Published: (2025)
by: Li, Yunwen, et al.
Published: (2025)
Aligning Instruction Tuning with Pre-training
by: Liang, Yiming, et al.
Published: (2025)
by: Liang, Yiming, et al.
Published: (2025)
Beyond Correctness: Evaluating Subjective Writing Preferences Across Cultures
by: Ying, Shuangshuang, et al.
Published: (2025)
by: Ying, Shuangshuang, et al.
Published: (2025)
Exceptional Wear Resistance and Unique Friction Behavior in High‐Entropy Rare‐Earth Dodecaboride Composite
by: Shixuan Gu, et al.
Published: (2026)
by: Shixuan Gu, et al.
Published: (2026)
LongEval: A Comprehensive Analysis of Long-Text Generation Through a Plan-based Paradigm
by: Wu, Siwei, et al.
Published: (2025)
by: Wu, Siwei, et al.
Published: (2025)
CoTJudger: A Graph-Driven Framework for Automatic Evaluation of Chain-of-Thought Efficiency and Redundancy in LRMs
by: Li, Siyi, et al.
Published: (2026)
by: Li, Siyi, et al.
Published: (2026)
Super‐Resolution Ghost Imaging Through Complex Scattering in Dynamic Media
by: Yang Peng, et al.
Published: (2025)
by: Yang Peng, et al.
Published: (2025)
CMDAG: A Chinese Metaphor Dataset with Annotated Grounds as CoT for Boosting Metaphor Generation
by: Shao, Yujie, et al.
Published: (2024)
by: Shao, Yujie, et al.
Published: (2024)
Avoiding Premature Collapse: Adaptive Annealing for Entropy-Regularized Structural Inference
by: Liu, Yizhi
Published: (2026)
by: Liu, Yizhi
Published: (2026)
CMMMU: A Chinese Massive Multi-discipline Multimodal Understanding Benchmark
by: Zhang, Ge, et al.
Published: (2024)
by: Zhang, Ge, et al.
Published: (2024)
A First Look at GPT Apps: Landscape and Vulnerability
by: Zhang, Zejun, et al.
Published: (2024)
by: Zhang, Zejun, et al.
Published: (2024)
Towards Prospective Medical Image Reconstruction via Knowledge-Informed Dynamic Optimal Transport
by: Zheng, Taoran, et al.
Published: (2025)
by: Zheng, Taoran, et al.
Published: (2025)
General-Reasoner: Advancing LLM Reasoning Across All Domains
by: Ma, Xueguang, et al.
Published: (2025)
by: Ma, Xueguang, et al.
Published: (2025)
Orbital angular momentum transmission in time-varying scattering media using dual orthogonal polarization channels
by: Li, Heshen, et al.
Published: (2026)
by: Li, Heshen, et al.
Published: (2026)
OmniBench: Towards The Future of Universal Omni-Language Models
by: Li, Yizhi, et al.
Published: (2024)
by: Li, Yizhi, et al.
Published: (2024)
Audio Contrastive-based Fine-tuning: Decoupling Representation Learning and Classification
by: Wang, Yang, et al.
Published: (2023)
by: Wang, Yang, et al.
Published: (2023)
Target Return Strategy
by: Ying Xue, et al.
Published: (2025)
by: Ying Xue, et al.
Published: (2025)
CodeEditorBench: Evaluating Code Editing Capability of Large Language Models
by: Guo, Jiawei, et al.
Published: (2024)
by: Guo, Jiawei, et al.
Published: (2024)
ConceptMoE: Adaptive Token-to-Concept Compression for Implicit Compute Allocation
by: Huang, Zihao, et al.
Published: (2026)
by: Huang, Zihao, et al.
Published: (2026)
CIF-Bench: A Chinese Instruction-Following Benchmark for Evaluating the Generalizability of Large Language Models
by: LI, Yizhi, et al.
Published: (2024)
by: LI, Yizhi, et al.
Published: (2024)
MMTE: Corpus and Metrics for Evaluating Machine Translation Quality of Metaphorical Language
by: Wang, Shun, et al.
Published: (2024)
by: Wang, Shun, et al.
Published: (2024)
Re:Form -- Reducing Human Priors in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny
by: Yan, Chuanhao, et al.
Published: (2025)
by: Yan, Chuanhao, et al.
Published: (2025)
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
by: Guo, Jarvis, et al.
Published: (2024)
by: Guo, Jarvis, et al.
Published: (2024)
LIME: Less Is More for MLLM Evaluation
by: Zhu, King, et al.
Published: (2024)
by: Zhu, King, et al.
Published: (2024)
SAGA: Summarization-Guided Assert Statement Generation
by: Zhang, Yuwei, et al.
Published: (2023)
by: Zhang, Yuwei, et al.
Published: (2023)
SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval
by: Wu, Siwei, et al.
Published: (2024)
by: Wu, Siwei, et al.
Published: (2024)
Optimal Design for Human Preference Elicitation
by: Mukherjee, Subhojyoti, et al.
Published: (2024)
by: Mukherjee, Subhojyoti, et al.
Published: (2024)
GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
by: Yu, Jiahao, et al.
Published: (2023)
by: Yu, Jiahao, et al.
Published: (2023)
Similar Items
-
TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
by: Li, Yizhi, et al.
Published: (2025) -
Overview of the NLPCC 2024 Shared Task on Chinese Metaphor Generation
by: Qu, Xingwei, et al.
Published: (2024) -
Steel-LLM:From Scratch to Open Source -- A Personal Journey in Building a Chinese-Centric LLM
by: Gu, Qingshui, et al.
Published: (2025) -
Kun: Answer Polishment for Chinese Self-Alignment with Instruction Back-Translation
by: Zheng, Tianyu, et al.
Published: (2024) -
I-SHEEP: Self-Alignment of LLM from Scratch through an Iterative Self-Enhancement Paradigm
by: Liang, Yiming, et al.
Published: (2024)