d-TreeRPO: Towards More Reliable Policy Optimization for Diffusion Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Pan, Leyi, Tao, Shuchang, Zhai, Yunpeng, Fu, Zheyu, Fang, Liancheng, He, Minghua, Zhang, Lingzhe, Liu, Zhaoyang, Ding, Bolin, Liu, Aiwei, Wen, Lijie |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models
di: Pan, Leyi, et al.
Pubblicazione: (2025)
di: Pan, Leyi, et al.
Pubblicazione: (2025)
MicroRemed: Benchmarking LLMs in Microservices Remediation
di: Zhang, Lingzhe, et al.
Pubblicazione: (2025)
di: Zhang, Lingzhe, et al.
Pubblicazione: (2025)
A Survey on Parallel Text Generation: From Parallel Decoding to Diffusion Language Models
di: Zhang, Lingzhe, et al.
Pubblicazione: (2025)
di: Zhang, Lingzhe, et al.
Pubblicazione: (2025)
Can Watermarked LLMs be Identified by Users via Crafted Prompts?
di: Liu, Aiwei, et al.
Pubblicazione: (2024)
di: Liu, Aiwei, et al.
Pubblicazione: (2024)
A Semantic Invariant Robust Watermark for Large Language Models
di: Liu, Aiwei, et al.
Pubblicazione: (2023)
di: Liu, Aiwei, et al.
Pubblicazione: (2023)
MarkDiffusion: An Open-Source Toolkit for Generative Watermarking of Latent Diffusion Models
di: Pan, Leyi, et al.
Pubblicazione: (2025)
di: Pan, Leyi, et al.
Pubblicazione: (2025)
An Unforgeable Publicly Verifiable Watermark for Large Language Models
di: Liu, Aiwei, et al.
Pubblicazione: (2023)
di: Liu, Aiwei, et al.
Pubblicazione: (2023)
Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation?
di: Pan, Leyi, et al.
Pubblicazione: (2025)
di: Pan, Leyi, et al.
Pubblicazione: (2025)
WaterSeeker: Pioneering Efficient Detection of Watermarked Segments in Large Documents
di: Pan, Leyi, et al.
Pubblicazione: (2024)
di: Pan, Leyi, et al.
Pubblicazione: (2024)
A Survey of Text Watermarking in the Era of Large Language Models
di: Liu, Aiwei, et al.
Pubblicazione: (2023)
di: Liu, Aiwei, et al.
Pubblicazione: (2023)
WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference
di: Liu, Aiwei, et al.
Pubblicazione: (2025)
di: Liu, Aiwei, et al.
Pubblicazione: (2025)
MarkLLM: An Open-Source Toolkit for LLM Watermarking
di: Pan, Leyi, et al.
Pubblicazione: (2024)
di: Pan, Leyi, et al.
Pubblicazione: (2024)
On the Robustness of Document-Level Relation Extraction Models to Entity Name Variations
di: Meng, Shiao, et al.
Pubblicazione: (2024)
di: Meng, Shiao, et al.
Pubblicazione: (2024)
ChatCite: LLM Agent with Human Workflow Guidance for Comparative Literature Summary
di: Li, Yutong, et al.
Pubblicazione: (2024)
di: Li, Yutong, et al.
Pubblicazione: (2024)
Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt Distillation
di: Liu, Aiwei, et al.
Pubblicazione: (2024)
di: Liu, Aiwei, et al.
Pubblicazione: (2024)
TIS-DPO: Token-level Importance Sampling for Direct Preference Optimization With Estimated Weights
di: Liu, Aiwei, et al.
Pubblicazione: (2024)
di: Liu, Aiwei, et al.
Pubblicazione: (2024)
TreeRPO: Tree Relative Policy Optimization
di: Yang, Zhicheng, et al.
Pubblicazione: (2025)
di: Yang, Zhicheng, et al.
Pubblicazione: (2025)
Tokenization Is More Than Compression
di: Schmidt, Craig W., et al.
Pubblicazione: (2024)
di: Schmidt, Craig W., et al.
Pubblicazione: (2024)
Dynamic-TreeRPO: Breaking the Independent Trajectory Bottleneck with Structured Sampling
di: Fu, Xiaolong, et al.
Pubblicazione: (2025)
di: Fu, Xiaolong, et al.
Pubblicazione: (2025)
Optimizing What We Trust: Reliability-Guided QUBO Selection of Multi-Agent Weak Framing Signals for Arabic Sentiment Prediction
di: Alkhalifa, Rabab
Pubblicazione: (2026)
di: Alkhalifa, Rabab
Pubblicazione: (2026)
Adaptive Steering and Remasking for Safe Generation in Diffusion Language Models
di: Lee, Yejin, et al.
Pubblicazione: (2026)
di: Lee, Yejin, et al.
Pubblicazione: (2026)
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
di: Orgad, Hadas, et al.
Pubblicazione: (2024)
di: Orgad, Hadas, et al.
Pubblicazione: (2024)
Multi-chain Graph Refinement and Selection for Reliable Reasoning in Large Language Models
di: Yang, Yujiao, et al.
Pubblicazione: (2025)
di: Yang, Yujiao, et al.
Pubblicazione: (2025)
E2E-REME: Towards End-to-End Microservices Auto-Remediation via Experience-Simulation Reinforcement Fine-Tuning
di: Zhang, Lingzhe, et al.
Pubblicazione: (2026)
di: Zhang, Lingzhe, et al.
Pubblicazione: (2026)
GATE: Graph-based Adaptive Tool Evolution Across Diverse Tasks
di: Luo, Jianwen, et al.
Pubblicazione: (2025)
di: Luo, Jianwen, et al.
Pubblicazione: (2025)
Parametric Social Identity Injection and Diversification in Public Opinion Simulation
di: Wang, Hexi, et al.
Pubblicazione: (2026)
di: Wang, Hexi, et al.
Pubblicazione: (2026)
Fast Quiet-STaR: Thinking Without Thought Tokens
di: Huang, Wei, et al.
Pubblicazione: (2025)
di: Huang, Wei, et al.
Pubblicazione: (2025)
Math Natural Language Inference: this should be easy!
di: de Paiva, Valeria, et al.
Pubblicazione: (2025)
di: de Paiva, Valeria, et al.
Pubblicazione: (2025)
Semantic Synergy: Unlocking Policy Insights and Learning Pathways Through Advanced Skill Mapping
di: Koundouri, Phoebe, et al.
Pubblicazione: (2025)
di: Koundouri, Phoebe, et al.
Pubblicazione: (2025)
Prompt Engineering and the Effectiveness of Large Language Models in Enhancing Human Productivity
di: Anam, Rizal Khoirul
Pubblicazione: (2025)
di: Anam, Rizal Khoirul
Pubblicazione: (2025)
Language Models, Graph Searching, and Supervision Adulteration: When More Supervision is Less and How to Make More More
di: Frydenlund, Arvid
Pubblicazione: (2025)
di: Frydenlund, Arvid
Pubblicazione: (2025)
Homogenization of Non-homogeneous Incompressible Navier-Stokes System in Critically Perforated Domains
di: Pan, Jiaojiao
Pubblicazione: (2024)
di: Pan, Jiaojiao
Pubblicazione: (2024)
LayerTracer: A Joint Task-Particle and Vulnerable-Layer Analysis framework for Arbitrary Large Language Model Architectures
di: Wu, Yuhang, et al.
Pubblicazione: (2026)
di: Wu, Yuhang, et al.
Pubblicazione: (2026)
Towards Reliable Retrieval in RAG Systems for Large Legal Datasets
di: Reuter, Markus, et al.
Pubblicazione: (2025)
di: Reuter, Markus, et al.
Pubblicazione: (2025)
The Knesset Corpus: An Annotated Corpus of Hebrew Parliamentary Proceedings
di: Goldin, Gili, et al.
Pubblicazione: (2024)
di: Goldin, Gili, et al.
Pubblicazione: (2024)
New Skills or Sharper Primitives? A Probabilistic Perspective on the Emergence of Reasoning in RLVR
di: Wang, Zhilin, et al.
Pubblicazione: (2026)
di: Wang, Zhilin, et al.
Pubblicazione: (2026)
Pitfalls in Evaluating Interpretability Agents
di: Haklay, Tal, et al.
Pubblicazione: (2026)
di: Haklay, Tal, et al.
Pubblicazione: (2026)
Towards Effective and Efficient Continual Pre-training of Large Language Models
di: Chen, Jie, et al.
Pubblicazione: (2024)
di: Chen, Jie, et al.
Pubblicazione: (2024)
The Unlikely Duel: Evaluating Creative Writing in LLMs through a Unique Scenario
di: Gómez-Rodríguez, Carlos, et al.
Pubblicazione: (2024)
di: Gómez-Rodríguez, Carlos, et al.
Pubblicazione: (2024)
SentiCSE: A Sentiment-aware Contrastive Sentence Embedding Framework with Sentiment-guided Textual Similarity
di: Kim, Jaemin, et al.
Pubblicazione: (2024)
di: Kim, Jaemin, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models
di: Pan, Leyi, et al.
Pubblicazione: (2025) -
MicroRemed: Benchmarking LLMs in Microservices Remediation
di: Zhang, Lingzhe, et al.
Pubblicazione: (2025) -
A Survey on Parallel Text Generation: From Parallel Decoding to Diffusion Language Models
di: Zhang, Lingzhe, et al.
Pubblicazione: (2025) -
Can Watermarked LLMs be Identified by Users via Crafted Prompts?
di: Liu, Aiwei, et al.
Pubblicazione: (2024) -
A Semantic Invariant Robust Watermark for Large Language Models
di: Liu, Aiwei, et al.
Pubblicazione: (2023)