POLCA: Stochastic Generative Optimization with LLM
Fuente:
arXiv
Salvato in:
| Autori principali: | Ren, Xuanfei, Nie, Allen, Xie, Tengyang, Cheng, Ching-An |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Trace is the Next AutoDiff: Generative Optimization with Rich Feedback, Execution Traces, and LLMs
di: Cheng, Ching-An, et al.
Pubblicazione: (2024)
di: Cheng, Ching-An, et al.
Pubblicazione: (2024)
Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
di: Rosset, Corby, et al.
Pubblicazione: (2024)
di: Rosset, Corby, et al.
Pubblicazione: (2024)
Reinforce LLM Reasoning through Multi-Agent Reflection
di: Yuan, Yurun, et al.
Pubblicazione: (2025)
di: Yuan, Yurun, et al.
Pubblicazione: (2025)
Breaking the Capability Ceiling of LLM Post-Training by Reintroducing Markov States
di: Yuan, Yurun, et al.
Pubblicazione: (2026)
di: Yuan, Yurun, et al.
Pubblicazione: (2026)
Principal Orthogonal Latent Components Analysis (POLCA Net)
di: H., Jose Antonio Martin, et al.
Pubblicazione: (2024)
di: H., Jose Antonio Martin, et al.
Pubblicazione: (2024)
Offline Reinforcement Learning in Large State Spaces: Algorithms and Guarantees
di: Jiang, Nan, et al.
Pubblicazione: (2025)
di: Jiang, Nan, et al.
Pubblicazione: (2025)
Do We Need to Verify Step by Step? Rethinking Process Supervision from a Theoretical Perspective
di: Jia, Zeyu, et al.
Pubblicazione: (2025)
di: Jia, Zeyu, et al.
Pubblicazione: (2025)
Learning Game-Playing Agents with Generative Code Optimization
di: Kuang, Zhiyi, et al.
Pubblicazione: (2025)
di: Kuang, Zhiyi, et al.
Pubblicazione: (2025)
Understanding the Challenges in Iterative Generative Optimization with LLMs
di: Nie, Allen, et al.
Pubblicazione: (2026)
di: Nie, Allen, et al.
Pubblicazione: (2026)
Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM Reasoning
di: Yuan, Yurun, et al.
Pubblicazione: (2025)
di: Yuan, Yurun, et al.
Pubblicazione: (2025)
Self-Play with Adversarial Critic: Provable and Scalable Offline Alignment for Language Models
di: Ji, Xiang, et al.
Pubblicazione: (2024)
di: Ji, Xiang, et al.
Pubblicazione: (2024)
Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits
di: Chen, Fan, et al.
Pubblicazione: (2025)
di: Chen, Fan, et al.
Pubblicazione: (2025)
Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
di: Xie, Tengyang, et al.
Pubblicazione: (2024)
di: Xie, Tengyang, et al.
Pubblicazione: (2024)
Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
di: Huang, Audrey, et al.
Pubblicazione: (2024)
di: Huang, Audrey, et al.
Pubblicazione: (2024)
Mano: Restriking Manifold Optimization for LLM Training
di: Gu, Yufei, et al.
Pubblicazione: (2026)
di: Gu, Yufei, et al.
Pubblicazione: (2026)
Fine-grained Analysis of Stability and Generalization for Stochastic Bilevel Optimization
di: Zhang, Xuelin, et al.
Pubblicazione: (2026)
di: Zhang, Xuelin, et al.
Pubblicazione: (2026)
Doubly Stochastic Adaptive Neighbors Clustering via the Marcus Mapping
di: Yuan, Jinghui, et al.
Pubblicazione: (2024)
di: Yuan, Jinghui, et al.
Pubblicazione: (2024)
Second-Order Convergence in Private Stochastic Non-Convex Optimization
di: Tao, Youming, et al.
Pubblicazione: (2025)
di: Tao, Youming, et al.
Pubblicazione: (2025)
Towards Principled Representation Learning from Videos for Reinforcement Learning
di: Misra, Dipendra, et al.
Pubblicazione: (2024)
di: Misra, Dipendra, et al.
Pubblicazione: (2024)
From Stochastic Answers to Verifiable Reasoning: Interpretable Decision-Making with LLM-Generated Code
di: Mahesh, Anirudh Jaidev, et al.
Pubblicazione: (2026)
di: Mahesh, Anirudh Jaidev, et al.
Pubblicazione: (2026)
RosettaSearch: Multi-Objective Inference-Time Search for Protein Sequence Design
di: Kshirsagar, Meghana, et al.
Pubblicazione: (2026)
di: Kshirsagar, Meghana, et al.
Pubblicazione: (2026)
Uncertainty-Aware Decision Transformer for Stochastic Driving Environments
di: Li, Zenan, et al.
Pubblicazione: (2023)
di: Li, Zenan, et al.
Pubblicazione: (2023)
Stochastic Optimization of Inventory at Large-scale Supply Chains
di: Jin, Zhaoyang Larry, et al.
Pubblicazione: (2025)
di: Jin, Zhaoyang Larry, et al.
Pubblicazione: (2025)
CounterCurate: Enhancing Physical and Semantic Visio-Linguistic Compositional Reasoning via Counterfactual Examples
di: Zhang, Jianrui, et al.
Pubblicazione: (2024)
di: Zhang, Jianrui, et al.
Pubblicazione: (2024)
Paying Less Generalization Tax: A Cross-Domain Generalization Study of RL Training for LLM Agents
di: Liu, Zhihan, et al.
Pubblicazione: (2026)
di: Liu, Zhihan, et al.
Pubblicazione: (2026)
Certificate-Guided Pruning for Stochastic Lipschitz Optimization
di: Shihab, Ibne Farabi, et al.
Pubblicazione: (2026)
di: Shihab, Ibne Farabi, et al.
Pubblicazione: (2026)
Beyond Sharp Minima: Robust LLM Unlearning via Feedback-Guided Multi-Point Optimization
di: Wu, Wenhan, et al.
Pubblicazione: (2025)
di: Wu, Wenhan, et al.
Pubblicazione: (2025)
SE(3)-Stochastic Flow Matching for Protein Backbone Generation
di: Bose, Avishek Joey, et al.
Pubblicazione: (2023)
di: Bose, Avishek Joey, et al.
Pubblicazione: (2023)
A Universal Banach--Bregman Framework for Stochastic Iterations: Unifying Stochastic Mirror Descent, Learning and LLM Training
di: Zhang, Johnny R., et al.
Pubblicazione: (2025)
di: Zhang, Johnny R., et al.
Pubblicazione: (2025)
OAT-Rephrase: Optimization-Aware Training Data Rephrasing for Zeroth-Order LLM Fine-Tuning
di: Long, Jikai, et al.
Pubblicazione: (2025)
di: Long, Jikai, et al.
Pubblicazione: (2025)
Fairness Aware Reward Optimization
di: Choi, Ching Lam, et al.
Pubblicazione: (2026)
di: Choi, Ching Lam, et al.
Pubblicazione: (2026)
How Strategic Agents Respond: Comparing Analytical Models with LLM-Generated Responses in Strategic Classification
di: Xie, Tian, et al.
Pubblicazione: (2025)
di: Xie, Tian, et al.
Pubblicazione: (2025)
Deterministic Decomposition of Stochastic Generative Dynamics
di: Song, Xingyu, et al.
Pubblicazione: (2026)
di: Song, Xingyu, et al.
Pubblicazione: (2026)
Generative Modeling with Phase Stochastic Bridges
di: Chen, Tianrong, et al.
Pubblicazione: (2023)
di: Chen, Tianrong, et al.
Pubblicazione: (2023)
A Unified Gaussian Process for Branching and Nested Hyperparameter Optimization
di: Zhang, Jiazhao, et al.
Pubblicazione: (2024)
di: Zhang, Jiazhao, et al.
Pubblicazione: (2024)
Unifying Bayesian Flow Networks and Diffusion Models through Stochastic Differential Equations
di: Xue, Kaiwen, et al.
Pubblicazione: (2024)
di: Xue, Kaiwen, et al.
Pubblicazione: (2024)
SAGE: Sequence-level Adaptive Gradient Evolution for Generative Recommendation
di: Xie, Yu, et al.
Pubblicazione: (2026)
di: Xie, Yu, et al.
Pubblicazione: (2026)
E^2-LLM: Bridging Neural Signals and Interpretable Affective Analysis
di: Ma, Fei, et al.
Pubblicazione: (2026)
di: Ma, Fei, et al.
Pubblicazione: (2026)
KernelBand: Steering LLM-based Kernel Optimization via Hardware-Aware Multi-Armed Bandits
di: Ran, Dezhi, et al.
Pubblicazione: (2025)
di: Ran, Dezhi, et al.
Pubblicazione: (2025)
Predicting Long Term Sequential Policy Value Using Softer Surrogates
di: Nam, Hyunji, et al.
Pubblicazione: (2024)
di: Nam, Hyunji, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Trace is the Next AutoDiff: Generative Optimization with Rich Feedback, Execution Traces, and LLMs
di: Cheng, Ching-An, et al.
Pubblicazione: (2024) -
Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
di: Rosset, Corby, et al.
Pubblicazione: (2024) -
Reinforce LLM Reasoning through Multi-Agent Reflection
di: Yuan, Yurun, et al.
Pubblicazione: (2025) -
Breaking the Capability Ceiling of LLM Post-Training by Reintroducing Markov States
di: Yuan, Yurun, et al.
Pubblicazione: (2026) -
Principal Orthogonal Latent Components Analysis (POLCA Net)
di: H., Jose Antonio Martin, et al.
Pubblicazione: (2024)