POLCA: Stochastic Generative Optimization with LLM
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ren, Xuanfei, Nie, Allen, Xie, Tengyang, Cheng, Ching-An |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Trace is the Next AutoDiff: Generative Optimization with Rich Feedback, Execution Traces, and LLMs
von: Cheng, Ching-An, et al.
Veröffentlicht: (2024)
von: Cheng, Ching-An, et al.
Veröffentlicht: (2024)
Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
von: Rosset, Corby, et al.
Veröffentlicht: (2024)
von: Rosset, Corby, et al.
Veröffentlicht: (2024)
Reinforce LLM Reasoning through Multi-Agent Reflection
von: Yuan, Yurun, et al.
Veröffentlicht: (2025)
von: Yuan, Yurun, et al.
Veröffentlicht: (2025)
Breaking the Capability Ceiling of LLM Post-Training by Reintroducing Markov States
von: Yuan, Yurun, et al.
Veröffentlicht: (2026)
von: Yuan, Yurun, et al.
Veröffentlicht: (2026)
Principal Orthogonal Latent Components Analysis (POLCA Net)
von: H., Jose Antonio Martin, et al.
Veröffentlicht: (2024)
von: H., Jose Antonio Martin, et al.
Veröffentlicht: (2024)
Offline Reinforcement Learning in Large State Spaces: Algorithms and Guarantees
von: Jiang, Nan, et al.
Veröffentlicht: (2025)
von: Jiang, Nan, et al.
Veröffentlicht: (2025)
Do We Need to Verify Step by Step? Rethinking Process Supervision from a Theoretical Perspective
von: Jia, Zeyu, et al.
Veröffentlicht: (2025)
von: Jia, Zeyu, et al.
Veröffentlicht: (2025)
Learning Game-Playing Agents with Generative Code Optimization
von: Kuang, Zhiyi, et al.
Veröffentlicht: (2025)
von: Kuang, Zhiyi, et al.
Veröffentlicht: (2025)
Understanding the Challenges in Iterative Generative Optimization with LLMs
von: Nie, Allen, et al.
Veröffentlicht: (2026)
von: Nie, Allen, et al.
Veröffentlicht: (2026)
Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM Reasoning
von: Yuan, Yurun, et al.
Veröffentlicht: (2025)
von: Yuan, Yurun, et al.
Veröffentlicht: (2025)
Self-Play with Adversarial Critic: Provable and Scalable Offline Alignment for Language Models
von: Ji, Xiang, et al.
Veröffentlicht: (2024)
von: Ji, Xiang, et al.
Veröffentlicht: (2024)
Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits
von: Chen, Fan, et al.
Veröffentlicht: (2025)
von: Chen, Fan, et al.
Veröffentlicht: (2025)
Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
von: Xie, Tengyang, et al.
Veröffentlicht: (2024)
von: Xie, Tengyang, et al.
Veröffentlicht: (2024)
Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
von: Huang, Audrey, et al.
Veröffentlicht: (2024)
von: Huang, Audrey, et al.
Veröffentlicht: (2024)
Mano: Restriking Manifold Optimization for LLM Training
von: Gu, Yufei, et al.
Veröffentlicht: (2026)
von: Gu, Yufei, et al.
Veröffentlicht: (2026)
Fine-grained Analysis of Stability and Generalization for Stochastic Bilevel Optimization
von: Zhang, Xuelin, et al.
Veröffentlicht: (2026)
von: Zhang, Xuelin, et al.
Veröffentlicht: (2026)
Doubly Stochastic Adaptive Neighbors Clustering via the Marcus Mapping
von: Yuan, Jinghui, et al.
Veröffentlicht: (2024)
von: Yuan, Jinghui, et al.
Veröffentlicht: (2024)
Second-Order Convergence in Private Stochastic Non-Convex Optimization
von: Tao, Youming, et al.
Veröffentlicht: (2025)
von: Tao, Youming, et al.
Veröffentlicht: (2025)
Towards Principled Representation Learning from Videos for Reinforcement Learning
von: Misra, Dipendra, et al.
Veröffentlicht: (2024)
von: Misra, Dipendra, et al.
Veröffentlicht: (2024)
From Stochastic Answers to Verifiable Reasoning: Interpretable Decision-Making with LLM-Generated Code
von: Mahesh, Anirudh Jaidev, et al.
Veröffentlicht: (2026)
von: Mahesh, Anirudh Jaidev, et al.
Veröffentlicht: (2026)
RosettaSearch: Multi-Objective Inference-Time Search for Protein Sequence Design
von: Kshirsagar, Meghana, et al.
Veröffentlicht: (2026)
von: Kshirsagar, Meghana, et al.
Veröffentlicht: (2026)
Uncertainty-Aware Decision Transformer for Stochastic Driving Environments
von: Li, Zenan, et al.
Veröffentlicht: (2023)
von: Li, Zenan, et al.
Veröffentlicht: (2023)
Stochastic Optimization of Inventory at Large-scale Supply Chains
von: Jin, Zhaoyang Larry, et al.
Veröffentlicht: (2025)
von: Jin, Zhaoyang Larry, et al.
Veröffentlicht: (2025)
CounterCurate: Enhancing Physical and Semantic Visio-Linguistic Compositional Reasoning via Counterfactual Examples
von: Zhang, Jianrui, et al.
Veröffentlicht: (2024)
von: Zhang, Jianrui, et al.
Veröffentlicht: (2024)
Paying Less Generalization Tax: A Cross-Domain Generalization Study of RL Training for LLM Agents
von: Liu, Zhihan, et al.
Veröffentlicht: (2026)
von: Liu, Zhihan, et al.
Veröffentlicht: (2026)
Certificate-Guided Pruning for Stochastic Lipschitz Optimization
von: Shihab, Ibne Farabi, et al.
Veröffentlicht: (2026)
von: Shihab, Ibne Farabi, et al.
Veröffentlicht: (2026)
Beyond Sharp Minima: Robust LLM Unlearning via Feedback-Guided Multi-Point Optimization
von: Wu, Wenhan, et al.
Veröffentlicht: (2025)
von: Wu, Wenhan, et al.
Veröffentlicht: (2025)
SE(3)-Stochastic Flow Matching for Protein Backbone Generation
von: Bose, Avishek Joey, et al.
Veröffentlicht: (2023)
von: Bose, Avishek Joey, et al.
Veröffentlicht: (2023)
A Universal Banach--Bregman Framework for Stochastic Iterations: Unifying Stochastic Mirror Descent, Learning and LLM Training
von: Zhang, Johnny R., et al.
Veröffentlicht: (2025)
von: Zhang, Johnny R., et al.
Veröffentlicht: (2025)
OAT-Rephrase: Optimization-Aware Training Data Rephrasing for Zeroth-Order LLM Fine-Tuning
von: Long, Jikai, et al.
Veröffentlicht: (2025)
von: Long, Jikai, et al.
Veröffentlicht: (2025)
Fairness Aware Reward Optimization
von: Choi, Ching Lam, et al.
Veröffentlicht: (2026)
von: Choi, Ching Lam, et al.
Veröffentlicht: (2026)
How Strategic Agents Respond: Comparing Analytical Models with LLM-Generated Responses in Strategic Classification
von: Xie, Tian, et al.
Veröffentlicht: (2025)
von: Xie, Tian, et al.
Veröffentlicht: (2025)
Deterministic Decomposition of Stochastic Generative Dynamics
von: Song, Xingyu, et al.
Veröffentlicht: (2026)
von: Song, Xingyu, et al.
Veröffentlicht: (2026)
Generative Modeling with Phase Stochastic Bridges
von: Chen, Tianrong, et al.
Veröffentlicht: (2023)
von: Chen, Tianrong, et al.
Veröffentlicht: (2023)
A Unified Gaussian Process for Branching and Nested Hyperparameter Optimization
von: Zhang, Jiazhao, et al.
Veröffentlicht: (2024)
von: Zhang, Jiazhao, et al.
Veröffentlicht: (2024)
Unifying Bayesian Flow Networks and Diffusion Models through Stochastic Differential Equations
von: Xue, Kaiwen, et al.
Veröffentlicht: (2024)
von: Xue, Kaiwen, et al.
Veröffentlicht: (2024)
SAGE: Sequence-level Adaptive Gradient Evolution for Generative Recommendation
von: Xie, Yu, et al.
Veröffentlicht: (2026)
von: Xie, Yu, et al.
Veröffentlicht: (2026)
E^2-LLM: Bridging Neural Signals and Interpretable Affective Analysis
von: Ma, Fei, et al.
Veröffentlicht: (2026)
von: Ma, Fei, et al.
Veröffentlicht: (2026)
KernelBand: Steering LLM-based Kernel Optimization via Hardware-Aware Multi-Armed Bandits
von: Ran, Dezhi, et al.
Veröffentlicht: (2025)
von: Ran, Dezhi, et al.
Veröffentlicht: (2025)
Predicting Long Term Sequential Policy Value Using Softer Surrogates
von: Nam, Hyunji, et al.
Veröffentlicht: (2024)
von: Nam, Hyunji, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Trace is the Next AutoDiff: Generative Optimization with Rich Feedback, Execution Traces, and LLMs
von: Cheng, Ching-An, et al.
Veröffentlicht: (2024) -
Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
von: Rosset, Corby, et al.
Veröffentlicht: (2024) -
Reinforce LLM Reasoning through Multi-Agent Reflection
von: Yuan, Yurun, et al.
Veröffentlicht: (2025) -
Breaking the Capability Ceiling of LLM Post-Training by Reintroducing Markov States
von: Yuan, Yurun, et al.
Veröffentlicht: (2026) -
Principal Orthogonal Latent Components Analysis (POLCA Net)
von: H., Jose Antonio Martin, et al.
Veröffentlicht: (2024)