OptPO: Optimal Rollout Allocation for Test-time Policy Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Youkang, Wang, Jian, Chen, Rubing, Zeng, Tianyi, Wei, Xiao-Yong, Li, Qing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Pay Attention to What You Need
von: Gao, Yifei, et al.
Veröffentlicht: (2023)
von: Gao, Yifei, et al.
Veröffentlicht: (2023)
Self-Attention as Transport: Limits of Symmetric Spectral Diagnostics
von: Dahlem, Dominik, et al.
Veröffentlicht: (2026)
von: Dahlem, Dominik, et al.
Veröffentlicht: (2026)
Unraveling Media Perspectives: A Comprehensive Methodology Combining Large Language Models, Topic Modeling, Sentiment Analysis, and Ontology Learning to Analyse Media Bias
von: Jähde, Orlando, et al.
Veröffentlicht: (2025)
von: Jähde, Orlando, et al.
Veröffentlicht: (2025)
OPENXRD: A Comprehensive Benchmark Framework for LLM/MLLM XRD Question Answering
von: Vosoughi, Ali, et al.
Veröffentlicht: (2025)
von: Vosoughi, Ali, et al.
Veröffentlicht: (2025)
Bridging the Language Gap: Enhancing Multilingual Prompt-Based Code Generation in LLMs via Zero-Shot Cross-Lingual Transfer
von: Li, Mingda, et al.
Veröffentlicht: (2024)
von: Li, Mingda, et al.
Veröffentlicht: (2024)
Uncertainty Estimation and Quantification for LLMs: A Simple Supervised Approach
von: Liu, Linyu, et al.
Veröffentlicht: (2024)
von: Liu, Linyu, et al.
Veröffentlicht: (2024)
Mechanistic Analysis of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
von: Imanov, Olaf Yunus Laitinen
Veröffentlicht: (2026)
von: Imanov, Olaf Yunus Laitinen
Veröffentlicht: (2026)
Latent Object Permanence: Topological Phase Transitions, Free-Energy Principles, and Renormalization Group Flows in Deep Transformer Manifolds
von: Alpay, Faruk, et al.
Veröffentlicht: (2026)
von: Alpay, Faruk, et al.
Veröffentlicht: (2026)
Inference acceleration for large language models using "stairs" assisted greedy generation
von: Grigaliūnas, Domas, et al.
Veröffentlicht: (2024)
von: Grigaliūnas, Domas, et al.
Veröffentlicht: (2024)
Strategic Doctrine Language Models (sdLM): A Learning-System Framework for Doctrinal Consistency and Geopolitical Forecasting
von: Imanov, Olaf Yunus Laitinen, et al.
Veröffentlicht: (2026)
von: Imanov, Olaf Yunus Laitinen, et al.
Veröffentlicht: (2026)
Forging GEMs: Advancing Greek NLP through Quality-Based Corpus Curation
von: Apostolopoulou, Alexandra, et al.
Veröffentlicht: (2025)
von: Apostolopoulou, Alexandra, et al.
Veröffentlicht: (2025)
Can Out-of-Distribution Evaluations Uncover Reliance on Shortcuts? A Case Study in Question Answering
von: Štefánik, Michal, et al.
Veröffentlicht: (2025)
von: Štefánik, Michal, et al.
Veröffentlicht: (2025)
Harnessing non-adversarial robustness in large language models
von: Zhou, Qinghua, et al.
Veröffentlicht: (2026)
von: Zhou, Qinghua, et al.
Veröffentlicht: (2026)
mHC-SSM: Manifold-Constrained Hyper-Connections for State Space Language Models with Stream-Specialized Adapters
von: Mutlu, Abdulvahap, et al.
Veröffentlicht: (2026)
von: Mutlu, Abdulvahap, et al.
Veröffentlicht: (2026)
Extracting Sentence Embeddings from Pretrained Transformer Models
von: Stankevičius, Lukas, et al.
Veröffentlicht: (2024)
von: Stankevičius, Lukas, et al.
Veröffentlicht: (2024)
Sentiment Analysis of Lithuanian Online Reviews Using Large Language Models
von: Vileikytė, Brigita, et al.
Veröffentlicht: (2024)
von: Vileikytė, Brigita, et al.
Veröffentlicht: (2024)
ReFactor GNNs: Revisiting Factorisation-based Models from a Message-Passing Perspective
von: Chen, Yihong, et al.
Veröffentlicht: (2022)
von: Chen, Yihong, et al.
Veröffentlicht: (2022)
A Language Model-Driven Semi-Supervised Ensemble Framework for Illicit Market Detection Across Deep/Dark Web and Social Platforms
von: Yazdanjue, Navid, et al.
Veröffentlicht: (2025)
von: Yazdanjue, Navid, et al.
Veröffentlicht: (2025)
Clustering in pure-attention hardmax transformers and its role in sentiment analysis
von: Alcalde, Albert, et al.
Veröffentlicht: (2024)
von: Alcalde, Albert, et al.
Veröffentlicht: (2024)
Cache-to-Cache: Direct Semantic Communication Between Large Language Models
von: Fu, Tianyu, et al.
Veröffentlicht: (2025)
von: Fu, Tianyu, et al.
Veröffentlicht: (2025)
HInter: Exposing Hidden Intersectional Bias in Large Language Models
von: Souani, Badr, et al.
Veröffentlicht: (2025)
von: Souani, Badr, et al.
Veröffentlicht: (2025)
LLMs as Deceptive Agents: How Role-Based Prompting Induces Semantic Ambiguity in Puzzle Tasks
von: Yoo, Seunghyun
Veröffentlicht: (2025)
von: Yoo, Seunghyun
Veröffentlicht: (2025)
How Well Do Vision-Language Models Understand Sequential Driving Scenes? A Sensitivity Study
von: Brusnicki, Roberto, et al.
Veröffentlicht: (2026)
von: Brusnicki, Roberto, et al.
Veröffentlicht: (2026)
ExpliCa: Evaluating Explicit Causal Reasoning in Large Language Models
von: Miliani, Martina, et al.
Veröffentlicht: (2025)
von: Miliani, Martina, et al.
Veröffentlicht: (2025)
Context Aware Lemmatization and Morphological Tagging Method in Turkish
von: Sayallar, Cagri
Veröffentlicht: (2025)
von: Sayallar, Cagri
Veröffentlicht: (2025)
MetaCheckGPT -- A Multi-task Hallucination Detector Using LLM Uncertainty and Meta-models
von: Mehta, Rahul, et al.
Veröffentlicht: (2024)
von: Mehta, Rahul, et al.
Veröffentlicht: (2024)
Prompting Encoder Models for Zero-Shot Classification: A Cross-Domain Study in Italian
von: Auriemma, Serena, et al.
Veröffentlicht: (2024)
von: Auriemma, Serena, et al.
Veröffentlicht: (2024)
A Generalization Bound for a Family of Implicit Networks
von: Fung, Samy Wu, et al.
Veröffentlicht: (2024)
von: Fung, Samy Wu, et al.
Veröffentlicht: (2024)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
von: Pather, Kaviraj, et al.
Veröffentlicht: (2025)
von: Pather, Kaviraj, et al.
Veröffentlicht: (2025)
HEFT: A Coarse-to-Fine Hierarchy for Enhancing the Efficiency and Accuracy of Language Model Reasoning
von: Hill, Brennen
Veröffentlicht: (2025)
von: Hill, Brennen
Veröffentlicht: (2025)
Differentiable Neural Networks with RePU Activation: with Applications to Score Estimation and Isotonic Regression
von: Shen, Guohao, et al.
Veröffentlicht: (2023)
von: Shen, Guohao, et al.
Veröffentlicht: (2023)
Retrieval-augmented code completion for local projects using large language models
von: Hostnik, Marko, et al.
Veröffentlicht: (2024)
von: Hostnik, Marko, et al.
Veröffentlicht: (2024)
Exact Sequence Interpolation with Transformers
von: Alcalde, Albert, et al.
Veröffentlicht: (2025)
von: Alcalde, Albert, et al.
Veröffentlicht: (2025)
Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models
von: Fu, Tianyu, et al.
Veröffentlicht: (2025)
von: Fu, Tianyu, et al.
Veröffentlicht: (2025)
A Survey on Collaborating Small and Large Language Models for Performance, Cost-effectiveness, Cloud-edge Privacy, and Trustworthiness
von: Wang, Fali, et al.
Veröffentlicht: (2025)
von: Wang, Fali, et al.
Veröffentlicht: (2025)
Who's Asking? Investigating Bias Through the Lens of Disability Framed Queries in LLMs
von: Hari, Vishnu, et al.
Veröffentlicht: (2025)
von: Hari, Vishnu, et al.
Veröffentlicht: (2025)
Z-Order Transformer for Feed-Forward Gaussian Splatting
von: Wang, Can, et al.
Veröffentlicht: (2026)
von: Wang, Can, et al.
Veröffentlicht: (2026)
Parameter-Efficient Transformer Embeddings
von: Ndubuaku, Henry, et al.
Veröffentlicht: (2025)
von: Ndubuaku, Henry, et al.
Veröffentlicht: (2025)
NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution
von: Breneur, Oleksandr Marchenko, et al.
Veröffentlicht: (2026)
von: Breneur, Oleksandr Marchenko, et al.
Veröffentlicht: (2026)
DYNAMAX: Dynamic computing for Transformers and Mamba based architectures
von: Nogales, Miguel, et al.
Veröffentlicht: (2025)
von: Nogales, Miguel, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Pay Attention to What You Need
von: Gao, Yifei, et al.
Veröffentlicht: (2023) -
Self-Attention as Transport: Limits of Symmetric Spectral Diagnostics
von: Dahlem, Dominik, et al.
Veröffentlicht: (2026) -
Unraveling Media Perspectives: A Comprehensive Methodology Combining Large Language Models, Topic Modeling, Sentiment Analysis, and Ontology Learning to Analyse Media Bias
von: Jähde, Orlando, et al.
Veröffentlicht: (2025) -
OPENXRD: A Comprehensive Benchmark Framework for LLM/MLLM XRD Question Answering
von: Vosoughi, Ali, et al.
Veröffentlicht: (2025) -
Bridging the Language Gap: Enhancing Multilingual Prompt-Based Code Generation in LLMs via Zero-Shot Cross-Lingual Transfer
von: Li, Mingda, et al.
Veröffentlicht: (2024)