SCOPE: Self-Play via Co-Evolving Policies for Open-Ended Tasks
Fuente:
arXiv
Guardado en:
| Autores principales: | Kwan, Wai-Chung, Gema, Aryo Pradipta, Leang, Joshua Ong Jun, Minervini, Pasquale |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
OpenSIR: Open-Ended Self-Improving Reasoner
por: Kwan, Wai-Chung, et al.
Publicado: (2025)
por: Kwan, Wai-Chung, et al.
Publicado: (2025)
CoMAT: Chain of Mathematically Annotated Thought Improves Mathematical Reasoning
por: Leang, Joshua Ong Jun, et al.
Publicado: (2024)
por: Leang, Joshua Ong Jun, et al.
Publicado: (2024)
Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs
por: Saxena, Rohit, et al.
Publicado: (2025)
por: Saxena, Rohit, et al.
Publicado: (2025)
PiCSAR: Probabilistic Confidence Selection And Ranking for Reasoning Chains
por: Leang, Joshua Ong Jun, et al.
Publicado: (2025)
por: Leang, Joshua Ong Jun, et al.
Publicado: (2025)
Edinburgh Clinical NLP at SemEval-2024 Task 2: Fine-tune your model unless you have access to GPT-4
por: Gema, Aryo Pradipta, et al.
Publicado: (2024)
por: Gema, Aryo Pradipta, et al.
Publicado: (2024)
Parameter-Efficient Fine-Tuning of LLaMA for the Clinical Domain
por: Gema, Aryo Pradipta, et al.
Publicado: (2023)
por: Gema, Aryo Pradipta, et al.
Publicado: (2023)
Self-Training Large Language Models for Tool-Use Without Demonstrations
por: Luo, Ne, et al.
Publicado: (2025)
por: Luo, Ne, et al.
Publicado: (2025)
Edinburgh Clinical NLP at MEDIQA-CORR 2024: Guiding Large Language Models with Hints
por: Gema, Aryo Pradipta, et al.
Publicado: (2024)
por: Gema, Aryo Pradipta, et al.
Publicado: (2024)
DeCoRe: Decoding by Contrasting Retrieval Heads to Mitigate Hallucinations
por: Gema, Aryo Pradipta, et al.
Publicado: (2024)
por: Gema, Aryo Pradipta, et al.
Publicado: (2024)
Noiser: Bounded Input Perturbations for Attributing Large Language Models
por: Madani, Mohammad Reza Ghasemi, et al.
Publicado: (2025)
por: Madani, Mohammad Reza Ghasemi, et al.
Publicado: (2025)
Do Composed Image Retrieval Benchmarks Require Multimodal Composition?
por: Attimonelli, Matteo, et al.
Publicado: (2026)
por: Attimonelli, Matteo, et al.
Publicado: (2026)
Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering
por: Zhao, Yu, et al.
Publicado: (2024)
por: Zhao, Yu, et al.
Publicado: (2024)
An Analysis of Decoding Methods for LLM-based Agents for Faithful Multi-Hop Question Answering
por: Murphy, Alexander, et al.
Publicado: (2025)
por: Murphy, Alexander, et al.
Publicado: (2025)
The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models
por: Hong, Giwon, et al.
Publicado: (2024)
por: Hong, Giwon, et al.
Publicado: (2024)
Analysing the Residual Stream of Language Models Under Knowledge Conflicts
por: Zhao, Yu, et al.
Publicado: (2024)
por: Zhao, Yu, et al.
Publicado: (2024)
Are We Done with MMLU?
por: Gema, Aryo Pradipta, et al.
Publicado: (2024)
por: Gema, Aryo Pradipta, et al.
Publicado: (2024)
Analyzing LLM Instruction Optimization for Tabular Fact Verification
por: Du, Xiaotang, et al.
Publicado: (2026)
por: Du, Xiaotang, et al.
Publicado: (2026)
Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them
por: Rajani, Neel, et al.
Publicado: (2025)
por: Rajani, Neel, et al.
Publicado: (2025)
Same Answer, Different Representations: Hidden instability in VLMs
por: Wani, Farooq Ahmad, et al.
Publicado: (2026)
por: Wani, Farooq Ahmad, et al.
Publicado: (2026)
Answerability in Retrieval-Augmented Open-Domain Question Answering
por: Abdumalikov, Rustam, et al.
Publicado: (2024)
por: Abdumalikov, Rustam, et al.
Publicado: (2024)
EvoRubric: Self-Evolving Rubric-Driven RL for Open-Ended Generation
por: Guan, Xin, et al.
Publicado: (2026)
por: Guan, Xin, et al.
Publicado: (2026)
AEL: Agent Evolving Learning for Open-Ended Environments
por: Xu, Wujiang, et al.
Publicado: (2026)
por: Xu, Wujiang, et al.
Publicado: (2026)
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
por: Chen, Jiaqi, et al.
Publicado: (2025)
por: Chen, Jiaqi, et al.
Publicado: (2025)
CausalEvolve: Towards Open-Ended Discovery with Causal Scratchpad
por: Chen, Yongqiang, et al.
Publicado: (2026)
por: Chen, Yongqiang, et al.
Publicado: (2026)
Inverse Scaling in Test-Time Compute
por: Gema, Aryo Pradipta, et al.
Publicado: (2025)
por: Gema, Aryo Pradipta, et al.
Publicado: (2025)
G-Zero: Self-Play for Open-Ended Generation from Zero Data
por: Huang, Chengsong, et al.
Publicado: (2026)
por: Huang, Chengsong, et al.
Publicado: (2026)
Bagpiper: Solving Open-Ended Audio Tasks via Rich Captions
por: Tian, Jinchuan, et al.
Publicado: (2026)
por: Tian, Jinchuan, et al.
Publicado: (2026)
ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program Evolution
por: Lange, Robert Tjarko, et al.
Publicado: (2025)
por: Lange, Robert Tjarko, et al.
Publicado: (2025)
Can GPT-3.5 Generate and Code Discharge Summaries?
por: Falis, Matúš, et al.
Publicado: (2024)
por: Falis, Matúš, et al.
Publicado: (2024)
CheckEmbed: Effective Verification of LLM Solutions to Open-Ended Tasks
por: Besta, Maciej, et al.
Publicado: (2024)
por: Besta, Maciej, et al.
Publicado: (2024)
Enhancing Long Document Long Form Summarisation with Self-Planning
por: Du, Xiaotang, et al.
Publicado: (2025)
por: Du, Xiaotang, et al.
Publicado: (2025)
A Comparative Study on Patient Language across Therapeutic Domains for Effective Patient Voice Classification in Online Health Discussions
por: Lysandrou, Giorgos, et al.
Publicado: (2024)
por: Lysandrou, Giorgos, et al.
Publicado: (2024)
GRADA: Graph-based Reranking against Adversarial Documents Attack
por: Zheng, Jingjie, et al.
Publicado: (2025)
por: Zheng, Jingjie, et al.
Publicado: (2025)
REAL Sampling: Boosting Factuality and Diversity of Open-Ended Generation via Asymptotic Entropy
por: Chang, Haw-Shiuan, et al.
Publicado: (2024)
por: Chang, Haw-Shiuan, et al.
Publicado: (2024)
PosterSum: A Multimodal Benchmark for Scientific Poster Summarization
por: Saxena, Rohit, et al.
Publicado: (2025)
por: Saxena, Rohit, et al.
Publicado: (2025)
Improving Open-Ended Text Generation via Adaptive Decoding
por: Zhu, Wenhong, et al.
Publicado: (2024)
por: Zhu, Wenhong, et al.
Publicado: (2024)
OpenEP: Open-Ended Future Event Prediction
por: Guan, Yong, et al.
Publicado: (2024)
por: Guan, Yong, et al.
Publicado: (2024)
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
por: Xie, Tianbao, et al.
Publicado: (2024)
por: Xie, Tianbao, et al.
Publicado: (2024)
SPELL: Self-Play Reinforcement Learning for Evolving Long-Context Language Models
por: Yang, Ziyi, et al.
Publicado: (2025)
por: Yang, Ziyi, et al.
Publicado: (2025)
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training
por: Wang, Pengkai, et al.
Publicado: (2025)
por: Wang, Pengkai, et al.
Publicado: (2025)
Ejemplares similares
-
OpenSIR: Open-Ended Self-Improving Reasoner
por: Kwan, Wai-Chung, et al.
Publicado: (2025) -
CoMAT: Chain of Mathematically Annotated Thought Improves Mathematical Reasoning
por: Leang, Joshua Ong Jun, et al.
Publicado: (2024) -
Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs
por: Saxena, Rohit, et al.
Publicado: (2025) -
PiCSAR: Probabilistic Confidence Selection And Ranking for Reasoning Chains
por: Leang, Joshua Ong Jun, et al.
Publicado: (2025) -
Edinburgh Clinical NLP at SemEval-2024 Task 2: Fine-tune your model unless you have access to GPT-4
por: Gema, Aryo Pradipta, et al.
Publicado: (2024)