Learning to Play 7 Wonders Duel Without Human Supervision
Fuente:
arXiv
Saved in:
| Main Authors: | Paolini, Giovanni, Moreschini, Lorenzo, Veneziano, Francesco, Iraci, Alessandro |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning To Play Atari Games Using Dueling Q-Learning and Hebbian Plasticity
by: Salehin, Md Ashfaq
Published: (2024)
by: Salehin, Md Ashfaq
Published: (2024)
Neural Dueling Bandits: Preference-Based Optimization with Human Feedback
by: Verma, Arun, et al.
Published: (2024)
by: Verma, Arun, et al.
Published: (2024)
Online Clustering of Dueling Bandits
by: Wang, Zhiyong, et al.
Published: (2025)
by: Wang, Zhiyong, et al.
Published: (2025)
Fusing Reward and Dueling Feedback in Stochastic Bandits
by: Wang, Xuchuang, et al.
Published: (2025)
by: Wang, Xuchuang, et al.
Published: (2025)
Linear and Neural Dueling Bandits with Delayed Feedback
by: Wang, Xiangyi, et al.
Published: (2026)
by: Wang, Xiangyi, et al.
Published: (2026)
Mapping Uncharted Symmetries: Machine Discovery in Combinatorics
by: Cainelli, Eugenio, et al.
Published: (2026)
by: Cainelli, Eugenio, et al.
Published: (2026)
Learning Tennis Strategy Through Curriculum-Based Dueling Double Deep Q-Networks
by: Mohan, Vishnu
Published: (2025)
by: Mohan, Vishnu
Published: (2025)
PFML: Self-Supervised Learning of Time-Series Data Without Representation Collapse
by: Vaaras, Einari, et al.
Published: (2024)
by: Vaaras, Einari, et al.
Published: (2024)
Expected Possession Value of Control and Duel Actions for Soccer Player's Skills Estimation
by: Shelopugin, Andrei
Published: (2024)
by: Shelopugin, Andrei
Published: (2024)
A Controlled Study of Double DQN and Dueling DQN Under Cross-Environment Transfer
by: Nasir, Azkaa, et al.
Published: (2026)
by: Nasir, Azkaa, et al.
Published: (2026)
Wonderful Matrices: More Efficient and Effective Architecture for Language Modeling Tasks
by: Shi, Jingze, et al.
Published: (2024)
by: Shi, Jingze, et al.
Published: (2024)
Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents
by: Xia, Fanzeng, et al.
Published: (2024)
by: Xia, Fanzeng, et al.
Published: (2024)
How to Square Tensor Networks and Circuits Without Squaring Them
by: Loconte, Lorenzo, et al.
Published: (2025)
by: Loconte, Lorenzo, et al.
Published: (2025)
Towards a Learning Theory of Representation Alignment
by: Insulla, Francesco, et al.
Published: (2025)
by: Insulla, Francesco, et al.
Published: (2025)
Self-Harmony: Learning to Harmonize Self-Supervision and Self-Play in Test-Time Reinforcement Learning
by: Wang, Ru, et al.
Published: (2025)
by: Wang, Ru, et al.
Published: (2025)
From Demonstrations to Rewards: Alignment Without Explicit Human Preferences
by: Zeng, Siliang, et al.
Published: (2025)
by: Zeng, Siliang, et al.
Published: (2025)
Discovering Latent Knowledge in Language Models Without Supervision
by: Burns, Collin, et al.
Published: (2022)
by: Burns, Collin, et al.
Published: (2022)
XQSV: A Structurally Variable Network to Imitate Human Play in Xiangqi
by: Zhou, Chenliang
Published: (2024)
by: Zhou, Chenliang
Published: (2024)
Training Plug-n-Play Knowledge Modules with Deep Context Distillation
by: Caccia, Lucas, et al.
Published: (2025)
by: Caccia, Lucas, et al.
Published: (2025)
Learning to Play Blackjack: A Curriculum Learning Perspective
by: Alasti, Amirreza, et al.
Published: (2026)
by: Alasti, Amirreza, et al.
Published: (2026)
TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment
by: Tan, Zhewen, et al.
Published: (2026)
by: Tan, Zhewen, et al.
Published: (2026)
Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences
by: Karlekar, Sweta, et al.
Published: (2026)
by: Karlekar, Sweta, et al.
Published: (2026)
HINTS: Extraction of Human Insights from Time-Series Without External Sources
by: Jhin, Sheo Yon, et al.
Published: (2025)
by: Jhin, Sheo Yon, et al.
Published: (2025)
Semi-Supervised Graph Representation Learning with Human-centric Explanation for Predicting Fatty Liver Disease
by: Kim, So Yeon, et al.
Published: (2024)
by: Kim, So Yeon, et al.
Published: (2024)
LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
by: Balestriero, Randall, et al.
Published: (2025)
by: Balestriero, Randall, et al.
Published: (2025)
Learning Game-Playing Agents with Generative Code Optimization
by: Kuang, Zhiyi, et al.
Published: (2025)
by: Kuang, Zhiyi, et al.
Published: (2025)
Aligning Human and Machine Attention for Enhanced Supervised Learning
by: Chriqui, Avihay, et al.
Published: (2025)
by: Chriqui, Avihay, et al.
Published: (2025)
Auto-ICL: In-Context Learning without Human Supervision
by: Yang, Jinghan, et al.
Published: (2023)
by: Yang, Jinghan, et al.
Published: (2023)
The Importance of Being Lazy: Scaling Limits of Continual Learning
by: Graldi, Jacopo, et al.
Published: (2025)
by: Graldi, Jacopo, et al.
Published: (2025)
RLZero: Direct Policy Inference from Language Without In-Domain Supervision
by: Sikchi, Harshit, et al.
Published: (2024)
by: Sikchi, Harshit, et al.
Published: (2024)
Wonderful Matrices: Combining for a More Efficient and Effective Foundation Model Architecture
by: Shi, Jingze, et al.
Published: (2024)
by: Shi, Jingze, et al.
Published: (2024)
Learning to Play Piano in the Real World
by: Zeulner, Yves-Simon, et al.
Published: (2025)
by: Zeulner, Yves-Simon, et al.
Published: (2025)
Learning Quantifiable Visual Explanations Without Ground-Truth
by: Singh, Amritpal, et al.
Published: (2026)
by: Singh, Amritpal, et al.
Published: (2026)
Power Plays: Unleashing Machine Learning Magic in Smart Grids
by: Rashid, Abdur, et al.
Published: (2024)
by: Rashid, Abdur, et al.
Published: (2024)
Cogito, Ergo Ludo: An Agent that Learns to Play by Reasoning and Planning
by: Wang, Sai, et al.
Published: (2025)
by: Wang, Sai, et al.
Published: (2025)
Learning Robust Reasoning through Guided Adversarial Self-Play
by: Li, Shuozhe, et al.
Published: (2026)
by: Li, Shuozhe, et al.
Published: (2026)
On the Universality of Self-Supervised Learning
by: Qiang, Wenwen, et al.
Published: (2024)
by: Qiang, Wenwen, et al.
Published: (2024)
Fewer Truncations Improve Language Modeling
by: Ding, Hantian, et al.
Published: (2024)
by: Ding, Hantian, et al.
Published: (2024)
There are no Champions in Supervised Long-Term Time Series Forecasting
by: Brigato, Lorenzo, et al.
Published: (2025)
by: Brigato, Lorenzo, et al.
Published: (2025)
Evidence of Learned Look-Ahead in a Chess-Playing Neural Network
by: Jenner, Erik, et al.
Published: (2024)
by: Jenner, Erik, et al.
Published: (2024)
Similar Items
-
Learning To Play Atari Games Using Dueling Q-Learning and Hebbian Plasticity
by: Salehin, Md Ashfaq
Published: (2024) -
Neural Dueling Bandits: Preference-Based Optimization with Human Feedback
by: Verma, Arun, et al.
Published: (2024) -
Online Clustering of Dueling Bandits
by: Wang, Zhiyong, et al.
Published: (2025) -
Fusing Reward and Dueling Feedback in Stochastic Bandits
by: Wang, Xuchuang, et al.
Published: (2025) -
Linear and Neural Dueling Bandits with Delayed Feedback
by: Wang, Xiangyi, et al.
Published: (2026)