Identity Concealment Games: How I Learned to Stop Revealing and Love the Coincidences

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Karabag, Mustafa O., Ornik, Melkior, Topcu, Ufuk
Format: Preprint
Published: 2021
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914700235636736
author Karabag, Mustafa O.
Ornik, Melkior
Topcu, Ufuk
author_facet Karabag, Mustafa O.
Ornik, Melkior
Topcu, Ufuk
contents In an adversarial environment, a hostile player performing a task may behave like a non-hostile one in order not to reveal its identity to an opponent. To model such a scenario, we define identity concealment games: zero-sum stochastic reachability games with a zero-sum objective of identity concealment. To measure the identity concealment of the player, we introduce the notion of an average player. The average player's policy represents the expected behavior of a non-hostile player. We show that there exists an equilibrium policy pair for every identity concealment game and give the optimality equations to synthesize an equilibrium policy pair. If the player's opponent follows a non-equilibrium policy, the player can hide its identity better. For this reason, we study how the hostile player may learn the opponent's policy. Since learning via exploration policies would quickly reveal the hostile player's identity to the opponent, we consider the problem of learning a near-optimal policy for the hostile player using the game runs collected under the average player's policy. Consequently, we propose an algorithm that provably learns a near-optimal policy and give an upper bound on the number of sample runs to be collected.
format Preprint
id arxiv_https___arxiv_org_abs_2105_05377
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle Identity Concealment Games: How I Learned to Stop Revealing and Love the Coincidences
Karabag, Mustafa O.
Ornik, Melkior
Topcu, Ufuk
Computer Science and Game Theory
Multiagent Systems
Optimization and Control
In an adversarial environment, a hostile player performing a task may behave like a non-hostile one in order not to reveal its identity to an opponent. To model such a scenario, we define identity concealment games: zero-sum stochastic reachability games with a zero-sum objective of identity concealment. To measure the identity concealment of the player, we introduce the notion of an average player. The average player's policy represents the expected behavior of a non-hostile player. We show that there exists an equilibrium policy pair for every identity concealment game and give the optimality equations to synthesize an equilibrium policy pair. If the player's opponent follows a non-equilibrium policy, the player can hide its identity better. For this reason, we study how the hostile player may learn the opponent's policy. Since learning via exploration policies would quickly reveal the hostile player's identity to the opponent, we consider the problem of learning a near-optimal policy for the hostile player using the game runs collected under the average player's policy. Consequently, we propose an algorithm that provably learns a near-optimal policy and give an upper bound on the number of sample runs to be collected.
title Identity Concealment Games: How I Learned to Stop Revealing and Love the Coincidences
topic Computer Science and Game Theory
Multiagent Systems
Optimization and Control
url https://arxiv.org/abs/2105.05377