> For the complete documentation index, see [llms.txt](https://kilmya1.gitbook.io/deep-multi-agent-reinforcement-learning/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://kilmya1.gitbook.io/deep-multi-agent-reinforcement-learning/iii-learning-to-reciprocate/9-dice-the-infinitely-differentiable-monte-carlo-estimator/9.4-correct-gradient-estimators-with-dice/9.4.1-implement-of-dice.md).

# 9.4.1 Implement of DiCE

&#x20;DiCE의 처음부터 실용적인 논문임을 강조했는데, 이는 단순한 구현방법에 있습니다. MagicBox는 두가지 특성을 만족하면 됐었는데, 이는 다음과 같이 정의함으로써 두가지 성질을 다 가져갈 수 있습니다. 아래선 위를 확인하도록 합니다.

&#x20;                                                                $$\square (\mathcal{W}) = \exp(\tau - \bot(\tau))$$

&#x20;                                                                $$\tau = \sum\_{w \in \mathcal{W}} \log(p(w;\theta))$$

$$\bot$$은 $$\nabla\_x \bot(x) = 0$$이 되도록 하는 gradient를 안흐르도록 하는 operator로 pytorch의 detach같은 역할을 합니다. $$\bot(x) \rightarrow x$$이므로, $$\square (\mathcal{W}) \rightarrow 1$$임이 자명합니다. 이로써 첫번째 성질이 증명되었습니다. 바로 두번째 성질을 증명하면 다음과 같습니다.

&#x20;                                                        $$\nabla\_\theta \square (\mathcal{W}) = \nabla\_\theta\exp(\tau-\bot(\tau))$$

&#x20;                                                                            $$= \exp(\tau-\bot(\tau))\nabla\_\theta(\tau-\bot(\tau))$$

&#x20;                                                                            $$=\square(\mathcal{W})(\nabla\_\theta\tau-0)$$

&#x20;                                                                            $$=\square(\mathcal{W})\sum\_{w \in \mathcal{W}}\nabla\_\theta \log(p(w;\theta))$$

&#x20;그리고 magicbox operator를 구현하게되면, 주로 objective와 바로 연관지어 구현하는게 가장 간단한데, 일반적인 RL에선 $$J  = \mathbb{E}\[\sum r\_t]$$로 나타낼 때, DiCE의 objective는 $$J\_\square= \sum\_t \square({a\_{t'},t'\leq t})r\_t$$로 나타내야합니다. (이는 이전의 action에 따라 reward에 stochastic하게 영향을 주기 때문입니다.
