In give_way, I get rewards around 17 after training, while Figure 5 in the paper shows episode rewards mean near 1000. Is this due to different scaling, or different version of vmas (i'm using vmas 1.2.12)?
In give_way, I get rewards around 17 after training, while Figure 5 in the paper shows episode rewards mean near 1000.
Is this due to different scaling, or different version of vmas (i'm using vmas 1.2.12)?