跳到主要內容

Distributed tensorflow

ref:
https://www.youtube.com/watch?v=bRMGoPqsn20




#multi GPUs
#multi machines



#multi GPUs:
or say, 1 machine with multi devices






Mirrored strategy


#multi machines
use the Estimators' train_and_evaluate API.
async Parameter Server approach

https://www.tensorflow.org/api_docs/python/tfestimator/train_and_evaluate






留言

這個網誌中的熱門文章

DDPG in Torcs within Docker Container

Dockerfile: FROM tensorflow/tensorflow:0.10.0-gpu WORKDIR /home/frank/old_gym_torcs ADD . /home/old_gym_torcs  RUN apt update RUN apt install -y vim xautomation torcs RUN apt-get install -y libjpeg-dev cmake swig python-pyglet python3-opengl libboost-all-dev \         libsdl2-2.0.0 libsdl2-dev libglu1-mesa libglu1-mesa-dev libgles2-mesa-dev \         freeglut3 xvfb libav-tools RUN pip install gym RUN pip install keras==1.1.0 ENV PATH="/usr/games:${PATH}" CMD ["/bin/bash"] viper1 $ docker run --runtime=nvidia -it -e DISPLAY=$DISPLAY -v /tmp/.X11-unix:/tmp/.X11-unix -v /home/frank/old_gym_torcs:/home/old_gym_torcs -v /var/run/docker.sock:/var/run/docker.sock -v /home/frank/gym_torcs:/home/gym_torcs -v /home/frank/gym:/home/gym -p 3101:3101 --workdir /home/old_gym_torcs -p 8888:8888 ddpgfrk:tf0.10 /bin/bash (grey part is not necessary) After  $ docker commit id ddpg:tf0.10.0-gpu viper1 $ docker r...

增強式學習

   迴力球遊戲-ATARI     賽車遊戲DQN-ATARI 賽車遊戲-TORCS Ref:     李宏毅老師 YOUTUBE DRL 1-3 On-policy VS Off-policy On-policy     The agent learned and the agent interacting with the environment is the same     阿光自已下棋學習 Off-policy     The agent learned and the agent interacting with the environment is different     佐助下棋,阿光在旁邊看 Add a baseline:     It is possible that R is always positive     So R subtract a expectation value Policy in " Policy Gradient" means output action, like left/right/fire gamma-discounted rewards: 時間愈遠的貢獻,降低其權重 Reward Function & Action is defined in prior to training MC v.s. TD MC 蒙弟卡羅: critic after episode end : larger variance(cuz conditions differ a lot in every episode), unbiased (judge until episode end, more fair) TD: Temporal-difference approach: critic during episode :smaller variance, biased maybe atari : a3c  ...

Loss function

why cross entroy is more suitable for categorical classcification? ref: https://www.youtube.com/watch?v=Li5sVEXTIJw