This paper aims to reduce GPU memory usage during DNN training. Capuchin achieves this goal though swapping and recomputation , using tensor as unit of operation. The major question is how to balance …
This paper aims to reduce GPU memory usage during DNN training. Capuchin achieves this goal though swapping and recomputation , using tensor as unit of operation. The major question is how to balance …
讨论
登录后参与讨论
还没有评论,来说第一句吧。