polish(pu): polish the nstep_return_ngu and null_padding action in NGU #116

puyuan1996 · 2021-11-01T14:09:01Z

Description

polish the nstep_return_ngu and null_padding action, and change the name of reward model: rnd to rnd-ngu, to avoid the confusion between the rnd reward model in NGU and the conventional RND paper.

Related Issue

TODO

Check List

merge the latest version source branch/repo, and resolve all the conflicts
pass style check
pass all the tests

… usage in setup.py

separate doc from main repo to doc repo

* test(nyz): comment subprocess env manager and parallel entry unittest * fix(nyz): try to fix test_ppg and flask_fs_collector test close * test(nyz): modify unittest worker number * test(nyz) fix unittest worker and ignore 1v1 collector test * test(nyz): test different range for unittest(env, rl_utils, entry, interaction) * test(nyz): test different range for unittest(env, rl_utils, entry, interaction, league, model) and add execution timeout * test(nyz): test different range for unittest(env, rl_utils, entry, interaction, league, model) * test(nyz): test different range for unittest(env, rl_utils, entry, interaction, league, model, torch_utils) * test(nyz): fix test DingEnvWrapper unittest env bug * test(nyz): add utils unittest and disable dataloader unittest * test(nyz): simplify reward model unittest * test(nyz): enable all the unittest except dataloader * test(nyz): enable parallel entry and dataloader unittest * test(nyz): fix test ppg rerun bug * test(nyz): enable windows test * test(nyz): disable subprocess env manager unittest * test(nyz): fix test auto checkpoint bug * test(nyz): disable test dataloader * test(nyz): enable subprocess env manager unittest * test(nyz): update coveragerc * test(nyz): add coverage upload workflow * test(nyz): disable test_block in subprocess env manager * test(nyz): enable rerun in test demo buffer

…md && update code coverage badge (opendilab#8)

* refactor(nyz): refactor read_config to 3 different function interface * feature(nyz): enable env_setting param in entry * polish(nyz): remove redundant code and global declaration * polish(nyz): remove flag in import_helper * polish(nyz): remove unused import * style(nyz): correct format

…pole (opendilab#114) * added gail entry * added lunarlander and cartpole config * added gail mujoco config * added mujoco exp * update22-10 * added third exp * added metric to evaluate policies * added GAIL entry and config for Cartpole and Walker2d * checked style and unittest * restored lunarlander env * style problems * bug correction * Delete expert_data_train.pkl * changed loss of GAIL * Update walker2d_ddpg_gail_config.py * changed gail reward from -D(s, a) to -log(D(s, a)) * added small constant to reward function * added comment to clarify config * Update walker2d_ddpg_gail_config.py * added lunarlander entry + config * Added Atari discriminator + Pong entry config * Update gail_irl_model.py * Update gail_irl_model.py * added gail serial pipeline and onehot actions for gail atari * related to previous commit * removed main files * removed old comment

…v-polish-ngu Conflicts: dizoo/gym_hybrid/config/gym_hybrid_ddpg_config.py

into dev-polish-ngu

…v-polish-ngu

into dev-polish-ngu

…v-polish-ngu

into dev-polish-ngu

codecov · 2022-03-21T03:36:08Z

Codecov Report

Merging #116 (1ca8b10) into main (bb356ed) will decrease coverage by 0.02%.
The diff coverage is 67.85%.

@@            Coverage Diff             @@
##             main     #116      +/-   ##
==========================================
- Coverage   85.65%   85.62%   -0.03%     
==========================================
  Files         466      466              
  Lines       36019    35958      -61     
==========================================
- Hits        30851    30790      -61     
  Misses       5168     5168

Flag	Coverage Δ
unittests	`85.62% <67.85%> (-0.03%)`	⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

Impacted Files	Coverage Δ
ding/entry/serial_entry_reward_model_ngu.py	`21.25% <0.00%> (+0.26%)`	⬆️
ding/reward_model/ngu_reward_model.py	`15.22% <0.00%> (+0.05%)`	⬆️
ding/rl_utils/__init__.py	`100.00% <ø> (ø)`
ding/worker/collector/base_serial_evaluator_ngu.py	`18.75% <0.00%> (-0.36%)`	⬇️
...g/worker/collector/interaction_serial_evaluator.py	`93.80% <ø> (ø)`
...ng/worker/collector/sample_serial_collector_ngu.py	`13.77% <0.00%> (-0.09%)`	⬇️
ding/policy/ngu.py	`15.67% <16.66%> (ø)`
ding/rl_utils/adder.py	`89.10% <50.00%> (-0.90%)`	⬇️
ding/entry/serial_entry.py	`94.91% <100.00%> (ø)`
ding/model/wrapper/model_wrappers.py	`90.35% <100.00%> (-0.36%)`	⬇️
... and 12 more

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update bb356ed...1ca8b10. Read the comment docs.

puyuan1996 · 2022-04-14T14:21:45Z

open a new PR to polish NGU.

* fix/fix_submodule_err (opendilab#61) * fix/fix_submodule_err --------- Co-authored-by: ChenQiaoling00 <qiaoling_chen@u.nus.edu> * fix issue templates (opendilab#65) * fix(tokenizer): refactor tokenizer and update usage in readme (opendilab#51) * update tokenizer example * fix(readme, requirements): fix typo at Chinese readme and select a lower version of transformers (opendilab#73) * fix a typo in readme * in order to find InternLMTokenizer, select a lower version of Transformers --------- Co-authored-by: gouhchangjiang <gouhchangjiang@gmail.com> * [Doc] Add wechat and discord link in readme (opendilab#78) * Doc：add wechat and discord link * Doc：update wechat and discord link * Doc：update wechat and discord link * Doc：update wechat and discord link * Doc：update wechat and discord link * Doc：update wechat and discord link * Doc：update wechat and discord link * Doc：update wechat and discord link * Doc：update wechat and discord link * Doc：update wechat and discord link * Doc：update wechat and discord link * [Docs]: add Japanese README (opendilab#43) * Add Japanese README * Update README-ja-JP.md replace message * Update README-ja-JP.md * add repetition_penalty in GenerationConfig in web_demo.py (opendilab#48) Co-authored-by: YWMditto <862779238@qq.com> * use fp16 in instruction (opendilab#80) * [Enchancement] add more options for issue template (opendilab#77) * [Enchancement] add more options for issue template * update qustion icon * fix link * Use tempfile for convert2hf.py (opendilab#23) Fix InternLM/InternLM#50 * delete torch_dtype of README's example code (opendilab#100) * set the value of repetition_penalty to 1.0 to avoid random outputs (opendilab#99) * Update web_demo.py (opendilab#97) Remove meaningless log. * [Fix]Fix wrong string cutoff in the script for sft text tokenizing (opendilab#106) * docs(install.md): update dependency package transformers version to >= 4.28.0 (opendilab#124) Co-authored-by: 黄婷 <huangting3@CN0014010744M.local> * docs(LICENSE): add license (opendilab#125) * add license of colossalai and flash-attn * fix lint * modify the name * fix AutoModel map in convert2hf.py (opendilab#116) * variables are not printly as expect (opendilab#114) * feat(solver): fix code to adapt to torch2.0 and provide docker images (opendilab#128) * feat(solver): fix code to adapt to torch2.0 * docs(install.md): publish internlm environment image * docs(install.md): update dependency packages version * docs(install.md): update default image --------- Co-authored-by: 黄婷 <huangting3@CN0014010744M.local> * add demo test (opendilab#132) Co-authored-by: qa-caif-cicd <qa-caif-cicd@pjlab.org.cn> * fix web_demo cache accelerate (opendilab#133) * fix(hybrid_zero_optim.py): delete math import * Update embedding.py --------- Co-authored-by: ChenQiaoling00 <qiaoling_chen@u.nus.edu> Co-authored-by: Kai Chen <chenkaidev@gmail.com> Co-authored-by: Yang Gao <Gary1546308416AL@gmail.com> Co-authored-by: Changjiang GOU <gouchangjiang@gmail.com> Co-authored-by: gouhchangjiang <gouhchangjiang@gmail.com> Co-authored-by: vansin <msnode@163.com> Co-authored-by: Ikko Eltociear Ashimine <eltociear@gmail.com> Co-authored-by: YWMditto <46778265+YWMditto@users.noreply.github.com> Co-authored-by: YWMditto <862779238@qq.com> Co-authored-by: WRH <12756472+wangruohui@users.noreply.github.com> Co-authored-by: liukuikun <24622904+Harold-lkk@users.noreply.github.com> Co-authored-by: x54-729 <45304952+x54-729@users.noreply.github.com> Co-authored-by: Shuo Zhang <zhangshuolove@live.com> Co-authored-by: Miao Zheng <76149310+MeowZheng@users.noreply.github.com> Co-authored-by: huangting4201 <1538303371@qq.com> Co-authored-by: 黄婷 <huangting3@CN0014010744M.local> Co-authored-by: ytxiong <45058324+yingtongxiong@users.noreply.github.com> Co-authored-by: Zaida Zhou <58739961+zhouzaida@users.noreply.github.com> Co-authored-by: kkscilife <126147887+kkscilife@users.noreply.github.com> Co-authored-by: qa-caif-cicd <qa-caif-cicd@pjlab.org.cn> Co-authored-by: hw <45089338+MorningForest@users.noreply.github.com>

* fix/fix_submodule_err (opendilab#61) * fix/fix_submodule_err --------- Co-authored-by: ChenQiaoling00 <qiaoling_chen@u.nus.edu> * fix issue templates (opendilab#65) * fix(tokenizer): refactor tokenizer and update usage in readme (opendilab#51) * update tokenizer example * fix(readme, requirements): fix typo at Chinese readme and select a lower version of transformers (opendilab#73) * fix a typo in readme * in order to find InternLMTokenizer, select a lower version of Transformers --------- Co-authored-by: gouhchangjiang <gouhchangjiang@gmail.com> * [Doc] Add wechat and discord link in readme (opendilab#78) * Doc：add wechat and discord link * Doc：update wechat and discord link * Doc：update wechat and discord link * Doc：update wechat and discord link * Doc：update wechat and discord link * Doc：update wechat and discord link * Doc：update wechat and discord link * Doc：update wechat and discord link * Doc：update wechat and discord link * Doc：update wechat and discord link * Doc：update wechat and discord link * [Docs]: add Japanese README (opendilab#43) * Add Japanese README * Update README-ja-JP.md replace message * Update README-ja-JP.md * add repetition_penalty in GenerationConfig in web_demo.py (opendilab#48) Co-authored-by: YWMditto <862779238@qq.com> * use fp16 in instruction (opendilab#80) * [Enchancement] add more options for issue template (opendilab#77) * [Enchancement] add more options for issue template * update qustion icon * fix link * Use tempfile for convert2hf.py (opendilab#23) Fix InternLM/InternLM#50 * delete torch_dtype of README's example code (opendilab#100) * set the value of repetition_penalty to 1.0 to avoid random outputs (opendilab#99) * Update web_demo.py (opendilab#97) Remove meaningless log. * [Fix]Fix wrong string cutoff in the script for sft text tokenizing (opendilab#106) * docs(install.md): update dependency package transformers version to >= 4.28.0 (opendilab#124) Co-authored-by: 黄婷 <huangting3@CN0014010744M.local> * docs(LICENSE): add license (opendilab#125) * add license of colossalai and flash-attn * fix lint * modify the name * fix AutoModel map in convert2hf.py (opendilab#116) * variables are not printly as expect (opendilab#114) * feat(solver): fix code to adapt to torch2.0 and provide docker images (opendilab#128) * feat(solver): fix code to adapt to torch2.0 * docs(install.md): publish internlm environment image * docs(install.md): update dependency packages version * docs(install.md): update default image --------- Co-authored-by: 黄婷 <huangting3@CN0014010744M.local> * add demo test (opendilab#132) Co-authored-by: qa-caif-cicd <qa-caif-cicd@pjlab.org.cn> * fix web_demo cache accelerate (opendilab#133) * Doc: add twitter link (opendilab#141) * Feat add checkpoint fraction (opendilab#151) * feat(config): add checkpoint_fraction into config * feat: remove checkpoint_fraction from configs/7B_sft.py --------- Co-authored-by: wangguoteng.p <wangguoteng925@qq.com> * [Doc] update deployment guide to keep consistency with lmdeploy (opendilab#136) * update deployment guide * fix error * use llm partition (opendilab#159) Co-authored-by: qa-caif-cicd <qa-caif-cicd@pjlab.org.cn> * test(ci_scripts): clean test data after test, remove unnecessary global variables, and other optimizations (opendilab#165) * test: optimization of ci scripts(variables, test data cleaning, etc). * chore(workflows): disable ci job on push. * fix: update partition * test(ci_scripts): add install requirements automaticlly,trigger event about lint check and other optimizations (opendilab#174) * add pull_request in lint check * use default variables in ci_scripts * fix format * check and install requirements automaticlly * fix format --------- Co-authored-by: qa-caif-cicd <qa-caif-cicd@pjlab.org.cn> * feat(profiling): add a simple memory profiler (opendilab#89) * feat(profiling): add simple memory profiler * feat(profiling): add profiling argument * feat(CI_workflow): Add PR & Issue auto remove workflow (opendilab#184) * feat(ci_workflow): Add PR & Issue auto remove workflow Add a workflow for stale PR & Issue auto remove - pr & issue well be labeled as stale for inactive in 7 days - staled PR & Issue well be remove in 7 days - run this workflow every day on 1:30 a.m. * Update stale.yml * feat(bot): Create .owners.yml for Auto Assign (opendilab#176) * Create .owners.yml: for issue/pr assign automatically * Update .owners.yml * Update .owners.yml fix typo * [feat]: add pal reasoning script (opendilab#163) * [Feat] Add PAL inference script * Update README.md * Update tools/README.md Co-authored-by: BigDong <yudongwang1226@gmail.com> * Update tools/pal_inference.py Co-authored-by: BigDong <yudongwang1226@gmail.com> * Update pal script * Update README.md * restore .ore-commit-config.yaml * Update tools/README.md Co-authored-by: BigDong <yudongwang1226@gmail.com> * Update tools/README.md Co-authored-by: BigDong <yudongwang1226@gmail.com> * Update pal inference script * Update READMD.md * Update internlm/utils/interface.py Co-authored-by: Wenwei Zhang <40779233+ZwwWayne@users.noreply.github.com> * Update pal script * Update pal script * Update script * Add docstring * Update format * Update script * Update script * Update script --------- Co-authored-by: BigDong <yudongwang1226@gmail.com> Co-authored-by: Wenwei Zhang <40779233+ZwwWayne@users.noreply.github.com> * test(ci_scripts): add timeout settings and clean work after the slurm job (opendilab#185) * restore pr test on develop branch * add mask * add post action to cancel slurm job * remove readonly attribute on job log * add debug info * debug job log * try stdin * use stdin * set default value avoid error * try setting readonly on job log * performance echo * remove debug info * use squeue to check slurm job status * restore the lossed parm * litmit retry times * use exclusive to avoid port already in use * optimize loop body * remove partition * add {} for variables * set env variable for slurm partition --------- Co-authored-by: qa-caif-cicd <qa-caif-cicd@pjlab.org.cn> * refactor(tools): move interface.py and import it to web_demo (opendilab#195) * move interface.py and import it to web_demo * typo * fix(ci): fix lint error * fix(ci): fix lint error --------- Co-authored-by: Sun Peng <sunpengsdu@gmail.com> Co-authored-by: ChenQiaoling00 <qiaoling_chen@u.nus.edu> Co-authored-by: Kai Chen <chenkaidev@gmail.com> Co-authored-by: Yang Gao <Gary1546308416AL@gmail.com> Co-authored-by: Changjiang GOU <gouchangjiang@gmail.com> Co-authored-by: gouhchangjiang <gouhchangjiang@gmail.com> Co-authored-by: vansin <msnode@163.com> Co-authored-by: Ikko Eltociear Ashimine <eltociear@gmail.com> Co-authored-by: YWMditto <46778265+YWMditto@users.noreply.github.com> Co-authored-by: YWMditto <862779238@qq.com> Co-authored-by: WRH <12756472+wangruohui@users.noreply.github.com> Co-authored-by: liukuikun <24622904+Harold-lkk@users.noreply.github.com> Co-authored-by: x54-729 <45304952+x54-729@users.noreply.github.com> Co-authored-by: Shuo Zhang <zhangshuolove@live.com> Co-authored-by: Miao Zheng <76149310+MeowZheng@users.noreply.github.com> Co-authored-by: 黄婷 <huangting3@CN0014010744M.local> Co-authored-by: ytxiong <45058324+yingtongxiong@users.noreply.github.com> Co-authored-by: Zaida Zhou <58739961+zhouzaida@users.noreply.github.com> Co-authored-by: kkscilife <126147887+kkscilife@users.noreply.github.com> Co-authored-by: qa-caif-cicd <qa-caif-cicd@pjlab.org.cn> Co-authored-by: hw <45089338+MorningForest@users.noreply.github.com> Co-authored-by: Guoteng <32697156+SolenoidWGT@users.noreply.github.com> Co-authored-by: wangguoteng.p <wangguoteng925@qq.com> Co-authored-by: lvhan028 <lvhan_028@163.com> Co-authored-by: zachtzy <141206206+zachtzy@users.noreply.github.com> Co-authored-by: cx <759046501@qq.com> Co-authored-by: Jaylin Lee <61487970+APX103@users.noreply.github.com> Co-authored-by: del-zhenwu <dele.zhenwu@gmail.com> Co-authored-by: Shaoyuan Xie <66255889+Daniel-xsy@users.noreply.github.com> Co-authored-by: BigDong <yudongwang1226@gmail.com> Co-authored-by: Wenwei Zhang <40779233+ZwwWayne@users.noreply.github.com> Co-authored-by: huangting4201 <huangting3@sensetime.com>

opendilab and others added 30 commits July 8, 2021 16:23

style(nyz): update badges with actions and issues

90b2797

hotfix(nyz): fix dist entry disable-flask-log typo

e95b01a

doc(nyz): remove doc file in main repo, modify doc workflow(enable doc)

3b88f70

doc(nyz): modify doc workflow on condition and remove install workflow

1567939

doc(nyz): modify doc workflow on condition

539e8ac

doc(nyz): fix doc workflow install cmd bug

cc9f682

doc(nyz): fix doc workflow public path bug

0ee3636

doc(nyz): fix doc workflow html path bug

f15d444

style(nyz): update conda install cmd

43fab4e

hotfix(nyz): fix subprocess env manager state transition bug and exec…

5adf800

… usage in setup.py

Merge branch 'main' of https://github.com/opendilab/DI-engine

98b938a

create badge.json

240d013

delete badges.json

814ab70

style(nyz): update badges in readme

ce54633

doc(nyz): fix quick start file path problem

d5fe8a6

doc(nyz): indicate doc repo branch name in doc workflow

742d513

Merge remote-tracking branch 'origin/main' into doc/separate

7f9b8e3

Merge pull request opendilab#4 from PaParaZz1/doc/separate

25e43c2

separate doc from main repo to doc repo

style(nyz): update logo svg path

71a7ec5

style(nyz): update issue template (ci skip)

e41864d

style(nyz): update coverage badge and colab quick start

e322083

style(nyz): update requests and urllib3 version

df81bee

badge(hansbug): add LoC and Documentation Percentage badge to README.…

f92680a

…md && update code coverage badge (opendilab#8)

support layer_num==0 in MLP layer.

966a6d7

add on policy ppo; modify ddpg/td3 config.

7accb4e

style(nyz): add chinese version doc link

557a44c

style(nyz): add 3 minutes kickoff chinese version

582f707

refactor(nyz): refactor output file structure

1849da7

davide97l and others added 12 commits November 19, 2021 22:33

style(nyz): add PDQN/MAPPO link, DQN doc zh link and correct format

f9f92a8

Merge branch 'main' of https://github.com/opendilab/DI-engine into de…

babf050

…v-polish-ngu Conflicts: dizoo/gym_hybrid/config/gym_hybrid_ddpg_config.py

Merge branch 'main' of https://github.com/opendilab/DI-engine into de…

a87f4bb

…v-polish-ngu Conflicts: dizoo/gym_hybrid/config/gym_hybrid_ddpg_config.py

style(pu): yapf format

a97a772

style(pu): yapf format

883ddcf

Merge branch 'dev-polish-ngu' of https://github.com/puyuan1996/DI-engine

a1b0f89

into dev-polish-ngu

Merge branch 'dev-polish-ngu' of https://github.com/puyuan1996/DI-engine

4f9cca2

into dev-polish-ngu

polish(pu): polish ngu update_per_collect and eps end decay para

0bd47fd

polish(pu): polish ngu update_per_collect and eps end decay para

cb3aa6b

polish(pu): polish ngu update_per_collect and eps end decay para

fe05389

polish(pu): polish ngu update_per_collect and eps end decay para

a63494d

PaParaZz1 force-pushed the main branch 2 times, most recently from ee876e0 to c6947cd Compare January 4, 2022 06:27

puyuan1996 changed the title ~~Polish(pu): polish the nstep_return_ngu and null_padding action in NGU~~ polish(pu): polish the nstep_return_ngu and null_padding action in NGU Jan 19, 2022

puyuan1996 added 8 commits March 18, 2022 00:32

Merge branch 'main' of https://github.com/opendilab/DI-engine into de…

7de0ce0

…v-polish-ngu

Merge branch 'dev-polish-ngu' of https://github.com/puyuan1996/DI-engine

7141eb2

into dev-polish-ngu

style(pu): yapf format

ba4d45a

Merge branch 'main' of https://github.com/opendilab/DI-engine into de…

363003b

…v-polish-ngu

fix(pu): fix import error

d552a69

fix(pu): fix unittest

2bdd053

Merge branch 'dev-polish-ngu' of https://github.com/puyuan1996/DI-engine

5df07c2

into dev-polish-ngu

fix(pu): fix model_wrappers unittest

1ca8b10

puyuan1996 closed this Apr 14, 2022

puyuan1996 deleted the dev-polish-ngu branch April 20, 2022 08:13

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

polish(pu): polish the nstep_return_ngu and null_padding action in NGU #116

polish(pu): polish the nstep_return_ngu and null_padding action in NGU #116

puyuan1996 commented Nov 1, 2021

codecov bot commented Mar 21, 2022 •

edited

Loading

puyuan1996 commented Apr 14, 2022

polish(pu): polish the nstep_return_ngu and null_padding action in NGU #116

polish(pu): polish the nstep_return_ngu and null_padding action in NGU #116

Conversation

puyuan1996 commented Nov 1, 2021

Description

Related Issue

TODO

Check List

codecov bot commented Mar 21, 2022 • edited Loading

Codecov Report

puyuan1996 commented Apr 14, 2022

codecov bot commented Mar 21, 2022 •

edited

Loading