I try to run RLHF pipeline for StackLLama with your pretrained models and it failed.
It seems that the reward model has improper output layer.
I'm following this README
To got the pretrained models I do:
-
Download the original llama-7b model
-
To get LLAMA_SE_MODEL, I downloaded your adapter: trl-lib/llama-7b-se-peft and merged it with llama-7b via your merge_peft_adapter.py script:
python merge_peft_adapter.py --adapter_model_name=/home/user/llama-7b-se-peft/ --base_model_name=/home/user/llama7b/ --output_name=/home/user/llama-7b-se-ckpt/
As I understand now LLAMA_SE_MODEL is saved in /home/user/llama-7b-se-ckpt/
-
After that to get LLAMA_SE_RM_MODEL, I downloaded your adapter: trl-lib/llama-7b-se-rm-peft and merged it with LLAMA_SE_MODEL (step 2) via your merge_peft_adapter.py script:
python merge_peft_adapter.py --adapter_model_name=/home/user/llama-7b-se-rm-peft --base_model_name=/home/user/llama-7b-se-ckpt/ --output_name=/home/user/llama-7b-se-rm-ckpt
As I understand now LLAMA_SE_RM_MODEL is placed in /home/user/llama-7b-se-rm-ckpt
After that I tried to start rl_training.py script:
python examples/stack_llama/scripts/rl_training.py --model_name=/home/user/llama-7b-se-ckpt --reward_model_name=/home/user/llama-7b-se-rm-ckpt --adafactor=False --tokenizer_name=/home/user/llama-7b-se-ckpt --save_freq=100 --output_max_length=128 --batch_size=1 --gradient_accumulation_steps=8 --batched_gen=True --ppo_epochs=4 --seed=0 --learning_rate=1.4e-5 --early_stopping=True --output_dir=llama-se-rl-finetune-128-8-8-1.4e-5_adam
And get the following error related to reward model:
File "examples/stack_llama/scripts/rl_training.py", line 255, in
rewards = [torch.tensor(output[0]["score"] - script_args.reward_baseline) for output in pipe_outputs]
KeyError: 0
It seems that the /home/user/llama-7b-se-rm-ckpt has not 'score' output and was not merged properly.
This similar issue is described here: #297
So, could you please clarify this issue and check how pretrained LLAMA_SE_RM_MODEL can be got from your adapter(s).
I try to run RLHF pipeline for StackLLama with your pretrained models and it failed.
It seems that the reward model has improper output layer.
I'm following this README
To got the pretrained models I do:
Download the original llama-7b model
To get LLAMA_SE_MODEL, I downloaded your adapter: trl-lib/llama-7b-se-peft and merged it with llama-7b via your merge_peft_adapter.py script:
python merge_peft_adapter.py --adapter_model_name=/home/user/llama-7b-se-peft/ --base_model_name=/home/user/llama7b/ --output_name=/home/user/llama-7b-se-ckpt/As I understand now LLAMA_SE_MODEL is saved in /home/user/llama-7b-se-ckpt/
After that to get LLAMA_SE_RM_MODEL, I downloaded your adapter: trl-lib/llama-7b-se-rm-peft and merged it with LLAMA_SE_MODEL (step 2) via your merge_peft_adapter.py script:
python merge_peft_adapter.py --adapter_model_name=/home/user/llama-7b-se-rm-peft --base_model_name=/home/user/llama-7b-se-ckpt/ --output_name=/home/user/llama-7b-se-rm-ckptAs I understand now LLAMA_SE_RM_MODEL is placed in /home/user/llama-7b-se-rm-ckpt
After that I tried to start rl_training.py script:
python examples/stack_llama/scripts/rl_training.py --model_name=/home/user/llama-7b-se-ckpt --reward_model_name=/home/user/llama-7b-se-rm-ckpt --adafactor=False --tokenizer_name=/home/user/llama-7b-se-ckpt --save_freq=100 --output_max_length=128 --batch_size=1 --gradient_accumulation_steps=8 --batched_gen=True --ppo_epochs=4 --seed=0 --learning_rate=1.4e-5 --early_stopping=True --output_dir=llama-se-rl-finetune-128-8-8-1.4e-5_adamAnd get the following error related to reward model:
File "examples/stack_llama/scripts/rl_training.py", line 255, in
rewards = [torch.tensor(output[0]["score"] - script_args.reward_baseline) for output in pipe_outputs]
KeyError: 0
It seems that the /home/user/llama-7b-se-rm-ckpt has not 'score' output and was not merged properly.
This similar issue is described here: #297
So, could you please clarify this issue and check how pretrained LLAMA_SE_RM_MODEL can be got from your adapter(s).