After running a training session on multiple GPUs, batch_size is read wrongly from opt.yaml causing an error.
python -m torch.distributed.launch --nproc_per_node 2 train.py --batch-size 30 --data coco128.yaml --weights yolov5s.pt --device 1,2
Press Ctrl+C to stop the session.
Now, try to resume:
python -m torch.distributed.launch --nproc_per_node 2 train.py --batch-size 30 --data coco128.yaml --weights yolov5s.pt --device 1,2 --resume
*****************************************
Setting OMP_NUM_THREADS environment variable for each process to be 1 in default, to avoid your system being overloaded, please further tune the variable for optimal performance in your application as needed.
*****************************************
github: Traceback (most recent call last):
File "train.py", line 492, in <module>
device = select_device(opt.device, batch_size=opt.batch_size)
File "/user/detection/tests/yolov5/utils/torch_utils.py", line 68, in select_device
assert batch_size % n == 0, f'batch-size {batch_size} not multiple of GPU count {n}'
AssertionError: batch-size 15 not multiple of GPU count 2
up to date with https://github.com/ultralytics/yolov5 â
Resuming training from ./runs/train/exp/weights/last.pt
Traceback (most recent call last):
File "train.py", line 492, in <module>
device = select_device(opt.device, batch_size=opt.batch_size)
File "/user/detection/tests/yolov5/utils/torch_utils.py", line 68, in select_device
assert batch_size % n == 0, f'batch-size {batch_size} not multiple of GPU count {n}'
AssertionError: batch-size 15 not multiple of GPU count 2
Traceback (most recent call last):
File "/usr/lib/python3.8/runpy.py", line 194, in _run_module_as_main
return _run_code(code, main_globals, None,
File "/usr/lib/python3.8/runpy.py", line 87, in _run_code
exec(code, run_globals)
File "/user/venv/yolov5/lib/python3.8/site-packages/torch/distributed/launch.py", line 260, in <module>
main()
File "/user/venv/yolov5/lib/python3.8/site-packages/torch/distributed/launch.py", line 255, in main
raise subprocess.CalledProcessError(returncode=process.returncode,
subprocess.CalledProcessError: Command '['/user/venv/yolov5/bin/python', '-u', 'train.py', '--local_rank=1', '--batch-size', '30', '--data', 'coco128.yaml', '--weights', 'yolov5s.pt', '--device', '1,2', '--resume']' returned non-zero exit status 1.
The training should resume correctly, with the right batch size.
ð Bug
After running a training session on multiple GPUs, batch_size is read wrongly from opt.yaml causing an error.
To Reproduce (REQUIRED)
Run this line:
Press Ctrl+C to stop the session.
Now, try to resume:
Output:
Expected behavior
The training should resume correctly, with the right batch size.
Environment
Additional context
--