Repository navigation
Caching image path #2121
Description
Activity
ð Hello @train255, thank you for your interest in ð YOLOv5! Please visit our âïļ Tutorials to get started, where you can find quickstart guides for simple tasks like Custom Data Training all the way to advanced concepts like Hyperparameter Evolution.
If this is a ð Bug Report, please provide screenshots and minimum viable code to reproduce your issue, otherwise we can not help you.
If this is a custom training â Question, please provide as much information as possible, including dataset images, training logs, screenshots, and a public link to online W&B logging if available.
For business inquiries or professional support requests please visit https://www.ultralytics.com or email Glenn Jocher at glenn.jocher@ultralytics.com.
Requirements
Python 3.8 or later with all requirements.txt dependencies installed, including torch>=1.7. To install run:
$ pip install -r requirements.txtEnvironments
YOLOv5 may be run in any of the following up-to-date verified environments (with all dependencies including CUDA/CUDNN, Python and PyTorch preinstalled):
- Google Colab and Kaggle notebooks with free GPU:
- Google Cloud Deep Learning VM. See GCP Quickstart Guide
- Amazon Deep Learning AMI. See AWS Quickstart Guide
- Docker Image. See Docker Quickstart Guide
Status
If this badge is green, all YOLOv5 GitHub Actions Continuous Integration (CI) tests are currently passing. CI tests verify correct operation of YOLOv5 training (train.py), testing (test.py), inference (detect.py) and export (export.py) on MacOS, Windows, and Ubuntu every 24 hours and on every commit.
@train255 you can specify your separate train and test directories or text files in your dataset.yaml file here:
Lines 12 to 16 in 73a0669
| # train and val data as 1) directory: path/images/, 2) file: path/images.txt, or 3) list: [path1/images/, path2/images/] | |
| train: ../coco/train2017.txt # 118287 images | |
| val: ../coco/val2017.txt # 5000 images | |
| test: ../coco/test-dev2017.txt # 20288 of 40670 images, submit to https://competitions.codalab.org/competitions/20794 | |
@train255 you can specify your separate train and test directories or text files in your dataset.yaml file here:
Lines 12 to 16 in 73a0669
# train and val data as 1) directory: path/images/, 2) file: path/images.txt, or 3) list: [path1/images/, path2/images/] train: ../coco/train2017.txt # 118287 images val: ../coco/val2017.txt # 5000 images test: ../coco/test-dev2017.txt # 20288 of 40670 images, submit to https://competitions.codalab.org/competitions/20794
@glenn-jocher I think it is not the point of @train255. Sometimes it is inconvenient to separate directories if you already made the workflow that uses this dataset and also is being used in another framework. Why don't you just add a prefix or a suffix (e.g. train, val, test) to the cache file name?
@rcg12387 I don't understand. Can you provide a notebook that reproduces the original issue? Also a PR with a proposed fix would make parsing it simpler.
@rcg12387 I don't understand. Can you provide a notebook that reproduces the original issue? Also a PR with a proposed fix would make parsing it simpler.
For example, when all train, val, test sets are in the same directory, cashe file names of each cases conflict with each other. Especially, in the training case this is bad because training and validation are being performed successively.
# Check cache
cache_path = Path(self.label_files[0]).parent.with_suffix('.cache') # cached labels
if cache_path.is_file(): # sometimes train and val caches conflict
# load or re-cache
else:
# create cache
In some workflow, it is compelled not to separate these sets, e.g. owing to collaboration.
Anyway, PR(#2134) of @train255 can be one of ways to solve the issue.
@rcg12387 ah, yes I see! Does this result in actual incorrect dataloading in your case?
The cache system is designed to be foolproof in the sense that your underlying data is compared to the *.cache contents on every training run (once for train, once for val), and if the two hashes don't match, the *.cache file is ignored and the new data is re-cached. So it could be that you are only seeing one cache file, which would be the latest created (the val set).
You can check the dataloader tqdm messages to see if the correct number of images were loaded for each part (train and val).
@rcg12387 the cache checking and re-caching on different hashes happens here. L376 compares the *.cache hash against the requested data hash, and in case of disparity the *.cache file is ignored.
Lines 371 to 380 in 73a0669
| # Check cache | |
| self.label_files = img2label_paths(self.img_files) # labels | |
| cache_path = Path(self.label_files[0]).parent.with_suffix('.cache') # cached labels | |
| if cache_path.is_file(): | |
| cache = torch.load(cache_path) # load | |
| if cache['hash'] != get_hash(self.label_files + self.img_files) or 'results' not in cache: # changed | |
| cache = self.cache_labels(cache_path, prefix) # re-cache | |
| else: | |
| cache = self.cache_labels(cache_path, prefix) # cache | |
@rcg12387 ah, yes I see! Does this result in actual incorrect dataloading in your case?
Yes. In my case it causes exception.
cache_path = Path(self.label_files[0]).parent.with_suffix('.cache') # cached labels
# cache_path = Path(str(Path(self.label_files[0]).parent) + '_' + cache_affix).with_suffix('.cache') # cached labels
@rcg12387 I see. I think you may be running into a separate issue though, as I see from your screenshots train and val sets are defined differently (i.e. train has 237 batches, val only has 60 in both cases), so there does not appear to be any confusion between the two.
If your error is reproducible can you create a notebook to allow us to run it and debug?
@rcg12387 I created a small dataset called coco6, which uses the first 4 images of COCO128 for training, and then images 5 and 6 for testing, with *.txt files pointing to the two different sets of images (all 6 images are in the /images directory). Everything appears to work correctly. A single cache file remains after the dataloaders are defined, which corresponds to the val set, and training proceeds without issue.
Yeah. I also tried using coco128 (train: 80, val: 48) but cannot reproduce the issue. However, using my own dataset it crashes. Let me see why.
Sorry I will explain my issue more clearly.
This is my data structure
content
| -- data
| -- 1.jpg
| -- 1.txt
| -- 2.jpg
| -- 2.txt
| -- 3.jpg
| -- 3.txt
| -- 4.jpg
| -- 4.txt
| -- config.yaml
| -- train.txt
| -- val.txt
train.txt file content
/content/data/1.jpg
/content/data/2.jpg
val.txt file content
/content/data/3.jpg
/content/data/4.jpg
config.yaml file content
train: /content/train.txt
val: /content/val.txt
nc: 2
names: ['Class1','Class2']
The training proceeds without issue in the first time.
python train.py --img 640 --batch 16 --epochs 100 --data /content/config.yaml --cfg ./models/yolov5x.yaml --weights ''
...
Scanning images: 100% 14915/14915 [6:48:26<00:00, 1.64s/it]
Scanning labels /content/train.cache (14915 found, 0 missing, 0 empty, 0 duplicate, for 14915 images): 14915it [00:00, 15900.80it/s]
Scanning images: 100% 8657/8657 [02:26<00:00, 59.22it/s]
Scanning labels /content/train.cache (8657 found, 0 missing, 0 empty, 0 duplicate, for 8657 images): 8657it [00:00, 16335.68it/s]
...
The data is too big so I must train many times.
python train.py --img 640 --batch 16 --epochs 100 --data /content/config.yaml --cfg ./models/yolov5x.yaml --weights 'checkpoints/last.pt'
....
Scanning labels /content/train.cache (0 found, 0 missing, 0 empty, 0 duplicate, for ....
...
I have a problem here. The scanning image go back to a beginning because the train.cache was overwritten by the val caching file. We will have to spend a lot of time with the big data.
4 remaining items
@rcg12387 @train255 good news!! The re-caching issue is now completely solved in PR #2134. As long as the underlying images and labels do not change, you will never need to recache your dataset. The change was simply modifying the cache name to the *.txt filename when *.txt files are used to define the dataset:
Please git pull to receive this update and let us know if you run into any other issues!
Cool
@train255 I see, so the difference is that your labels are in the same directory as your images. This should also work, I will run another coco6 test in this configuration. Your code is very out of date by the way, you should git pull to update or reclone the repo, as there have been many changes since the dataloading process you are showing.
I found out the real reason. There is a delay when I use google colab to load caching file in My Drive. Thank you for your support.
@train255 oh great! You should never train from remote data, mounted drives etc. Always copy your dataset to local storage first.





I have 2 files train.txt and valid.txt. The train images and valid images are in the same directory "/content/train"
When we train a model, the train and val caching file is the same name
train.cache