Skip to content

Caching image path #2121

Description

@train255

I have 2 files train.txt and valid.txt. The train images and valid images are in the same directory "/content/train"

# train.txt file
/content/train/1.jpg
/content/train/2.jpg
# val.txt file
/content/train/3.jpg
/content/train/4.jpg

When we train a model, the train and val caching file is the same name train.cache

...
Scanning images: 100% 14915/14915 [6:48:26<00:00,  1.64s/it]
Scanning labels /content/train.cache (14915 found, 0 missing, 0 empty, 0 duplicate, for 14915 images): 14915it [00:00, 15900.80it/s]
Scanning images: 100% 8657/8657 [02:26<00:00, 59.22it/s]
Scanning labels /content/train.cache (8657 found, 0 missing, 0 empty, 0 duplicate, for 8657 images): 8657it [00:00, 16335.68it/s]
...

Activity

github-actions commented on Feb 3, 2021

@github-actions
Contributor

👋 Hello @train255, thank you for your interest in 🚀 YOLOv5! Please visit our ⭐ïļ Tutorials to get started, where you can find quickstart guides for simple tasks like Custom Data Training all the way to advanced concepts like Hyperparameter Evolution.

If this is a 🐛 Bug Report, please provide screenshots and minimum viable code to reproduce your issue, otherwise we can not help you.

If this is a custom training ❓ Question, please provide as much information as possible, including dataset images, training logs, screenshots, and a public link to online W&B logging if available.

For business inquiries or professional support requests please visit https://www.ultralytics.com or email Glenn Jocher at glenn.jocher@ultralytics.com.

Requirements

Python 3.8 or later with all requirements.txt dependencies installed, including torch>=1.7. To install run:

$ pip install -r requirements.txt

Environments

YOLOv5 may be run in any of the following up-to-date verified environments (with all dependencies including CUDA/CUDNN, Python and PyTorch preinstalled):

Status

CI CPU testing

If this badge is green, all YOLOv5 GitHub Actions Continuous Integration (CI) tests are currently passing. CI tests verify correct operation of YOLOv5 training (train.py), testing (test.py), inference (detect.py) and export (export.py) on MacOS, Windows, and Ubuntu every 24 hours and on every commit.

glenn-jocher commented on Feb 3, 2021

@glenn-jocher
Member

@train255 you can specify your separate train and test directories or text files in your dataset.yaml file here:

yolov5/data/coco.yaml

Lines 12 to 16 in 73a0669

# train and val data as 1) directory: path/images/, 2) file: path/images.txt, or 3) list: [path1/images/, path2/images/]
train: ../coco/train2017.txt # 118287 images
val: ../coco/val2017.txt # 5000 images
test: ../coco/test-dev2017.txt # 20288 of 40670 images, submit to https://competitions.codalab.org/competitions/20794

rcg12387 commented on Feb 4, 2021

@rcg12387

@train255 you can specify your separate train and test directories or text files in your dataset.yaml file here:

yolov5/data/coco.yaml

Lines 12 to 16 in 73a0669

# train and val data as 1) directory: path/images/, 2) file: path/images.txt, or 3) list: [path1/images/, path2/images/]
train: ../coco/train2017.txt # 118287 images
val: ../coco/val2017.txt # 5000 images
test: ../coco/test-dev2017.txt # 20288 of 40670 images, submit to https://competitions.codalab.org/competitions/20794

@glenn-jocher I think it is not the point of @train255. Sometimes it is inconvenient to separate directories if you already made the workflow that uses this dataset and also is being used in another framework. Why don't you just add a prefix or a suffix (e.g. train, val, test) to the cache file name?

glenn-jocher commented on Feb 4, 2021

@glenn-jocher
Member

@rcg12387 I don't understand. Can you provide a notebook that reproduces the original issue? Also a PR with a proposed fix would make parsing it simpler.

rcg12387 commented on Feb 5, 2021

@rcg12387

@rcg12387 I don't understand. Can you provide a notebook that reproduces the original issue? Also a PR with a proposed fix would make parsing it simpler.

For example, when all train, val, test sets are in the same directory, cashe file names of each cases conflict with each other. Especially, in the training case this is bad because training and validation are being performed successively.

# Check cache
cache_path = Path(self.label_files[0]).parent.with_suffix('.cache')  # cached labels
if cache_path.is_file(): # sometimes train and val caches conflict
    # load or re-cache
else:
    # create cache

In some workflow, it is compelled not to separate these sets, e.g. owing to collaboration.
Anyway, PR(#2134) of @train255 can be one of ways to solve the issue.

glenn-jocher commented on Feb 5, 2021

@glenn-jocher
Member

@rcg12387 ah, yes I see! Does this result in actual incorrect dataloading in your case?

The cache system is designed to be foolproof in the sense that your underlying data is compared to the *.cache contents on every training run (once for train, once for val), and if the two hashes don't match, the *.cache file is ignored and the new data is re-cached. So it could be that you are only seeing one cache file, which would be the latest created (the val set).

You can check the dataloader tqdm messages to see if the correct number of images were loaded for each part (train and val).

glenn-jocher commented on Feb 5, 2021

@glenn-jocher
Member

@rcg12387 the cache checking and re-caching on different hashes happens here. L376 compares the *.cache hash against the requested data hash, and in case of disparity the *.cache file is ignored.

yolov5/utils/datasets.py

Lines 371 to 380 in 73a0669

# Check cache
self.label_files = img2label_paths(self.img_files) # labels
cache_path = Path(self.label_files[0]).parent.with_suffix('.cache') # cached labels
if cache_path.is_file():
cache = torch.load(cache_path) # load
if cache['hash'] != get_hash(self.label_files + self.img_files) or 'results' not in cache: # changed
cache = self.cache_labels(cache_path, prefix) # re-cache
else:
cache = self.cache_labels(cache_path, prefix) # cache

rcg12387 commented on Feb 5, 2021

@rcg12387

@rcg12387 ah, yes I see! Does this result in actual incorrect dataloading in your case?

Yes. In my case it causes exception.

cache_path = Path(self.label_files[0]).parent.with_suffix('.cache')  # cached labels
# cache_path = Path(str(Path(self.label_files[0]).parent) + '_' + cache_affix).with_suffix('.cache')  # cached labels

bug

rcg12387 commented on Feb 5, 2021

@rcg12387

If I separate two cache files it goes well.

# cache_path = Path(self.label_files[0]).parent.with_suffix('.cache')  # cached labels
cache_path = Path(str(Path(self.label_files[0]).parent) + '_' + cache_affix).with_suffix('.cache')  # cached labels

normal

glenn-jocher commented on Feb 5, 2021

@glenn-jocher
Member

@rcg12387 I see. I think you may be running into a separate issue though, as I see from your screenshots train and val sets are defined differently (i.e. train has 237 batches, val only has 60 in both cases), so there does not appear to be any confusion between the two.

If your error is reproducible can you create a notebook to allow us to run it and debug?

glenn-jocher commented on Feb 5, 2021

@glenn-jocher
Member

@rcg12387 I created a small dataset called coco6, which uses the first 4 images of COCO128 for training, and then images 5 and 6 for testing, with *.txt files pointing to the two different sets of images (all 6 images are in the /images directory). Everything appears to work correctly. A single cache file remains after the dataloaders are defined, which corresponds to the val set, and training proceeds without issue.

Screen Shot 2021-02-04 at 7 16 32 PM

Screen Shot 2021-02-04 at 7 18 22 PM

rcg12387 commented on Feb 5, 2021

@rcg12387

Yeah. I also tried using coco128 (train: 80, val: 48) but cannot reproduce the issue. However, using my own dataset it crashes. Let me see why.

train255 commented on Feb 5, 2021

@train255
ContributorAuthor

Sorry I will explain my issue more clearly.

This is my data structure

content
| -- data
     | -- 1.jpg
     | -- 1.txt
     | -- 2.jpg
     | -- 2.txt
     | -- 3.jpg
     | -- 3.txt
     | -- 4.jpg
     | -- 4.txt
| -- config.yaml
| -- train.txt
| -- val.txt

train.txt file content

/content/data/1.jpg
/content/data/2.jpg

val.txt file content

/content/data/3.jpg
/content/data/4.jpg

config.yaml file content

train: /content/train.txt
val: /content/val.txt
nc: 2
names: ['Class1','Class2']

The training proceeds without issue in the first time.

python train.py --img 640 --batch 16 --epochs 100 --data /content/config.yaml --cfg ./models/yolov5x.yaml --weights ''
...
Scanning images: 100% 14915/14915 [6:48:26<00:00,  1.64s/it]
Scanning labels /content/train.cache (14915 found, 0 missing, 0 empty, 0 duplicate, for 14915 images): 14915it [00:00, 15900.80it/s]
Scanning images: 100% 8657/8657 [02:26<00:00, 59.22it/s]
Scanning labels /content/train.cache (8657 found, 0 missing, 0 empty, 0 duplicate, for 8657 images): 8657it [00:00, 16335.68it/s]
...

The data is too big so I must train many times.

python train.py --img 640 --batch 16 --epochs 100 --data /content/config.yaml --cfg ./models/yolov5x.yaml --weights 'checkpoints/last.pt'
....
Scanning labels /content/train.cache (0 found, 0 missing, 0 empty, 0 duplicate, for ....
...

I have a problem here. The scanning image go back to a beginning because the train.cache was overwritten by the val caching file. We will have to spend a lot of time with the big data.

4 remaining items

linked a pull request that will close this issueUnique *.cache filenames fix #2134on Feb 5, 2021

glenn-jocher commented on Feb 5, 2021

@glenn-jocher
Member

@rcg12387 @train255 good news!! The re-caching issue is now completely solved in PR #2134. As long as the underlying images and labels do not change, you will never need to recache your dataset. The change was simply modifying the cache name to the *.txt filename when *.txt files are used to define the dataset:

Screen Shot 2021-02-05 at 11 12 34 AM

Please git pull to receive this update and let us know if you run into any other issues!

rcg12387 commented on Feb 6, 2021

@rcg12387

Cool

train255 commented on Feb 6, 2021

@train255
ContributorAuthor

@train255 I see, so the difference is that your labels are in the same directory as your images. This should also work, I will run another coco6 test in this configuration. Your code is very out of date by the way, you should git pull to update or reclone the repo, as there have been many changes since the dataloading process you are showing.

I found out the real reason. There is a delay when I use google colab to load caching file in My Drive. Thank you for your support.

glenn-jocher commented on Feb 6, 2021

@glenn-jocher
Member

@train255 oh great! You should never train from remote data, mounted drives etc. Always copy your dataset to local storage first.

added a commit that references this issue on Feb 26, 2021
cbd55da
added 2 commits that reference this issue on May 12, 2021
6394a2a
7a6bba7
added a commit that references this issue on May 17, 2021
9402cd1
added 2 commits that reference this issue on Oct 12, 2021
e69f773
8caa475
added 2 commits that reference this issue on Aug 26, 2022
1922bb9
2c6baf7
added a commit that references this issue on May 7, 2024
e9b3de4
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions