Feeding images / tensors of different size using PyTorch dataloader classes
Struggled to do this properly on DS Bowl (I resorted to random crops there for training and 1-image sized batches for validation).
Suppose your dataset has some internal structure in it.
For example - you may have images of vastly different aspect ratios (3x1, 1x3 and 1x1) and you would like to squeeze every bit of performance from your pipeline.
Of course, you may pad your images / center-crop them / random crop them - but in this case you will lose some of the information.
I played with this on some tasks - sometimes force-resize works better than crops, but trying to apply your model convolutionally worked really good on SemSeg challenges.
So it may work very well on plain classification as well.
So, if you apply your model convolutionally, you will end up with differently-sized feature maps for each cluster of images.
Within the model, it can be fixed with:
(0) Adaptive avg pooling layers
(1) Some simple logic in .forward statement of the model
But anyway you end up with a small technical issue - PyTorch cannot concatenate tensors of different sizes using standard collation function.
Theoretically, there are several ways to fix this:
(0) Stupid solution - create N datasets, train on them sequentially.
In practice I tried that on DS Bowl - it worked poorly - the model overfitted to each cluster, and then performed poorly on next one;
(1) Crop / pad / resize images (suppose you deliberately want to avoid that);
(2) Insert some custom logic into PyTorch collattion function, i.e. resize there;
(3) Just sample images so that only images of one size end up within each batch;
(0) and (1) I would like to avoid intentionally.
(2) seems a bit stupid as well, because resizing should be done as a pre-processing step (collation function deals with normalized tensors, not images) and it is better not to mix purposes of your modules
Ofc, you can try to produce N tensors in (2) - i.e. tensor for each image size, but that would require additional loop downstream.
In the end, I decided that (3) is the best approach - because it can be easily transferred to other datasets / domains / tasks.
Long story short - here is my solution - I just extended their sampling function:
https://github.com/pytorch/pytorch/issues/1512#issuecomment-405015099
Maybe it is worth a PR on Github?
What do you think?
#deep_learning
#data_science
Like this post or have something to say => tell us more in the comments or donate!
Struggled to do this properly on DS Bowl (I resorted to random crops there for training and 1-image sized batches for validation).
Suppose your dataset has some internal structure in it.
For example - you may have images of vastly different aspect ratios (3x1, 1x3 and 1x1) and you would like to squeeze every bit of performance from your pipeline.
Of course, you may pad your images / center-crop them / random crop them - but in this case you will lose some of the information.
I played with this on some tasks - sometimes force-resize works better than crops, but trying to apply your model convolutionally worked really good on SemSeg challenges.
So it may work very well on plain classification as well.
So, if you apply your model convolutionally, you will end up with differently-sized feature maps for each cluster of images.
Within the model, it can be fixed with:
(0) Adaptive avg pooling layers
(1) Some simple logic in .forward statement of the model
But anyway you end up with a small technical issue - PyTorch cannot concatenate tensors of different sizes using standard collation function.
Theoretically, there are several ways to fix this:
(0) Stupid solution - create N datasets, train on them sequentially.
In practice I tried that on DS Bowl - it worked poorly - the model overfitted to each cluster, and then performed poorly on next one;
(1) Crop / pad / resize images (suppose you deliberately want to avoid that);
(2) Insert some custom logic into PyTorch collattion function, i.e. resize there;
(3) Just sample images so that only images of one size end up within each batch;
(0) and (1) I would like to avoid intentionally.
(2) seems a bit stupid as well, because resizing should be done as a pre-processing step (collation function deals with normalized tensors, not images) and it is better not to mix purposes of your modules
Ofc, you can try to produce N tensors in (2) - i.e. tensor for each image size, but that would require additional loop downstream.
In the end, I decided that (3) is the best approach - because it can be easily transferred to other datasets / domains / tasks.
Long story short - here is my solution - I just extended their sampling function:
https://github.com/pytorch/pytorch/issues/1512#issuecomment-405015099
Maybe it is worth a PR on Github?
What do you think?
#deep_learning
#data_science
Like this post or have something to say => tell us more in the comments or donate!