We need to rework distributed hdf5 access once more, which appears to have stopped working.
TypeError: h5py objects cannot be pickle
There exists h5pickle, but this does not currently work as drop-in replacement. The issue appears to be related to DaanVanVugt/h5pickle#14.
A naive strategy would to provide only the filepath and hdf5 path to the dataset, opening/reading/closing the file for every dask chunk. The repeated opening and closing introduces a performance penalty. See discussion here
We need to rework distributed hdf5 access once more, which appears to have stopped working.
There exists h5pickle, but this does not currently work as drop-in replacement. The issue appears to be related to DaanVanVugt/h5pickle#14.
A naive strategy would to provide only the filepath and hdf5 path to the dataset, opening/reading/closing the file for every dask chunk. The repeated opening and closing introduces a performance penalty. See discussion here