[BUG]: HDF5Storage fails to collect #198
No reviewers
Labels
No labels
CRITICAL
Stale
WIP
bug
concept
coordinate
dataset
dependencies
documentation
duplicate
enhancement
github_actions
good first issue
help wanted
invalid
maintenance
maps
marker
mask
on hold
parcellation
preprocess
question
ready
storage
template-space
triage
wontfix
No milestone
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
juaml/junifer!198
Loading…
Reference in a new issue
No description provided.
Delete branch "fix/hdf5_collect"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Is there an existing issue for this?
Current Behavior
HCP results (4214 files) are not colected:
Closer inspection at the issue, it seems that the logic of chunking the collect is not correct:
github.com/juaml/junifer@85b5ff538d/junifer/storage/hdf5.py (L858-L874)The problem is that the number of files is not multiple of the chunk_size, so the last chunk is not complete. This raises an error were
iis 4214.And, since we are fixing this logic, the nested progress bars are not correct. Indeed, this is how it looks:
This is because for the "chunk" progress bar, is not possible to know the size in an enumerate:
I would even suggest that we do not need separate progressbars due to the chunks. Indeed what we care is to show how the files are being included. So I would only keep 2 bars, one for the feature and one for the file.
Expected Behavior
Collect to run without issues.
Steps To Reproduce
Might need access to this specific project directory.
Environment
Relevant log output
No response
Anything else?
No response
Codecov Report
93.22% <91.17%> (-0.03%)Flags with carried forward coverage won't be shown. Click here to find out more.
93.17% <91.17%> (-0.38%)@ -24,0 +41,4 @@The data to be chunked.kind : strThe kind of data to be chunked.element_count : intMissing docstring here.
Would be a safe measure to
del to_writeto further reduce memory usage.@ -785,7 +791,91 @@ def test_store_timeseries(tmp_path: Path) -> None:assert_array_equal(read_df.values, data)Missing type annotations.
Let the GC do it when it considers it necessary. For small datasets, this might increase the computation time.
Thanks for the fix!