[DAT]: Add ENKI dataset #47

Open
opened 2022-09-12 11:18:10 +00:00 by fraimondo · 13 comments
fraimondo commented 2022-09-12 11:18:10 +00:00 (Migrated from github.com)

Which dataset is it?

Enki dataset for Juseless

Implementation

Did not implement anything

Your implementation

No response

Dataset access restrictions

  • Public and open access (no registration)
  • Public and open access (registration required)
  • Restricted access (needs approuval)
  • Available only in specific locations (Juseless, Jureca)

Anything else to say?

No response

### Which dataset is it? Enki dataset for Juseless ### Implementation Did not implement anything ### Your implementation _No response_ ### Dataset access restrictions - [ ] Public and open access (no registration) - [ ] Public and open access (registration required) - [X] Restricted access (needs approuval) - [X] Available only in specific locations (Juseless, Jureca) ### Anything else to say? _No response_
synchon commented 2022-10-12 06:37:35 +00:00 (Migrated from github.com)

@fraimondo Is this the eNKI one or the eNKI-pheno one?

@fraimondo Is this the eNKI one or the eNKI-pheno one?
LeSasse commented 2022-10-12 06:41:21 +00:00 (Migrated from github.com)

We interpreted it as the eNKI processed data on juseless: https://github.com/juaml/junifer/tree/add_enki_datagrabber (datagrabber here)

We interpreted it as the eNKI processed data on juseless: https://github.com/juaml/junifer/tree/add_enki_datagrabber ([datagrabber here](https://github.com/juaml/junifer/blob/add_enki_datagrabber/junifer/configs/juseless/datagrabbers/enki.py))
synchon commented 2022-10-12 06:43:28 +00:00 (Migrated from github.com)

@LeSasse Ah great, you have already started working on it. 👍

@LeSasse Ah great, you have already started working on it. 👍
verakye commented 2022-10-12 14:41:04 +00:00 (Migrated from github.com)

So far there is the anatomical and the BOLD data from the fmriprep output included. Do we also want to add the Freesurfer data?

So far there is the anatomical and the BOLD data from the fmriprep output included. Do we also want to add the Freesurfer data?
LeSasse commented 2022-10-12 15:29:07 +00:00 (Migrated from github.com)

One problem that we have identified also with the dataset at /data/project/enki/processed is that some files do not actually exist. They are shown by the file system, so the PatternDataGrabber thinks they are there, but the file cannot be accessed. One example is at /data/project/enki/processed/fmriprep/sub-A00054581/ses-CLG5/func/sub-A00054581_ses-CLG5_task-checkerboard_acq-1400_space-MNI152NLin6Asym_desc-preproc_bold.nii.gz.

Its some symlink that is shown by the file system so the datagrabber seems to find it, but then the test:
assert out["BOLD"]["path"].exists()
fails, because the file does not in fact exist?

One problem that we have identified also with the dataset at ```/data/project/enki/processed``` is that some files do not actually exist. They are shown by the file system, so the PatternDataGrabber thinks they are there, but the file cannot be accessed. One example is at ```/data/project/enki/processed/fmriprep/sub-A00054581/ses-CLG5/func/sub-A00054581_ses-CLG5_task-checkerboard_acq-1400_space-MNI152NLin6Asym_desc-preproc_bold.nii.gz```. Its some symlink that is shown by the file system so the datagrabber seems to find it, but then the test: ```assert out["BOLD"]["path"].exists()``` fails, because the file does not in fact exist?
fraimondo commented 2022-10-26 13:14:13 +00:00 (Migrated from github.com)

@LeSasse: is there a subdataset in this dataset? Is that why the file is not there?

@LeSasse: is there a subdataset in this dataset? Is that why the file is not there?
verakye commented 2022-10-26 14:34:41 +00:00 (Migrated from github.com)

As far as I know it's not because of a sub dataset but because some files were removed because the brain scans didn't pass some quality controls. However the symlinks still exist. There was a discussion with the datalad people if the dataset should be "cleaned" but for now it sounded as if this wasn't planned for the near future.

As far as I know it's not because of a sub dataset but because some files were removed because the brain scans didn't pass some quality controls. However the symlinks still exist. There was a discussion with the datalad people if the dataset should be "cleaned" but for now it sounded as if this wasn't planned for the near future.
fraimondo commented 2022-10-26 15:04:23 +00:00 (Migrated from github.com)

Then there is not much to do, unless those subjects are manually excluded. Any project dealing with this dataset needs to exclude this subjects manually. You just can't go through all the data and check if a file exists (symlink points to the right place)

Then there is not much to do, unless those subjects are manually excluded. Any project dealing with this dataset needs to exclude this subjects manually. You just can't go through all the data and check if a file exists (symlink points to the right place)
LeSasse commented 2022-11-02 13:50:39 +00:00 (Migrated from github.com)

in that case i will leave the grabber as is and make a pull request for now.

in that case i will leave the grabber as is and make a pull request for now.
LeSasse commented 2022-11-02 14:10:03 +00:00 (Migrated from github.com)

As far as I know it's not because of a sub dataset but because some files were removed because the brain scans didn't pass some quality controls. However the symlinks still exist. There was a discussion with the datalad people if the dataset should be "cleaned" but for now it sounded as if this wasn't planned for the near future.

I believe the plan to "remove" the enki dataset has been followed through as I can not find the path /data/project/enki/processed anymore. @fraimondo @verakye

> As far as I know it's not because of a sub dataset but because some files were removed because the brain scans didn't pass some quality controls. However the symlinks still exist. There was a discussion with the datalad people if the dataset should be "cleaned" but for now it sounded as if this wasn't planned for the near future. I believe the plan to "remove" the enki dataset has been followed through as I can not find the path ```/data/project/enki/processed``` anymore. @fraimondo @verakye
verakye commented 2022-11-02 14:16:27 +00:00 (Migrated from github.com)

I also can't find it anymore. I am currently waiting for Alex' reply about it.

I also can't find it anymore. I am currently waiting for Alex' reply about it.
fraimondo commented 2022-11-02 14:19:54 +00:00 (Migrated from github.com)

Then this issue will need to be on hold for the moment.

Then this issue will need to be on hold for the moment.
verakye commented 2022-11-02 17:30:28 +00:00 (Migrated from github.com)

According to Alex the future of the dataset depends "on stakeholders putting in effort to declare what they want". He hasn't heard anyone mention anything concrete recently, so for now it seems to be unclear. Here (https://jugit.fz-juelich.de/inm7/datasets/datasets_repo/-/issues/71) is the history of the dataset where people would/should report what they want for the future, so we should be able to see there if things change/upcoming plans for this dataset for the future.

According to Alex the future of the dataset depends "on stakeholders putting in effort to declare what they want". He hasn't heard anyone mention anything concrete recently, so for now it seems to be unclear. Here (https://jugit.fz-juelich.de/inm7/datasets/datasets_repo/-/issues/71) is the history of the dataset where people would/should report what they want for the future, so we should be able to see there if things change/upcoming plans for this dataset for the future.
Sign in to join this conversation.
No milestone
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
juaml/junifer#47
No description provided.