[ENH]: Add support for subject(s) file for junifer run #182
No reviewers
Labels
No labels
CRITICAL
Stale
WIP
bug
concept
coordinate
dataset
dependencies
documentation
duplicate
enhancement
github_actions
good first issue
help wanted
invalid
maintenance
maps
marker
mask
on hold
parcellation
preprocess
question
ready
storage
template-space
triage
wontfix
No milestone
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
juaml/junifer!182
Loading…
Reference in a new issue
No description provided.
Delete branch "update/allow-element-via-file"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Are you requiring a new dataset or marker?
Which feature do you want to include?
I think in most data processing projects it is desirable to be able to select specific subsets of all subjects. Sometimes one just wants to get all available data, but sometimes one may want only specific subsets (like for example only unrelated subjects or subjects matched on some other variable). In other cases one may want to just preprocess a specific subset for some initial exploration or testing before getting the full sample.
How do you imagine this integrated in junifer?
Add a subject parameter to in-built data grabbers. Ideally in the full pipeline one can add a list of subjects to the yaml directly, or for example by providing a "subject.txt" file that lists all desired subjects. Something like that.
Do you have a sample code that implements this outside of junifer?
No response
Anything else to say?
No response
This was thought already:
github.com/juaml/junifer@ace98da340/junifer/api/cli.py (L54)The only thing we need to define is the file format.
Main issue is elements as tuples.
For example, using HCP1200 datagrabber, we can set in the datagrabber
task="REST1"and then the iterator of the datagrabber will go through all the element tuples, but only for the specified task:Setting the elements on the CLI or the YAML is just a way of bypassing the iterator. So in this case, the .CSV file with the elements to process should contain all the keys of the element (
"subject","task","phase_encoding").Possible solutions:
filterfunction:then these lines in the run function:
github.com/juaml/junifer@ace98da340/junifer/api/functions.py (L165-L172)would change for something like:
What you say makes sense, but I really only mean like setting one parameter of the element, i.e. mostly for subjects, to a list. The rest of the element can still be constructed. So for example, as a user I dont want to construct all elements, but just provide a list of subjects, for example of the HCP in a txt file. Junifer can then construct the elements, for example by using the default values for all other parameters (i.e. [LR, RL] and [REST1..., LAST_TASK]), without me having to make an iterator/file over all elements myself. The reasoning is that subject lists can be quite long and are therefore more difficult to set in the YAML file than other parameters. Having said that, making a csv file of all elements is also easy enough, but in my opinion puts more burden on the user (also in terms of data discovery).
Passing only the subject IDs makes sense to me as this will make it simple for users and also enable one to filter what the
DataGrabberiterates over in combination with other parameters that oneDataGrabberprovides, for example,types,tasksand so on. From the implementation perspective, we already have a good way as @fraimondo shows.Then let's move on that way.
Codecov Report
100.00% <ø> (ø)93.04% <96.42%> (+0.01%)Flags with carried forward coverage won't be shown. Click here to find out more.
71.28% <100.00%> (+2.13%)96.23% <100.00%> (ø)98.36% <93.75%> (-1.64%)While this works, it's not quite robust. Splitting elements by
,should be done only if there's no file involved.@ -61,19 +80,49 @@ def _parse_elements(element: str, config: Dict) -> Union[List, None]:"over the configuration file. That is, the elements specified ""in the command line will be used. The elements specified in ""the configuration file will be ignored. To remove this warning, "This function should parse the file completely, giving the list of tuples required.
Why this?
Well, because we can have things like:
sub-01,ses-01orsub-01, ses-01orsub-01, ses-01.Using pandas to parse the CSV file would be easier.
@ -61,19 +80,49 @@ def _parse_elements(element: str, config: Dict) -> Union[List, None]:"over the configuration file. That is, the elements specified ""in the command line will be used. The elements specified in ""the configuration file will be ignored. To remove this warning, "This worked but I anyway changed it to use
pandasas you suggested.Addressed.
@ -77,0 +117,4 @@contents["workdir"] = str(workdir.resolve())# Output directoryoutdir = tmp_path / "outdir"# StorageThis is not fully testing all the options.
What if the user wants to specify
subjectandsession? Can you test that, it should be a comma-separated list.@ -77,0 +117,4 @@contents["workdir"] = str(workdir.resolve())# Output directoryoutdir = tmp_path / "outdir"# StorageShould be addressed now.