refactor: consistent use of DataGrabber #226

Merged
synchon merged 32 commits from chore/dg-cleanup into main 2023-06-20 10:27:52 +00:00
48 changed files with 442 additions and 410 deletions

View file

@ -0,0 +1 @@
Adopt ``DataGrabber`` consistently throughout codebase to match with the documentation by `Synchon Mandal`_

View file

@ -22,7 +22,7 @@ data type including source and previous transformation steps.
The :ref:`Data Grabber <datagrabber>` step adds the ``path`` second-level key The :ref:`Data Grabber <datagrabber>` step adds the ``path`` second-level key
which gives the path to the file containing the data. The ``meta`` key in this which gives the path to the file containing the data. The ``meta`` key in this
step only contains information about the datagrabber used. step only contains information about the DataGrabber used.
.. code-block:: python .. code-block:: python

View file

@ -9,29 +9,29 @@ Description
----------- -----------
The ``DataGrabber`` is an object that can provide an interface to datasets you The ``DataGrabber`` is an object that can provide an interface to datasets you
want to work with in junifer. Every concrete implementation of a datagrabber is want to work with in junifer. Every concrete implementation of a DataGrabber is
aware of a particular dataset's structure and thus allows you to fetch specific aware of a particular dataset's structure and thus allows you to fetch specific
elements of interest from the dataset. It adds the ``path`` key to each elements of interest from the dataset. It adds the ``path`` key to each
:ref:`data type <data_types>` in the :ref:`Data object <data_object>`. :ref:`data type <data_types>` in the :ref:`Data object <data_object>`.
Datagrabbers are intended to be used as context managers. When used within a DataGrabbers are intended to be used as context managers. When used within a
context, a datagrabber takes care of any pre and post steps for interacting with context, a DataGrabber takes care of any pre and post steps for interacting with
the dataset, for example, downloading and cleaning up. As the interface the dataset, for example, downloading and cleaning up. As the interface
is consistent, you always use the same procedure to interact with the datagrabber. is consistent, you always use the same procedure to interact with the DataGrabber.
For example, a concrete implementation of :class:`.DataladDataGrabber` can For example, a concrete implementation of :class:`.DataladDataGrabber` can
provide junifer with data from a Datalad dataset. Of course, datagrabbers are not provide junifer with data from a Datalad dataset. Of course, DataGrabbers are not
only meant to work with Datalad datasets but any dataset. only meant to work with Datalad datasets but any dataset.
If you are interested in using already provided datagrabbers, please go to If you are interested in using already provided DataGrabbers, please go to
:doc:`../builtin`. And, if you want to implement your own datagrabber, you need :doc:`../builtin`. And, if you want to implement your own DataGrabber, you need
to provide concrete implementations of base classes already provided. to provide concrete implementations of base classes already provided.
Base Classes Base Classes
------------ ------------
In this section, we showcase different abstract base classes you might want to In this section, we showcase different abstract base classes you might want to
use to implement your own datagrabber. use to implement your own DataGrabber.
.. list-table:: .. list-table::
:widths: auto :widths: auto
@ -41,16 +41,16 @@ use to implement your own datagrabber.
- Description - Description
* - :class:`.BaseDataGrabber` * - :class:`.BaseDataGrabber`
- | The abstract base class providing you an interface to implement your - | The abstract base class providing you an interface to implement your
| own datagrabber. You should try to avoid using this directly and | own DataGrabber. You should try to avoid using this directly and
| instead use :class:`.PatternDataGrabber` or | instead use :class:`.PatternDataGrabber` or
| :class:`.DataladDataGrabber`. To build your own custom *low-level* | :class:`.DataladDataGrabber`. To build your own custom *low-level*
| datagrabber, you need to override the ``get_elements_keys``, | DataGrabber, you need to override the ``get_elements_keys``,
| ``get_elements`` and ``get_item`` methods, and most of the time you | ``get_elements`` and ``get_item`` methods, and most of the time you
| should also override other existing methods like ``__enter__`` and | should also override other existing methods like ``__enter__`` and
| ``__exit__``. | ``__exit__``.
* - :class:`.PatternDataGrabber` * - :class:`.PatternDataGrabber`
- | It implements functionality to help you define the pattern of the - | It implements functionality to help you define the pattern of the
| dataset you want to get. For example, you know that T1 images are | dataset you want to get. For example, you know that T1w images are
| found in a directory following the pattern: | found in a directory following the pattern:
| ``{subject}/anat/{subject}_T1w.nii.gz`` inside of the dataset. Now you | ``{subject}/anat/{subject}_T1w.nii.gz`` inside of the dataset. Now you
| can provide this to the :class:`.PatternDataGrabber` and it will be | can provide this to the :class:`.PatternDataGrabber` and it will be

View file

@ -75,10 +75,10 @@ Data Grabber
^^^^^^^^^^^^ ^^^^^^^^^^^^
The ``datagrabber`` section must be configured using the ``kind`` key to specify The ``datagrabber`` section must be configured using the ``kind`` key to specify
the datagrabber to use. Additional keys correspond to the parameters of the the DataGrabber to use. Additional keys correspond to the parameters of the
datagrabber. DataGrabber constructor.
For example, to use the :class:`.DataladAOMICPIOP1` datagrabber, we just need to For example, to use the :class:`.DataladAOMICPIOP1` DataGrabber, we just need to
specify its name as the ``kind`` key. specify its name as the ``kind`` key.
.. code-block:: yaml .. code-block:: yaml
@ -86,8 +86,9 @@ specify its name as the ``kind`` key.
datagrabber: datagrabber:
kind: DataladAOMICPIOP1 kind: DataladAOMICPIOP1
However, it is also possible to pass parameters to the datagrabber. In this case, However, it is also possible to pass parameters to the DataGrabber constructor.
we can restrict the datagrabber to fetch only the ``restingstate`` task. In this case, we can restrict the DataGrabber to fetch only the ``restingstate``
task.
.. code-block:: yaml .. code-block:: yaml

View file

@ -1,8 +1,8 @@
""" """
Generic BIDS datagrabber for datalad. Generic BIDS DataGrabber for datalad.
===================================== =====================================
This example uses a generic BIDS datagraber to get the data from a BIDS dataset This example uses a generic BIDS DataGraber to get the data from a BIDS dataset
store in a datalad remote sibling. store in a datalad remote sibling.
Authors: Federico Raimondo Authors: Federico Raimondo
@ -20,9 +20,9 @@ configure_logging(level="INFO")
############################################################################### ###############################################################################
# The BIDS datagrabber requires three parameters: the types of data we want, # The BIDS DataGrabber requires three parameters: the types of data we want,
# the specific pattern that matches each type, and the variables that will be # the specific pattern that matches each type, and the variables that will be
# replaced int he patterns. # replaced in the patterns.
types = ["T1w", "BOLD"] types = ["T1w", "BOLD"]
patterns = { patterns = {
"T1w": "{subject}/anat/{subject}_T1w.nii.gz", "T1w": "{subject}/anat/{subject}_T1w.nii.gz",
@ -30,15 +30,15 @@ patterns = {
} }
replacements = ["subject"] replacements = ["subject"]
############################################################################### ###############################################################################
# Additionally, a datalad datagrabber requires the URI of the remote sibling # Additionally, a datalad-based DataGrabber requires the URI of the remote
# and the location of the dataset within the remote sibling. # sibling and the location of the dataset within the remote sibling.
repo_uri = "https://gin.g-node.org/juaml/datalad-example-bids" repo_uri = "https://gin.g-node.org/juaml/datalad-example-bids"
rootdir = "example_bids" rootdir = "example_bids"
############################################################################### ###############################################################################
# Now we can use the datagrabber within a `with` context # Now we can use the DataGrabber within a `with` context.
# One thing we can do with any datagrabber is iterate over the elements. # One thing we can do with any DataGrabber is iterate over the elements.
# In this case, each element of the datagrabber is one session. # In this case, each element of the DataGrabber is one session.
with PatternDataladDataGrabber( with PatternDataladDataGrabber(
rootdir=rootdir, rootdir=rootdir,
types=types, types=types,
@ -50,9 +50,9 @@ with PatternDataladDataGrabber(
print(elem) print(elem)
############################################################################### ###############################################################################
# Another feature of the datagrabber is the ability to get a specific # Another feature of the DataGrabber is the ability to get a specific
# element by its name. In this case, we index `sub-01` and we get the file # element by its name. In this case, we index `sub-01` and we get the file
# paths for the two types of data we want (T1w and bold). # paths for the two types of data we want (T1w and BOLD).
with PatternDataladDataGrabber( with PatternDataladDataGrabber(
rootdir=rootdir, rootdir=rootdir,
types=types, types=types,

View file

@ -25,7 +25,7 @@ from junifer.utils import configure_logging
configure_logging(level="INFO") configure_logging(level="INFO")
############################################################################## ##############################################################################
# Define the datagrabber interface # Define the DataGrabber interface
datagrabber = { datagrabber = {
"kind": "SPMAuditoryTestingDataGrabber", "kind": "SPMAuditoryTestingDataGrabber",
} }

View file

@ -11,20 +11,24 @@ from ..pipeline.registry import register
def register_datagrabber(klass: Type) -> Type: def register_datagrabber(klass: Type) -> Type:
"""Datagrabber registration decorator. """Register DataGrabber.
Registers the datagrabber so it can be used by name. Registers the DataGrabber so it can be used by name.
Parameters Parameters
---------- ----------
klass: class klass: class
The class of the datagrabber to register. The class of the DataGrabber to register.
Returns Returns
------- -------
klass: class klass: class
The unmodified input class. The unmodified input class.
Notes
-----
It should only be used as a decorator.
""" """
register( register(
step="datagrabber", step="datagrabber",

View file

@ -24,17 +24,17 @@ from .utils import yaml
def _get_datagrabber(datagrabber_config: Dict) -> BaseDataGrabber: def _get_datagrabber(datagrabber_config: Dict) -> BaseDataGrabber:
"""Get datagrabber. """Get DataGrabber.
Parameters Parameters
---------- ----------
datagrabber_config : dict datagrabber_config : dict
The config to get the datagrabber using. The config to get the DataGrabber using.
Returns Returns
------- -------
dict object
The datagrabber. The DataGrabber.
""" """
datagrabber_params = datagrabber_config.copy() datagrabber_params = datagrabber_config.copy()
@ -90,8 +90,8 @@ def run(
workdir : str or pathlib.Path workdir : str or pathlib.Path
Directory where the pipeline will be executed. Directory where the pipeline will be executed.
datagrabber : dict datagrabber : dict
Datagrabber to use. Must have a key ``kind`` with the kind of DataGrabber to use. Must have a key ``kind`` with the kind of
datagrabber to use. All other keys are passed to the datagrabber DataGrabber to use. All other keys are passed to the DataGrabber
init function. init function.
markers : list of dict markers : list of dict
List of markers to extract. Each marker is a dict with at least two List of markers to extract. Each marker is a dict with at least two
@ -108,7 +108,7 @@ def run(
preprocessor to use. All other keys are passed to the preprocessor preprocessor to use. All other keys are passed to the preprocessor
init function (default None). init function (default None).
elements : str or tuple or list of str or tuple, optional elements : str or tuple or list of str or tuple, optional
Element(s) to process. Will be used to index the datagrabber Element(s) to process. Will be used to index the DataGrabber
(default None). (default None).
""" """
@ -220,7 +220,7 @@ def queue(
overwrite : bool, optional overwrite : bool, optional
Whether to overwrite if job directory already exists (default False). Whether to overwrite if job directory already exists (default False).
elements : str or tuple or list of str or tuple, optional elements : str or tuple or list of str or tuple, optional
Element(s) to process. Will be used to index the datagrabber Element(s) to process. Will be used to index the DataGrabber
(default None). (default None).
**kwargs : dict **kwargs : dict
The keyword arguments to pass to the job queue system. The keyword arguments to pass to the job queue system.
@ -339,7 +339,7 @@ def _queue_condor(
yaml_config : pathlib.Path yaml_config : pathlib.Path
The path to the YAML config file. The path to the YAML config file.
elements : list of str or tuple elements : list of str or tuple
Element(s) to process. Will be used to index the datagrabber. Element(s) to process. Will be used to index the DataGrabber.
config : dict config : dict
The configuration to be used for queueing the job. The configuration to be used for queueing the job.
env : dict, optional env : dict, optional
@ -559,7 +559,7 @@ def _queue_slurm(
yaml_config : pathlib.Path yaml_config : pathlib.Path
The path to the YAML config file. The path to the YAML config file.
elements : str or tuple or list[str or tuple], optional elements : str or tuple or list[str or tuple], optional
Element(s) to process. Will be used to index the datagrabber Element(s) to process. Will be used to index the DataGrabber
(default None). (default None).
config : dict config : dict
The configuration to be used for queueing the job. The configuration to be used for queueing the job.

View file

@ -1,4 +1,4 @@
"""Provide class for AOMIC1000 VBM juseless datalad datagrabber.""" """Provide concrete implementation for AOMIC ID1000 VBM DataGrabber."""
# Authors: Felix Hoffstaedter <f.hoffstaedter@fz-juelich.de> # Authors: Felix Hoffstaedter <f.hoffstaedter@fz-juelich.de>
# Synchon Mandal <s.mandal@fz-juelich.de> # Synchon Mandal <s.mandal@fz-juelich.de>
@ -13,13 +13,13 @@ from ....datagrabber import PatternDataladDataGrabber
@register_datagrabber @register_datagrabber
class JuselessDataladAOMICID1000VBM(PatternDataladDataGrabber): class JuselessDataladAOMICID1000VBM(PatternDataladDataGrabber):
"""Juseless AOMICID1000 VBM Data Grabber class. """Concrete implementation for Juseless AOMIC ID1000 VBM data fetching.
Implements a Data Grabber to access the AOMICID1000 VBM data in Juseless. Implements a DataGrabber to access the AOMIC ID1000 VBM data in Juseless.
Parameters Parameters
---------- ----------
datadir : str or pathlib.Path, optional datadir : str or pathlib.Path or None, optional
The directory where the datalad dataset will be cloned. If None, The directory where the datalad dataset will be cloned. If None,
the datalad dataset will be cloned into a temporary directory the datalad dataset will be cloned into a temporary directory
(default None). (default None).

View file

@ -1,4 +1,4 @@
"""Provide class for CamCAN VBM juseless datalad datagrabber.""" """Provide concrete implementation for CamCAN VBM DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de> # Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de> # Leonard Sasse <l.sasse@fz-juelich.de>
@ -14,13 +14,13 @@ from ....datagrabber import PatternDataladDataGrabber
@register_datagrabber @register_datagrabber
class JuselessDataladCamCANVBM(PatternDataladDataGrabber): class JuselessDataladCamCANVBM(PatternDataladDataGrabber):
"""Juseless CamCAN VBM Data Grabber class. """Concrete implementation for Juseless CamCAN VBM data fetching.
Implements a Data Grabber to access the CamCAN VBM data in Juseless. Implements a DataGrabber to access the CamCAN VBM data in Juseless.
Parameters Parameters
---------- ----------
datadir : str or pathlib.Path, optional datadir : str or pathlib.Path or None, optional
The directory where the datalad dataset will be cloned. If None, The directory where the datalad dataset will be cloned. If None,
the datalad dataset will be cloned into a temporary directory the datalad dataset will be cloned into a temporary directory
(default None). (default None).

View file

@ -1,4 +1,4 @@
"""Provide class for IXI VBM juseless datalad datagrabber.""" """Provide concrete implementation for IXI VBM DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de> # Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de> # Leonard Sasse <l.sasse@fz-juelich.de>
@ -15,19 +15,20 @@ from ....utils import raise_error
@register_datagrabber @register_datagrabber
class JuselessDataladIXIVBM(PatternDataladDataGrabber): class JuselessDataladIXIVBM(PatternDataladDataGrabber):
"""Juseless IXI VBM Data Grabber class. """Concrete implementation for Juseless IXI VBM data fetching.
Implements a Data Grabber to access the IXI VBM data in Juseless. Implements a DataGrabber to access the IXI VBM data in Juseless.
Parameters Parameters
---------- ----------
datadir : str or pathlib.Path, optional datadir : str or pathlib.Path or None, optional
The directory where the datalad dataset will be cloned. If None, The directory where the datalad dataset will be cloned. If None,
the datalad dataset will be cloned into a temporary directory the datalad dataset will be cloned into a temporary directory
(default None). (default None).
sites : {"Guys", "HH", "IOP"} or list of the options, optional. sites : {"Guys", "HH", "IOP"} or list of the options or None, optional
Which sites to access data from. If None, all available sites are Which sites to access data from. If None, all available sites are
selected (default None). selected (default None).
""" """
def __init__( def __init__(

View file

@ -1,4 +1,4 @@
"""Provide tests for AOMICID1000 VBM juseless datagrabber.""" """Provide tests for JuselessDataladAOMICID1000VBM."""
# Authors: Felix Hoffstaedter <f.hoffstaedter@fz-juelich.de> # Authors: Felix Hoffstaedter <f.hoffstaedter@fz-juelich.de>
# Synchon Mandal <s.mandal@fz-juelich.de> # Synchon Mandal <s.mandal@fz-juelich.de>
@ -19,8 +19,8 @@ if socket.gethostname() != "juseless":
configure_logging(level="DEBUG") configure_logging(level="DEBUG")
def test_juselessdataladaomicid1000vbm_datagrabber() -> None: def test_JuselessDataladAOMICID1000VBM() -> None:
"""Test datalad AOMICID1000VBM datagrabber.""" """Test JuselessDataladAOMICID1000VBM."""
with JuselessDataladAOMICID1000VBM() as dg: with JuselessDataladAOMICID1000VBM() as dg:
all_elements = dg.get_elements() all_elements = dg.get_elements()
test_element = all_elements[0] test_element = all_elements[0]

View file

@ -1,4 +1,4 @@
"""Provide tests for CamCAN VBM juseless datagrabber.""" """Provide tests for JuselessDataladCamCANVBM."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de> # Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de> # Leonard Sasse <l.sasse@fz-juelich.de>
@ -20,8 +20,8 @@ if socket.gethostname() != "juseless":
configure_logging(level="DEBUG") configure_logging(level="DEBUG")
def test_juselessdataladcamcanvbm_datagrabber() -> None: def test_JuselessDataladCamCANVBM() -> None:
"""Test datalad CamCANVBM datagrabber.""" """Test JuselessDataladCamCANVBM."""
with JuselessDataladCamCANVBM() as dg: with JuselessDataladCamCANVBM() as dg:
all_elements = dg.get_elements() all_elements = dg.get_elements()
test_element = all_elements[0] test_element = all_elements[0]

View file

@ -1,4 +1,4 @@
"""Provide tests for IXI VBM juseless datagrabber.""" """Provide tests for JuselessDataladIXIVBM."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de> # Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de> # Leonard Sasse <l.sasse@fz-juelich.de>
@ -20,8 +20,8 @@ if socket.gethostname() != "juseless":
configure_logging(level="DEBUG") configure_logging(level="DEBUG")
def test_juselessdataladixivbm_datagrabber() -> None: def test_JuselessDataladIXIVBM() -> None:
"""Test datalad IXIVBM datagrabber.""" """Test JuselessDataladIXIVBM."""
with JuselessDataladIXIVBM() as dg: with JuselessDataladIXIVBM() as dg:
all_elements = dg.get_elements() all_elements = dg.get_elements()
test_element = all_elements[0] test_element = all_elements[0]
@ -33,8 +33,8 @@ def test_juselessdataladixivbm_datagrabber() -> None:
assert out["VBM_GM"]["path"].exists() assert out["VBM_GM"]["path"].exists()
def test_juselessdataladixivbm_datagrabber_invalid_site() -> None: def test_JuselessDataladIXIVBM_invalid_site() -> None:
"""Test datalad IXIVBM datagrabber with invalid site.""" """Test JuselessDataladIXIVBM with invalid site."""
with pytest.raises(ValueError, match="notavalidsite not a valid site"): with pytest.raises(ValueError, match="notavalidsite not a valid site"):
with JuselessDataladIXIVBM(sites="notavalidsite"): with JuselessDataladIXIVBM(sites="notavalidsite"):
pass pass

View file

@ -1,4 +1,4 @@
"""Provide tests for UCLA juseless datagrabber.""" """Provide tests for JuselessUCLA."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de> # Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de> # Leonard Sasse <l.sasse@fz-juelich.de>
@ -21,8 +21,8 @@ if socket.gethostname() != "juseless":
configure_logging(level="DEBUG") configure_logging(level="DEBUG")
def test_juseless_ucla_datagrabber() -> None: def test_JuselessUCLA() -> None:
"""Test juseless ucla datagrabber.""" """Test JuselessUCLA."""
with JuselessUCLA() as dg: with JuselessUCLA() as dg:
all_elements = dg.get_elements() all_elements = dg.get_elements()
test_element = all_elements[0] test_element = all_elements[0]
@ -46,8 +46,8 @@ def test_juseless_ucla_datagrabber() -> None:
"tasks", "tasks",
[None, "rest", ["rest", "stopsignal"]], [None, "rest", ["rest", "stopsignal"]],
) )
def test_juseless_ucla_datagrabber_task_params(tasks: Optional[str]) -> None: def test_JuselessUCLA_task_params(tasks: Optional[str]) -> None:
"""Test juseless ucla datagrabber with different task parameters. """Test JuselessUCLA with different task parameters.
Parameters Parameters
---------- ----------
@ -81,8 +81,8 @@ def test_juseless_ucla_datagrabber_task_params(tasks: Optional[str]) -> None:
assert el[1] in ["rest", "stopsignal"] assert el[1] in ["rest", "stopsignal"]
def test_juseless_ucla_datagrabber_invalid_tasks() -> None: def test_JuselessUCLA_invalid_tasks() -> None:
"""Test juseless ucla datagrabber with invalid task parameters.""" """Test JuselessUCLA with invalid task parameters."""
with pytest.raises( with pytest.raises(
ValueError, match="invalid is not a valid task in the UCLA" ValueError, match="invalid is not a valid task in the UCLA"
): ):

View file

@ -1,4 +1,4 @@
"""Provide tests for juseless datagrabber.""" """Provide tests for JuselessDataladUKBVBM."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de> # Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de> # Leonard Sasse <l.sasse@fz-juelich.de>
@ -20,8 +20,8 @@ if socket.gethostname() != "juseless":
configure_logging(level="DEBUG") configure_logging(level="DEBUG")
def test_juselessdataladukbvbm_datagrabber() -> None: def test_JuselessDataladUKBVBM() -> None:
"""Test datalad UKBVBM datagrabber.""" """Test JuselessDataladUKBVBM."""
with JuselessDataladUKBVBM() as dg: with JuselessDataladUKBVBM() as dg:
all_elements = dg.get_elements() all_elements = dg.get_elements()
test_element = all_elements[0] test_element = all_elements[0]

View file

@ -1,4 +1,4 @@
"""Provide a concrete implementation for UCLA dataset.""" """Provide concrete implementation for UCLA DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de> # Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de> # Leonard Sasse <l.sasse@fz-juelich.de>
@ -14,16 +14,18 @@ from ....utils import raise_error
@register_datagrabber @register_datagrabber
class JuselessUCLA(PatternDataGrabber): class JuselessUCLA(PatternDataGrabber):
"""Concrete implementation for pattern-based data fetching of UCLA data. """Concrete implementation for Juseless UCLA data fetching.
Implements a DataGrabber to access the UCLA data in Juseless.
Parameters Parameters
---------- ----------
datadir : str or Path, optional datadir : str or pathlib.Path, optional
The directory where the dataset is stored The directory where the dataset is stored
(default '/data/project/psychosis_thalamus/data/fmriprep'). (default "/data/project/psychosis_thalamus/data/fmriprep").
tasks : {"rest", "bart", "bht", "pamenc", "pamret", \ tasks : {"rest", "bart", "bht", "pamenc", "pamret", \
"scap", "taskswitch", "stopsignal"} or \ "scap", "taskswitch", "stopsignal"} or \
list of the options, optional list of the options or None, optional
UCLA task sessions. If None, all available task sessions are UCLA task sessions. If None, all available task sessions are
selected (default None). selected (default None).
@ -118,5 +120,6 @@ class JuselessUCLA(PatternDataGrabber):
elements : list elements : list
The list of elements that can be grabbed in the dataset after The list of elements that can be grabbed in the dataset after
imposing constraints based on specified tasks. imposing constraints based on specified tasks.
""" """
return [x for x in super().get_elements() if x[1] in self.tasks] return [x for x in super().get_elements() if x[1] in self.tasks]

View file

@ -1,4 +1,4 @@
"""Provide class for juseless datalad datagrabber.""" """Provide concrete implementation for UKB VBM DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de> # Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de> # Leonard Sasse <l.sasse@fz-juelich.de>
@ -14,13 +14,13 @@ from ....datagrabber import PatternDataladDataGrabber
@register_datagrabber @register_datagrabber
class JuselessDataladUKBVBM(PatternDataladDataGrabber): class JuselessDataladUKBVBM(PatternDataladDataGrabber):
"""Juseless UKB VBM Data Grabber class. """Concrete implementation for Juseless UKB VBM data fetching.
Implements a Data Grabber to access the UKB VBM data in Juseless. Implements a DataGrabber to access the UKB VBM data in Juseless.
Parameters Parameters
---------- ----------
datadir : str or pathlib.Path, optional datadir : str or pathlib.Path or None, optional
The directory where the datalad dataset will be cloned. If None, The directory where the datalad dataset will be cloned. If None,
the datalad dataset will be cloned into a temporary directory the datalad dataset will be cloned into a temporary directory
(default None). (default None).

View file

@ -1,32 +0,0 @@
"""Provide tests for juseless datagrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de>
# Synchon Mandal <s.mandal@fz-juelich.de>
# License: AGPL
import socket
import pytest
from junifer.datagrabber.hcp import DataladHCP1200
from junifer.utils.logging import configure_logging
# Check if the test is running on juseless
if socket.gethostname() != "juseless":
pytest.skip("These tests are only for juseless", allow_module_level=True)
configure_logging(level="DEBUG")
def test_juselessdataladhcp_datagrabber() -> None:
"""Test datalad HCP datagrabber."""
with DataladHCP1200() as dg:
all_elements = dg.get_elements()
test_element = all_elements[0]
out = dg[test_element]
assert out["BOLD"]["path"].exists()
assert out["BOLD"]["path"].isfile()

View file

@ -2,6 +2,7 @@
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de> # Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de> # Leonard Sasse <l.sasse@fz-juelich.de>
# Synchon Mandal <s.mandal@fz-juelich.de>
# License: AGPL # License: AGPL
@ -12,5 +13,5 @@ from .pattern import PatternDataGrabber
from .pattern_datalad import PatternDataladDataGrabber from .pattern_datalad import PatternDataladDataGrabber
from .aomic import DataladAOMICID1000, DataladAOMICPIOP1, DataladAOMICPIOP2 from .aomic import DataladAOMICID1000, DataladAOMICPIOP1, DataladAOMICPIOP2
from .hcp import HCP1200, DataladHCP1200 from .hcp1200 import HCP1200, DataladHCP1200
from .multiple import MultipleDataGrabber from .multiple import MultipleDataGrabber

View file

@ -1,4 +1,4 @@
"""Provide imports for datagrabber 'aomic' sub-package.""" """Provide imports for aomic sub-package."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de> # Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de> # Leonard Sasse <l.sasse@fz-juelich.de>

View file

@ -1,4 +1,4 @@
"""Provide concrete implementations for AOMIC1000 data access.""" """Provide concrete implementation for AOMIC ID1000 DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de> # Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Vera Komeyer <v.komeyer@fz-juelich.de> # Vera Komeyer <v.komeyer@fz-juelich.de>
@ -9,21 +9,21 @@
from pathlib import Path from pathlib import Path
from typing import Dict, Union from typing import Dict, Union
from junifer.datagrabber import PatternDataladDataGrabber
from ...api.decorators import register_datagrabber from ...api.decorators import register_datagrabber
from ..pattern_datalad import PatternDataladDataGrabber
@register_datagrabber @register_datagrabber
class DataladAOMICID1000(PatternDataladDataGrabber): class DataladAOMICID1000(PatternDataladDataGrabber):
"""Concrete implementation for pattern-based data fetching of AOMICID1000. """Concrete implementation for datalad-based data fetching of AOMIC ID1000.
Parameters Parameters
---------- ----------
datadir : str or Path, optional datadir : str or Path or None, optional
The directory where the datalad dataset will be cloned. If None, The directory where the datalad dataset will be cloned. If None,
the datalad dataset will be cloned into a temporary directory the datalad dataset will be cloned into a temporary directory
(default None). (default None).
""" """
def __init__( def __init__(
@ -115,6 +115,7 @@ class DataladAOMICID1000(PatternDataladDataGrabber):
out : dict out : dict
Dictionary of paths for each type of data required for the Dictionary of paths for each type of data required for the
specified element. specified element.
""" """
out = super().get_item(subject=subject) out = super().get_item(subject=subject)
out["BOLD"]["mask_item"] = "BOLD_mask" out["BOLD"]["mask_item"] = "BOLD_mask"

View file

@ -1,4 +1,4 @@
"""Provide concrete implementations for AOMICPIOP1 data access.""" """Provide concrete implementation for AOMIC PIOP1 DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de> # Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Vera Komeyer <v.komeyer@fz-juelich.de> # Vera Komeyer <v.komeyer@fz-juelich.de>
@ -10,19 +10,18 @@ from itertools import product
from pathlib import Path from pathlib import Path
from typing import Dict, List, Union from typing import Dict, List, Union
from junifer.datagrabber import PatternDataladDataGrabber
from ...api.decorators import register_datagrabber from ...api.decorators import register_datagrabber
from ...utils import raise_error from ...utils import raise_error
from ..pattern_datalad import PatternDataladDataGrabber
@register_datagrabber @register_datagrabber
class DataladAOMICPIOP1(PatternDataladDataGrabber): class DataladAOMICPIOP1(PatternDataladDataGrabber):
"""Concrete implementation for pattern-based data fetching of AOMICPIOP1. """Concrete implementation for pattern-based data fetching of AOMIC PIOP1.
Parameters Parameters
---------- ----------
datadir : str or Path, optional datadir : str or Path or None, optional
The directory where the datalad dataset will be cloned. If None, The directory where the datalad dataset will be cloned. If None,
the datalad dataset will be cloned into a temporary directory the datalad dataset will be cloned into a temporary directory
(default None). (default None).
@ -30,6 +29,7 @@ class DataladAOMICPIOP1(PatternDataladDataGrabber):
"gstroop", "workingmemory"} or list of the options, optional "gstroop", "workingmemory"} or list of the options, optional
AOMIC PIOP1 task sessions. If None, all available task sessions are AOMIC PIOP1 task sessions. If None, all available task sessions are
selected (default None). selected (default None).
""" """
def __init__( def __init__(
@ -147,6 +147,7 @@ class DataladAOMICPIOP1(PatternDataladDataGrabber):
out : dict out : dict
Dictionary of paths for each type of data required for the Dictionary of paths for each type of data required for the
specified element. specified element.
""" """
task_acqs = { task_acqs = {

View file

@ -1,4 +1,4 @@
"""Provide concrete implementations for AOMICPIOP2 data access.""" """Provide concrete implementation for AOMIC PIOP2 DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de> # Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Vera Komeyer <v.komeyer@fz-juelich.de> # Vera Komeyer <v.komeyer@fz-juelich.de>
@ -9,19 +9,18 @@
from pathlib import Path from pathlib import Path
from typing import Dict, List, Union from typing import Dict, List, Union
from junifer.datagrabber import PatternDataladDataGrabber
from ...api.decorators import register_datagrabber from ...api.decorators import register_datagrabber
from ...utils import raise_error from ...utils import raise_error
from ..pattern_datalad import PatternDataladDataGrabber
@register_datagrabber @register_datagrabber
class DataladAOMICPIOP2(PatternDataladDataGrabber): class DataladAOMICPIOP2(PatternDataladDataGrabber):
"""Concrete implementation for pattern-based data fetching of AOMICPIOP2. """Concrete implementation for pattern-based data fetching of AOMIC PIOP2.
Parameters Parameters
---------- ----------
datadir : str or Path, optional datadir : str or Path or None, optional
The directory where the datalad dataset will be cloned. If None, The directory where the datalad dataset will be cloned. If None,
the datalad dataset will be cloned into a temporary directory the datalad dataset will be cloned into a temporary directory
(default None). (default None).
@ -29,6 +28,7 @@ class DataladAOMICPIOP2(PatternDataladDataGrabber):
or list of the options, optional or list of the options, optional
AOMIC PIOP2 task sessions. If None, all available task sessions are AOMIC PIOP2 task sessions. If None, all available task sessions are
selected (default None). selected (default None).
""" """
def __init__( def __init__(
@ -136,6 +136,7 @@ class DataladAOMICPIOP2(PatternDataladDataGrabber):
list list
The list of elements that can be grabbed in the dataset after The list of elements that can be grabbed in the dataset after
imposing constraints based on specified tasks. imposing constraints based on specified tasks.
""" """
all_elements = super().get_elements() all_elements = super().get_elements()
return [x for x in all_elements if x[1] in self.tasks] return [x for x in all_elements if x[1] in self.tasks]
@ -156,6 +157,7 @@ class DataladAOMICPIOP2(PatternDataladDataGrabber):
out : dict out : dict
Dictionary of paths for each type of data required for the Dictionary of paths for each type of data required for the
specified element. specified element.
""" """
out = super().get_item(subject=subject, task=task) out = super().get_item(subject=subject, task=task)
out["BOLD"]["mask_item"] = "BOLD_mask" out["BOLD"]["mask_item"] = "BOLD_mask"

View file

@ -1,4 +1,4 @@
"""Provide tests for aomicid1000.""" """Provide tests for DataladAOMICID1000 DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de> # Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Vera Komeyer <v.komeyer@fz-juelich.de> # Vera Komeyer <v.komeyer@fz-juelich.de>
@ -6,12 +6,12 @@
# Leonard Sasse <l.sasse@fz-juelich.de> # Leonard Sasse <l.sasse@fz-juelich.de>
# License: AGPL # License: AGPL
from junifer.datagrabber.aomic.id1000 import DataladAOMICID1000 from junifer.datagrabber import DataladAOMICID1000
from junifer.utils import configure_logging from junifer.utils import configure_logging
def test_aomic1000_datagrabber() -> None: def test_DataladAOMICID1000() -> None:
"""Test datalad AOMIC1000 datagrabber.""" """Test DataladAOMICID1000 DataGrabber."""
uri_ID1000 = "https://gin.g-node.org/juaml/datalad-example-aomic1000" uri_ID1000 = "https://gin.g-node.org/juaml/datalad-example-aomic1000"
configure_logging(level="DEBUG") configure_logging(level="DEBUG")

View file

@ -1,4 +1,4 @@
"""Provide tests for aomic piop1.""" """Provide tests for DataladAOMICPIOP1 DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de> # Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Vera Komeyer <v.komeyer@fz-juelich.de> # Vera Komeyer <v.komeyer@fz-juelich.de>
@ -8,12 +8,12 @@
import pytest import pytest
from junifer.datagrabber.aomic.piop1 import DataladAOMICPIOP1 from junifer.datagrabber import DataladAOMICPIOP1
from junifer.utils import configure_logging from junifer.utils import configure_logging
def test_aomic_piop1_datagrabber() -> None: def test_DataladAOMICPIOP1() -> None:
"""Test datalad AOMICPIOP1 datagrabber.""" """Test DataladAOMICPIOP1 DataGrabber."""
configure_logging(level="DEBUG") configure_logging(level="DEBUG")
uri_PIOP1 = "https://gin.g-node.org/juaml/datalad-example-aomicpiop1" uri_PIOP1 = "https://gin.g-node.org/juaml/datalad-example-aomicpiop1"
@ -138,7 +138,7 @@ def test_aomic_piop1_datagrabber() -> None:
assert sub == meta["element"]["subject"] assert sub == meta["element"]["subject"]
def test_piop1_invalid_tasks(): def test_DataladAOMICPIOP1_invalid_tasks():
"""Test whether invalid task fails.""" """Test whether invalid task fails."""
with pytest.raises( with pytest.raises(
ValueError, ValueError,

View file

@ -1,4 +1,4 @@
"""Provide tests for aomic piop2.""" """Provide tests DataladAOMICPIOP2 DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de> # Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Vera Komeyer <v.komeyer@fz-juelich.de> # Vera Komeyer <v.komeyer@fz-juelich.de>
@ -8,12 +8,12 @@
import pytest import pytest
from junifer.datagrabber.aomic import DataladAOMICPIOP2 from junifer.datagrabber import DataladAOMICPIOP2
from junifer.utils import configure_logging from junifer.utils import configure_logging
def test_aomic_piop2_datagrabber() -> None: def test_DataladAOMICPIOP2() -> None:
"""Test datalad AOMICPIOP2 datagrabber.""" """Test DataladAOMICPIOP2 DataGrabber."""
configure_logging(level="DEBUG") configure_logging(level="DEBUG")
uri_PIOP2 = "https://gin.g-node.org/juaml/datalad-example-aomicpiop2" uri_PIOP2 = "https://gin.g-node.org/juaml/datalad-example-aomicpiop2"
@ -132,7 +132,7 @@ def test_aomic_piop2_datagrabber() -> None:
assert sub == meta["element"]["subject"] assert sub == meta["element"]["subject"]
def test_piop2_invalid_tasks(): def test_DataladAOMICPIOP2_invalid_tasks():
"""Test whether invalid task fails.""" """Test whether invalid task fails."""
with pytest.raises( with pytest.raises(
ValueError, ValueError,

View file

@ -1,4 +1,4 @@
"""Provide abstract base class for datagrabber.""" """Provide abstract base class for DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de> # Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de> # Leonard Sasse <l.sasse@fz-juelich.de>
@ -15,7 +15,7 @@ from .utils import validate_types
class BaseDataGrabber(ABC, UpdateMetaMixin): class BaseDataGrabber(ABC, UpdateMetaMixin):
"""Abstract base class for datagrabber. """Abstract base class for DataGrabber.
For every interface that is required, one needs to provide a concrete For every interface that is required, one needs to provide a concrete
implementation of this abstract class. implementation of this abstract class.
@ -31,6 +31,7 @@ class BaseDataGrabber(ABC, UpdateMetaMixin):
---------- ----------
datadir : pathlib.Path datadir : pathlib.Path
The directory where the data is / will be stored. The directory where the data is / will be stored.
""" """
def __init__(self, types: List[str], datadir: Union[str, Path]) -> None: def __init__(self, types: List[str], datadir: Union[str, Path]) -> None:
@ -52,7 +53,7 @@ class BaseDataGrabber(ABC, UpdateMetaMixin):
Yields Yields
------ ------
object object
An element that can be indexed by the datagrabber. An element that can be indexed by the DataGrabber.
""" """
for elem in self.get_elements(): for elem in self.get_elements():
@ -126,7 +127,7 @@ class BaseDataGrabber(ABC, UpdateMetaMixin):
Returns Returns
------- -------
str list of str
The element keys. The element keys.
""" """
@ -144,7 +145,7 @@ class BaseDataGrabber(ABC, UpdateMetaMixin):
list list
List of elements that can be grabbed. The elements can be strings, List of elements that can be grabbed. The elements can be strings,
tuples or any object that will be then used as a key to index the tuples or any object that will be then used as a key to index the
datagrabber. DataGrabber.
""" """
raise_error( raise_error(

View file

@ -1,4 +1,4 @@
"""Provide abstract base class for datalad datagrabber.""" """Provide abstract base class for datalad-based DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de> # Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de> # Leonard Sasse <l.sasse@fz-juelich.de>
@ -24,20 +24,20 @@ from .base import BaseDataGrabber
@register_datagrabber @register_datagrabber
class DataladDataGrabber(BaseDataGrabber): class DataladDataGrabber(BaseDataGrabber):
"""Abstract base class for data fetching via Datalad. """Abstract base class for datalad-based data fetching.
Defines a Data Grabber that gets data from a datalad sibling. Defines a DataGrabber that gets data from a datalad sibling.
Parameters Parameters
---------- ----------
rootdir : str or Path, optional rootdir : str or pathlib.Path, optional
The path within the datalad dataset to the root directory The path within the datalad dataset to the root directory
(default "."). (default ".").
datadir : str or Path, optional datadir : str or pathlib.Path or None, optional
That directory where the datalad dataset will be cloned. If None, That directory where the datalad dataset will be cloned. If None,
the datalad dataset will be cloned into a temporary directory the datalad dataset will be cloned into a temporary directory
(default None). (default None).
uri : str, optional uri : str or None, optional
URI of the datalad sibling (default None). URI of the datalad sibling (default None).
**kwargs **kwargs
Keyword arguments passed to superclass. Keyword arguments passed to superclass.
@ -45,23 +45,26 @@ class DataladDataGrabber(BaseDataGrabber):
Methods Methods
------- -------
install: install:
Installs (clones) the datalad dataset into the `datadir`. This method Installs (clones) the datalad dataset into the ``datadir``. This method
is called automatically when the datagrabber is used within a context. is called automatically when the datagrabber is used within a context.
remove: remove:
Removes the datalad dataset from the `datadir`. This method is called Removes the datalad dataset from the ``datadir``. This method is called
automatically when the datagrabber is used within a context. automatically when the datagrabber is used within a context.
See Also See Also
-------- --------
BaseDataGrabber BaseDataGrabber:
For the abstract base class of any datagrabber. Abstract base class for DataGrabber.
PatternDataGrabber:
Concrete implementation for pattern-based data fetching.
PatternDataladDataGrabber:
Concrete implementation for pattern and datalad based data fetching.
Notes Notes
----- -----
By itself, this class is still abstract as the `__getitem__` method relies This class is intended to be used as a superclass of a subclass
on the parent class `BaseDataGrabber.__getitem__` which is not yet
implemented. This class is intended to be used as a superclass of a class
with multiple inheritance. with multiple inheritance.
""" """
def __init__( def __init__(
@ -129,16 +132,23 @@ class DataladDataGrabber(BaseDataGrabber):
@property @property
def datadir(self) -> Path: def datadir(self) -> Path:
"""Get data directory path.""" """Get data directory path.
Returns
-------
pathlib.Path
Path to the data directory.
"""
return super().datadir / self._rootdir return super().datadir / self._rootdir
def _get_dataset_id_remote(self) -> str: def _get_dataset_id_remote(self) -> str:
"""Get the dataset id from the remote. """Get the dataset ID from the remote.
Returns Returns
------- -------
str str
The dataset id. The dataset ID.
""" """
remote_id = None remote_id = None
@ -155,7 +165,7 @@ class DataladDataGrabber(BaseDataGrabber):
return remote_id return remote_id
def _dataset_get(self, out: Dict) -> Dict: def _dataset_get(self, out: Dict) -> Dict:
"""Get the dataset found from the path in `out`. """Get the dataset found from the path in ``out``.
Parameters Parameters
---------- ----------
@ -165,7 +175,7 @@ class DataladDataGrabber(BaseDataGrabber):
Returns Returns
------- -------
dict dict
The modified dictionary with meta updated. The unmodified input dictionary.
""" """
to_get = [v["path"] for v in out.values() if "path" in v] to_get = [v["path"] for v in out.values() if "path" in v]
@ -197,12 +207,13 @@ class DataladDataGrabber(BaseDataGrabber):
return out return out
def install(self) -> None: def install(self) -> None:
"""Install the datalad dataset into the datadir. """Install the datalad dataset into the ``datadir``.
Raises Raises
------ ------
ValueError ValueError
If the dataset is already installed but with a different ID. If the dataset is already installed but with a different ID.
""" """
isinstalled = dl.Dataset(self._datadir).is_installed() isinstalled = dl.Dataset(self._datadir).is_installed()
if isinstalled: if isinstalled:
@ -256,9 +267,7 @@ class DataladDataGrabber(BaseDataGrabber):
"""Implement single element indexing in the Datalad database. """Implement single element indexing in the Datalad database.
It will first obtain the paths from the parent class and then It will first obtain the paths from the parent class and then
`datalad get` each of the files. ``datalad get`` each of the files.
This method only works with multiple inheritance.
Parameters Parameters
---------- ----------
@ -279,12 +288,12 @@ class DataladDataGrabber(BaseDataGrabber):
out = self._dataset_get(out) out = self._dataset_get(out)
return out return out
def __enter__(self): def __enter__(self) -> "DataladDataGrabber":
"""Implement context entry.""" """Implement context entry."""
self.install() self.install()
return self return self
def __exit__(self, exc_type, exc_value, exc_traceback): def __exit__(self, exc_type, exc_value, exc_traceback) -> None:
"""Implement context exit.""" """Implement context exit."""
logger.debug("Cleaning up dataset") logger.debug("Cleaning up dataset")
self.cleanup() self.cleanup()

View file

@ -0,0 +1,7 @@
"""Provide imports for hcp1200 sub-package."""
# Authors: Synchon Mandal <s.mandal@fz-juelich.de>
# License: AGPL
from .hcp1200 import HCP1200
from .datalad_hcp1200 import DataladHCP1200

View file

@ -0,0 +1,68 @@
"""Provide concrete implementation for datalad-based HCP1200 DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de>
# Synchon Mandal <s.mandal@fz-juelich.de>
# License: AGPL
from pathlib import Path
from typing import List, Union
from junifer.datagrabber.datalad_base import DataladDataGrabber
from ...api.decorators import register_datagrabber
from .hcp1200 import HCP1200
@register_datagrabber
class DataladHCP1200(DataladDataGrabber, HCP1200):
"""Concrete implementation for datalad-based data fetching of HCP1200.
Parameters
----------
datadir : str or Path or None, optional
The directory where the datalad dataset will be cloned. If None,
the datalad dataset will be cloned into a temporary directory
(default None).
tasks : {"REST1", "REST2", "SOCIAL", "WM", "RELATIONAL", "EMOTION", \
"LANGUAGE", "GAMBLING", "MOTOR"} or list of the options or None \
, optional
HCP task sessions. If None, all available task sessions are selected
(default None).
phase_encodings : {"LR", "RL"} or list of the options or None, optional
HCP phase encoding directions. If None, both will be used
(default None).
ica_fix : bool, optional
Whether to retrieve data that was processed with ICA+FIX.
Only "REST1" and "REST2" tasks are available with ICA+FIX (default
False).
"""
def __init__(
self,
datadir: Union[str, Path, None] = None,
tasks: Union[str, List[str], None] = None,
phase_encodings: Union[str, List[str], None] = None,
ica_fix: bool = False,
) -> None:
uri = (
"https://github.com/datalad-datasets/"
"human-connectome-project-openaccess.git"
)
rootdir = "HCP1200"
super().__init__(
datadir=datadir,
tasks=tasks,
phase_encodings=phase_encodings,
uri=uri,
rootdir=rootdir,
ica_fix=ica_fix,
)
# Needed here as HCP1200's subjects are sub-datasets, so will not be
# found when elements are checked.
@property
def skip_file_check(self) -> bool:
"""Skip file check existence."""
return True

View file

@ -1,14 +1,17 @@
"""Provide concrete implementations for HCP data access.""" """Provide concrete implementation for pattern-based HCP1200 DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de>
# Synchon Mandal <s.mandal@fz-juelich.de>
# License: AGPL
from itertools import product from itertools import product
from pathlib import Path from pathlib import Path
from typing import Dict, List, Union from typing import Dict, List, Union
from junifer.datagrabber.datalad_base import DataladDataGrabber from ...api.decorators import register_datagrabber
from ..pattern import PatternDataGrabber
from ..api.decorators import register_datagrabber
from ..utils import raise_error from ..utils import raise_error
from .pattern import PatternDataGrabber
@register_datagrabber @register_datagrabber
@ -18,20 +21,22 @@ class HCP1200(PatternDataGrabber):
Parameters Parameters
---------- ----------
datadir : str or Path, optional datadir : str or Path, optional
The directory where the datalad dataset will be cloned. The directory where the data is / will be stored.
tasks : {"REST1", "REST2", "SOCIAL", "WM", "RELATIONAL", "EMOTION", \ tasks : {"REST1", "REST2", "SOCIAL", "WM", "RELATIONAL", "EMOTION", \
"LANGUAGE", "GAMBLING", "MOTOR"} or list of the options, optional "LANGUAGE", "GAMBLING", "MOTOR"} or list of the options or None \
, optional
HCP task sessions. If None, all available task sessions are selected HCP task sessions. If None, all available task sessions are selected
(default None). (default None).
phase_encodings : {"LR", "RL"} or list of the options, optional phase_encodings : {"LR", "RL"} or list of the options or None, optional
HCP phase encoding directions. If None, both will be used HCP phase encoding directions. If None, both will be used
(default None). (default None).
ica_fix : bool, optional ica_fix : bool, optional
Whether to retrieve data that was processed with ICA+FIX. Whether to retrieve data that was processed with ICA+FIX.
Only 'REST1' and 'REST2' tasks are available with ICA+FIX (default Only "REST1" and "REST2" tasks are available with ICA+FIX (default
False). False).
**kwargs **kwargs
Keyword arguments passed to superclass. Keyword arguments passed to superclass.
""" """
def __init__( def __init__(
@ -114,7 +119,7 @@ class HCP1200(PatternDataGrabber):
self.phase_encodings = phase_encodings self.phase_encodings = phase_encodings
def get_item(self, subject: str, task: str, phase_encoding: str) -> Dict: def get_item(self, subject: str, task: str, phase_encoding: str) -> Dict:
"""Index one element in the dataset. """Implement single element indexing in the database.
Parameters Parameters
---------- ----------
@ -128,8 +133,8 @@ class HCP1200(PatternDataGrabber):
Returns Returns
------- -------
out : dict dict
Dictionary of paths for each type of data required for the Dictionary of dictionaries for each type of data required for the
specified element. specified element.
""" """
@ -150,7 +155,7 @@ class HCP1200(PatternDataGrabber):
Returns Returns
------- -------
list list
The list of elements in the dataset. The list of elements that can be grabbed in the dataset.
""" """
subjects = [ subjects = [
@ -165,53 +170,3 @@ class HCP1200(PatternDataGrabber):
elems.append((subject, task, phase_encoding)) elems.append((subject, task, phase_encoding))
return elems return elems
@register_datagrabber
class DataladHCP1200(DataladDataGrabber, HCP1200):
"""Concrete implementation for datalad-based data fetching of HCP1200.
Parameters
----------
datadir : str or Path, optional
The directory where the datalad dataset will be cloned. If None,
the datalad dataset will be cloned into a temporary directory
(default None).
tasks : {"REST1", "REST2", "SOCIAL", "WM", "RELATIONAL", "EMOTION", \
"LANGUAGE", "GAMBLING", "MOTOR"} or list of the options, optional
HCP task sessions. If None, all available task sessions are selected
(default None).
phase_encodings : {"LR", "RL"} or list of the options, optional
HCP phase encoding directions. If None, both will be used
(default None).
ica_fix : bool, optional
Whether to retrieve data that was processed with ICA+FIX.
Only 'REST1' and 'REST2' tasks are available with ICA+FIX (default
False).
"""
def __init__(
self,
datadir: Union[str, Path, None] = None,
tasks: Union[str, List[str], None] = None,
phase_encodings: Union[str, List[str], None] = None,
ica_fix: bool = False,
) -> None:
uri = (
"https://github.com/datalad-datasets/"
"human-connectome-project-openaccess.git"
)
rootdir = "HCP1200"
super().__init__(
datadir=datadir,
tasks=tasks,
phase_encodings=phase_encodings,
uri=uri,
rootdir=rootdir,
ica_fix=ica_fix,
)
@property
def skip_file_check(self) -> bool:
"""Skip file check existence."""
return True

View file

@ -7,7 +7,7 @@ from typing import Iterable, Optional
import pytest import pytest
from junifer.datagrabber.hcp import HCP1200, DataladHCP1200 from junifer.datagrabber import HCP1200, DataladHCP1200
from junifer.utils import configure_logging from junifer.utils import configure_logging
@ -16,7 +16,7 @@ URI = "https://gin.g-node.org/juaml/datalad-example-hcp1200"
@pytest.fixture(scope="module") @pytest.fixture(scope="module")
def hcpdg() -> Iterable[DataladHCP1200]: def hcpdg() -> Iterable[DataladHCP1200]:
"""Return a HCP1200 datagrabber.""" """Return a HCP1200 DataGrabber."""
dg = DataladHCP1200() dg = DataladHCP1200()
# Set URI to Gin # Set URI to Gin
dg.uri = URI dg.uri = URI
@ -56,19 +56,19 @@ def hcpdg() -> Iterable[DataladHCP1200]:
("REST2", "RL", True, "rfMRI_REST2_RL_hp2000_clean.nii.gz"), ("REST2", "RL", True, "rfMRI_REST2_RL_hp2000_clean.nii.gz"),
], ],
) )
def test_hcp1200_datagrabber( def test_HCP1200(
hcpdg: DataladHCP1200, hcpdg: DataladHCP1200,
tasks: Optional[str], tasks: Optional[str],
phase_encodings: Optional[str], phase_encodings: Optional[str],
ica_fix: bool, ica_fix: bool,
expected_path_name: str, expected_path_name: str,
) -> None: ) -> None:
"""Test HCP1200 datagrabber. """Test HCP1200 DataGrabber.
Parameters Parameters
---------- ----------
hcpdg : DataladHCP1200 hcpdg : DataladHCP1200
The Datalad version of the datagrabber with the first subject The Datalad version of the DataGrabber with the first subject
already cloned. already cloned.
tasks : str tasks : str
The parametrized tasks. The parametrized tasks.
@ -132,17 +132,17 @@ def test_hcp1200_datagrabber(
("MOTOR", "RL"), ("MOTOR", "RL"),
], ],
) )
def test_hcp1200_datagrabber_single_access( def test_HCP1200_single_access(
hcpdg: DataladHCP1200, hcpdg: DataladHCP1200,
tasks: Optional[str], tasks: Optional[str],
phase_encodings: Optional[str], phase_encodings: Optional[str],
) -> None: ) -> None:
"""Test HCP1200 datagrabber single access. """Test HCP1200 DataGrabber single access.
Parameters Parameters
---------- ----------
hcpdg : DataladHCP1200 hcpdg : DataladHCP1200
The Datalad version of the datagrabber with the first subject The Datalad version of the DataGrabber with the first subject
already cloned. already cloned.
tasks : str tasks : str
The parametrized tasks. The parametrized tasks.
@ -172,17 +172,17 @@ def test_hcp1200_datagrabber_single_access(
(["REST1", "REST2"], None), (["REST1", "REST2"], None),
], ],
) )
def test_hcp1200_datagrabber_multi_access( def test_HCP1200_multi_access(
hcpdg: DataladHCP1200, hcpdg: DataladHCP1200,
tasks: Optional[str], tasks: Optional[str],
phase_encodings: Optional[str], phase_encodings: Optional[str],
) -> None: ) -> None:
"""Test HCP1200 datagrabber multiple access. """Test HCP1200 DataGrabber multiple access.
Parameters Parameters
---------- ----------
hcpdg : DataladHCP1200 hcpdg : DataladHCP1200
The Datalad version of the datagrabber with the first subject The Datalad version of the DataGrabber with the first subject
already cloned. already cloned.
tasks : str tasks : str
The parametrized tasks. The parametrized tasks.
@ -205,16 +205,17 @@ def test_hcp1200_datagrabber_multi_access(
assert element[2] in ["LR", "RL"] assert element[2] in ["LR", "RL"]
def test_hcp1200_datagrabber_multi_access_task_simple( def test_HCP1200_multi_access_task_simple(
hcpdg: DataladHCP1200, hcpdg: DataladHCP1200,
) -> None: ) -> None:
"""Test HCP1200 datagrabber simple multiple access for task. """Test HCP1200 DataGrabber simple multiple access for task.
Parameters Parameters
---------- ----------
hcpdg : DataladHCP1200 hcpdg : DataladHCP1200
The Datalad version of the datagrabber with the first subject The Datalad version of the DataGrabber with the first subject
already cloned. already cloned.
""" """
configure_logging(level="DEBUG") configure_logging(level="DEBUG")
dg = HCP1200( dg = HCP1200(
@ -231,16 +232,17 @@ def test_hcp1200_datagrabber_multi_access_task_simple(
assert element[2] in ["LR", "RL"] assert element[2] in ["LR", "RL"]
def test_hcp1200_datagrabber_multi_access_phase_simple( def test_HCP1200_multi_access_phase_simple(
hcpdg: DataladHCP1200, hcpdg: DataladHCP1200,
) -> None: ) -> None:
"""Test HCP1200 datagrabber simple multiple access for phase. """Test HCP1200 DataGrabber simple multiple access for phase.
Parameters Parameters
---------- ----------
hcpdg : DataladHCP1200 hcpdg : DataladHCP1200
The Datalad version of the datagrabber with the first subject The Datalad version of the DataGrabber with the first subject
already cloned. already cloned.
""" """
configure_logging(level="DEBUG") configure_logging(level="DEBUG")
dg = HCP1200( dg = HCP1200(
@ -266,11 +268,11 @@ def test_hcp1200_datagrabber_multi_access_phase_simple(
(["FOO", "BAR"], "LR"), (["FOO", "BAR"], "LR"),
], ],
) )
def test_hcp1200_datagrabber_incorrect_access_task( def test_HCP1200_incorrect_access_task(
tasks: Optional[str], tasks: Optional[str],
phase_encodings: Optional[str], phase_encodings: Optional[str],
) -> None: ) -> None:
"""Test HCP1200 datagrabber incorrect access for task. """Test HCP1200 DataGrabber incorrect access for task.
Parameters Parameters
---------- ----------
@ -298,11 +300,11 @@ def test_hcp1200_datagrabber_incorrect_access_task(
(["REST1", "REST2"], "BAR"), (["REST1", "REST2"], "BAR"),
], ],
) )
def test_hcp1200_datagrabber_incorrect_access_phase( def test_HCP1200_incorrect_access_phase(
tasks: Optional[str], tasks: Optional[str],
phase_encodings: Optional[str], phase_encodings: Optional[str],
) -> None: ) -> None:
"""Test HCP1200 datagrabber incorrect access for phase. """Test HCP1200 DataGrabber incorrect access for phase.
Parameters Parameters
---------- ----------
@ -321,16 +323,17 @@ def test_hcp1200_datagrabber_incorrect_access_phase(
) )
def test_hcp1200_datagrabber_elements( def test_HCP1200_elements(
hcpdg: DataladHCP1200, hcpdg: DataladHCP1200,
) -> None: ) -> None:
"""Test HCP1200 datagrabber elements. """Test HCP1200 DataGrabber elements.
Parameters Parameters
---------- ----------
hcpdg : DataladHCP1200 hcpdg : DataladHCP1200
The Datalad version of the datagrabber with the first subject The Datalad version of the DataGrabber with the first subject
already cloned. already cloned.
""" """
configure_logging(level="DEBUG") configure_logging(level="DEBUG")
dg = HCP1200( dg = HCP1200(
@ -364,10 +367,10 @@ def test_hcp1200_datagrabber_elements(
("MOTOR", True), ("MOTOR", True),
], ],
) )
def test_hcp1200_datagrabber_incorrect_access_icafix( def test_HCP1200_incorrect_access_icafix(
tasks: Optional[str], ica_fix: bool tasks: Optional[str], ica_fix: bool
) -> None: ) -> None:
"""Test HCP1200 datagrabber incorrect access for icafix. """Test HCP1200 DataGrabber incorrect access for icafix.
Parameters Parameters
---------- ----------

View file

@ -1,4 +1,4 @@
"""Provide abstract base class for multiple source datagrabber.""" """Provide concrete implementation for multi sourced DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de> # Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de> # Leonard Sasse <l.sasse@fz-juelich.de>
@ -7,21 +7,23 @@
from typing import Dict, List, Tuple, Union from typing import Dict, List, Tuple, Union
from ..utils import raise_error
from .base import BaseDataGrabber from .base import BaseDataGrabber
class MultipleDataGrabber(BaseDataGrabber): class MultipleDataGrabber(BaseDataGrabber):
"""Data Grabber class for data fetching from multiple sources. """Concrete implementation for multi sourced data fetching.
Defines a Data Grabber which can be used to fetch data from multiple Implements a DataGrabber which can be used to fetch data from multiple
datagrabbers. DataGrabbers.
Parameters Parameters
---------- ----------
datagrabbers : list of datagrabber-like objects datagrabbers : list of DataGrabber-like objects
The datagrabbers to use to fetch data using. The DataGrabbers to use for fetching data.
**kwargs **kwargs
Keyword arguments passed to superclass. Keyword arguments passed to superclass.
""" """
def __init__(self, datagrabbers: List[BaseDataGrabber], **kwargs) -> None: def __init__(self, datagrabbers: List[BaseDataGrabber], **kwargs) -> None:
@ -30,11 +32,11 @@ class MultipleDataGrabber(BaseDataGrabber):
first_keys = datagrabbers[0].get_element_keys() first_keys = datagrabbers[0].get_element_keys()
for dg in datagrabbers[1:]: for dg in datagrabbers[1:]:
if dg.get_element_keys() != first_keys: if dg.get_element_keys() != first_keys:
raise ValueError("Datagrabbers have different element keys.") raise_error("DataGrabbers have different element keys.")
# 2) no overlapping types # 2) no overlapping types
types = [x for dg in datagrabbers for x in dg.get_types()] types = [x for dg in datagrabbers for x in dg.get_types()]
if len(types) != len(set(types)): if len(types) != len(set(types)):
raise ValueError("Datagrabbers have overlapping types.") raise_error("DataGrabbers have overlapping types.")
self._datagrabbers = datagrabbers self._datagrabbers = datagrabbers
def __getitem__(self, element: Union[str, Tuple]) -> Dict: def __getitem__(self, element: Union[str, Tuple]) -> Dict:
@ -73,8 +75,64 @@ class MultipleDataGrabber(BaseDataGrabber):
out[kind]["meta"]["datagrabber"]["datagrabbers"] = metas out[kind]["meta"]["datagrabber"]["datagrabbers"] = metas
return out return out
def get_item(self, **element: Dict) -> Dict[str, Dict]: def __enter__(self) -> "MultipleDataGrabber":
"""Get item. """Implement context entry."""
for dg in self._datagrabbers:
dg.__enter__()
return self
def __exit__(self, exc_type, exc_value, exc_traceback) -> None:
"""Implement context exit."""
for dg in self._datagrabbers:
dg.__exit__(exc_type, exc_value, exc_traceback)
# TODO: return type should be List[List[str]], but base type is List[str]
def get_types(self) -> List[str]:
"""Get types.
Returns
-------
list of list of str
The types of data to be grabbed.
"""
types = [x for dg in self._datagrabbers for x in dg.get_types()]
return types
def get_element_keys(self) -> List[str]:
"""Get element keys.
For each item in the ``element`` tuple passed to ``__getitem__()``,
this method returns the corresponding key(s).
Returns
-------
list of str
The element keys.
"""
return self._datagrabbers[0].get_element_keys()
def get_elements(self) -> List:
"""Get elements.
Returns
-------
list
List of elements that can be grabbed. The elements can be strings,
tuples or any object that will be then used as a key to index the
the DataGrabber. The element should be present in all of the
related DataGrabbers.
"""
all_elements = [dg.get_elements() for dg in self._datagrabbers]
elements = set(all_elements[0])
for s in all_elements[1:]:
elements.intersection_update(s)
return list(elements)
def get_item(self, **_: Dict) -> Dict[str, Dict]:
"""Get the specified item from the dataset.
Parameters Parameters
---------- ----------
@ -90,59 +148,12 @@ class MultipleDataGrabber(BaseDataGrabber):
Notes Notes
----- -----
This function is not implemented for this class as it is useless. This function is not implemented for this class as it is useless.
""" """
raise NotImplementedError( raise_error(
"get_item() is not useful for this class, hence not implemented." msg=(
"get_item() is not useful for this class, hence not "
"implemented."
),
klass=NotImplementedError,
) )
def __enter__(self) -> "BaseDataGrabber":
"""Implement context entry."""
for dg in self._datagrabbers:
dg.__enter__()
return self
def __exit__(self, exc_type, exc_value, exc_traceback) -> None:
"""Implement context exit."""
for dg in self._datagrabbers:
dg.__exit__(exc_type, exc_value, exc_traceback)
def get_elements(self) -> List:
"""Get elements.
Returns
-------
list
The list of elements that can be grabbed in the dataset. It
corresponds to the elements that are present in all the
related datagrabbers.
"""
all_elements = [dg.get_elements() for dg in self._datagrabbers]
elements = set(all_elements[0])
for s in all_elements[1:]:
elements.intersection_update(s)
return list(elements)
def get_element_keys(self) -> List[str]:
"""Get element keys.
For each item in the ``element`` tuple passed to ``__getitem__()``,
this method returns the corresponding key(s).
Returns
-------
list of str
The element keys.
"""
return self._datagrabbers[0].get_element_keys()
def get_types(self) -> List[str]:
"""Get types.
Returns
-------
list of list of str
The types of data to be grabbed.
"""
types = [x for dg in self._datagrabbers for x in dg.get_types()]
return types

View file

@ -1,4 +1,4 @@
"""Provide concrete implementation for pattern-based datagrabber.""" """Provide concrete implementation for pattern-based DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de> # Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de> # Leonard Sasse <l.sasse@fz-juelich.de>
@ -23,9 +23,9 @@ _CONFOUNDS_FORMATS = ("fmriprep", "adhoc")
@register_datagrabber @register_datagrabber
class PatternDataGrabber(BaseDataGrabber): class PatternDataGrabber(BaseDataGrabber):
"""Concrete implementation for data grabbing using patterns. """Concrete implementation for pattern-based data fetching.
Implements a Data Grabber that understands patterns to grab data. Implements a DataGrabber that understands patterns to grab data.
Parameters Parameters
---------- ----------
@ -34,12 +34,12 @@ class PatternDataGrabber(BaseDataGrabber):
patterns : dict patterns : dict
Patterns for each type of data as a dictionary. The keys are the types Patterns for each type of data as a dictionary. The keys are the types
and the values are the patterns. Each occurrence of the string and the values are the patterns. Each occurrence of the string
`{subject}` in the pattern will be replaced by the indexed element. ``{subject}`` in the pattern will be replaced by the indexed element.
replacements : list of str replacements : str or list of str
Replacements in the patterns for each item in the "element" tuple. Replacements in the patterns for each item in the "element" tuple.
datadir : str or pathlib.Path datadir : str or pathlib.Path
The directory where the data is / will be stored. The directory where the data is / will be stored.
confounds_format : {"fmriprep", "adhoc"}, optional confounds_format : {"fmriprep", "adhoc"} or None, optional
The format of the confounds for the dataset (default None). The format of the confounds for the dataset (default None).
""" """
@ -83,7 +83,7 @@ class PatternDataGrabber(BaseDataGrabber):
def _replace_patterns_regex( def _replace_patterns_regex(
self, pattern: str self, pattern: str
) -> Tuple[str, str, List[str]]: ) -> Tuple[str, str, List[str]]:
"""Replace the patterns in `pattern` with the named groups. """Replace the patterns in ``pattern`` with the named groups.
It allows elements to be obtained from the filesystem. It allows elements to be obtained from the filesystem.
@ -122,7 +122,7 @@ class PatternDataGrabber(BaseDataGrabber):
return re_pattern, glob_pattern, t_replacements return re_pattern, glob_pattern, t_replacements
def _replace_patterns_glob(self, element: Dict, pattern: str) -> str: def _replace_patterns_glob(self, element: Dict, pattern: str) -> str:
"""Replace patterns with the element so it can be globbed. """Replace ``pattern`` with the ``element`` so it can be globbed.
Parameters Parameters
---------- ----------
@ -209,7 +209,7 @@ class PatternDataGrabber(BaseDataGrabber):
raise_error( raise_error(
"`confounds_format` needs to be one of " "`confounds_format` needs to be one of "
f"{_CONFOUNDS_FORMATS}, None provided. " f"{_CONFOUNDS_FORMATS}, None provided. "
"As the datagrabber used specifies " "As the DataGrabber used specifies "
"'BOLD_confounds', None is invalid." "'BOLD_confounds', None is invalid."
) )
# Set the format # Set the format
@ -228,6 +228,7 @@ class PatternDataGrabber(BaseDataGrabber):
------- -------
list list
The list of elements that can be grabbed in the dataset. The list of elements that can be grabbed in the dataset.
""" """
elements = None elements = None

View file

@ -1,4 +1,4 @@
"""Provide base class for pattern-based datalad datagrabber.""" """Provide concrete implementation for pattern + datalad based DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de> # Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de> # Leonard Sasse <l.sasse@fz-juelich.de>
@ -15,29 +15,29 @@ from .utils import validate_patterns
@register_datagrabber @register_datagrabber
class PatternDataladDataGrabber(DataladDataGrabber, PatternDataGrabber): class PatternDataladDataGrabber(DataladDataGrabber, PatternDataGrabber):
"""Base class for pattern-based data fetching via Datalad. """Concrete implementation for pattern and datalad based data fetching.
Defines a DataGrabber that gets data from a datalad sibling, Implements a DataGrabber that gets data from a datalad sibling,
interpreting patterns. interpreting patterns.
Parameters Parameters
---------- ----------
types : list of str types : list of str
The types of data to be grabbed. The types of data to be grabbed.
patterns : dict, optional patterns : dict
Patterns for each type of data as a dictionary. The keys are the types Patterns for each type of data as a dictionary. The keys are the types
and the values are the patterns. Each occurrence of the string and the values are the patterns. Each occurrence of the string
`{subject}` in the pattern will be replaced by the indexed element ``{subject}`` in the pattern will be replaced by the indexed element.
(default None).
**kwargs **kwargs
Keyword arguments passed to superclass. Keyword arguments passed to superclass.
See Also See Also
-------- --------
DataladDataGrabber: DataladDataGrabber:
Base class for data fetching via Datalad. Abstract base class for datalad-based data fetching.
PatternDataGrabber: PatternDataGrabber:
Base class for pattern-based data fetching. Concrete implementation for pattern-based data fetching.
""" """
def __init__( def __init__(

View file

@ -1,4 +1,4 @@
"""Provide tests for base.""" """Provide tests for BaseDataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de> # Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de> # Leonard Sasse <l.sasse@fz-juelich.de>
@ -9,7 +9,7 @@ from pathlib import Path
import pytest import pytest
from junifer.datagrabber.base import BaseDataGrabber from junifer.datagrabber import BaseDataGrabber
def test_BaseDataGrabber_abstractness() -> None: def test_BaseDataGrabber_abstractness() -> None:

View file

@ -1,4 +1,4 @@
"""Provide tests for datalad_base.""" """Provide tests for DataladDataGrabber."""
# Authors: Synchon Mandal <s.mandal@fz-juelich.de> # Authors: Synchon Mandal <s.mandal@fz-juelich.de>
# License: AGPL # License: AGPL
@ -9,7 +9,7 @@ from typing import Type
import datalad.api as dl import datalad.api as dl
import pytest import pytest
from junifer.datagrabber.datalad_base import DataladDataGrabber from junifer.datagrabber import DataladDataGrabber
_testing_dataset = { _testing_dataset = {
@ -26,20 +26,20 @@ _testing_dataset = {
} }
def test_datalad_base_abstractness() -> None: def test_DataladDataGrabber_abstractness() -> None:
"""Test datalad base is abstract.""" """Test DataladDataGrabber is abstract base class."""
with pytest.raises(TypeError, match=r"abstract"): with pytest.raises(TypeError, match=r"abstract"):
DataladDataGrabber() # type: ignore DataladDataGrabber() # type: ignore
@pytest.fixture @pytest.fixture
def concrete_datagrabber() -> Type[DataladDataGrabber]: def concrete_datagrabber() -> Type[DataladDataGrabber]:
"""Return a concrete datagrabber class. """Return a concrete datalad-based DataGrabber.
Returns Returns
------- -------
DataladDataGrabber DataladDataGrabber
A concrete datagrabber class. A concrete datalad-based DataGrabber.
""" """
@ -74,17 +74,18 @@ def concrete_datagrabber() -> Type[DataladDataGrabber]:
return MyDataGrabber return MyDataGrabber
def test_datalad_install_errors( def test_DataladDataGrabber_install_errors(
tmp_path: Path, concrete_datagrabber: Type tmp_path: Path, concrete_datagrabber: Type
) -> None: ) -> None:
"""Test datalad base install errors / warnings. """Test DataladDataGrabber install errors / warnings.
Parameters Parameters
---------- ----------
tmp_path : pathlib.Path tmp_path : pathlib.Path
The path to the test directory. The path to the test directory.
concrete_datagrabber : DataladDataGrabber concrete_datagrabber : DataladDataGrabber
A concrete datagrabber class to use. A concrete datalad-based DataGrabber class to use.
""" """
# Dataset cloned outside of datagrabber # Dataset cloned outside of datagrabber
@ -112,17 +113,18 @@ def test_datalad_install_errors(
pass pass
def test_datalad_clone_cleanup( def test_DataladDataGrabber_clone_cleanup(
tmp_path: Path, concrete_datagrabber: Type tmp_path: Path, concrete_datagrabber: Type
) -> None: ) -> None:
"""Test datalad base clone and remove. """Test DataladDataGrabber clone and remove.
Parameters Parameters
---------- ----------
tmp_path : pathlib.Path tmp_path : pathlib.Path
The path to the test directory. The path to the test directory.
concrete_datagrabber : DataladDataGrabber concrete_datagrabber : DataladDataGrabber
A concrete datagrabber class to use. A concrete datalad-based DataGrabber class to use.
""" """
# Clone whole dataset # Clone whole dataset
@ -160,13 +162,16 @@ def test_datalad_clone_cleanup(
assert len(list(datadir.glob("*"))) == 0 assert len(list(datadir.glob("*"))) == 0
def test_datalad_clone_create_cleanup(concrete_datagrabber: Type) -> None: def test_DataladDataGrabber_clone_create_cleanup(
"""Test datalad base tempdir clone and remove. concrete_datagrabber: Type,
) -> None:
"""Test DataladDataGrabber tempdir clone and remove.
Parameters Parameters
---------- ----------
concrete_datagrabber : DataladDataGrabber concrete_datagrabber : DataladDataGrabber
A concrete datagrabber class to use. A concrete datalad-based DataGrabber class to use.
""" """
# Clone whole dataset # Clone whole dataset
@ -203,17 +208,18 @@ def test_datalad_clone_create_cleanup(concrete_datagrabber: Type) -> None:
assert len(list(datadir.glob("*"))) == 0 assert len(list(datadir.glob("*"))) == 0
def test_datalad_previously_cloned( def test_DataladDataGrabber_previously_cloned(
tmp_path: Path, concrete_datagrabber: Type tmp_path: Path, concrete_datagrabber: Type
) -> None: ) -> None:
"""Test datalad base on cloned dataset. """Test DataladDataGrabber on cloned dataset.
Parameters Parameters
---------- ----------
tmp_path : pathlib.Path tmp_path : pathlib.Path
The path to the test directory. The path to the test directory.
concrete_datagrabber : DataladDataGrabber concrete_datagrabber : DataladDataGrabber
A concrete datagrabber class to use. A concrete datalad-based DataGrabber class to use.
""" """
# Dataset cloned outside of datagrabber # Dataset cloned outside of datagrabber
@ -271,17 +277,18 @@ def test_datalad_previously_cloned(
assert len(list(datadir.glob("*"))) > 0 assert len(list(datadir.glob("*"))) > 0
def test_datalad_previously_cloned_and_get( def test_DataladDataGrabber_previously_cloned_and_get(
tmp_path: Path, concrete_datagrabber: Type tmp_path: Path, concrete_datagrabber: Type
) -> None: ) -> None:
"""Test datalad base on cloned dataset with files present. """Test DataladDataGrabber on cloned dataset with files present.
Parameters Parameters
---------- ----------
tmp_path : pathlib.Path tmp_path : pathlib.Path
The path to the test directory. The path to the test directory.
concrete_datagrabber : DataladDataGrabber concrete_datagrabber : DataladDataGrabber
A concrete datagrabber class to use. A concrete datalad-based DataGrabber class to use.
""" """
# Dataset cloned outside of datagrabber with some files present # Dataset cloned outside of datagrabber with some files present
@ -353,17 +360,18 @@ def test_datalad_previously_cloned_and_get(
assert elem1_t1w.is_file() is True assert elem1_t1w.is_file() is True
def test_datalad_previously_cloned_and_get_dirty( def test_DataladDataGrabber_previously_cloned_and_get_dirty(
tmp_path: Path, concrete_datagrabber: Type tmp_path: Path, concrete_datagrabber: Type
) -> None: ) -> None:
"""Test datalad base on a dirty cloned dataset. """Test DataladDataGrabber on a dirty cloned dataset.
Parameters Parameters
---------- ----------
tmp_path : pathlib.Path tmp_path : pathlib.Path
The path to the test directory. The path to the test directory.
concrete_datagrabber : DataladDataGrabber concrete_datagrabber : DataladDataGrabber
A concrete datagrabber class to use. A concrete datalad-based DataGrabber class to use.
""" """
# Dataset cloned outside of datagrabber with some files present and dirty # Dataset cloned outside of datagrabber with some files present and dirty

View file

@ -1,4 +1,4 @@
"""Provide tests for multiple.""" """Provide tests for MultipleDataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de> # Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# License: AGPL # License: AGPL
@ -20,8 +20,8 @@ _testing_dataset = {
} }
def test_multiple() -> None: def test_MultipleDataGrabber() -> None:
"""Test a multiple datagrabber.""" """Test MultipleDataGrabber."""
repo_uri = _testing_dataset["example_bids_ses"]["uri"] repo_uri = _testing_dataset["example_bids_ses"]["uri"]
rootdir = "example_bids_ses" rootdir = "example_bids_ses"
replacements = ["subject", "session"] replacements = ["subject", "session"]
@ -77,8 +77,8 @@ def test_multiple() -> None:
assert meta["datagrabbers"][1]["class"] == "PatternDataladDataGrabber" assert meta["datagrabbers"][1]["class"] == "PatternDataladDataGrabber"
def test_multiple_no_intersection() -> None: def test_MultipleDataGrabber_no_intersection() -> None:
"""Test a multiple datagrabber without intersection (0 elements).""" """Test MultipleDataGrabber without intersection (0 elements)."""
repo_uri1 = _testing_dataset["example_bids"]["uri"] repo_uri1 = _testing_dataset["example_bids"]["uri"]
repo_uri2 = _testing_dataset["example_bids_ses"]["uri"] repo_uri2 = _testing_dataset["example_bids_ses"]["uri"]
rootdir = "example_bids_ses" rootdir = "example_bids_ses"
@ -113,8 +113,8 @@ def test_multiple_no_intersection() -> None:
assert set(subs) == set(expected_subs) assert set(subs) == set(expected_subs)
def test_multiple_get_item() -> None: def test_MultipleDataGrabber_get_item() -> None:
"""Test a multiple datagrabber get_item error.""" """Test MultipleDataGrabber get_item() error."""
repo_uri1 = _testing_dataset["example_bids"]["uri"] repo_uri1 = _testing_dataset["example_bids"]["uri"]
rootdir = "example_bids_ses" rootdir = "example_bids_ses"
replacements = ["subject", "session"] replacements = ["subject", "session"]
@ -134,8 +134,8 @@ def test_multiple_get_item() -> None:
dg.get_item(subject="sub-01") # type: ignore dg.get_item(subject="sub-01") # type: ignore
def test_multiple_validation() -> None: def test_MultipleDataGrabber_validation() -> None:
"""Test a multiple datagrabber init validation.""" """Test MultipleDataGrabber init validation."""
repo_uri1 = _testing_dataset["example_bids"]["uri"] repo_uri1 = _testing_dataset["example_bids"]["uri"]
repo_uri2 = _testing_dataset["example_bids_ses"]["uri"] repo_uri2 = _testing_dataset["example_bids_ses"]["uri"]
rootdir = "example_bids_ses" rootdir = "example_bids_ses"

View file

@ -1,4 +1,4 @@
"""Provide tests for pattern.""" """Provide tests for PatternDataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de> # Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de> # Leonard Sasse <l.sasse@fz-juelich.de>
@ -9,7 +9,7 @@ from pathlib import Path
import pytest import pytest
from junifer.datagrabber.pattern import PatternDataGrabber from junifer.datagrabber import PatternDataGrabber
def test_PatternDataGrabber_errors(tmp_path: Path) -> None: def test_PatternDataGrabber_errors(tmp_path: Path) -> None:
@ -261,7 +261,7 @@ def test_PatternDataGrabber(tmp_path: Path) -> None:
assert out1["vbm"]["path"] != out2["vbm"]["path"] assert out1["vbm"]["path"] != out2["vbm"]["path"]
def test_pattern_data_grabber_confounds_format_error_on_init() -> None: def test_PatternDataGrabber_confounds_format_error_on_init() -> None:
"""Test PatterDataGrabber confounds format error on initialisation.""" """Test PatterDataGrabber confounds format error on initialisation."""
with pytest.raises( with pytest.raises(
ValueError, match="Invalid value for `confounds_format`" ValueError, match="Invalid value for `confounds_format`"
@ -275,7 +275,7 @@ def test_pattern_data_grabber_confounds_format_error_on_init() -> None:
) )
def test_pattern_data_grabber_confounds_format_error_on_fetch( def test_PatternDataGrabber_confounds_format_error_on_fetch(
tmp_path: Path, tmp_path: Path,
) -> None: ) -> None:
"""Test PatterDataGrabber confounds format error on fetching. """Test PatterDataGrabber confounds format error on fetching.
@ -301,6 +301,6 @@ def test_pattern_data_grabber_confounds_format_error_on_fetch(
) )
# Check error on fetch # Check error on fetch
with pytest.raises( with pytest.raises(
ValueError, match="As the datagrabber used specifies 'BOLD_confounds'" ValueError, match="As the DataGrabber used specifies 'BOLD_confounds'"
): ):
datagrabber.get_item(subject="sub-001") datagrabber.get_item(subject="sub-001")

View file

@ -1,4 +1,4 @@
"""Provide tests for pattern_datalad.""" """Provide tests for PatternDataladDataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de> # Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de> # Leonard Sasse <l.sasse@fz-juelich.de>
@ -9,7 +9,7 @@ from pathlib import Path
import pytest import pytest
from junifer.datagrabber.pattern_datalad import PatternDataladDataGrabber from junifer.datagrabber import PatternDataladDataGrabber
_testing_dataset = { _testing_dataset = {
@ -26,8 +26,8 @@ _testing_dataset = {
} }
def test_bids_pattern_datalad_datagrabber_missing_uri() -> None: def test_bids_PatternDataladDataGrabber_missing_uri() -> None:
"""Test check of missing URI in pattern datalad datagrabber.""" """Test check of missing URI in PatternDataladDataGrabber."""
with pytest.raises(ValueError, match=r"`uri` must be provided"): with pytest.raises(ValueError, match=r"`uri` must be provided"):
PatternDataladDataGrabber( PatternDataladDataGrabber(
datadir=None, datadir=None,
@ -37,15 +37,8 @@ def test_bids_pattern_datalad_datagrabber_missing_uri() -> None:
) )
def test_bids_PatternDataladDataGrabber(tmp_path: Path) -> None: def test_bids_PatternDataladDataGrabber() -> None:
"""Test a subject-based BIDS datalad datagrabber. """Test subject-based BIDS PatternDataladDataGrabber."""
Parameters
----------
tmp_path : pathlib.Path
The path to the test directory.
"""
# Define types # Define types
types = ["T1w", "BOLD"] types = ["T1w", "BOLD"]
# Define patterns # Define patterns
@ -97,15 +90,8 @@ def test_bids_PatternDataladDataGrabber(tmp_path: Path) -> None:
assert f.readlines()[0].startswith("placeholder") assert f.readlines()[0].startswith("placeholder")
def test_bids_PatternDataladDataGrabber_datadir(tmp_path: Path) -> None: def test_bids_PatternDataladDataGrabber_datadir() -> None:
"""Test a datalad datagrabber with a datadir set to a relative path. """Test PatternDataladDataGrabber with a datadir set to a relative path."""
Parameters
----------
tmp_path : pathlib.Path
The path to the test directory.
"""
# Define types # Define types
types = ["T1w", "BOLD"] types = ["T1w", "BOLD"]
# Define patterns # Define patterns
@ -144,7 +130,7 @@ def test_bids_PatternDataladDataGrabber_datadir(tmp_path: Path) -> None:
def test_bids_PatternDataladDataGrabber_session(): def test_bids_PatternDataladDataGrabber_session():
"""Test a subject and session-based BIDS datalad datagrabber.""" """Test a subject and session-based BIDS PatternDataladDataGrabber."""
types = ["T1w", "BOLD"] types = ["T1w", "BOLD"]
patterns = { patterns = {
"T1w": "{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz", "T1w": "{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz",

View file

@ -104,8 +104,8 @@ class MarkerCollection:
Parameters Parameters
---------- ----------
datagrabber : datagrabber-like datagrabber : DataGrabber-like
The datagrabber to validate. The DataGrabber to validate.
""" """
logger.info("Validating Marker Collection") logger.info("Validating Marker Collection")

View file

@ -1,4 +1,4 @@
"""Provide testing datagrabbers.""" """Provide testing DataGrabbers."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de> # Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Synchon Mandal <s.mandal@fz-juelich.de> # Synchon Mandal <s.mandal@fz-juelich.de>
@ -15,9 +15,9 @@ from ..datagrabber.base import BaseDataGrabber
class OasisVBMTestingDataGrabber(BaseDataGrabber): class OasisVBMTestingDataGrabber(BaseDataGrabber):
"""Data Grabber for Oasis VBM testing data. """DataGrabber for Oasis VBM testing data.
Wrapper for :func:`nilearn.datasets.fetch_oasis_vbm` Wrapper for :func:`nilearn.datasets.fetch_oasis_vbm`.
""" """
@ -83,7 +83,7 @@ class OasisVBMTestingDataGrabber(BaseDataGrabber):
class SPMAuditoryTestingDataGrabber(BaseDataGrabber): class SPMAuditoryTestingDataGrabber(BaseDataGrabber):
"""Data Grabber for SPM Auditory dataset. """DataGrabber for SPM Auditory dataset.
Wrapper for :func:`nilearn.datasets.fetch_spm_auditory`. Wrapper for :func:`nilearn.datasets.fetch_spm_auditory`.
@ -147,9 +147,9 @@ class SPMAuditoryTestingDataGrabber(BaseDataGrabber):
class PartlyCloudyTestingDataGrabber(BaseDataGrabber): class PartlyCloudyTestingDataGrabber(BaseDataGrabber):
"""Data Grabber for Partly Cloudy dataset. """DataGrabber for Partly Cloudy dataset.
Wrapper for :func:`nilearn.datasets.fetch_development_fmri` Wrapper for :func:`nilearn.datasets.fetch_development_fmri`.
Parameters Parameters
---------- ----------

View file

@ -12,7 +12,7 @@ from .datagrabbers import (
) )
# Register testing datagrabber # Register testing DataGrabbers
register( register(
step="datagrabber", step="datagrabber",
name="OasisVBMTestingDataGrabber", name="OasisVBMTestingDataGrabber",

View file

@ -1,4 +1,4 @@
"""Provide tests for Oasis VBM Testing datagrabber.""" """Provide tests for OasisVBMTestingDataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de> # Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# License: AGPL # License: AGPL
@ -7,7 +7,7 @@ from junifer.testing.datagrabbers import OasisVBMTestingDataGrabber
def test_OasisVBMTestingDataGrabber() -> None: def test_OasisVBMTestingDataGrabber() -> None:
"""Test Oasis VBM Testing datagrabber.""" """Test OasisVBMTestingDataGrabber."""
expected_elements = [ expected_elements = [
"sub-01", "sub-01",
"sub-02", "sub-02",

View file

@ -1,4 +1,4 @@
"""Provide tests for PartlyCloudy datagrabber.""" """Provide tests for PartlyCloudyTestingDataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de> # Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# License: AGPL # License: AGPL
@ -7,7 +7,7 @@ from junifer.testing.datagrabbers import PartlyCloudyTestingDataGrabber
def test_PartlyCloudyTestingDataGrabber() -> None: def test_PartlyCloudyTestingDataGrabber() -> None:
"""Test PartlyCloudy datagrabber.""" """Test PartlyCloudyTestingDataGrabber."""
expected_elements = [ expected_elements = [
"sub-01", "sub-01",
"sub-02", "sub-02",

View file

@ -1,4 +1,4 @@
"""Provide tests for SPM Auditory datagrabber.""" """Provide tests for SPMAuditoryTestingDataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de> # Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# License: AGPL # License: AGPL
@ -7,7 +7,7 @@ from junifer.testing.datagrabbers import SPMAuditoryTestingDataGrabber
def test_SPMAuditoryTestingDataGrabber() -> None: def test_SPMAuditoryTestingDataGrabber() -> None:
"""Test SPM Auditory datagrabber.""" """Test SPMAuditoryTestingDataGrabber."""
expected_elements = [ expected_elements = [
"sub001", "sub001",
"sub002", "sub002",

View file

@ -1,4 +1,4 @@
"""Create a testing dataset for the DataladAOMICID1000 pattern datagrabber.""" """Create a testing dataset for DataladAOMICID1000."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de> # Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Vera Komeyer <v.komeyer@fz-juelich.de> # Vera Komeyer <v.komeyer@fz-juelich.de>