refactor: consistent use of DataGrabber #226

Merged
synchon merged 32 commits from chore/dg-cleanup into main 2023-06-20 10:27:52 +00:00
48 changed files with 442 additions and 410 deletions

View file

@ -0,0 +1 @@
Adopt ``DataGrabber`` consistently throughout codebase to match with the documentation by `Synchon Mandal`_

View file

@ -22,7 +22,7 @@ data type including source and previous transformation steps.
The :ref:`Data Grabber <datagrabber>` step adds the ``path`` second-level key
which gives the path to the file containing the data. The ``meta`` key in this
step only contains information about the datagrabber used.
step only contains information about the DataGrabber used.
.. code-block:: python

View file

@ -9,29 +9,29 @@ Description
-----------
The ``DataGrabber`` is an object that can provide an interface to datasets you
want to work with in junifer. Every concrete implementation of a datagrabber is
want to work with in junifer. Every concrete implementation of a DataGrabber is
aware of a particular dataset's structure and thus allows you to fetch specific
elements of interest from the dataset. It adds the ``path`` key to each
:ref:`data type <data_types>` in the :ref:`Data object <data_object>`.
Datagrabbers are intended to be used as context managers. When used within a
context, a datagrabber takes care of any pre and post steps for interacting with
DataGrabbers are intended to be used as context managers. When used within a
context, a DataGrabber takes care of any pre and post steps for interacting with
the dataset, for example, downloading and cleaning up. As the interface
is consistent, you always use the same procedure to interact with the datagrabber.
is consistent, you always use the same procedure to interact with the DataGrabber.
For example, a concrete implementation of :class:`.DataladDataGrabber` can
provide junifer with data from a Datalad dataset. Of course, datagrabbers are not
provide junifer with data from a Datalad dataset. Of course, DataGrabbers are not
only meant to work with Datalad datasets but any dataset.
If you are interested in using already provided datagrabbers, please go to
:doc:`../builtin`. And, if you want to implement your own datagrabber, you need
If you are interested in using already provided DataGrabbers, please go to
:doc:`../builtin`. And, if you want to implement your own DataGrabber, you need
to provide concrete implementations of base classes already provided.
Base Classes
------------
In this section, we showcase different abstract base classes you might want to
use to implement your own datagrabber.
use to implement your own DataGrabber.
.. list-table::
:widths: auto
@ -41,16 +41,16 @@ use to implement your own datagrabber.
- Description
* - :class:`.BaseDataGrabber`
- | The abstract base class providing you an interface to implement your
| own datagrabber. You should try to avoid using this directly and
| own DataGrabber. You should try to avoid using this directly and
| instead use :class:`.PatternDataGrabber` or
| :class:`.DataladDataGrabber`. To build your own custom *low-level*
| datagrabber, you need to override the ``get_elements_keys``,
| DataGrabber, you need to override the ``get_elements_keys``,
| ``get_elements`` and ``get_item`` methods, and most of the time you
| should also override other existing methods like ``__enter__`` and
| ``__exit__``.
* - :class:`.PatternDataGrabber`
- | It implements functionality to help you define the pattern of the
| dataset you want to get. For example, you know that T1 images are
| dataset you want to get. For example, you know that T1w images are
| found in a directory following the pattern:
| ``{subject}/anat/{subject}_T1w.nii.gz`` inside of the dataset. Now you
| can provide this to the :class:`.PatternDataGrabber` and it will be

View file

@ -75,10 +75,10 @@ Data Grabber
^^^^^^^^^^^^
The ``datagrabber`` section must be configured using the ``kind`` key to specify
the datagrabber to use. Additional keys correspond to the parameters of the
datagrabber.
the DataGrabber to use. Additional keys correspond to the parameters of the
DataGrabber constructor.
For example, to use the :class:`.DataladAOMICPIOP1` datagrabber, we just need to
For example, to use the :class:`.DataladAOMICPIOP1` DataGrabber, we just need to
specify its name as the ``kind`` key.
.. code-block:: yaml
@ -86,8 +86,9 @@ specify its name as the ``kind`` key.
datagrabber:
kind: DataladAOMICPIOP1
However, it is also possible to pass parameters to the datagrabber. In this case,
we can restrict the datagrabber to fetch only the ``restingstate`` task.
However, it is also possible to pass parameters to the DataGrabber constructor.
In this case, we can restrict the DataGrabber to fetch only the ``restingstate``
task.
.. code-block:: yaml

View file

@ -1,8 +1,8 @@
"""
Generic BIDS datagrabber for datalad.
Generic BIDS DataGrabber for datalad.
=====================================
This example uses a generic BIDS datagraber to get the data from a BIDS dataset
This example uses a generic BIDS DataGraber to get the data from a BIDS dataset
store in a datalad remote sibling.
Authors: Federico Raimondo
@ -20,9 +20,9 @@ configure_logging(level="INFO")
###############################################################################
# The BIDS datagrabber requires three parameters: the types of data we want,
# The BIDS DataGrabber requires three parameters: the types of data we want,
# the specific pattern that matches each type, and the variables that will be
# replaced int he patterns.
# replaced in the patterns.
types = ["T1w", "BOLD"]
patterns = {
"T1w": "{subject}/anat/{subject}_T1w.nii.gz",
@ -30,15 +30,15 @@ patterns = {
}
replacements = ["subject"]
###############################################################################
# Additionally, a datalad datagrabber requires the URI of the remote sibling
# and the location of the dataset within the remote sibling.
# Additionally, a datalad-based DataGrabber requires the URI of the remote
# sibling and the location of the dataset within the remote sibling.
repo_uri = "https://gin.g-node.org/juaml/datalad-example-bids"
rootdir = "example_bids"
###############################################################################
# Now we can use the datagrabber within a `with` context
# One thing we can do with any datagrabber is iterate over the elements.
# In this case, each element of the datagrabber is one session.
# Now we can use the DataGrabber within a `with` context.
# One thing we can do with any DataGrabber is iterate over the elements.
# In this case, each element of the DataGrabber is one session.
with PatternDataladDataGrabber(
rootdir=rootdir,
types=types,
@ -50,9 +50,9 @@ with PatternDataladDataGrabber(
print(elem)
###############################################################################
# Another feature of the datagrabber is the ability to get a specific
# Another feature of the DataGrabber is the ability to get a specific
# element by its name. In this case, we index `sub-01` and we get the file
# paths for the two types of data we want (T1w and bold).
# paths for the two types of data we want (T1w and BOLD).
with PatternDataladDataGrabber(
rootdir=rootdir,
types=types,

View file

@ -25,7 +25,7 @@ from junifer.utils import configure_logging
configure_logging(level="INFO")
##############################################################################
# Define the datagrabber interface
# Define the DataGrabber interface
datagrabber = {
"kind": "SPMAuditoryTestingDataGrabber",
}

View file

@ -11,20 +11,24 @@ from ..pipeline.registry import register
def register_datagrabber(klass: Type) -> Type:
"""Datagrabber registration decorator.
"""Register DataGrabber.
Registers the datagrabber so it can be used by name.
Registers the DataGrabber so it can be used by name.
Parameters
----------
klass: class
The class of the datagrabber to register.
The class of the DataGrabber to register.
Returns
-------
klass: class
The unmodified input class.
Notes
-----
It should only be used as a decorator.
"""
register(
step="datagrabber",

View file

@ -24,17 +24,17 @@ from .utils import yaml
def _get_datagrabber(datagrabber_config: Dict) -> BaseDataGrabber:
"""Get datagrabber.
"""Get DataGrabber.
Parameters
----------
datagrabber_config : dict
The config to get the datagrabber using.
The config to get the DataGrabber using.
Returns
-------
dict
The datagrabber.
object
The DataGrabber.
"""
datagrabber_params = datagrabber_config.copy()
@ -90,8 +90,8 @@ def run(
workdir : str or pathlib.Path
Directory where the pipeline will be executed.
datagrabber : dict
Datagrabber to use. Must have a key ``kind`` with the kind of
datagrabber to use. All other keys are passed to the datagrabber
DataGrabber to use. Must have a key ``kind`` with the kind of
DataGrabber to use. All other keys are passed to the DataGrabber
init function.
markers : list of dict
List of markers to extract. Each marker is a dict with at least two
@ -108,7 +108,7 @@ def run(
preprocessor to use. All other keys are passed to the preprocessor
init function (default None).
elements : str or tuple or list of str or tuple, optional
Element(s) to process. Will be used to index the datagrabber
Element(s) to process. Will be used to index the DataGrabber
(default None).
"""
@ -220,7 +220,7 @@ def queue(
overwrite : bool, optional
Whether to overwrite if job directory already exists (default False).
elements : str or tuple or list of str or tuple, optional
Element(s) to process. Will be used to index the datagrabber
Element(s) to process. Will be used to index the DataGrabber
(default None).
**kwargs : dict
The keyword arguments to pass to the job queue system.
@ -339,7 +339,7 @@ def _queue_condor(
yaml_config : pathlib.Path
The path to the YAML config file.
elements : list of str or tuple
Element(s) to process. Will be used to index the datagrabber.
Element(s) to process. Will be used to index the DataGrabber.
config : dict
The configuration to be used for queueing the job.
env : dict, optional
@ -559,7 +559,7 @@ def _queue_slurm(
yaml_config : pathlib.Path
The path to the YAML config file.
elements : str or tuple or list[str or tuple], optional
Element(s) to process. Will be used to index the datagrabber
Element(s) to process. Will be used to index the DataGrabber
(default None).
config : dict
The configuration to be used for queueing the job.

View file

@ -1,4 +1,4 @@
"""Provide class for AOMIC1000 VBM juseless datalad datagrabber."""
"""Provide concrete implementation for AOMIC ID1000 VBM DataGrabber."""
# Authors: Felix Hoffstaedter <f.hoffstaedter@fz-juelich.de>
# Synchon Mandal <s.mandal@fz-juelich.de>
@ -13,13 +13,13 @@ from ....datagrabber import PatternDataladDataGrabber
@register_datagrabber
class JuselessDataladAOMICID1000VBM(PatternDataladDataGrabber):
"""Juseless AOMICID1000 VBM Data Grabber class.
"""Concrete implementation for Juseless AOMIC ID1000 VBM data fetching.
Implements a Data Grabber to access the AOMICID1000 VBM data in Juseless.
Implements a DataGrabber to access the AOMIC ID1000 VBM data in Juseless.
Parameters
----------
datadir : str or pathlib.Path, optional
datadir : str or pathlib.Path or None, optional
The directory where the datalad dataset will be cloned. If None,
the datalad dataset will be cloned into a temporary directory
(default None).

View file

@ -1,4 +1,4 @@
"""Provide class for CamCAN VBM juseless datalad datagrabber."""
"""Provide concrete implementation for CamCAN VBM DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de>
@ -14,13 +14,13 @@ from ....datagrabber import PatternDataladDataGrabber
@register_datagrabber
class JuselessDataladCamCANVBM(PatternDataladDataGrabber):
"""Juseless CamCAN VBM Data Grabber class.
"""Concrete implementation for Juseless CamCAN VBM data fetching.
Implements a Data Grabber to access the CamCAN VBM data in Juseless.
Implements a DataGrabber to access the CamCAN VBM data in Juseless.
Parameters
----------
datadir : str or pathlib.Path, optional
datadir : str or pathlib.Path or None, optional
The directory where the datalad dataset will be cloned. If None,
the datalad dataset will be cloned into a temporary directory
(default None).

View file

@ -1,4 +1,4 @@
"""Provide class for IXI VBM juseless datalad datagrabber."""
"""Provide concrete implementation for IXI VBM DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de>
@ -15,19 +15,20 @@ from ....utils import raise_error
@register_datagrabber
class JuselessDataladIXIVBM(PatternDataladDataGrabber):
"""Juseless IXI VBM Data Grabber class.
"""Concrete implementation for Juseless IXI VBM data fetching.
Implements a Data Grabber to access the IXI VBM data in Juseless.
Implements a DataGrabber to access the IXI VBM data in Juseless.
Parameters
----------
datadir : str or pathlib.Path, optional
datadir : str or pathlib.Path or None, optional
The directory where the datalad dataset will be cloned. If None,
the datalad dataset will be cloned into a temporary directory
(default None).
sites : {"Guys", "HH", "IOP"} or list of the options, optional.
sites : {"Guys", "HH", "IOP"} or list of the options or None, optional
Which sites to access data from. If None, all available sites are
selected (default None).
"""
def __init__(

View file

@ -1,4 +1,4 @@
"""Provide tests for AOMICID1000 VBM juseless datagrabber."""
"""Provide tests for JuselessDataladAOMICID1000VBM."""
# Authors: Felix Hoffstaedter <f.hoffstaedter@fz-juelich.de>
# Synchon Mandal <s.mandal@fz-juelich.de>
@ -19,8 +19,8 @@ if socket.gethostname() != "juseless":
configure_logging(level="DEBUG")
def test_juselessdataladaomicid1000vbm_datagrabber() -> None:
"""Test datalad AOMICID1000VBM datagrabber."""
def test_JuselessDataladAOMICID1000VBM() -> None:
"""Test JuselessDataladAOMICID1000VBM."""
with JuselessDataladAOMICID1000VBM() as dg:
all_elements = dg.get_elements()
test_element = all_elements[0]

View file

@ -1,4 +1,4 @@
"""Provide tests for CamCAN VBM juseless datagrabber."""
"""Provide tests for JuselessDataladCamCANVBM."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de>
@ -20,8 +20,8 @@ if socket.gethostname() != "juseless":
configure_logging(level="DEBUG")
def test_juselessdataladcamcanvbm_datagrabber() -> None:
"""Test datalad CamCANVBM datagrabber."""
def test_JuselessDataladCamCANVBM() -> None:
"""Test JuselessDataladCamCANVBM."""
with JuselessDataladCamCANVBM() as dg:
all_elements = dg.get_elements()
test_element = all_elements[0]

View file

@ -1,4 +1,4 @@
"""Provide tests for IXI VBM juseless datagrabber."""
"""Provide tests for JuselessDataladIXIVBM."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de>
@ -20,8 +20,8 @@ if socket.gethostname() != "juseless":
configure_logging(level="DEBUG")
def test_juselessdataladixivbm_datagrabber() -> None:
"""Test datalad IXIVBM datagrabber."""
def test_JuselessDataladIXIVBM() -> None:
"""Test JuselessDataladIXIVBM."""
with JuselessDataladIXIVBM() as dg:
all_elements = dg.get_elements()
test_element = all_elements[0]
@ -33,8 +33,8 @@ def test_juselessdataladixivbm_datagrabber() -> None:
assert out["VBM_GM"]["path"].exists()
def test_juselessdataladixivbm_datagrabber_invalid_site() -> None:
"""Test datalad IXIVBM datagrabber with invalid site."""
def test_JuselessDataladIXIVBM_invalid_site() -> None:
"""Test JuselessDataladIXIVBM with invalid site."""
with pytest.raises(ValueError, match="notavalidsite not a valid site"):
with JuselessDataladIXIVBM(sites="notavalidsite"):
pass

View file

@ -1,4 +1,4 @@
"""Provide tests for UCLA juseless datagrabber."""
"""Provide tests for JuselessUCLA."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de>
@ -21,8 +21,8 @@ if socket.gethostname() != "juseless":
configure_logging(level="DEBUG")
def test_juseless_ucla_datagrabber() -> None:
"""Test juseless ucla datagrabber."""
def test_JuselessUCLA() -> None:
"""Test JuselessUCLA."""
with JuselessUCLA() as dg:
all_elements = dg.get_elements()
test_element = all_elements[0]
@ -46,8 +46,8 @@ def test_juseless_ucla_datagrabber() -> None:
"tasks",
[None, "rest", ["rest", "stopsignal"]],
)
def test_juseless_ucla_datagrabber_task_params(tasks: Optional[str]) -> None:
"""Test juseless ucla datagrabber with different task parameters.
def test_JuselessUCLA_task_params(tasks: Optional[str]) -> None:
"""Test JuselessUCLA with different task parameters.
Parameters
----------
@ -81,8 +81,8 @@ def test_juseless_ucla_datagrabber_task_params(tasks: Optional[str]) -> None:
assert el[1] in ["rest", "stopsignal"]
def test_juseless_ucla_datagrabber_invalid_tasks() -> None:
"""Test juseless ucla datagrabber with invalid task parameters."""
def test_JuselessUCLA_invalid_tasks() -> None:
"""Test JuselessUCLA with invalid task parameters."""
with pytest.raises(
ValueError, match="invalid is not a valid task in the UCLA"
):

View file

@ -1,4 +1,4 @@
"""Provide tests for juseless datagrabber."""
"""Provide tests for JuselessDataladUKBVBM."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de>
@ -20,8 +20,8 @@ if socket.gethostname() != "juseless":
configure_logging(level="DEBUG")
def test_juselessdataladukbvbm_datagrabber() -> None:
"""Test datalad UKBVBM datagrabber."""
def test_JuselessDataladUKBVBM() -> None:
"""Test JuselessDataladUKBVBM."""
with JuselessDataladUKBVBM() as dg:
all_elements = dg.get_elements()
test_element = all_elements[0]

View file

@ -1,4 +1,4 @@
"""Provide a concrete implementation for UCLA dataset."""
"""Provide concrete implementation for UCLA DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de>
@ -14,16 +14,18 @@ from ....utils import raise_error
@register_datagrabber
class JuselessUCLA(PatternDataGrabber):
"""Concrete implementation for pattern-based data fetching of UCLA data.
"""Concrete implementation for Juseless UCLA data fetching.
Implements a DataGrabber to access the UCLA data in Juseless.
Parameters
----------
datadir : str or Path, optional
datadir : str or pathlib.Path, optional
The directory where the dataset is stored
(default '/data/project/psychosis_thalamus/data/fmriprep').
(default "/data/project/psychosis_thalamus/data/fmriprep").
tasks : {"rest", "bart", "bht", "pamenc", "pamret", \
"scap", "taskswitch", "stopsignal"} or \
list of the options, optional
list of the options or None, optional
UCLA task sessions. If None, all available task sessions are
selected (default None).
@ -118,5 +120,6 @@ class JuselessUCLA(PatternDataGrabber):
elements : list
The list of elements that can be grabbed in the dataset after
imposing constraints based on specified tasks.
"""
return [x for x in super().get_elements() if x[1] in self.tasks]

View file

@ -1,4 +1,4 @@
"""Provide class for juseless datalad datagrabber."""
"""Provide concrete implementation for UKB VBM DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de>
@ -14,13 +14,13 @@ from ....datagrabber import PatternDataladDataGrabber
@register_datagrabber
class JuselessDataladUKBVBM(PatternDataladDataGrabber):
"""Juseless UKB VBM Data Grabber class.
"""Concrete implementation for Juseless UKB VBM data fetching.
Implements a Data Grabber to access the UKB VBM data in Juseless.
Implements a DataGrabber to access the UKB VBM data in Juseless.
Parameters
----------
datadir : str or pathlib.Path, optional
datadir : str or pathlib.Path or None, optional
The directory where the datalad dataset will be cloned. If None,
the datalad dataset will be cloned into a temporary directory
(default None).

View file

@ -1,32 +0,0 @@
"""Provide tests for juseless datagrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de>
# Synchon Mandal <s.mandal@fz-juelich.de>
# License: AGPL
import socket
import pytest
from junifer.datagrabber.hcp import DataladHCP1200
from junifer.utils.logging import configure_logging
# Check if the test is running on juseless
if socket.gethostname() != "juseless":
pytest.skip("These tests are only for juseless", allow_module_level=True)
configure_logging(level="DEBUG")
def test_juselessdataladhcp_datagrabber() -> None:
"""Test datalad HCP datagrabber."""
with DataladHCP1200() as dg:
all_elements = dg.get_elements()
test_element = all_elements[0]
out = dg[test_element]
assert out["BOLD"]["path"].exists()
assert out["BOLD"]["path"].isfile()

View file

@ -2,6 +2,7 @@
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de>
# Synchon Mandal <s.mandal@fz-juelich.de>
# License: AGPL
@ -12,5 +13,5 @@ from .pattern import PatternDataGrabber
from .pattern_datalad import PatternDataladDataGrabber
from .aomic import DataladAOMICID1000, DataladAOMICPIOP1, DataladAOMICPIOP2
from .hcp import HCP1200, DataladHCP1200
from .hcp1200 import HCP1200, DataladHCP1200
from .multiple import MultipleDataGrabber

View file

@ -1,4 +1,4 @@
"""Provide imports for datagrabber 'aomic' sub-package."""
"""Provide imports for aomic sub-package."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de>

View file

@ -1,4 +1,4 @@
"""Provide concrete implementations for AOMIC1000 data access."""
"""Provide concrete implementation for AOMIC ID1000 DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Vera Komeyer <v.komeyer@fz-juelich.de>
@ -9,21 +9,21 @@
from pathlib import Path
from typing import Dict, Union
from junifer.datagrabber import PatternDataladDataGrabber
from ...api.decorators import register_datagrabber
from ..pattern_datalad import PatternDataladDataGrabber
@register_datagrabber
class DataladAOMICID1000(PatternDataladDataGrabber):
"""Concrete implementation for pattern-based data fetching of AOMICID1000.
"""Concrete implementation for datalad-based data fetching of AOMIC ID1000.
Parameters
----------
datadir : str or Path, optional
datadir : str or Path or None, optional
The directory where the datalad dataset will be cloned. If None,
the datalad dataset will be cloned into a temporary directory
(default None).
"""
def __init__(
@ -115,6 +115,7 @@ class DataladAOMICID1000(PatternDataladDataGrabber):
out : dict
Dictionary of paths for each type of data required for the
specified element.
"""
out = super().get_item(subject=subject)
out["BOLD"]["mask_item"] = "BOLD_mask"

View file

@ -1,4 +1,4 @@
"""Provide concrete implementations for AOMICPIOP1 data access."""
"""Provide concrete implementation for AOMIC PIOP1 DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Vera Komeyer <v.komeyer@fz-juelich.de>
@ -10,19 +10,18 @@ from itertools import product
from pathlib import Path
from typing import Dict, List, Union
from junifer.datagrabber import PatternDataladDataGrabber
from ...api.decorators import register_datagrabber
from ...utils import raise_error
from ..pattern_datalad import PatternDataladDataGrabber
@register_datagrabber
class DataladAOMICPIOP1(PatternDataladDataGrabber):
"""Concrete implementation for pattern-based data fetching of AOMICPIOP1.
"""Concrete implementation for pattern-based data fetching of AOMIC PIOP1.
Parameters
----------
datadir : str or Path, optional
datadir : str or Path or None, optional
The directory where the datalad dataset will be cloned. If None,
the datalad dataset will be cloned into a temporary directory
(default None).
@ -30,6 +29,7 @@ class DataladAOMICPIOP1(PatternDataladDataGrabber):
"gstroop", "workingmemory"} or list of the options, optional
AOMIC PIOP1 task sessions. If None, all available task sessions are
selected (default None).
"""
def __init__(
@ -147,6 +147,7 @@ class DataladAOMICPIOP1(PatternDataladDataGrabber):
out : dict
Dictionary of paths for each type of data required for the
specified element.
"""
task_acqs = {

View file

@ -1,4 +1,4 @@
"""Provide concrete implementations for AOMICPIOP2 data access."""
"""Provide concrete implementation for AOMIC PIOP2 DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Vera Komeyer <v.komeyer@fz-juelich.de>
@ -9,19 +9,18 @@
from pathlib import Path
from typing import Dict, List, Union
from junifer.datagrabber import PatternDataladDataGrabber
from ...api.decorators import register_datagrabber
from ...utils import raise_error
from ..pattern_datalad import PatternDataladDataGrabber
@register_datagrabber
class DataladAOMICPIOP2(PatternDataladDataGrabber):
"""Concrete implementation for pattern-based data fetching of AOMICPIOP2.
"""Concrete implementation for pattern-based data fetching of AOMIC PIOP2.
Parameters
----------
datadir : str or Path, optional
datadir : str or Path or None, optional
The directory where the datalad dataset will be cloned. If None,
the datalad dataset will be cloned into a temporary directory
(default None).
@ -29,6 +28,7 @@ class DataladAOMICPIOP2(PatternDataladDataGrabber):
or list of the options, optional
AOMIC PIOP2 task sessions. If None, all available task sessions are
selected (default None).
"""
def __init__(
@ -136,6 +136,7 @@ class DataladAOMICPIOP2(PatternDataladDataGrabber):
list
The list of elements that can be grabbed in the dataset after
imposing constraints based on specified tasks.
"""
all_elements = super().get_elements()
return [x for x in all_elements if x[1] in self.tasks]
@ -156,6 +157,7 @@ class DataladAOMICPIOP2(PatternDataladDataGrabber):
out : dict
Dictionary of paths for each type of data required for the
specified element.
"""
out = super().get_item(subject=subject, task=task)
out["BOLD"]["mask_item"] = "BOLD_mask"

View file

@ -1,4 +1,4 @@
"""Provide tests for aomicid1000."""
"""Provide tests for DataladAOMICID1000 DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Vera Komeyer <v.komeyer@fz-juelich.de>
@ -6,12 +6,12 @@
# Leonard Sasse <l.sasse@fz-juelich.de>
# License: AGPL
from junifer.datagrabber.aomic.id1000 import DataladAOMICID1000
from junifer.datagrabber import DataladAOMICID1000
from junifer.utils import configure_logging
def test_aomic1000_datagrabber() -> None:
"""Test datalad AOMIC1000 datagrabber."""
def test_DataladAOMICID1000() -> None:
"""Test DataladAOMICID1000 DataGrabber."""
uri_ID1000 = "https://gin.g-node.org/juaml/datalad-example-aomic1000"
configure_logging(level="DEBUG")

View file

@ -1,4 +1,4 @@
"""Provide tests for aomic piop1."""
"""Provide tests for DataladAOMICPIOP1 DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Vera Komeyer <v.komeyer@fz-juelich.de>
@ -8,12 +8,12 @@
import pytest
from junifer.datagrabber.aomic.piop1 import DataladAOMICPIOP1
from junifer.datagrabber import DataladAOMICPIOP1
from junifer.utils import configure_logging
def test_aomic_piop1_datagrabber() -> None:
"""Test datalad AOMICPIOP1 datagrabber."""
def test_DataladAOMICPIOP1() -> None:
"""Test DataladAOMICPIOP1 DataGrabber."""
configure_logging(level="DEBUG")
uri_PIOP1 = "https://gin.g-node.org/juaml/datalad-example-aomicpiop1"
@ -138,7 +138,7 @@ def test_aomic_piop1_datagrabber() -> None:
assert sub == meta["element"]["subject"]
def test_piop1_invalid_tasks():
def test_DataladAOMICPIOP1_invalid_tasks():
"""Test whether invalid task fails."""
with pytest.raises(
ValueError,

View file

@ -1,4 +1,4 @@
"""Provide tests for aomic piop2."""
"""Provide tests DataladAOMICPIOP2 DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Vera Komeyer <v.komeyer@fz-juelich.de>
@ -8,12 +8,12 @@
import pytest
from junifer.datagrabber.aomic import DataladAOMICPIOP2
from junifer.datagrabber import DataladAOMICPIOP2
from junifer.utils import configure_logging
def test_aomic_piop2_datagrabber() -> None:
"""Test datalad AOMICPIOP2 datagrabber."""
def test_DataladAOMICPIOP2() -> None:
"""Test DataladAOMICPIOP2 DataGrabber."""
configure_logging(level="DEBUG")
uri_PIOP2 = "https://gin.g-node.org/juaml/datalad-example-aomicpiop2"
@ -132,7 +132,7 @@ def test_aomic_piop2_datagrabber() -> None:
assert sub == meta["element"]["subject"]
def test_piop2_invalid_tasks():
def test_DataladAOMICPIOP2_invalid_tasks():
"""Test whether invalid task fails."""
with pytest.raises(
ValueError,

View file

@ -1,4 +1,4 @@
"""Provide abstract base class for datagrabber."""
"""Provide abstract base class for DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de>
@ -15,7 +15,7 @@ from .utils import validate_types
class BaseDataGrabber(ABC, UpdateMetaMixin):
"""Abstract base class for datagrabber.
"""Abstract base class for DataGrabber.
For every interface that is required, one needs to provide a concrete
implementation of this abstract class.
@ -31,6 +31,7 @@ class BaseDataGrabber(ABC, UpdateMetaMixin):
----------
datadir : pathlib.Path
The directory where the data is / will be stored.
"""
def __init__(self, types: List[str], datadir: Union[str, Path]) -> None:
@ -52,7 +53,7 @@ class BaseDataGrabber(ABC, UpdateMetaMixin):
Yields
------
object
An element that can be indexed by the datagrabber.
An element that can be indexed by the DataGrabber.
"""
for elem in self.get_elements():
@ -126,7 +127,7 @@ class BaseDataGrabber(ABC, UpdateMetaMixin):
Returns
-------
str
list of str
The element keys.
"""
@ -144,7 +145,7 @@ class BaseDataGrabber(ABC, UpdateMetaMixin):
list
List of elements that can be grabbed. The elements can be strings,
tuples or any object that will be then used as a key to index the
datagrabber.
DataGrabber.
"""
raise_error(

View file

@ -1,4 +1,4 @@
"""Provide abstract base class for datalad datagrabber."""
"""Provide abstract base class for datalad-based DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de>
@ -24,20 +24,20 @@ from .base import BaseDataGrabber
@register_datagrabber
class DataladDataGrabber(BaseDataGrabber):
"""Abstract base class for data fetching via Datalad.
"""Abstract base class for datalad-based data fetching.
Defines a Data Grabber that gets data from a datalad sibling.
Defines a DataGrabber that gets data from a datalad sibling.
Parameters
----------
rootdir : str or Path, optional
rootdir : str or pathlib.Path, optional
The path within the datalad dataset to the root directory
(default ".").
datadir : str or Path, optional
datadir : str or pathlib.Path or None, optional
That directory where the datalad dataset will be cloned. If None,
the datalad dataset will be cloned into a temporary directory
(default None).
uri : str, optional
uri : str or None, optional
URI of the datalad sibling (default None).
**kwargs
Keyword arguments passed to superclass.
@ -45,23 +45,26 @@ class DataladDataGrabber(BaseDataGrabber):
Methods
-------
install:
Installs (clones) the datalad dataset into the `datadir`. This method
Installs (clones) the datalad dataset into the ``datadir``. This method
is called automatically when the datagrabber is used within a context.
remove:
Removes the datalad dataset from the `datadir`. This method is called
Removes the datalad dataset from the ``datadir``. This method is called
automatically when the datagrabber is used within a context.
See Also
--------
BaseDataGrabber
For the abstract base class of any datagrabber.
BaseDataGrabber:
Abstract base class for DataGrabber.
PatternDataGrabber:
Concrete implementation for pattern-based data fetching.
PatternDataladDataGrabber:
Concrete implementation for pattern and datalad based data fetching.
Notes
-----
By itself, this class is still abstract as the `__getitem__` method relies
on the parent class `BaseDataGrabber.__getitem__` which is not yet
implemented. This class is intended to be used as a superclass of a class
This class is intended to be used as a superclass of a subclass
with multiple inheritance.
"""
def __init__(
@ -129,16 +132,23 @@ class DataladDataGrabber(BaseDataGrabber):
@property
def datadir(self) -> Path:
"""Get data directory path."""
"""Get data directory path.
Returns
-------
pathlib.Path
Path to the data directory.
"""
return super().datadir / self._rootdir
def _get_dataset_id_remote(self) -> str:
"""Get the dataset id from the remote.
"""Get the dataset ID from the remote.
Returns
-------
str
The dataset id.
The dataset ID.
"""
remote_id = None
@ -155,7 +165,7 @@ class DataladDataGrabber(BaseDataGrabber):
return remote_id
def _dataset_get(self, out: Dict) -> Dict:
"""Get the dataset found from the path in `out`.
"""Get the dataset found from the path in ``out``.
Parameters
----------
@ -165,7 +175,7 @@ class DataladDataGrabber(BaseDataGrabber):
Returns
-------
dict
The modified dictionary with meta updated.
The unmodified input dictionary.
"""
to_get = [v["path"] for v in out.values() if "path" in v]
@ -197,12 +207,13 @@ class DataladDataGrabber(BaseDataGrabber):
return out
def install(self) -> None:
"""Install the datalad dataset into the datadir.
"""Install the datalad dataset into the ``datadir``.
Raises
------
ValueError
If the dataset is already installed but with a different ID.
"""
isinstalled = dl.Dataset(self._datadir).is_installed()
if isinstalled:
@ -256,9 +267,7 @@ class DataladDataGrabber(BaseDataGrabber):
"""Implement single element indexing in the Datalad database.
It will first obtain the paths from the parent class and then
`datalad get` each of the files.
This method only works with multiple inheritance.
``datalad get`` each of the files.
Parameters
----------
@ -279,12 +288,12 @@ class DataladDataGrabber(BaseDataGrabber):
out = self._dataset_get(out)
return out
def __enter__(self):
def __enter__(self) -> "DataladDataGrabber":
"""Implement context entry."""
self.install()
return self
def __exit__(self, exc_type, exc_value, exc_traceback):
def __exit__(self, exc_type, exc_value, exc_traceback) -> None:
"""Implement context exit."""
logger.debug("Cleaning up dataset")
self.cleanup()

View file

@ -0,0 +1,7 @@
"""Provide imports for hcp1200 sub-package."""
# Authors: Synchon Mandal <s.mandal@fz-juelich.de>
# License: AGPL
from .hcp1200 import HCP1200
from .datalad_hcp1200 import DataladHCP1200

View file

@ -0,0 +1,68 @@
"""Provide concrete implementation for datalad-based HCP1200 DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de>
# Synchon Mandal <s.mandal@fz-juelich.de>
# License: AGPL
from pathlib import Path
from typing import List, Union
from junifer.datagrabber.datalad_base import DataladDataGrabber
from ...api.decorators import register_datagrabber
from .hcp1200 import HCP1200
@register_datagrabber
class DataladHCP1200(DataladDataGrabber, HCP1200):
"""Concrete implementation for datalad-based data fetching of HCP1200.
Parameters
----------
datadir : str or Path or None, optional
The directory where the datalad dataset will be cloned. If None,
the datalad dataset will be cloned into a temporary directory
(default None).
tasks : {"REST1", "REST2", "SOCIAL", "WM", "RELATIONAL", "EMOTION", \
"LANGUAGE", "GAMBLING", "MOTOR"} or list of the options or None \
, optional
HCP task sessions. If None, all available task sessions are selected
(default None).
phase_encodings : {"LR", "RL"} or list of the options or None, optional
HCP phase encoding directions. If None, both will be used
(default None).
ica_fix : bool, optional
Whether to retrieve data that was processed with ICA+FIX.
Only "REST1" and "REST2" tasks are available with ICA+FIX (default
False).
"""
def __init__(
self,
datadir: Union[str, Path, None] = None,
tasks: Union[str, List[str], None] = None,
phase_encodings: Union[str, List[str], None] = None,
ica_fix: bool = False,
) -> None:
uri = (
"https://github.com/datalad-datasets/"
"human-connectome-project-openaccess.git"
)
rootdir = "HCP1200"
super().__init__(
datadir=datadir,
tasks=tasks,
phase_encodings=phase_encodings,
uri=uri,
rootdir=rootdir,
ica_fix=ica_fix,
)
# Needed here as HCP1200's subjects are sub-datasets, so will not be
# found when elements are checked.
@property
def skip_file_check(self) -> bool:
"""Skip file check existence."""
return True

View file

@ -1,14 +1,17 @@
"""Provide concrete implementations for HCP data access."""
"""Provide concrete implementation for pattern-based HCP1200 DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de>
# Synchon Mandal <s.mandal@fz-juelich.de>
# License: AGPL
from itertools import product
from pathlib import Path
from typing import Dict, List, Union
from junifer.datagrabber.datalad_base import DataladDataGrabber
from ..api.decorators import register_datagrabber
from ...api.decorators import register_datagrabber
from ..pattern import PatternDataGrabber
from ..utils import raise_error
from .pattern import PatternDataGrabber
@register_datagrabber
@ -18,20 +21,22 @@ class HCP1200(PatternDataGrabber):
Parameters
----------
datadir : str or Path, optional
The directory where the datalad dataset will be cloned.
The directory where the data is / will be stored.
tasks : {"REST1", "REST2", "SOCIAL", "WM", "RELATIONAL", "EMOTION", \
"LANGUAGE", "GAMBLING", "MOTOR"} or list of the options, optional
"LANGUAGE", "GAMBLING", "MOTOR"} or list of the options or None \
, optional
HCP task sessions. If None, all available task sessions are selected
(default None).
phase_encodings : {"LR", "RL"} or list of the options, optional
phase_encodings : {"LR", "RL"} or list of the options or None, optional
HCP phase encoding directions. If None, both will be used
(default None).
ica_fix : bool, optional
Whether to retrieve data that was processed with ICA+FIX.
Only 'REST1' and 'REST2' tasks are available with ICA+FIX (default
Only "REST1" and "REST2" tasks are available with ICA+FIX (default
False).
**kwargs
Keyword arguments passed to superclass.
"""
def __init__(
@ -114,7 +119,7 @@ class HCP1200(PatternDataGrabber):
self.phase_encodings = phase_encodings
def get_item(self, subject: str, task: str, phase_encoding: str) -> Dict:
"""Index one element in the dataset.
"""Implement single element indexing in the database.
Parameters
----------
@ -128,8 +133,8 @@ class HCP1200(PatternDataGrabber):
Returns
-------
out : dict
Dictionary of paths for each type of data required for the
dict
Dictionary of dictionaries for each type of data required for the
specified element.
"""
@ -150,7 +155,7 @@ class HCP1200(PatternDataGrabber):
Returns
-------
list
The list of elements in the dataset.
The list of elements that can be grabbed in the dataset.
"""
subjects = [
@ -165,53 +170,3 @@ class HCP1200(PatternDataGrabber):
elems.append((subject, task, phase_encoding))
return elems
@register_datagrabber
class DataladHCP1200(DataladDataGrabber, HCP1200):
"""Concrete implementation for datalad-based data fetching of HCP1200.
Parameters
----------
datadir : str or Path, optional
The directory where the datalad dataset will be cloned. If None,
the datalad dataset will be cloned into a temporary directory
(default None).
tasks : {"REST1", "REST2", "SOCIAL", "WM", "RELATIONAL", "EMOTION", \
"LANGUAGE", "GAMBLING", "MOTOR"} or list of the options, optional
HCP task sessions. If None, all available task sessions are selected
(default None).
phase_encodings : {"LR", "RL"} or list of the options, optional
HCP phase encoding directions. If None, both will be used
(default None).
ica_fix : bool, optional
Whether to retrieve data that was processed with ICA+FIX.
Only 'REST1' and 'REST2' tasks are available with ICA+FIX (default
False).
"""
def __init__(
self,
datadir: Union[str, Path, None] = None,
tasks: Union[str, List[str], None] = None,
phase_encodings: Union[str, List[str], None] = None,
ica_fix: bool = False,
) -> None:
uri = (
"https://github.com/datalad-datasets/"
"human-connectome-project-openaccess.git"
)
rootdir = "HCP1200"
super().__init__(
datadir=datadir,
tasks=tasks,
phase_encodings=phase_encodings,
uri=uri,
rootdir=rootdir,
ica_fix=ica_fix,
)
@property
def skip_file_check(self) -> bool:
"""Skip file check existence."""
return True

View file

@ -7,7 +7,7 @@ from typing import Iterable, Optional
import pytest
from junifer.datagrabber.hcp import HCP1200, DataladHCP1200
from junifer.datagrabber import HCP1200, DataladHCP1200
from junifer.utils import configure_logging
@ -16,7 +16,7 @@ URI = "https://gin.g-node.org/juaml/datalad-example-hcp1200"
@pytest.fixture(scope="module")
def hcpdg() -> Iterable[DataladHCP1200]:
"""Return a HCP1200 datagrabber."""
"""Return a HCP1200 DataGrabber."""
dg = DataladHCP1200()
# Set URI to Gin
dg.uri = URI
@ -56,19 +56,19 @@ def hcpdg() -> Iterable[DataladHCP1200]:
("REST2", "RL", True, "rfMRI_REST2_RL_hp2000_clean.nii.gz"),
],
)
def test_hcp1200_datagrabber(
def test_HCP1200(
hcpdg: DataladHCP1200,
tasks: Optional[str],
phase_encodings: Optional[str],
ica_fix: bool,
expected_path_name: str,
) -> None:
"""Test HCP1200 datagrabber.
"""Test HCP1200 DataGrabber.
Parameters
----------
hcpdg : DataladHCP1200
The Datalad version of the datagrabber with the first subject
The Datalad version of the DataGrabber with the first subject
already cloned.
tasks : str
The parametrized tasks.
@ -132,17 +132,17 @@ def test_hcp1200_datagrabber(
("MOTOR", "RL"),
],
)
def test_hcp1200_datagrabber_single_access(
def test_HCP1200_single_access(
hcpdg: DataladHCP1200,
tasks: Optional[str],
phase_encodings: Optional[str],
) -> None:
"""Test HCP1200 datagrabber single access.
"""Test HCP1200 DataGrabber single access.
Parameters
----------
hcpdg : DataladHCP1200
The Datalad version of the datagrabber with the first subject
The Datalad version of the DataGrabber with the first subject
already cloned.
tasks : str
The parametrized tasks.
@ -172,17 +172,17 @@ def test_hcp1200_datagrabber_single_access(
(["REST1", "REST2"], None),
],
)
def test_hcp1200_datagrabber_multi_access(
def test_HCP1200_multi_access(
hcpdg: DataladHCP1200,
tasks: Optional[str],
phase_encodings: Optional[str],
) -> None:
"""Test HCP1200 datagrabber multiple access.
"""Test HCP1200 DataGrabber multiple access.
Parameters
----------
hcpdg : DataladHCP1200
The Datalad version of the datagrabber with the first subject
The Datalad version of the DataGrabber with the first subject
already cloned.
tasks : str
The parametrized tasks.
@ -205,16 +205,17 @@ def test_hcp1200_datagrabber_multi_access(
assert element[2] in ["LR", "RL"]
def test_hcp1200_datagrabber_multi_access_task_simple(
def test_HCP1200_multi_access_task_simple(
hcpdg: DataladHCP1200,
) -> None:
"""Test HCP1200 datagrabber simple multiple access for task.
"""Test HCP1200 DataGrabber simple multiple access for task.
Parameters
----------
hcpdg : DataladHCP1200
The Datalad version of the datagrabber with the first subject
The Datalad version of the DataGrabber with the first subject
already cloned.
"""
configure_logging(level="DEBUG")
dg = HCP1200(
@ -231,16 +232,17 @@ def test_hcp1200_datagrabber_multi_access_task_simple(
assert element[2] in ["LR", "RL"]
def test_hcp1200_datagrabber_multi_access_phase_simple(
def test_HCP1200_multi_access_phase_simple(
hcpdg: DataladHCP1200,
) -> None:
"""Test HCP1200 datagrabber simple multiple access for phase.
"""Test HCP1200 DataGrabber simple multiple access for phase.
Parameters
----------
hcpdg : DataladHCP1200
The Datalad version of the datagrabber with the first subject
The Datalad version of the DataGrabber with the first subject
already cloned.
"""
configure_logging(level="DEBUG")
dg = HCP1200(
@ -266,11 +268,11 @@ def test_hcp1200_datagrabber_multi_access_phase_simple(
(["FOO", "BAR"], "LR"),
],
)
def test_hcp1200_datagrabber_incorrect_access_task(
def test_HCP1200_incorrect_access_task(
tasks: Optional[str],
phase_encodings: Optional[str],
) -> None:
"""Test HCP1200 datagrabber incorrect access for task.
"""Test HCP1200 DataGrabber incorrect access for task.
Parameters
----------
@ -298,11 +300,11 @@ def test_hcp1200_datagrabber_incorrect_access_task(
(["REST1", "REST2"], "BAR"),
],
)
def test_hcp1200_datagrabber_incorrect_access_phase(
def test_HCP1200_incorrect_access_phase(
tasks: Optional[str],
phase_encodings: Optional[str],
) -> None:
"""Test HCP1200 datagrabber incorrect access for phase.
"""Test HCP1200 DataGrabber incorrect access for phase.
Parameters
----------
@ -321,16 +323,17 @@ def test_hcp1200_datagrabber_incorrect_access_phase(
)
def test_hcp1200_datagrabber_elements(
def test_HCP1200_elements(
hcpdg: DataladHCP1200,
) -> None:
"""Test HCP1200 datagrabber elements.
"""Test HCP1200 DataGrabber elements.
Parameters
----------
hcpdg : DataladHCP1200
The Datalad version of the datagrabber with the first subject
The Datalad version of the DataGrabber with the first subject
already cloned.
"""
configure_logging(level="DEBUG")
dg = HCP1200(
@ -364,10 +367,10 @@ def test_hcp1200_datagrabber_elements(
("MOTOR", True),
],
)
def test_hcp1200_datagrabber_incorrect_access_icafix(
def test_HCP1200_incorrect_access_icafix(
tasks: Optional[str], ica_fix: bool
) -> None:
"""Test HCP1200 datagrabber incorrect access for icafix.
"""Test HCP1200 DataGrabber incorrect access for icafix.
Parameters
----------

View file

@ -1,4 +1,4 @@
"""Provide abstract base class for multiple source datagrabber."""
"""Provide concrete implementation for multi sourced DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de>
@ -7,21 +7,23 @@
from typing import Dict, List, Tuple, Union
from ..utils import raise_error
from .base import BaseDataGrabber
class MultipleDataGrabber(BaseDataGrabber):
"""Data Grabber class for data fetching from multiple sources.
"""Concrete implementation for multi sourced data fetching.
Defines a Data Grabber which can be used to fetch data from multiple
datagrabbers.
Implements a DataGrabber which can be used to fetch data from multiple
DataGrabbers.
Parameters
----------
datagrabbers : list of datagrabber-like objects
The datagrabbers to use to fetch data using.
datagrabbers : list of DataGrabber-like objects
The DataGrabbers to use for fetching data.
**kwargs
Keyword arguments passed to superclass.
"""
def __init__(self, datagrabbers: List[BaseDataGrabber], **kwargs) -> None:
@ -30,11 +32,11 @@ class MultipleDataGrabber(BaseDataGrabber):
first_keys = datagrabbers[0].get_element_keys()
for dg in datagrabbers[1:]:
if dg.get_element_keys() != first_keys:
raise ValueError("Datagrabbers have different element keys.")
raise_error("DataGrabbers have different element keys.")
# 2) no overlapping types
types = [x for dg in datagrabbers for x in dg.get_types()]
if len(types) != len(set(types)):
raise ValueError("Datagrabbers have overlapping types.")
raise_error("DataGrabbers have overlapping types.")
self._datagrabbers = datagrabbers
def __getitem__(self, element: Union[str, Tuple]) -> Dict:
@ -73,8 +75,64 @@ class MultipleDataGrabber(BaseDataGrabber):
out[kind]["meta"]["datagrabber"]["datagrabbers"] = metas
return out
def get_item(self, **element: Dict) -> Dict[str, Dict]:
"""Get item.
def __enter__(self) -> "MultipleDataGrabber":
"""Implement context entry."""
for dg in self._datagrabbers:
dg.__enter__()
return self
def __exit__(self, exc_type, exc_value, exc_traceback) -> None:
"""Implement context exit."""
for dg in self._datagrabbers:
dg.__exit__(exc_type, exc_value, exc_traceback)
# TODO: return type should be List[List[str]], but base type is List[str]
def get_types(self) -> List[str]:
"""Get types.
Returns
-------
list of list of str
The types of data to be grabbed.
"""
types = [x for dg in self._datagrabbers for x in dg.get_types()]
return types
def get_element_keys(self) -> List[str]:
"""Get element keys.
For each item in the ``element`` tuple passed to ``__getitem__()``,
this method returns the corresponding key(s).
Returns
-------
list of str
The element keys.
"""
return self._datagrabbers[0].get_element_keys()
def get_elements(self) -> List:
"""Get elements.
Returns
-------
list
List of elements that can be grabbed. The elements can be strings,
tuples or any object that will be then used as a key to index the
the DataGrabber. The element should be present in all of the
related DataGrabbers.
"""
all_elements = [dg.get_elements() for dg in self._datagrabbers]
elements = set(all_elements[0])
for s in all_elements[1:]:
elements.intersection_update(s)
return list(elements)
def get_item(self, **_: Dict) -> Dict[str, Dict]:
"""Get the specified item from the dataset.
Parameters
----------
@ -90,59 +148,12 @@ class MultipleDataGrabber(BaseDataGrabber):
Notes
-----
This function is not implemented for this class as it is useless.
"""
raise NotImplementedError(
"get_item() is not useful for this class, hence not implemented."
raise_error(
msg=(
"get_item() is not useful for this class, hence not "
"implemented."
),
klass=NotImplementedError,
)
def __enter__(self) -> "BaseDataGrabber":
"""Implement context entry."""
for dg in self._datagrabbers:
dg.__enter__()
return self
def __exit__(self, exc_type, exc_value, exc_traceback) -> None:
"""Implement context exit."""
for dg in self._datagrabbers:
dg.__exit__(exc_type, exc_value, exc_traceback)
def get_elements(self) -> List:
"""Get elements.
Returns
-------
list
The list of elements that can be grabbed in the dataset. It
corresponds to the elements that are present in all the
related datagrabbers.
"""
all_elements = [dg.get_elements() for dg in self._datagrabbers]
elements = set(all_elements[0])
for s in all_elements[1:]:
elements.intersection_update(s)
return list(elements)
def get_element_keys(self) -> List[str]:
"""Get element keys.
For each item in the ``element`` tuple passed to ``__getitem__()``,
this method returns the corresponding key(s).
Returns
-------
list of str
The element keys.
"""
return self._datagrabbers[0].get_element_keys()
def get_types(self) -> List[str]:
"""Get types.
Returns
-------
list of list of str
The types of data to be grabbed.
"""
types = [x for dg in self._datagrabbers for x in dg.get_types()]
return types

View file

@ -1,4 +1,4 @@
"""Provide concrete implementation for pattern-based datagrabber."""
"""Provide concrete implementation for pattern-based DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de>
@ -23,9 +23,9 @@ _CONFOUNDS_FORMATS = ("fmriprep", "adhoc")
@register_datagrabber
class PatternDataGrabber(BaseDataGrabber):
"""Concrete implementation for data grabbing using patterns.
"""Concrete implementation for pattern-based data fetching.
Implements a Data Grabber that understands patterns to grab data.
Implements a DataGrabber that understands patterns to grab data.
Parameters
----------
@ -34,12 +34,12 @@ class PatternDataGrabber(BaseDataGrabber):
patterns : dict
Patterns for each type of data as a dictionary. The keys are the types
and the values are the patterns. Each occurrence of the string
`{subject}` in the pattern will be replaced by the indexed element.
replacements : list of str
``{subject}`` in the pattern will be replaced by the indexed element.
replacements : str or list of str
Replacements in the patterns for each item in the "element" tuple.
datadir : str or pathlib.Path
The directory where the data is / will be stored.
confounds_format : {"fmriprep", "adhoc"}, optional
confounds_format : {"fmriprep", "adhoc"} or None, optional
The format of the confounds for the dataset (default None).
"""
@ -83,7 +83,7 @@ class PatternDataGrabber(BaseDataGrabber):
def _replace_patterns_regex(
self, pattern: str
) -> Tuple[str, str, List[str]]:
"""Replace the patterns in `pattern` with the named groups.
"""Replace the patterns in ``pattern`` with the named groups.
It allows elements to be obtained from the filesystem.
@ -122,7 +122,7 @@ class PatternDataGrabber(BaseDataGrabber):
return re_pattern, glob_pattern, t_replacements
def _replace_patterns_glob(self, element: Dict, pattern: str) -> str:
"""Replace patterns with the element so it can be globbed.
"""Replace ``pattern`` with the ``element`` so it can be globbed.
Parameters
----------
@ -209,7 +209,7 @@ class PatternDataGrabber(BaseDataGrabber):
raise_error(
"`confounds_format` needs to be one of "
f"{_CONFOUNDS_FORMATS}, None provided. "
"As the datagrabber used specifies "
"As the DataGrabber used specifies "
"'BOLD_confounds', None is invalid."
)
# Set the format
@ -228,6 +228,7 @@ class PatternDataGrabber(BaseDataGrabber):
-------
list
The list of elements that can be grabbed in the dataset.
"""
elements = None

View file

@ -1,4 +1,4 @@
"""Provide base class for pattern-based datalad datagrabber."""
"""Provide concrete implementation for pattern + datalad based DataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de>
@ -15,29 +15,29 @@ from .utils import validate_patterns
@register_datagrabber
class PatternDataladDataGrabber(DataladDataGrabber, PatternDataGrabber):
"""Base class for pattern-based data fetching via Datalad.
"""Concrete implementation for pattern and datalad based data fetching.
Defines a DataGrabber that gets data from a datalad sibling,
Implements a DataGrabber that gets data from a datalad sibling,
interpreting patterns.
Parameters
----------
types : list of str
The types of data to be grabbed.
patterns : dict, optional
patterns : dict
Patterns for each type of data as a dictionary. The keys are the types
and the values are the patterns. Each occurrence of the string
`{subject}` in the pattern will be replaced by the indexed element
(default None).
``{subject}`` in the pattern will be replaced by the indexed element.
**kwargs
Keyword arguments passed to superclass.
See Also
--------
DataladDataGrabber:
Base class for data fetching via Datalad.
Abstract base class for datalad-based data fetching.
PatternDataGrabber:
Base class for pattern-based data fetching.
Concrete implementation for pattern-based data fetching.
"""
def __init__(

View file

@ -1,4 +1,4 @@
"""Provide tests for base."""
"""Provide tests for BaseDataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de>
@ -9,7 +9,7 @@ from pathlib import Path
import pytest
from junifer.datagrabber.base import BaseDataGrabber
from junifer.datagrabber import BaseDataGrabber
def test_BaseDataGrabber_abstractness() -> None:

View file

@ -1,4 +1,4 @@
"""Provide tests for datalad_base."""
"""Provide tests for DataladDataGrabber."""
# Authors: Synchon Mandal <s.mandal@fz-juelich.de>
# License: AGPL
@ -9,7 +9,7 @@ from typing import Type
import datalad.api as dl
import pytest
from junifer.datagrabber.datalad_base import DataladDataGrabber
from junifer.datagrabber import DataladDataGrabber
_testing_dataset = {
@ -26,20 +26,20 @@ _testing_dataset = {
}
def test_datalad_base_abstractness() -> None:
"""Test datalad base is abstract."""
def test_DataladDataGrabber_abstractness() -> None:
"""Test DataladDataGrabber is abstract base class."""
with pytest.raises(TypeError, match=r"abstract"):
DataladDataGrabber() # type: ignore
@pytest.fixture
def concrete_datagrabber() -> Type[DataladDataGrabber]:
"""Return a concrete datagrabber class.
"""Return a concrete datalad-based DataGrabber.
Returns
-------
DataladDataGrabber
A concrete datagrabber class.
A concrete datalad-based DataGrabber.
"""
@ -74,17 +74,18 @@ def concrete_datagrabber() -> Type[DataladDataGrabber]:
return MyDataGrabber
def test_datalad_install_errors(
def test_DataladDataGrabber_install_errors(
tmp_path: Path, concrete_datagrabber: Type
) -> None:
"""Test datalad base install errors / warnings.
"""Test DataladDataGrabber install errors / warnings.
Parameters
----------
tmp_path : pathlib.Path
The path to the test directory.
concrete_datagrabber : DataladDataGrabber
A concrete datagrabber class to use.
A concrete datalad-based DataGrabber class to use.
"""
# Dataset cloned outside of datagrabber
@ -112,17 +113,18 @@ def test_datalad_install_errors(
pass
def test_datalad_clone_cleanup(
def test_DataladDataGrabber_clone_cleanup(
tmp_path: Path, concrete_datagrabber: Type
) -> None:
"""Test datalad base clone and remove.
"""Test DataladDataGrabber clone and remove.
Parameters
----------
tmp_path : pathlib.Path
The path to the test directory.
concrete_datagrabber : DataladDataGrabber
A concrete datagrabber class to use.
A concrete datalad-based DataGrabber class to use.
"""
# Clone whole dataset
@ -160,13 +162,16 @@ def test_datalad_clone_cleanup(
assert len(list(datadir.glob("*"))) == 0
def test_datalad_clone_create_cleanup(concrete_datagrabber: Type) -> None:
"""Test datalad base tempdir clone and remove.
def test_DataladDataGrabber_clone_create_cleanup(
concrete_datagrabber: Type,
) -> None:
"""Test DataladDataGrabber tempdir clone and remove.
Parameters
----------
concrete_datagrabber : DataladDataGrabber
A concrete datagrabber class to use.
A concrete datalad-based DataGrabber class to use.
"""
# Clone whole dataset
@ -203,17 +208,18 @@ def test_datalad_clone_create_cleanup(concrete_datagrabber: Type) -> None:
assert len(list(datadir.glob("*"))) == 0
def test_datalad_previously_cloned(
def test_DataladDataGrabber_previously_cloned(
tmp_path: Path, concrete_datagrabber: Type
) -> None:
"""Test datalad base on cloned dataset.
"""Test DataladDataGrabber on cloned dataset.
Parameters
----------
tmp_path : pathlib.Path
The path to the test directory.
concrete_datagrabber : DataladDataGrabber
A concrete datagrabber class to use.
A concrete datalad-based DataGrabber class to use.
"""
# Dataset cloned outside of datagrabber
@ -271,17 +277,18 @@ def test_datalad_previously_cloned(
assert len(list(datadir.glob("*"))) > 0
def test_datalad_previously_cloned_and_get(
def test_DataladDataGrabber_previously_cloned_and_get(
tmp_path: Path, concrete_datagrabber: Type
) -> None:
"""Test datalad base on cloned dataset with files present.
"""Test DataladDataGrabber on cloned dataset with files present.
Parameters
----------
tmp_path : pathlib.Path
The path to the test directory.
concrete_datagrabber : DataladDataGrabber
A concrete datagrabber class to use.
A concrete datalad-based DataGrabber class to use.
"""
# Dataset cloned outside of datagrabber with some files present
@ -353,17 +360,18 @@ def test_datalad_previously_cloned_and_get(
assert elem1_t1w.is_file() is True
def test_datalad_previously_cloned_and_get_dirty(
def test_DataladDataGrabber_previously_cloned_and_get_dirty(
tmp_path: Path, concrete_datagrabber: Type
) -> None:
"""Test datalad base on a dirty cloned dataset.
"""Test DataladDataGrabber on a dirty cloned dataset.
Parameters
----------
tmp_path : pathlib.Path
The path to the test directory.
concrete_datagrabber : DataladDataGrabber
A concrete datagrabber class to use.
A concrete datalad-based DataGrabber class to use.
"""
# Dataset cloned outside of datagrabber with some files present and dirty

View file

@ -1,4 +1,4 @@
"""Provide tests for multiple."""
"""Provide tests for MultipleDataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# License: AGPL
@ -20,8 +20,8 @@ _testing_dataset = {
}
def test_multiple() -> None:
"""Test a multiple datagrabber."""
def test_MultipleDataGrabber() -> None:
"""Test MultipleDataGrabber."""
repo_uri = _testing_dataset["example_bids_ses"]["uri"]
rootdir = "example_bids_ses"
replacements = ["subject", "session"]
@ -77,8 +77,8 @@ def test_multiple() -> None:
assert meta["datagrabbers"][1]["class"] == "PatternDataladDataGrabber"
def test_multiple_no_intersection() -> None:
"""Test a multiple datagrabber without intersection (0 elements)."""
def test_MultipleDataGrabber_no_intersection() -> None:
"""Test MultipleDataGrabber without intersection (0 elements)."""
repo_uri1 = _testing_dataset["example_bids"]["uri"]
repo_uri2 = _testing_dataset["example_bids_ses"]["uri"]
rootdir = "example_bids_ses"
@ -113,8 +113,8 @@ def test_multiple_no_intersection() -> None:
assert set(subs) == set(expected_subs)
def test_multiple_get_item() -> None:
"""Test a multiple datagrabber get_item error."""
def test_MultipleDataGrabber_get_item() -> None:
"""Test MultipleDataGrabber get_item() error."""
repo_uri1 = _testing_dataset["example_bids"]["uri"]
rootdir = "example_bids_ses"
replacements = ["subject", "session"]
@ -134,8 +134,8 @@ def test_multiple_get_item() -> None:
dg.get_item(subject="sub-01") # type: ignore
def test_multiple_validation() -> None:
"""Test a multiple datagrabber init validation."""
def test_MultipleDataGrabber_validation() -> None:
"""Test MultipleDataGrabber init validation."""
repo_uri1 = _testing_dataset["example_bids"]["uri"]
repo_uri2 = _testing_dataset["example_bids_ses"]["uri"]
rootdir = "example_bids_ses"

View file

@ -1,4 +1,4 @@
"""Provide tests for pattern."""
"""Provide tests for PatternDataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de>
@ -9,7 +9,7 @@ from pathlib import Path
import pytest
from junifer.datagrabber.pattern import PatternDataGrabber
from junifer.datagrabber import PatternDataGrabber
def test_PatternDataGrabber_errors(tmp_path: Path) -> None:
@ -261,7 +261,7 @@ def test_PatternDataGrabber(tmp_path: Path) -> None:
assert out1["vbm"]["path"] != out2["vbm"]["path"]
def test_pattern_data_grabber_confounds_format_error_on_init() -> None:
def test_PatternDataGrabber_confounds_format_error_on_init() -> None:
"""Test PatterDataGrabber confounds format error on initialisation."""
with pytest.raises(
ValueError, match="Invalid value for `confounds_format`"
@ -275,7 +275,7 @@ def test_pattern_data_grabber_confounds_format_error_on_init() -> None:
)
def test_pattern_data_grabber_confounds_format_error_on_fetch(
def test_PatternDataGrabber_confounds_format_error_on_fetch(
tmp_path: Path,
) -> None:
"""Test PatterDataGrabber confounds format error on fetching.
@ -301,6 +301,6 @@ def test_pattern_data_grabber_confounds_format_error_on_fetch(
)
# Check error on fetch
with pytest.raises(
ValueError, match="As the datagrabber used specifies 'BOLD_confounds'"
ValueError, match="As the DataGrabber used specifies 'BOLD_confounds'"
):
datagrabber.get_item(subject="sub-001")

View file

@ -1,4 +1,4 @@
"""Provide tests for pattern_datalad."""
"""Provide tests for PatternDataladDataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Leonard Sasse <l.sasse@fz-juelich.de>
@ -9,7 +9,7 @@ from pathlib import Path
import pytest
from junifer.datagrabber.pattern_datalad import PatternDataladDataGrabber
from junifer.datagrabber import PatternDataladDataGrabber
_testing_dataset = {
@ -26,8 +26,8 @@ _testing_dataset = {
}
def test_bids_pattern_datalad_datagrabber_missing_uri() -> None:
"""Test check of missing URI in pattern datalad datagrabber."""
def test_bids_PatternDataladDataGrabber_missing_uri() -> None:
"""Test check of missing URI in PatternDataladDataGrabber."""
with pytest.raises(ValueError, match=r"`uri` must be provided"):
PatternDataladDataGrabber(
datadir=None,
@ -37,15 +37,8 @@ def test_bids_pattern_datalad_datagrabber_missing_uri() -> None:
)
def test_bids_PatternDataladDataGrabber(tmp_path: Path) -> None:
"""Test a subject-based BIDS datalad datagrabber.
Parameters
----------
tmp_path : pathlib.Path
The path to the test directory.
"""
def test_bids_PatternDataladDataGrabber() -> None:
"""Test subject-based BIDS PatternDataladDataGrabber."""
# Define types
types = ["T1w", "BOLD"]
# Define patterns
@ -97,15 +90,8 @@ def test_bids_PatternDataladDataGrabber(tmp_path: Path) -> None:
assert f.readlines()[0].startswith("placeholder")
def test_bids_PatternDataladDataGrabber_datadir(tmp_path: Path) -> None:
"""Test a datalad datagrabber with a datadir set to a relative path.
Parameters
----------
tmp_path : pathlib.Path
The path to the test directory.
"""
def test_bids_PatternDataladDataGrabber_datadir() -> None:
"""Test PatternDataladDataGrabber with a datadir set to a relative path."""
# Define types
types = ["T1w", "BOLD"]
# Define patterns
@ -144,7 +130,7 @@ def test_bids_PatternDataladDataGrabber_datadir(tmp_path: Path) -> None:
def test_bids_PatternDataladDataGrabber_session():
"""Test a subject and session-based BIDS datalad datagrabber."""
"""Test a subject and session-based BIDS PatternDataladDataGrabber."""
types = ["T1w", "BOLD"]
patterns = {
"T1w": "{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz",

View file

@ -104,8 +104,8 @@ class MarkerCollection:
Parameters
----------
datagrabber : datagrabber-like
The datagrabber to validate.
datagrabber : DataGrabber-like
The DataGrabber to validate.
"""
logger.info("Validating Marker Collection")

View file

@ -1,4 +1,4 @@
"""Provide testing datagrabbers."""
"""Provide testing DataGrabbers."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Synchon Mandal <s.mandal@fz-juelich.de>
@ -15,9 +15,9 @@ from ..datagrabber.base import BaseDataGrabber
class OasisVBMTestingDataGrabber(BaseDataGrabber):
"""Data Grabber for Oasis VBM testing data.
"""DataGrabber for Oasis VBM testing data.
Wrapper for :func:`nilearn.datasets.fetch_oasis_vbm`
Wrapper for :func:`nilearn.datasets.fetch_oasis_vbm`.
"""
@ -83,7 +83,7 @@ class OasisVBMTestingDataGrabber(BaseDataGrabber):
class SPMAuditoryTestingDataGrabber(BaseDataGrabber):
"""Data Grabber for SPM Auditory dataset.
"""DataGrabber for SPM Auditory dataset.
Wrapper for :func:`nilearn.datasets.fetch_spm_auditory`.
@ -147,9 +147,9 @@ class SPMAuditoryTestingDataGrabber(BaseDataGrabber):
class PartlyCloudyTestingDataGrabber(BaseDataGrabber):
"""Data Grabber for Partly Cloudy dataset.
"""DataGrabber for Partly Cloudy dataset.
Wrapper for :func:`nilearn.datasets.fetch_development_fmri`
Wrapper for :func:`nilearn.datasets.fetch_development_fmri`.
Parameters
----------

View file

@ -12,7 +12,7 @@ from .datagrabbers import (
)
# Register testing datagrabber
# Register testing DataGrabbers
register(
step="datagrabber",
name="OasisVBMTestingDataGrabber",

View file

@ -1,4 +1,4 @@
"""Provide tests for Oasis VBM Testing datagrabber."""
"""Provide tests for OasisVBMTestingDataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# License: AGPL
@ -7,7 +7,7 @@ from junifer.testing.datagrabbers import OasisVBMTestingDataGrabber
def test_OasisVBMTestingDataGrabber() -> None:
"""Test Oasis VBM Testing datagrabber."""
"""Test OasisVBMTestingDataGrabber."""
expected_elements = [
"sub-01",
"sub-02",

View file

@ -1,4 +1,4 @@
"""Provide tests for PartlyCloudy datagrabber."""
"""Provide tests for PartlyCloudyTestingDataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# License: AGPL
@ -7,7 +7,7 @@ from junifer.testing.datagrabbers import PartlyCloudyTestingDataGrabber
def test_PartlyCloudyTestingDataGrabber() -> None:
"""Test PartlyCloudy datagrabber."""
"""Test PartlyCloudyTestingDataGrabber."""
expected_elements = [
"sub-01",
"sub-02",

View file

@ -1,4 +1,4 @@
"""Provide tests for SPM Auditory datagrabber."""
"""Provide tests for SPMAuditoryTestingDataGrabber."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# License: AGPL
@ -7,7 +7,7 @@ from junifer.testing.datagrabbers import SPMAuditoryTestingDataGrabber
def test_SPMAuditoryTestingDataGrabber() -> None:
"""Test SPM Auditory datagrabber."""
"""Test SPMAuditoryTestingDataGrabber."""
expected_elements = [
"sub001",
"sub002",

View file

@ -1,4 +1,4 @@
"""Create a testing dataset for the DataladAOMICID1000 pattern datagrabber."""
"""Create a testing dataset for DataladAOMICID1000."""
# Authors: Federico Raimondo <f.raimondo@fz-juelich.de>
# Vera Komeyer <v.komeyer@fz-juelich.de>