[DOC] Section on extending junifer #104

Closed
fraimondo wants to merge 143 commits from docs_extending into main
13 changed files with 350 additions and 39 deletions

View file

@ -1,10 +1,17 @@
API Functions
=============
Main API functions
------------------
.. automodule:: junifer.api
:members:
:imported-members:
Decorators
----------
.. automodule:: junifer.api.decorators
:members:
:imported-members:

View file

@ -174,10 +174,10 @@ texts.
# The BIDS datagrabber requires three parameters: the types of data we want,
# the specific pattern that matches each type, and the variables that will be
# replaced int he patterns.
types = ["T1w", "bold"]
types = ["T1w", "BOLD"]
patterns = {
"T1w": "{subject}/anat/{subject}_T1w.nii.gz",
"bold": "{subject}/func/{subject}_task-rest_bold.nii.gz",
"BOLD": "{subject}/func/{subject}_task-rest_bold.nii.gz",
}
replacements = ["subject"]
###############################################################################
@ -203,7 +203,7 @@ texts.
###############################################################################
# Another feature of the datagrabber is the ability to get a specific
# element by its name. In this case, we index `sub-01` and we get the file
# paths for the two types of data we want (T1w and bold).
# paths for the two types of data we want (T1w and BOLD).
with PatternDataladDataGrabber(
rootdir=rootdir,
types=types,

View file

@ -0,0 +1,271 @@
.. include:: ../links.inc
.. _extending_datagrabbers:
Creating DataGrabbers
=====================
DataGrabbers are the first step of the pipeline. It's purpose is to interpret
the structure of a dataset and provide two specific functionalities:
1) Given an *element*, provide the path to each kind of data available for this
element (e.g. the path to the T1 image, the path to the T2 image, etc.)
2) Provide the list of *elements* available in the dataset.
In this section, we will see how to create a datagrabber for a dataset. Basic
aspects of datagrabbers are covedered in the
:ref:`Understanding DataGrabbers <datagrabber>` section.
.. _extending_datagrabbers_think:
Step 1: Think about the element
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
Like with any programming-related task, the first step is to think. When
creating a DataGrabber, we need to first define what an *element* is.
The *element* should be the smallest unit of data that can be processed. That
is, for each element, there should be a set of data that can be processed, but
only one of each *data type* (see :ref:`data_types`).
For example, if we have a dataset from an fMRI study in which:
a) both T1w and fMRI was acquired
b) 20 subjects went through a experiment twice
c) the experiment included resting-stage fMRI and a task named *stroop*
then the *element* should be composed of 3 items:
* ``subject``: The subject IDs, e.g. `sub001`, `sub002`, ... `sub020`
* ``session``: The sesion number, e.g. `ses1`, `ses2`
* ``task``: The task performed, e.g. `rest`, `stroop`
If any of this items were not part of the element, then we will have more than
one ``T1w`` and/or ``BOLD`` image for each subject, which is not allowed.
Importantly, nothing prevents that one image is part of two different elements.
For example, it is usually the case that the ``T1w`` image is not acquired for
each task, but once in the entire session. So in this case, the ``T1w`` image
for the element (``sub001``, ``ses1``, ``rest``) will be the same as the
``T1w`` image for the element (``sub001``, ``ses1``, ``stroop``).
We will now continue this section using as an example, a dataset in BIDS format
in which each 9 subjects (`sub-01` to `sub-09`) were scanned each in 3
sessions (`ses-01`, `ses-02`, `ses-03`) and each session included a `T1w` and
a `BOLD` image (resting-state), except for `ses-03` which was only anatomical.
Step 2: Think about the dataset's structure
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
Now that we have our element defined, we need to think about the structure of
the dataset. Mainly, because the structure of the dataset will determine how
the DataGrabber needs to be implemented.
Junifer provides an abstract class to deal with datasets that can be thought in
terms of *patterns*. A *pattern* is a string that contains placeholders that are
replaced by the actual values of the element. In our BIDS example, the path
to the T1w image of subject `sub-01` and session `ses-01`, relative to the
dataset location, is ``sub-01/ses-01/anat/sub-01_ses-01_T1w.nii.gz``. By
replacing ``sub-01`` with ``sub-02``, we can obtain the T1w image of the first
session of the second subject. Indeed, the path to the T1w images can be
expressed as a pattern:
``{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz``
where ``{subject}`` is the replacement for the subject id and ``{session}}``
is the replacement for the session id.
Since it is a BIDS dataaset, the same happens with the BOLD images. The path to
the BOLD images can be expressed as a pattern:
``{subject}/{session}/func/{subject}_{session}_task-rest_bold.nii.gz``
This will be the norm in most of the datasets. If your dataset can be expressed
in terms of patterns, then follow :ref:`extending_datagrabbers_pattern`.
Otherwise, we recommend that you take time to re-think about your dataset
structure and why it does not have clear *patterns*. Feel free to open a
discussion in the `junifer Discussions`_ page. Most probably we can help you
get your dataset in order.
If there is no other way, then you can follow :ref:`extending_datagrabbers_base`
to create a DataGrabber from scratch.
.. _extending_datagrabbers_pattern:
Option A: Extending from PatternDataGrabber
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
The :py:class:`~junifer.datagrabber.PatternDataGrabber` class is an
abstract class that has the functionality of understanding patterns embeded
in it.
Before creating the datagrabber, we need to define 3 variables:
* ``types``: A list with the available :ref:`data_types` in our dataset
* ``patterns``: A dictionary that specifies the pattern for each data type.
* ``replacements``: A list indicating which of the elements in the patterns
should be replaced by the values of the element.
For example, in our BIDS example, the variables will be:
.. code-block:: python
types = ["T1w", "BOLD"]
patterns = {
"T1w": "{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz",
"BOLD": "{subject}/{session}/func/{subject}_{session}_task-rest_bold.nii.gz",
}
replacements = ["subject", "session"]
An additional fourth variable is the ``datadir``, which should be the path to
where the dataset is located. For example, if the dataset is located in
``/data/project/test/data``, then ``datadir`` should be
``/data/project/test/data``. Or, if we want to allow the user to specify the
location of the dataset, we can expose the variable in the constructor, as in
this example
With this defined, we can now create our datagrabber, we will name it
``ExampleBIDSDataGrabber``:
.. code-block:: python
from junifer.datagrabber.pattern import PatternDataGrabber
class ExampleBIDSDataGrabber(PatternDataGrabber):
def __init__(self, datadir):
types = ["T1w", "BOLD"]
patterns = {
"T1w": "{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz",
"BOLD": "{subject}/{session}/func/{subject}_{session}_task-rest_bold.nii.gz",
}
replacements = ["subject", "session"]
super().__init__(
datadir=datadir,
types=types,
patterns=patterns,
replacements=replacements)
Our datagrabber is ready to be used by junifer. However, it is still unknown
to the library. We need to register it in the library. To do so, we need to
use the :py:func:`~junifer.api.decorators.register_datagrabber` decorator.
.. code-block:: python
from junifer.datagrabber.pattern import PatternDataGrabber
from junifer.api.decorators import register_datagrabber
@register_datagrabber
class ExampleBIDSDataGrabber(PatternDataGrabber):
def __init__(self, datadir):
types = ["T1w", "BOLD"]
patterns = {
"T1w": "{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz",
"BOLD": "{subject}/{session}/func/{subject}_{session}_task-rest_bold.nii.gz",
}
replacements = ["subject", "session"]
super().__init__(
datadir=datadir,
types=types,
patterns=patterns,
replacements=replacements)
Now, we can use our datagrabber in junifer, by setting the ``datagrabber`` kind
in the yaml file to ``ExampleBIDSDataGrabber``. Remember that we still need to
set the ``datadir``.
.. code-block:: yaml
datagrabber:
kind: ExampleBIDSDataGrabber
datadir: /data/project/test/data
Optional: Using datalad
-----------------------
If you are using datalad, you can use the
:py:class:`~junifer.datagrabber.PatternDataladDataGrabber` instead of the
:py:class:`~junifer.datagrabber.PatternDataGrabber`. This class will also
interpret patterns, but will also use datalad to `clone` and `get` the data.
The main difference between the two is that the ``datadir`` is not the actual
location of the dataset, but the location where the dataset will be cloned. It
can now be ``None``, which means that the data will be downloaded to a
temporary directory. To set the location of the dataset, you can use the
``uri`` argument in the constructor. Additionally, a ``rootdir`` argument can
be used to specify the path to the root directory of the dataset after doing
``datalad clone``
In the example, the dataset is hosted in gin
(``https://gin.g-node.org/juaml/datalad-example-bids``).
When we clone this dataset, we will see the following structure:
.. code-block::
.
└── example_bids_ses
├── sub-01
│ ├── ses-01
│ ├── ses-02
│ └── ses-03
├── sub-02
│ ├── ses-01
│ ├── ses-02
│ └── ses-03
├── sub-03
...
So the patterns will start after ``example_bids_ses``. This is our ``rootdir``.
Now we have our 2 additional variables:
.. code-block:: python
uri = "https://gin.g-node.org/juaml/datalad-example-bids"
rootdir = "example_bids_ses"
And we can create our datagrabber:
.. code-block:: python
from junifer.datagrabber.pattern import PatternDataladDataGrabber
from junifer.api.decorators import register_datagrabber
@register_datagrabber
class ExampleBIDSDataGrabber(PatternDataladDataGrabber):
def __init__(self):
types = ["T1w", "BOLD"]
patterns = {
"T1w": "{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz",
"BOLD": "{subject}/{session}/func/{subject}_{session}_task-rest_bold.nii.gz",
}
replacements = ["subject", "session"]
uri = "https://gin.g-node.org/juaml/datalad-example-bids"
rootdir = "example_bids_ses"
super().__init__(
datadir=None,
uri=uri,
rootdir=rootdir,
types=types,
patterns=patterns,
replacements=replacements)
.. _extending_datagrabbers_base:
Option B: Extending from BaseDataGrabber
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^

26
docs/extending/index.rst Normal file
View file

@ -0,0 +1,26 @@
.. include:: ../links.inc
.. _extending:
Extending junifer
=================
While we aim to provide as many datasets and markers as possible, we are also
interested in allowing users to extend the functionality with their own
datagrabbers, preprocessing, markers, etc.
This does not mean that the new functinality will have to be included in
junifer before the user can use them. Instead, the user can simply
create a new python file, code the desired functionality and use it with
junifer. This is the first step towards including the new functionality in
the junifer package.
In this section we will show how to extend junifer, by creating new
datagrabbers, preprocessing and markers, following the *junifer* way.
.. toctree::
:maxdepth: 1
:caption: Contents:
datagrabber

View file

@ -25,6 +25,7 @@ enabling others to extend it easily.
installation
understanding/index.rst
builtin
extending/index.rst
auto_examples/index.rst
api/index.rst
contribution

View file

@ -27,6 +27,7 @@
.. _`Github`: https://github.com/
.. _`junifer Github`: https://github.com/juaml/junifer
.. _`junifer Discussions`: https://github.com/juaml/junifer/discussions
.. _`Wikipedia`: https://wikipedia.org

View file

@ -42,6 +42,9 @@ Data types
* - ``BOLD``
- BOLD image (4D)
- Preprocessed/Denoised BOLD image (fmriprep output)
* - ``BOLD_confounds``
- BOLD image confounds (CSV/TSV file)
- Confounds that can be applied to the BOLD image.
* - ``VBM_GM``
- VBM Gray Matter segmentation (3D)
- CAT output (`m0wp1` images)

View file

@ -1,5 +1,7 @@
.. include:: ../links.inc
.. _understanding:
Understanding junifer
=====================

View file

@ -23,10 +23,10 @@ configure_logging(level="INFO")
# The BIDS datagrabber requires three parameters: the types of data we want,
# the specific pattern that matches each type, and the variables that will be
# replaced int he patterns.
types = ["T1w", "bold"]
types = ["T1w", "BOLD"]
patterns = {
"T1w": "{subject}/anat/{subject}_T1w.nii.gz",
"bold": "{subject}/func/{subject}_task-rest_bold.nii.gz",
"BOLD": "{subject}/func/{subject}_task-rest_bold.nii.gz",
}
replacements = ["subject"]
###############################################################################
@ -52,7 +52,7 @@ with PatternDataladDataGrabber(
###############################################################################
# Another feature of the datagrabber is the ability to get a specific
# element by its name. In this case, we index `sub-01` and we get the file
# paths for the two types of data we want (T1w and bold).
# paths for the two types of data we want (T1w and BOLD).
with PatternDataladDataGrabber(
rootdir=rootdir,
types=types,

View file

@ -19,7 +19,7 @@ def test_validate_types() -> None:
with pytest.raises(TypeError, match="must be a list of strings"):
validate_types([1]) # type: ignore
validate_types(["T1w", "bold"])
validate_types(["T1w", "BOLD"])
def test_validate_replacements() -> None:
@ -31,7 +31,7 @@ def test_validate_replacements() -> None:
patterns = {
"T1w": "{subject}/anat/{subject}_T1w.nii.gz",
"bold": "{subject}/func/{subject}_task-rest_bold.nii.gz",
"BOLD": "{subject}/func/{subject}_task-rest_bold.nii.gz",
}
with pytest.raises(TypeError, match="must be a list of strings"):
@ -42,7 +42,7 @@ def test_validate_replacements() -> None:
wrong_patterns = {
"T1w": "{subject}/anat/_T1w.nii.gz",
"bold": "{session}/func/_task-rest_bold.nii.gz",
"BOLD": "{session}/func/_task-rest_bold.nii.gz",
}
with pytest.raises(ValueError, match="At least one pattern"):
@ -53,7 +53,7 @@ def test_validate_replacements() -> None:
def test_validate_patterns() -> None:
"""Test validation of patterns."""
types = ["T1w", "bold"]
types = ["T1w", "BOLD"]
with pytest.raises(TypeError, match="must be a dict"):
validate_patterns(types, "wrong") # type: ignore
@ -74,12 +74,12 @@ def test_validate_patterns() -> None:
patterns = {
"T1w": "{subject}/anat/{subject}_T1w.nii.gz",
"bold": "{subject}/func/{subject}_task-rest_bold.nii.gz",
"BOLD": "{subject}/func/{subject}_task-rest_bold.nii.gz",
}
wrongpatterns = {
"T1w": "{subject}/anat/{subject}*.nii",
"bold": "{subject}/func/{subject}_task-rest_bold.nii.gz",
"BOLD": "{subject}/func/{subject}_task-rest_bold.nii.gz",
}
with pytest.raises(ValueError, match="following a replacement"):

View file

@ -27,7 +27,7 @@ def test_multiple() -> None:
"T1w": "{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz",
}
pattern2 = {
"bold": "{subject}/{session}/func/"
"BOLD": "{subject}/{session}/func/"
"{subject}_{session}_task-rest_bold.nii.gz",
}
dg1 = PatternDataladDataGrabber(
@ -41,7 +41,7 @@ def test_multiple() -> None:
dg2 = PatternDataladDataGrabber(
rootdir=rootdir,
uri=repo_uri,
types=["bold"],
types=["BOLD"],
patterns=pattern2,
replacements=replacements,
)
@ -50,7 +50,7 @@ def test_multiple() -> None:
types = dg.get_types()
assert "T1w" in types
assert "bold" in types
assert "BOLD" in types
expected_subs = [
(f"sub-{i:02d}", f"ses-{j:02d}")
@ -64,7 +64,7 @@ def test_multiple() -> None:
data = dg[("sub-01", "ses-01")]
assert "T1w" in data
assert "bold" in data
assert "BOLD" in data
meta = dg.get_meta()
assert "class" in meta
@ -85,7 +85,7 @@ def test_multiple_no_intersection() -> None:
"T1w": "{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz",
}
pattern2 = {
"bold": "{subject}/{session}/func/"
"BOLD": "{subject}/{session}/func/"
"{subject}_{session}_task-rest_bold.nii.gz",
}
dg1 = PatternDataladDataGrabber(
@ -99,7 +99,7 @@ def test_multiple_no_intersection() -> None:
dg2 = PatternDataladDataGrabber(
rootdir=rootdir,
uri=repo_uri2,
types=["bold"],
types=["BOLD"],
patterns=pattern2,
replacements=replacements,
)

View file

@ -45,11 +45,11 @@ def test_bids_PatternDataladDataGrabber(tmp_path: Path) -> None:
"""
# Define types
types = ["T1w", "bold"]
types = ["T1w", "BOLD"]
# Define patterns
patterns = {
"T1w": "{subject}/anat/{subject}_T1w.nii.gz",
"bold": "{subject}/func/{subject}_task-rest_bold.nii.gz",
"BOLD": "{subject}/func/{subject}_task-rest_bold.nii.gz",
}
# Define replacements
replacements = ["subject"]
@ -75,8 +75,8 @@ def test_bids_PatternDataladDataGrabber(tmp_path: Path) -> None:
assert t_sub["T1w"]["path"] == (
dg.datadir / f"{elem}/anat/{elem}_T1w.nii.gz"
)
assert "path" in t_sub["bold"]
assert t_sub["bold"]["path"] == (
assert "path" in t_sub["BOLD"]
assert t_sub["BOLD"]["path"] == (
dg.datadir / f"{elem}/func/{elem}_task-rest_bold.nii.gz"
)
@ -104,11 +104,11 @@ def test_bids_PatternDataladDataGrabber_datadir(tmp_path: Path) -> None:
"""
# Define types
types = ["T1w", "bold"]
types = ["T1w", "BOLD"]
# Define patterns
patterns = {
"T1w": "{subject}/anat/{subject}_T1w.nii.gz",
"bold": "{subject}/func/{subject}_task-rest_bold.nii.gz",
"BOLD": "{subject}/func/{subject}_task-rest_bold.nii.gz",
}
# Define replacements
replacements = ["subject"]
@ -118,7 +118,7 @@ def test_bids_PatternDataladDataGrabber_datadir(tmp_path: Path) -> None:
datadir = "dataset" # use string and not absolute path
patterns = {
"T1w": "example_bids/{subject}/anat/{subject}_T*w.nii.gz",
"bold": "example_bids/{subject}/func/{subject}_task-rest_*.nii.gz",
"BOLD": "example_bids/{subject}/func/{subject}_task-rest_*.nii.gz",
}
with PatternDataladDataGrabber(
uri=repo_uri,
@ -134,18 +134,18 @@ def test_bids_PatternDataladDataGrabber_datadir(tmp_path: Path) -> None:
assert t_sub["T1w"]["path"] == (
dg.datadir / f"{elem}/anat/{elem}_T1w.nii.gz"
)
assert "path" in t_sub["bold"]
assert t_sub["bold"]["path"] == (
assert "path" in t_sub["BOLD"]
assert t_sub["BOLD"]["path"] == (
dg.datadir / f"{elem}/func/{elem}_task-rest_bold.nii.gz"
)
def test_bids_PatternDataladDataGrabber_session():
"""Test a subject and session-based BIDS datalad datagrabber."""
types = ["T1w", "bold"]
types = ["T1w", "BOLD"]
patterns = {
"T1w": "{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz",
"bold": "{subject}/{session}/func/"
"BOLD": "{subject}/{session}/func/"
"{subject}_{session}_task-rest_bold.nii.gz",
}
replacements = ["subject", "session"]
@ -162,7 +162,7 @@ def test_bids_PatternDataladDataGrabber_session():
rootdir = "example_bids_ses"
# repo_commit = _testing_dataset['example_bids_ses']['id']
# With T1W and bold, only 2 sessions are available
# With T1W and BOLD, only 2 sessions are available
with PatternDataladDataGrabber(
rootdir=rootdir,
uri=repo_uri,

View file

@ -42,7 +42,7 @@ def test_meta() -> None:
nib_data_path = Path(nib_testing.data_path)
t_path = nib_data_path / "example4d.nii.gz"
input = {"bold": {"path": t_path}}
input = {"BOLD": {"path": t_path}}
output = reader.fit_transform(input)
assert "meta" in output
assert "datareader" in output["meta"]
@ -67,23 +67,23 @@ def test_read_nifti(fname: str) -> None:
t_path = nib_data_path / fname
input = {"bold": {"path": t_path}}
input = {"BOLD": {"path": t_path}}
output = reader.fit_transform(input)
assert isinstance(output, dict)
assert "bold" in output
assert isinstance(output["bold"], dict)
assert "path" in output["bold"]
assert "data" in output["bold"]
assert "BOLD" in output
assert isinstance(output["BOLD"], dict)
assert "path" in output["BOLD"]
assert "data" in output["BOLD"]
read_img = output["bold"]["data"]
read_img = output["BOLD"]["data"]
t_read_img = nib.load(t_path)
assert_array_equal(read_img.get_fdata(), t_read_img.get_fdata())
input = {"bold": {"path": t_path.as_posix()}}
input = {"BOLD": {"path": t_path.as_posix()}}
output2 = reader.fit_transform(input)
assert output["bold"]["path"] == output2["bold"]["path"]
assert output["BOLD"]["path"] == output2["BOLD"]["path"]
def test_read_unknown() -> None: