[DOC] Section on extending junifer (datagrabbers and markers) #124

Merged
fraimondo merged 22 commits from doc/extending into main 2022-11-23 16:52:13 +00:00
57 changed files with 1359 additions and 112 deletions

View file

@ -1,10 +1,17 @@
API Functions API Functions
============= =============
Main API functions
------------------
.. automodule:: junifer.api .. automodule:: junifer.api
:members: :members:
:imported-members: :imported-members:
Decorators
----------
.. automodule:: junifer.api.decorators .. automodule:: junifer.api.decorators
:members: :members:
:imported-members: :imported-members:

View file

@ -1,5 +1,5 @@
Data Grabbers Data Grabbers
============ =============
.. automodule:: junifer.datagrabber .. automodule:: junifer.datagrabber
:members: :members:

View file

@ -5,7 +5,7 @@ Built-in Pipeline steps and data
Data Grabbers Data Grabbers
------------ -------------
.. ..
Provide a list of the DataGrabbers that are implemented or planned. Provide a list of the DataGrabbers that are implemented or planned.

View file

@ -174,10 +174,10 @@ texts.
# The BIDS datagrabber requires three parameters: the types of data we want, # The BIDS datagrabber requires three parameters: the types of data we want,
synchon commented 2022-11-23 06:58:58 +00:00 (Migrated from github.com)

# replaced in the patterns.

`# replaced in the patterns.`
# the specific pattern that matches each type, and the variables that will be # the specific pattern that matches each type, and the variables that will be
# replaced in the patterns. # replaced in the patterns.
types = ["T1w", "bold"] types = ["T1w", "BOLD"]
patterns = { patterns = {
"T1w": "{subject}/anat/{subject}_T1w.nii.gz", "T1w": "{subject}/anat/{subject}_T1w.nii.gz",
"bold": "{subject}/func/{subject}_task-rest_bold.nii.gz", "BOLD": "{subject}/func/{subject}_task-rest_bold.nii.gz",
} }
replacements = ["subject"] replacements = ["subject"]
############################################################################### ###############################################################################

View file

@ -0,0 +1,424 @@
.. include:: ../links.inc
synchon commented 2022-11-23 07:09:32 +00:00 (Migrated from github.com)

Its

`Its`
synchon commented 2022-11-23 07:10:18 +00:00 (Migrated from github.com)

covered

`covered`
synchon commented 2022-11-23 07:11:20 +00:00 (Migrated from github.com)

an

`an`
synchon commented 2022-11-23 07:12:12 +00:00 (Migrated from github.com)

these

`these`
synchon commented 2022-11-23 07:17:01 +00:00 (Migrated from github.com)

... were scanned during 3 sessions ...

`... were scanned during 3 sessions ...`
synchon commented 2022-11-23 07:21:10 +00:00 (Migrated from github.com)

... ``{session}`` ...

```... ``{session}`` ...```
synchon commented 2022-11-23 07:25:01 +00:00 (Migrated from github.com)
    ...,
    replacements=replacements,
)
``` ..., replacements=replacements, ) ```
synchon commented 2022-11-23 07:26:24 +00:00 (Migrated from github.com)
    ...,
    replacements=replacements,
)
``` ..., replacements=replacements, ) ```
synchon commented 2022-11-23 07:29:47 +00:00 (Migrated from github.com)

Can we have a hyperlink reference for datalad?

Can we have a hyperlink reference for datalad?
synchon commented 2022-11-23 07:31:07 +00:00 (Migrated from github.com)

This class will not only interpret patterns but also use datalad to `clone` and `get` the data.

```This class will not only interpret patterns but also use datalad to `clone` and `get` the data.```
synchon commented 2022-11-23 07:31:39 +00:00 (Migrated from github.com)

... .

`... .`
synchon commented 2022-11-23 07:33:12 +00:00 (Migrated from github.com)
    ...,
    replacements=replacements,
)
``` ..., replacements=replacements, ) ```
synchon commented 2022-11-23 07:33:45 +00:00 (Migrated from github.com)

of

`of`
synchon commented 2022-11-23 07:35:05 +00:00 (Migrated from github.com)

... BOLD.

`... BOLD.`
synchon commented 2022-11-23 07:35:43 +00:00 (Migrated from github.com)

can

`can`
synchon commented 2022-11-23 07:36:40 +00:00 (Migrated from github.com)

represent each of the items ...

`represent each of the items ...`
synchon commented 2022-11-23 07:38:42 +00:00 (Migrated from github.com)

One the should be removed.

One `the` should be removed.
.. _extending_datagrabbers:
Creating Data Grabbers
======================
Data Grabbers are the first step of the pipeline. Its purpose is to interpret
the structure of a dataset and provide two specific functionalities:
1) Given an *element*, provide the path to each kind of data available for this
element (e.g. the path to the T1 image, the path to the T2 image, etc.)
2) Provide the list of *elements* available in the dataset.
In this section, we will see how to create a datagrabber for a dataset. Basic
aspects of datagrabbers are covered in the
:ref:`Understanding Data Grabbers <datagrabber>` section.
.. _extending_datagrabbers_think:
Step 1: Think about the element
-------------------------------
Like with any programming-related task, the first step is to think. When
creating a Data Grabber, we need to first define what an *element* is.
The *element* should be the smallest unit of data that can be processed. That
is, for each element, there should be a set of data that can be processed, but
only one of each *data type* (see :ref:`data_types`).
For example, if we have a dataset from an fMRI study in which:
a) both T1w and fMRI was acquired
b) 20 subjects went through an experiment twice
c) the experiment included resting-stage fMRI and a task named *stroop*
then the *element* should be composed of 3 items:
* ``subject``: The subject IDs, e.g. `sub001`, `sub002`, ... `sub020`
* ``session``: The sesion number, e.g. `ses1`, `ses2`
* ``task``: The task performed, e.g. `rest`, `stroop`
If any of these items were not part of the element, then we will have more than
one ``T1w`` and/or ``BOLD`` image for each subject, which is not allowed.
Importantly, nothing prevents that one image is part of two different elements.
For example, it is usually the case that the ``T1w`` image is not acquired for
each task, but once in the entire session. So in this case, the ``T1w`` image
for the element (``sub001``, ``ses1``, ``rest``) will be the same as the
synchon commented 2022-11-23 07:14:48 +00:00 (Migrated from github.com)

Maybe: ("sub001", "ses1", "rest")?

Maybe: ``("sub001", "ses1", "rest")``?
fraimondo commented 2022-11-23 14:18:21 +00:00 (Migrated from github.com)

not as strings. I like it like that.

not as strings. I like it like that.
synchon commented 2022-11-23 14:43:56 +00:00 (Migrated from github.com)

With strings, you directly link it to the code which IMO is simpler.

With strings, you directly link it to the code which IMO is simpler.
synchon commented 2022-11-23 14:45:56 +00:00 (Migrated from github.com)

Also, the rendering is not very pretty.

Also, the rendering is not very pretty.
synchon commented 2022-11-23 14:46:47 +00:00 (Migrated from github.com)

What I actually meant was putting the whole thing as monospace.

What I actually meant was putting the whole thing as monospace.
``T1w`` image for the element (``sub001``, ``ses1``, ``stroop``).
synchon commented 2022-11-23 07:15:15 +00:00 (Migrated from github.com)

And here maybe: ("sub001", "ses1", "stroop")?

And here maybe: ``("sub001", "ses1", "stroop")``?
We will now continue this section using as an example, a dataset in BIDS format
in which 9 subjects (`sub-01` to `sub-09`) were scanned each during 3
sessions (`ses-01`, `ses-02`, `ses-03`) and each session included a `T1w` and
a `BOLD` image (resting-state), except for `ses-03` which was only anatomical.
Step 2: Think about the dataset's structure
-------------------------------------------
Now that we have our element defined, we need to think about the structure of
the dataset. Mainly, because the structure of the dataset will determine how
the Data Grabber needs to be implemented.
Junifer provides an abstract class to deal with datasets that can be thought in
terms of *patterns*. A *pattern* is a string that contains placeholders that are
replaced by the actual values of the element. In our BIDS example, the path
to the T1w image of subject `sub-01` and session `ses-01`, relative to the
dataset location, is ``sub-01/ses-01/anat/sub-01_ses-01_T1w.nii.gz``. By
replacing ``sub-01`` with ``sub-02``, we can obtain the T1w image of the first
session of the second subject. Indeed, the path to the T1w images can be
expressed as a pattern:
``{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz``
where ``{subject}`` is the replacement for the subject id and ``{session}``
is the replacement for the session id.
Since it is a BIDS dataaset, the same happens with the BOLD images. The path to
the BOLD images can be expressed as a pattern:
``{subject}/{session}/func/{subject}_{session}_task-rest_bold.nii.gz``
This will be the norm in most of the datasets. If your dataset can be expressed
in terms of patterns, then follow :ref:`extending_datagrabbers_pattern`.
Otherwise, we recommend that you take time to re-think about your dataset
structure and why it does not have clear *patterns*. Feel free to open a
discussion in the `junifer Discussions`_ page. Most probably we can help you
get your dataset in order.
If there is no other way, then you can follow :ref:`extending_datagrabbers_base`
to create a Data Grabber from scratch.
.. _extending_datagrabbers_pattern:
Step 3: Create a Data Grabber
-----------------------------
Option A: Extending from PatternDataGrabber
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
The :py:class:`~junifer.datagrabber.PatternDataGrabber` class is an
abstract class that has the functionality of understanding patterns embeded
in it.
Before creating the datagrabber, we need to define 3 variables:
* ``types``: A list with the available :ref:`data_types` in our dataset
* ``patterns``: A dictionary that specifies the pattern for each data type.
* ``replacements``: A list indicating which of the elements in the patterns
should be replaced by the values of the element.
For example, in our BIDS example, the variables will be:
.. code-block:: python
types = ["T1w", "BOLD"]
patterns = {
"T1w": "{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz",
"BOLD": "{subject}/{session}/func/{subject}_{session}_task-rest_bold.nii.gz",
}
replacements = ["subject", "session"]
An additional fourth variable is the ``datadir``, which should be the path to
where the dataset is located. For example, if the dataset is located in
``/data/project/test/data``, then ``datadir`` should be
``/data/project/test/data``. Or, if we want to allow the user to specify the
location of the dataset, we can expose the variable in the constructor, as in
this example
With this defined, we can now create our datagrabber, we will name it
``ExampleBIDSDataGrabber``:
.. code-block:: python
from junifer.datagrabber.pattern import PatternDataGrabber
class ExampleBIDSDataGrabber(PatternDataGrabber):
def __init__(self, datadir):
types = ["T1w", "BOLD"]
patterns = {
"T1w": "{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz",
"BOLD": "{subject}/{session}/func/{subject}_{session}_task-rest_bold.nii.gz",
}
replacements = ["subject", "session"]
super().__init__(
datadir=datadir,
types=types,
patterns=patterns,
replacements=replacements,
)
Our datagrabber is ready to be used by junifer. However, it is still unknown
to the library. We need to register it in the library. To do so, we need to
use the :py:func:`~junifer.api.decorators.register_datagrabber` decorator.
.. code-block:: python
from junifer.datagrabber.pattern import PatternDataGrabber
from junifer.api.decorators import register_datagrabber
@register_datagrabber
class ExampleBIDSDataGrabber(PatternDataGrabber):
def __init__(self, datadir):
types = ["T1w", "BOLD"]
patterns = {
"T1w": "{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz",
"BOLD": "{subject}/{session}/func/{subject}_{session}_task-rest_bold.nii.gz",
}
replacements = ["subject", "session"]
super().__init__(
datadir=datadir,
types=types,
patterns=patterns,
replacements=replacements,
)
Now, we can use our datagrabber in junifer, by setting the ``datagrabber`` kind
in the yaml file to ``ExampleBIDSDataGrabber``. Remember that we still need to
set the ``datadir``.
.. code-block:: yaml
datagrabber:
kind: ExampleBIDSDataGrabber
datadir: /data/project/test/data
Optional: Using datalad
synchon commented 2022-11-23 15:39:05 +00:00 (Migrated from github.com)

..., but also ...

`..., but also ...`
"""""""""""""""""""""""
If you are using `datalad`_, you can use the
:py:class:`~junifer.datagrabber.PatternDataladDataGrabber` instead of the
:py:class:`~junifer.datagrabber.PatternDataGrabber`. This class will not only
interpret patterns, but also use `datalad`_ to `clone` and `get` the data.
The main difference between the two is that the ``datadir`` is not the actual
location of the dataset, but the location where the dataset will be cloned. It
can now be ``None``, which means that the data will be downloaded to a
temporary directory. To set the location of the dataset, you can use the
``uri`` argument in the constructor. Additionally, a ``rootdir`` argument can
be used to specify the path to the root directory of the dataset after doing
``datalad clone``.
In the example, the dataset is hosted in gin
(``https://gin.g-node.org/juaml/datalad-example-bids``).
When we clone this dataset, we will see the following structure:
.. code-block::
.
└── example_bids_ses
├── sub-01
│ ├── ses-01
│ ├── ses-02
│ └── ses-03
├── sub-02
│ ├── ses-01
│ ├── ses-02
│ └── ses-03
├── sub-03
...
So the patterns will start after ``example_bids_ses``. This is our ``rootdir``.
Now we have our 2 additional variables:
.. code-block:: python
uri = "https://gin.g-node.org/juaml/datalad-example-bids"
rootdir = "example_bids_ses"
And we can create our datagrabber:
.. code-block:: python
from junifer.datagrabber.pattern import PatternDataladDataGrabber
from junifer.api.decorators import register_datagrabber
@register_datagrabber
class ExampleBIDSDataGrabber(PatternDataladDataGrabber):
def __init__(self):
types = ["T1w", "BOLD"]
patterns = {
"T1w": "{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz",
"BOLD": "{subject}/{session}/func/{subject}_{session}_task-rest_bold.nii.gz",
}
replacements = ["subject", "session"]
uri = "https://gin.g-node.org/juaml/datalad-example-bids"
rootdir = "example_bids_ses"
super().__init__(
datadir=None,
uri=uri,
rootdir=rootdir,
types=types,
patterns=patterns,
replacements=replacements,
)
.. _extending_datagrabbers_base:
Option B: Extending from BaseDataGrabber
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
While we could not think of a use case in which the pattern-based data grabber would not be suitable, it is still
possible to create a datagrabber extending from the :py:class:`~junifer.datagrabber.base.BaseDataGrabber` class.
In order to create a datagrabber extending from :py:class:`~junifer.datagrabber.base.BaseDataGrabber`, we need to
implement the following methods:
- ``get_item``: to get a single item from the dataset.
- ``get_elements``: to get the list of all elements present in the dataset
- ``get_element_keys``: to get the keys of the elements in the dataset.
.. note::
The ``__init__`` method could also be implemented, but it is not mandatory. This is required if the datagrabber
requires any parameter.
We will now implement our BIDS example with this method.
The first method, ``get_item``, needs to obtain a single
item from the dataset. Since this dataset requires two variables, ``subject`` and ``session``, we will use them
as parameters of ``get_item``:
.. code-block:: python
def get_item(self, subject, session):
out = {
"T1w": f"{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz",
"BOLD": f"{subject}/{session}/func/{subject}_{session}_task-rest_bold.nii.gz",
}
return out
The second method, ``get_elements``, needs to return a list of all the elements in the dataset. In this case, we
know that the dataset contains 3 subjects and 3 sessions, so we can create a list of all the possible combinations.
However, we need to remember that for session *ses-03* there is no BOLD data.
.. code-block:: python
def get_elements(self):
subjects = ["sub-01", "sub-02", "sub-03"]
sessions = ["ses-01", "ses-02"]
# If we are not working on BOLD data, we can add "ses-03"
if "BOLD" not in self.types:
sessions.append("ses-03")
elements = []
for subject in subjects:
for session in sessions:
elements.append({"subject": subject, "session": session})
return elements
And finally, we can implement the ``get_element_keys`` method. This method needs to return a list of the keys that
represent each of the items in the element tuple. As a rule of thumb, they should be the parameters of the
``get_item`` method, in the same order.
.. code-block:: python
def get_element_keys(self):
return ["subject", "session"]
So, to summarize, our datagrabber will look like this:
.. code-block:: python
from junifer.datagrabber.base import BaseDataGrabber
from junifer.api.decorators import register_datagrabber
@register_datagrabber
class ExampleBIDSDataGrabber(BaseDataGrabber):
def get_item(self, subject, session):
out = {
"T1w": f"{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz",
"BOLD": f"{subject}/{session}/func/{subject}_{session}_task-rest_bold.nii.gz",
}
return out
def get_elements(self):
subjects = ["sub-01", "sub-02", "sub-03"]
sessions = ["ses-01", "ses-02"]
# If we are not working on BOLD data, we can add "ses-03"
if "BOLD" not in self.types:
sessions.append("ses-03")
elements = []
for subject in subjects:
for session in sessions:
elements.append({"subject": subject, "session": session})
return elements
def get_element_keys(self):
return ["subject", "session"]
Optional: Using datalad
"""""""""""""""""""""""
If this dataset is in a datalad dataset, we can extend from :class:`junifer.datagrabber.DataladDataGrabber` instead of
:class:`junifer.datagrabber.BaseDataGrabber`. This will allow us to use the datalad API to obtain the data.
Step 4: Optional: Adding *BOLD confounds*
-----------------------------------------
For some analyses, it is useful to have the confounds associated with the BOLD data. This corresponds to the
``BOLD_confounds`` item in the :ref:`Data Object <data_object>` (see :ref:`data_types`). However, the ``BOLD_confounds``
element does not only consists of a ``path``, but it requries more information about the format of the confounds file.
Thus, the ``BOLD_confounds`` element is a dictionary with the following keys:
- ``path``: the path to the confounds file.
- ``format``: the format of the confounds file. Currently, this can be either ``fmriprep`` or ``adhoc``.
The ``fmriprep`` format corresponds to the format of the confounds files generated by `fMRIPrep`_. The
``adhoc`` format corresponds to a format that is not standardized.
.. note::
The ``mappings`` key is only required if the ``format`` is ``adhoc``. If the ``format`` is ``fmriprep``, the
``mappings`` key is not required.
Currently, Junifer provides only one confound remover step
(:class:`junifer.preprocess.fMRIPrepConfoundRemover`), which relies entirely on the ``fmriprep`` confound
variable names. Thus, if the confounds are not in ``fmriprep`` format, the user will need to provide the mappings
between the *ad-hoc* variable names and the ``fmriprep`` variable names.
This is done by specifying the ``adhoc`` format and providing the mappings as a dictionary in the ``mappings`` key.
In the following example, the confounds file has 3 variables that are not in the ``fmriprep`` format. Thus, we will
provide the mappings for these variables to the ``fmriprep`` format.
.. code-block:: python
out["BOLD_confounds"]: {
"path": f"{subject}/{session}/func/{subject}_{session}_confounds.tsv",
"format": "adhoc",
"mappings": {
"fmriprep": {
"variable1": "rot_x",
"variable2": "rot_z",
"variable3": "rot_y",
}
},
}
.. note::
Not all of the mappings need to be provided. For the moment, this is used only by the
:class:`junifer.preprocess.fMRIPrepConfoundRemover` step, which requires variables based on the
strategy selected. However, it is recommended to provide all the mappings, as this will allow the user to
choose different strategies with the same dataset.

View file

@ -0,0 +1,28 @@
.. include:: ../links.inc
synchon commented 2022-11-23 07:02:53 +00:00 (Migrated from github.com)

functionality to junifer at runtime.

`functionality to junifer at runtime.`
.. _extending_extension:
Creating a Junifer extension
============================
Junifer is designed to be easily extensible. Through the use of a registry and decorators, we can easily add new
functionality to junifer on runtime. This is done by creating a new python module and importing it before running
junifer.
A special consideration has to be made when using the :ref:`code-less configuration<codeless>`. In this case, the
``with`` statement can be used to import a module or run a python ``.py`` file.
In the following example, we instruct junifer to first import ``my_module`` and then run the ``my_file.py`` file.
.. code-block:: yaml
with:
- my_module
- my_file.py
Thus, the code from ``my_file.py`` will be executed before running junifer. This is the ideal place to create junifer
extensions.
.. important:: Some junifer commands will not consider files imported from files included in the ``with`` statement.
That is, if ``my_file.py`` imports ``my_other_file.py``, some of the junifer commands will not consider
``my_other_file.py``. Either place all the code in one file or add multiple files to the ``with`` statement.

28
docs/extending/index.rst Normal file
View file

@ -0,0 +1,28 @@
.. include:: ../links.inc
synchon commented 2022-11-23 07:01:35 +00:00 (Migrated from github.com)

functionality

`functionality`
.. _extending:
Extending junifer
=================
While we aim to provide as many datasets and markers as possible, we are also
interested in allowing users to extend the functionality with their own
datagrabbers, preprocessing, markers, etc.
This does not mean that the new functionality will have to be included in
junifer before the user can use them. Instead, the user can simply
create a new python file, code the desired functionality and use it with
junifer. This is the first step towards including the new functionality in
the junifer package.
In this section we will show how to extend junifer, by creating new
datagrabbers, preprocessing and markers, following the *junifer* way.
.. toctree::
:maxdepth: 2
:caption: Contents:
extension
datagrabber
marker

276
docs/extending/marker.rst Normal file
View file

@ -0,0 +1,276 @@
.. include:: ../links.inc
synchon commented 2022-11-23 08:45:18 +00:00 (Migrated from github.com)

... with the data types that the marker ...

`... with the data types that the marker ... `
synchon commented 2022-11-23 08:48:21 +00:00 (Migrated from github.com)

In this example, the only parameter required for computation is the name of the parcellation to use.

`In this example, the only parameter required for computation is the name of the parcellation to use.`
synchon commented 2022-11-23 09:52:08 +00:00 (Migrated from github.com)

Rendering for this as well is weird.

Rendering for this as well is weird.
synchon commented 2022-11-23 09:53:08 +00:00 (Migrated from github.com)

useful

`useful`
synchon commented 2022-11-23 10:00:56 +00:00 (Migrated from github.com)
:ref:`data type <data_types>`
``` :ref:`data type <data_types>` ```
synchon commented 2022-11-23 10:01:47 +00:00 (Migrated from github.com)

The method ``store`` ...

```The method ``store`` ...```
synchon commented 2022-11-23 10:03:06 +00:00 (Migrated from github.com)

... simply ...

`... simply ...`
synchon commented 2022-11-23 10:04:05 +00:00 (Migrated from github.com)

Once all of the above steps are done, we ...

`Once all of the above steps are done, we ...`
synchon commented 2022-11-23 10:06:44 +00:00 (Migrated from github.com)

parcellation

`parcellation`
synchon commented 2022-11-23 10:07:31 +00:00 (Migrated from github.com)

parcellation_name

`parcellation_name`
synchon commented 2022-11-23 10:09:05 +00:00 (Migrated from github.com)

self.parcellation_name = parcellation_name

`self.parcellation_name = parcellation_name`
synchon commented 2022-11-23 10:09:47 +00:00 (Migrated from github.com)

parcellation_name

`parcellation_name`
synchon commented 2022-11-23 10:09:58 +00:00 (Migrated from github.com)

self.parcellation_name = parcellation_name

`self.parcellation_name = parcellation_name`
.. _extending_markers:
Creating Markers
================
Computing a marker (a.k.a. *feature*) is the main goal of junifer. While we aim to provide as many markers as possible,
it might be the case that the marker you are looking for is not available. In this case, you can create your own marker
by following this tutorial.
Most of the functionality of a junifer marker has been taken care by the :class:`junifer.markers.BaseMarker` class.
Thus, only a few methods are required:
1. ``get_valid_inputs``: a method to obtain the list of valid inputs for the marker. This is used to check that the
inputs provided by the user are valid. This method should return a list of strings, representing
:ref:`data types <data_types>`
2. ``get_output_kind``: a method to obtain the kind of output of the marker. This is used to check that the output
of the marker is compatible with the storage. This method should return a string, representing
:ref:`storage types <storage_types>`
3. ``compute``: the method that given the data, computes the marker.
4. ``store``: the method that stores the computed marker.
5. ``__init__``: the initialization method, where the marker is configured.
As an example, we will develop a Parcel Mean marker, that is, a marker that first applies a parcellation and
then computes the mean of the data in each parcel. This is a very simple example, but it will show you how to create
a new marker.
.. _extending_markers_input_output:
Step 1: Configure input and output
----------------------------------
This step is quite simple: we need to define the input and output of the marker. Based on the current
:ref:`data types <data_types>`, we can define as valid inputs ``BOLD``, ``VBM_WM`` and ``VBM_GM``.
.. code-block:: python
def get_valid_inputs(self):
return ['BOLD', 'VBM_WM', 'VBM_GM']
The output of the marker depends on the input. For ``BOLD``, it will be ``timeseries``, while for the rest of the inputs,
it will be ``table``. Thus, we can define the output as:
.. code-block:: python
def get_output_kind(self, input_kind):
if input_kind == 'BOLD':
return 'timeseries'
else:
return 'table'
.. _extending_markers_init:
Step 2: Initialize the marker
-----------------------------
In this step we need to define the parameters of the marker. That is, all the parameters that the user can provide
to configure how the marker will behave.
The parameters of the marker are defined in the ``__init__`` method. The :class:`junifer.markers.BaseMarker` class
requires two optional parameters:
1. ``name``: the name of the marker. This is used to identify the marker in the configuration file.
2. ``on``: a list or string with the data types that the marker will be applied to.
.. attention:: Only basic types (*int*, *bool* and *str*) as well as Lists, Tuples and Dictionaries are allowed as
parameters. This is because the parameters are stored in a JSON file, and JSON only supports these types.
In this example, the is only paramater required for the computation is the name of the parcellation to use. Thus, we can
define the ``__init__`` method as follows:
.. code-block:: python
def __init__(self, parcellation_name, on=None, name=None):
self.parcellation_name = parcellation_name
super().__init__(on=on, name=name)
.. caution:: Parameters of the marker must be stored as object attributes without using ``_`` as prefix. This is
because any attribute that starts with ``_`` will not be considered as a parameter and not stored as
part of the metadata of the marker.
.. _extending_markers_compute:
Step 3: Compute the marker
--------------------------
In this step, we will define the method that computes the marker. This method will be called by junifer when needed,
using the data provided by the datagrabber, as configured by the user. The function ``compute`` has two arguments:
* ``input``: a dictionary with the data to be used to compute the marker. This will be the corresponding element in the
synchon commented 2022-11-23 09:50:07 +00:00 (Migrated from github.com)

The indentation for this is a bit weird when rendered.

The indentation for this is a bit weird when rendered.
:ref:`Data Object<data_object>` alredy indexing. Thus, the dictionary has at least two keys: ``data`` and ``path``.
The first one contains the data, while the second one contains the path to the data. The dictionary can also contain
other keys, depending on the data type.
* ``extra_input``: the rest of the :ref:`Data Object<data_object>`. This is useful if you want to use other data to
compute the marker (e.g.: ``BOLD_confounds`` can be used to de-confound the ``BOLD`` data).
Following the example, we will compute the mean of the data in each parcel using
:class:`nilearn.maskers.NiftiLabelsMasker`. Importantly, the output of the compute function must be a dictionary.
This dictionary will later be passed onto the ``store`` method.
.. hint:: To simplify the ``store`` method, define keys of the dictionary based on the corresponding store functions
in the :ref:`storage types <storage_types>`. For example, if the output is a ``table``, the keys of the
dictionary should be ``data`` and ``columns``.
.. code-block:: python
from nilearn.maskers import NiftiLabelsMasker
from junifer.data import load_parcellation
def compute(self, input, extra_input):
# Get the data
data = input["data"]
# Get the min of the voxels sizes and use it as the resolution
resolution = np.min(data.header.get_zooms()[:3])
# Load the parcellation
t_parcellation, t_labels, _ = load_parcellation(
name=self.parcellation_name,
resolution=resolution,
)
# Create a masker
masker = NiftiLabelsMasker(
synchon commented 2022-11-23 09:57:01 +00:00 (Migrated from github.com)
masker = NiftiLabelsMasker(
    labels_img=t_parcellation,
    standardize=True,
    memory="nilearn_cache",
    verbose=5,
)
``` masker = NiftiLabelsMasker( labels_img=t_parcellation, standardize=True, memory="nilearn_cache", verbose=5, ) ```
labels_img=t_parcellation,
standardize=True,
memory='nilearn_cache',
verbose=5,
)
# mask the data
out_values = masker.fit_transform([data])
# Create the output dictionary
out = {"data": out_values, "columns": t_labels}
# If its 3D (BOLD), name each row as "scan"
if out_values.shape[0] > 1:
out["row_names"] = "scan"
return out
.. _extending_markers_store:
Step 4: Store the marker
------------------------
In this step, we will define the method that stores the marker. This method will be called by junifer when needed,
using the data provided by the ``compute`` method. The method ``store`` has three arguments:
* ``kind``: A string indicating the :ref:`data type <data_types>` that was used to compute the marker.
* ``out``: The output of the ``compute`` method.
* ``storage``: The storage object, that will be used to store the marker.
.. code-block:: python
def store(self, kind, out, storage):
if kind in ["VBM_GM", "VBM_WM"]:
storage.store(kind="table", **out)
elif kind in ["BOLD"]:
storage.store(kind="timeseries", **out)
.. hint:: Check the hint on :ref:`extending_markers_compute`. If the output of the ``compute`` method is a dictionary
with keys based on the :ref:`storage types <storage_types>`, the ``store`` method can simply call the right
storage function, based on the ``kind`` parameter, with ``**out``.
.. _extending_markers_finalize:
Step 5: Finalize the marker
---------------------------
Once all of the above steps are done, we just need to give our marker a name an register it using the
``@register_marker`` decorator:
.. code-block:: python
from nilearn.maskers import NiftiLabelsMasker
from junifer.data import load_parcellation
from junifer.api.decorators import register_marker
from junifer.markers.base import BaseMarker
@register_marker
class ParcelMean(BaseMarker):
def __init__(self, parcellation_name, on=None, name=None):
self.parcellation_name = parcellation_name
super().__init__(on=on, name=name)
def get_valid_inputs(self):
return ['BOLD', 'VBM_WM', 'VBM_GM']
def get_output_kind(self, input_kind):
if input_kind == 'BOLD':
return 'timeseries'
else:
return 'table'
def compute(self, input, extra_input):
# Get the data
data = input["data"]
# Get the min of the voxels sizes and use it as the resolution
resolution = np.min(data.header.get_zooms()[:3])
# Load the parcellation
t_parcellation, t_labels, _ = load_parcellation(
name=self.parcellation_name,
resolution=resolution,
)
# Create a masker
masker = NiftiLabelsMasker(
synchon commented 2022-11-23 10:10:30 +00:00 (Migrated from github.com)
masker = NiftiLabelsMasker(
    labels_img=t_parcellation,
    standardize=True,
    memory="nilearn_cache",
    verbose=5,
)
``` masker = NiftiLabelsMasker( labels_img=t_parcellation, standardize=True, memory="nilearn_cache", verbose=5, ) ```
labels_img=t_parcellation,
standardize=True,
memory='nilearn_cache',
verbose=5,
)
# mask the data
out_values = masker.fit_transform([data])
# Create the output dictionary
out = {"data": out_values, "columns": t_labels}
# If its 3D (BOLD), name each row as "scan"
if out_values.shape[0] > 1:
out["row_names"] = "scan"
return out
def store(self, kind, out, storage):
if kind in ["VBM_GM", "VBM_WM"]:
storage.store(kind="table", **out)
elif kind in ["BOLD"]:
storage.store(kind="timeseries", **out)
.. _extending_markers_template:
Template for a custom Marker
----------------------------
.. code-block:: python
from junifer.api.decorators import register_marker
from junifer.markers.base import BaseMarker
@register_marker
class TemplateMarker(BaseMarker):
def __init__(self, on=None, name=None):
# TODO: add marker-specific parameters
super().__init__(on=on, name=name)
def get_valid_inputs(self):
# TODO: Complete with the valid inputs
valid = []
return valid
def get_output_kind(self, input_kind):
# TODO: Return the valid output kind for each input kind
pass
def compute(self, input, extra_input):
# TODO: compute the marker and create the output dictionary
# Create the output dictionary
out = {"data": None, "columns": None}
return out
def store(self, kind, out, storage):
# TODO: store out using the storage object, based on the kind of data
pass

Binary file not shown.

After

Width:  |  Height:  |  Size: 28 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 103 KiB

View file

@ -24,7 +24,9 @@ enabling others to extend it easily.
installation installation
understanding/index.rst understanding/index.rst
using/index.rst
builtin builtin
extending/index.rst
auto_examples/index.rst auto_examples/index.rst
api/index.rst api/index.rst
contribution contribution

View file

@ -21,16 +21,24 @@
.. _`matplotlib`: https://matplotlib.org .. _`matplotlib`: https://matplotlib.org
.. _`fMRIPrep`: https://fmriprep.org .. _`fMRIPrep`: https://fmriprep.org
.. _`nilearn`: https://nilearn.github.io .. _`nilearn`: https://nilearn.github.io
.. _`nipype`: https://nipype.readthedocs.io
.. _`datalad`: https://datalad.org
.. _`venv`: https://docs.python.org/3/tutorial/venv.html .. _`venv`: https://docs.python.org/3/tutorial/venv.html
.. _`conda env`: https://docs.conda.io/projects/conda/en/latest/user-guide/tasks/manage-environments.html .. _`conda env`: https://docs.conda.io/projects/conda/en/latest/user-guide/tasks/manage-environments.html
.. _`Github`: https://github.com/ .. _`Github`: https://github.com/
.. _`junifer Github`: https://github.com/juaml/junifer .. _`junifer Github`: https://github.com/juaml/junifer
.. _`junifer Discussions`: https://github.com/juaml/junifer/discussions
.. _`Wikipedia`: https://wikipedia.org .. _`Wikipedia`: https://wikipedia.org
.. _`YAML`: https://yaml.org
.. _`setuptools_scm`: https://github.com/pypa/setuptools_scm/ .. _`setuptools_scm`: https://github.com/pypa/setuptools_scm/
.. _`sphinx gallery`: https://sphinx-gallery.github.io/stable/index.html .. _`sphinx gallery`: https://sphinx-gallery.github.io/stable/index.html
.. _`sphinx reST reference`: https://www.sphinx-doc.org/en/master/usage/restructuredtext/basics.html#inline-markup .. _`sphinx reST reference`: https://www.sphinx-doc.org/en/master/usage/restructuredtext/basics.html#inline-markup
.. _`HTCondor`: https://research.cs.wisc.edu/htcondor/
.. _`SLURM`: https://slurm.schedmd.com
.. _`GNU Parallel`: https://www.gnu.org/software/parallel/

View file

@ -42,9 +42,21 @@ Data types
* - ``BOLD`` * - ``BOLD``
synchon commented 2022-11-23 10:12:39 +00:00 (Migrated from github.com)

CONN toolbox?

`CONN toolbox`?
synchon commented 2022-11-23 10:12:47 +00:00 (Migrated from github.com)

CONN toolbox?

`CONN toolbox`?
- BOLD image (4D) - BOLD image (4D)
- Preprocessed/Denoised BOLD image (fmriprep output) - Preprocessed/Denoised BOLD image (fmriprep output)
* - ``BOLD_confounds``
- BOLD image confounds (CSV/TSV file)
- Confounds that can be applied to the BOLD image.
* - ``VBM_GM`` * - ``VBM_GM``
- VBM Gray Matter segmentation (3D) - VBM Gray Matter segmentation (3D)
- CAT output (`m0wp1` images) - CAT output (`m0wp1` images)
* - ``VBM_WM`` * - ``VBM_WM``
- VBM White Matter segmentation (3D) - VBM White Matter segmentation (3D)
- CAT output (`m0wp2` images) - CAT output (`m0wp2` images)
* - ``fALFF``
- Voxel-wise fALFF image (3D)
- fALFF computed with CONN toolbox
* - ``GCOR``
- Global Correlation image (3D)
- GCOR computed with CONN toolbox
* - ``LCOR``
- Local Correlation image (3D)
- LCOR computed with CONN toolbox

View file

@ -3,12 +3,12 @@
.. _datagrabber: .. _datagrabber:
Data Grabber Data Grabber
=========== ============
Description Description
----------- -----------
The ``DataGrabber`` is an object that can provide an interface to datasets you want to work with in junifer. The *Data Grabber* is an object that can provide an interface to datasets you want to work with in junifer.
Every concrete implementation of a datagrabber is aware of a particular dataset's structure and thus allows Every concrete implementation of a datagrabber is aware of a particular dataset's structure and thus allows
you to fetch specific elements of interest from the dataset. It adds the ``path`` key to each :ref:`data type <data_types>` you to fetch specific elements of interest from the dataset. It adds the ``path`` key to each :ref:`data type <data_types>`
in the :ref:`Data object <data_object>`. in the :ref:`Data object <data_object>`.

View file

@ -3,12 +3,12 @@
.. _datareader: .. _datareader:
Data Reader Data Reader
========== ===========
Description Description
----------- -----------
The ``DataReader`` is an object that is responsible for actually reading data files in junifer. The *Data Reader* is an object that is responsible for actually reading data files in junifer.
It reads the value of the key ``path`` for each :ref:`data type <data_types>` in the :ref:`Data object <data_object>` It reads the value of the key ``path`` for each :ref:`data type <data_types>` in the :ref:`Data object <data_object>`
and loads them to memory. After reading the data into memory, it adds the key ``data`` to the same level as ``path`` and loads them to memory. After reading the data into memory, it adds the key ``data`` to the same level as ``path``
and the value is the actual data in the memory. and the value is the actual data in the memory.
@ -16,7 +16,7 @@ and the value is the actual data in the memory.
Datareaders are meant to be used inside the datagrabber context but you can operate on them outside the context as long Datareaders are meant to be used inside the datagrabber context but you can operate on them outside the context as long
as the actual data is in the memory and the Python runtime has not garbage-collected it. as the actual data is in the memory and the Python runtime has not garbage-collected it.
For data formats not supported by junifer yet, you can either make your own ``DataReader`` or open an issue on For data formats not supported by junifer yet, you can either make your own *Data Reader* or open an issue on
`junifer Github`_ and we can help you out. `junifer Github`_ and we can help you out.
Currently supported file-formats Currently supported file-formats

View file

@ -1,5 +1,7 @@
.. include:: ../links.inc .. include:: ../links.inc
.. _understanding:
Understanding junifer Understanding junifer
===================== =====================
@ -16,11 +18,15 @@ structural MRI, diffusion MRI, etc.) and you want to extract features to
later use in statistical analyses or machine learning (for example, using later use in statistical analyses or machine learning (for example, using
julearn_). julearn_).
.. important:: Junifer is not a toolbox to create pipelines, but a tool to configure the junifer pipeline, which is
intended to be fixed and not to be changed. If you want to create a pipeline, you should use
other tools like nipype_.
.. toctree:: .. toctree::
:maxdepth: 2 :maxdepth: 2
:caption: Contents: :caption: Contents:
pipeline
data data
datagrabber datagrabber
datareader datareader

View file

@ -9,7 +9,7 @@ Description
----------- -----------
The ``Marker`` is an object that is responsible for feature extraction. It primarily operates on data loaded The ``Marker`` is an object that is responsible for feature extraction. It primarily operates on data loaded
in memory by :ref:`datareader <DataReader>` and stored in the ``data`` key of each :ref:`data type <data_types>` in memory by :ref:`Data Reader <datareader>` and stored in the ``data`` key of each :ref:`data type <data_types>`
in the :ref:`Data object <data_object>`. In some cases, it can also operate on pre-processed data as obtained in the :ref:`Data object <data_object>`. In some cases, it can also operate on pre-processed data as obtained
from the :ref:`Preprocess <preprocess>` step of the pipeline. It is important to note that this pre-process is from the :ref:`Preprocess <preprocess>` step of the pipeline. It is important to note that this pre-process is
not similar to pre-processing done by tools like FSL, SPM, AFNI, etc. . For example, one can perform confound not similar to pre-processing done by tools like FSL, SPM, AFNI, etc. . For example, one can perform confound

View file

@ -0,0 +1,30 @@
.. include:: ../links.inc
synchon commented 2022-11-23 10:15:09 +00:00 (Migrated from github.com)

DataGrabber

`DataGrabber`
synchon commented 2022-11-23 10:15:33 +00:00 (Migrated from github.com)

dataset instead of database?

`dataset` instead of `database`?
fraimondo commented 2022-11-23 14:25:44 +00:00 (Migrated from github.com)

It can be either. Given that we are coupled with datalad, I would keep it as dataset.

It can be either. Given that we are coupled with datalad, I would keep it as dataset.
fraimondo commented 2022-11-23 14:26:06 +00:00 (Migrated from github.com)

I stil prefer to use the two words to describe the concept and not the Class name.

I stil prefer to use the two words to describe the concept and not the Class name.
synchon commented 2022-11-23 14:43:14 +00:00 (Migrated from github.com)

In the Understanding section, we have DataGrabber for the concept as well. Again my reasoning is that if we keep it as one name throughout, users are not confused. And also, one can mentally link better to the name of the step being the class category's name.

In the Understanding section, we have `DataGrabber` for the concept as well. Again my reasoning is that if we keep it as one name throughout, users are not confused. And also, one can mentally link better to the name of the step being the class category's name.
.. _pipeline:
The Junifer Pipeline
====================
The junifer pipeline is the main execution path of junifer. It consists of five steps:
1. :ref:`Data Grabber <datagrabber>`: Interpret the dataset and provide a list of files.
2. :ref:`Data Reader <datareader>`: Read the files.
synchon commented 2022-11-23 10:15:42 +00:00 (Migrated from github.com)

DataReader

`DataReader`
fraimondo commented 2022-11-23 14:26:23 +00:00 (Migrated from github.com)

same as before

same as before
3. :ref:`Pre-processing <preprocess>`: Prepare the images for marker computation.
synchon commented 2022-11-23 10:16:21 +00:00 (Migrated from github.com)

Preprocess

`Preprocess`
fraimondo commented 2022-11-23 14:26:31 +00:00 (Migrated from github.com)

same

same
4. :ref:`Marker Computation <marker>`: Compute the marker.
synchon commented 2022-11-23 10:16:01 +00:00 (Migrated from github.com)

Marker

`Marker`
fraimondo commented 2022-11-23 14:26:42 +00:00 (Migrated from github.com)

same

same
5. :ref:`Storage <storage>`: Store the marker values.
The element that is passed accross the pipeline is called the :ref:`Data Object<data_object>`.
The following is a graphical representation of the pipeline:
.. image:: ../images/pipeline/pipeline.001.png
However, it is usually the case that several markers are computed for the same data. Thus, the *markers* step
of the pipeline is defined as a list of markers. The following is a graphical representation of the pipeline execution
on multiple markers:
.. image:: ../images/pipeline/pipeline.002.png
.. note:: To avoid keeping in memory all of the computed marker, the storage step is called after each marker
computation, releasing the memory used to compute each marker.

View file

@ -24,6 +24,35 @@ methods respectively.
For storage interfaces not supported by junifer yet, you can either make your own ``Storage`` by providing a concrete For storage interfaces not supported by junifer yet, you can either make your own ``Storage`` by providing a concrete
implementation of :class:`junifer.storage.BaseFeatureStorage` or open an issue on `junifer Github`_ and we can help you out. implementation of :class:`junifer.storage.BaseFeatureStorage` or open an issue on `junifer Github`_ and we can help you out.
.. _storage_types:
Currently supported storage types
---------------------------------
.. list-table::
:widths: auto
:header-rows: 1
* - Storage Type
- Description
- Options
- Reference
* - ``matrix``
- A 2D matrix with row and column names
- ``col_names``, ``row_names``, ``matrix_kind``, ``diagonal``
- :meth:`junifer.storage.BaseFeatureStorage.store_matrix`
* - ``table``
- A vector of values with column names
- ``columns``, ``row_names``
- :meth:`junifer.storage.BaseFeatureStorage.store_table`
* - ``timeseries``
- A 2D matrix of values with column names
- ``columns``, ``row_names``
- :meth:`junifer.storage.BaseFeatureStorage.store_timeseries`
.. _storage_interfaces:
Currently supported storage interfaces Currently supported storage interfaces
-------------------------------------- --------------------------------------
@ -36,6 +65,6 @@ Currently supported storage interfaces
- File type - File type
- Storage kinds - Storage kinds
* - :class:`junifer.storage.SQLiteFeatureStorage` * - :class:`junifer.storage.SQLiteFeatureStorage`
- ``.db`` - ``.sqlite``
- SQLite - SQLite
- ``matrix``, ``table``, ``timeseries`` - ``matrix``, ``table``, ``timeseries``

204
docs/using/codeless.rst Normal file
View file

@ -0,0 +1,204 @@
.. include:: ../links.inc
synchon commented 2022-11-23 10:19:47 +00:00 (Migrated from github.com)
``preprocess``
``` ``preprocess`` ```
synchon commented 2022-11-23 10:21:49 +00:00 (Migrated from github.com)

... datareader ...

`... datareader ...`
synchon commented 2022-11-23 10:24:34 +00:00 (Migrated from github.com)

I believe the YAML syntax is false and when we load YAML, it should automatically convert it to False.

I believe the YAML syntax is `false` and when we load YAML, it should automatically convert it to `False`.
synchon commented 2022-11-23 10:24:42 +00:00 (Migrated from github.com)

Same as above but for true and True.

Same as above but for `true` and `True`.
synchon commented 2022-11-23 10:25:35 +00:00 (Migrated from github.com)

compute

`compute`
fraimondo commented 2022-11-23 14:27:09 +00:00 (Migrated from github.com)

was not sure about this.

was not sure about this.
synchon commented 2022-11-23 14:41:13 +00:00 (Migrated from github.com)

I think it does, can you please check it?

I think it does, can you please check it?
.. _codeless:
Code-less configuration
=======================
On of the most important features of junifer is its capacity to run without writing a single line of code. This is
achieved by using a configuration file that is written in YAML_. In this file, we configure the different steps of
:ref:`pipeline`.
As a reminder, this is how the pipeline looks like:
.. image:: ../images/pipeline/pipeline.001.png
Thus, the configuration file must configure each of the sections of the pipeline, as well as some general parameters.
As an example, we will generate the configuration file for a pipeline that will extract the mean ``VBM_GM`` values
using two different parcellations and one set of coordinates, from the *Oasis VBM Testing dataset* included in
junifer.
General Parameters
------------------
The general parameters are the ones that are not specific to any of the sections of the pipeline, but configure
junifer as a whole. These parameters are:
* ``with``: A section used to specify modules and junifer extensions to use.
* ``workdir``: The working directory where junifer will store temporary files.
Since the example uses a specific datagrabber for testing, we need to add ``junifer.testing.registry`` to the
``with`` section. This will allow junifer to find the datagrabber. We will set the ``workdir`` to ``/tmp``.
.. code-block:: yaml
with: junifer.testing.registry
workdir: /tmp
Step-by-step configuration
--------------------------
In order to configure the pipeline, we need to configure each step:
* ``datagrabber``
* ``datareader``
* ``preprocess``
* ``markers``
* ``storage``
.. important:: The datareader step configuration is optional, as junifer only provides one datareader. Nevertheless,
it is possible to extend junifer with custom datareaders, and thus, it is also possible to configure this step.
Data Grabber
synchon commented 2022-11-23 10:21:19 +00:00 (Migrated from github.com)

DataGrabber

`DataGrabber`
fraimondo commented 2022-11-23 14:27:52 +00:00 (Migrated from github.com)

will keep it as concepts and not class names

will keep it as concepts and not class names
^^^^^^^^^^^^
The ``datagrabber`` section must be configured using the ``kind`` key to specify the datagrabber to use. Additional
keys correspond to the parameters of the datagrabber.
For example, to use the :class:`junifer.datagrabber.DataladAOMICPIOP1` datagrabber, we just need to
specify its name as the ``kind`` key.
.. code-block:: yaml
datagrabber:
kind: DataladAOMICPIOP1
However, it is also possible to pass parameters to the datagrabber. In this case, we can restrict the datagrabber to
fetch only the ``restingstate`` task.
.. code-block:: yaml
datagrabber:
kind: DataladAOMICPIOP1
tasks: restingstate
In the *Oasis VBM Testing dataset* example, the section will look like this:
.. code-block:: yaml
datagrabber:
kind: OasisVBMTesting
Data Reader
synchon commented 2022-11-23 10:21:29 +00:00 (Migrated from github.com)

DataReader

`DataReader`
^^^^^^^^^^^
As mentioned before, this section is entirely optional, as junifer only provides one data reader
(:class:`junifer.datareader.DefaultDataReader`), which is the default in case the section is not specified.
In any case, the syntax of the section is the same as for the ``datagrabber`` section, using the ``kind`` key to
specify the data reader to use, and additional keys to pass parameters to the data reader:
.. code-block:: yaml
datareader:
kind: DefaultDataReader
For the *Oasis VBM Testing dataset* example, we will not specify a ``datareader`` step.
Preprocessing
^^^^^^^^^^^^^
Preprocessing is also an optional step, as it might be the case that no pre-processing is needed. In the case that
preprocessing is needed, the section must be configured using the ``kind`` key to specify the preprocessor to use,
and additional keys to pass parameters to the preprocessor.
For example, to use the :class:`junifer.preprocess.fMRIPrepConfoundRemover` preprocessor, we just need to specify its
name as the ``kind`` key, as well as its parameters.
.. code-block:: yaml
preproces:
kind: fMRIPrepConfoundRemover
strategy:
motion: full
wm_csf: full
global_signal: basic
spike: 0.2
detrend: false
standardize: true
For the *Oasis VBM Testing dataset* example, we will not specify a preprocessing step.
Markers
^^^^^^^
The ``markers`` section diverges from the previous ones, as we need to specify a list of markers. Each marker has a
name that we can use to refer to it later, and a set of parameters that will be passed to the marker.
For the *Oasis VBM Testing dataset* example, we want to compute the mean ``VBM_GM`` value for each parcel using the
Schaefer parcellation (100 parcels, 7 networks), Schaefer parcellation (200 parcels, 7 networks), and the *DMNBuckner*
network, using 5mm spheres. Thus, we will configure the ``markers`` section as follows:
.. code-block:: yaml
markers:
- name: Schaefer100x7_mean
kind: ParcelAggregation
parcellation: Schaefer100x7
method: mean
- name: Schaefer200x7_mean
kind: ParcelAggregation
parcellation: Schaefer200x7
method: mean
- name: DMNBuckner_5mm_mean
kind: SphereAggregation
coords: DMNBuckner
radius: 5
method: mean
Storage
^^^^^^^
Finally, we need to define how and where the results will be stored. This is done using the ``storage`` section,
which must be configured using the ``kind`` key to specify the storage to use, and additional keys to pass parameters.
For example, to use the :class:`junifer.storage.SQLiteFeatureStorage` storage, we just need to specify where we want
to store the results:
.. code-block:: yaml
storage:
kind: SQLiteFeatureStorage
uri: /data/junifer/example/oasis_vbm_testing.sqlite
synchon commented 2022-11-23 10:28:52 +00:00 (Migrated from github.com)

In the storage types, we have .db extension for SQLite. Just to be consistent, maybe we use it here?

In the storage types, we have `.db` extension for SQLite. Just to be consistent, maybe we use it here?
fraimondo commented 2022-11-23 14:28:59 +00:00 (Migrated from github.com)

it does not matter, the user sets the name and extension. It should be sqlite.

it does not matter, the user sets the name and extension. It should be sqlite.
synchon commented 2022-11-23 14:40:54 +00:00 (Migrated from github.com)

It indeed does not matter, my argument is just for the sake of consistency. I can imagine it be confusing for users who are not familiar with SQLite in that detail.

It indeed does not matter, my argument is just for the sake of consistency. I can imagine it be confusing for users who are not familiar with SQLite in that detail.
The full example
----------------
This is how the full *Oasis VBM Testing dataset* example configuration file looks like:
.. code-block:: yaml
with: junifer.testing.registry
workdir: /tmp
datagrabber:
kind: OasisVBMTesting
markers:
- name: Schaefer100x7_mean
kind: ParcelAggregation
parcellation: Schaefer100x7
method: mean
- name: Schaefer200x7_mean
kind: ParcelAggregation
parcellation: Schaefer200x7
method: mean
- name: DMNBuckner_5mm_mean
kind: SphereAggregation
coords: DMNBuckner
radius: 5
method: mean
storage:
kind: SQLiteFeatureStorage
uri: /data/junifer/example/oasis_vbm_testing.sqlite

18
docs/using/index.rst Normal file
View file

@ -0,0 +1,18 @@
.. include:: ../links.inc
.. _using:
Using junifer
=============
In this section, we will cover the main aspects behind using junifer. We will first explain the basics behind junifer's
code-less configuration. Then we will show how to use the command line interface to ``run`` junifer and ``collect``
the results. Finally, we will show how to use the ``queue`` command to interact with HPC and HTC systems.
.. toctree::
:maxdepth: 2
:caption: Contents:
codeless
running
queueing

89
docs/using/queueing.rst Normal file
View file

@ -0,0 +1,89 @@
.. include:: ../links.inc
synchon commented 2022-11-23 10:33:27 +00:00 (Migrated from github.com)

If you are in immediate need of any of these ...

`If you are in immediate need of any of these ...`
synchon commented 2022-11-23 10:33:48 +00:00 (Migrated from github.com)

scheduler

`scheduler`
synchon commented 2022-11-23 10:34:15 +00:00 (Migrated from github.com)
``HTCondor``
``` ``HTCondor`` ```
synchon commented 2022-11-23 10:34:49 +00:00 (Migrated from github.com)

Python

`Python`
synchon commented 2022-11-23 10:36:24 +00:00 (Migrated from github.com)

Please check my argument for having true.

Please check my argument for having ``true``.
.. _queueing:
Queueing jobs (HPC, HTC)
========================
Yet another interesting feature of junifer is the ability to queue jobs on computational clusters. This is done by
adding the ``queue`` section in the :ref:`codeless` file and executing the ``junifer queue`` command.
While junifer is meant to support `HTCondor`_, `SLURM`_ and local queueing using `GNU Parallel`_, only HTCondor is
currently supported. This will be implemented in future relases of junifer. If you are in immediate need of any of these
schedulers, please create an issue on the `junifer github`_ repository.
The ``queue`` section of the :ref:`codeless` must start by defining the following general parameters:
* ``jobname``: name of the job to be queued. This will be used to name the folder where the job files will be created,
as well as any relevant file. Depending on the scheduler, it will also be listed in the queueing system with this
name.
* ``kind``: The kind of scheduler to be used. Currently, only ``HTCondor`` is supported.
Example:
.. code-block:: yaml
queue:
jobname: TestHTCondorQueue
kind: HTCondor
The rest of the parameters depend on the scheduler you are using.
.. _queueing_condor:
HTCondor
--------
When using HTCondor, junifer will use a DAG to queue one job per element (``junifer run``). As an option, the DAG can
include a final job (``junifer collect``) to collect the results once all of the individual element jobs are finished.
The following parameters are avilable for HTCondor:
* ``env``: Definition of the Python enviroment. It must provide two variables: ``kind`` and ``name``. The ``kind``
corresponds to the kind of virtual environment to use: ``conda``, ``virtualenv`` (not yet supported) or
``local`` (no virtual enviroment). The ``name`` is the name of the enviroment to use in case a virtual environment
is used.
* ``mem``: Memory to be used by the job. It must be provided as a string with the units (e.g. ``2GB``).
* ``cpus``: Number of CPUs to be used by the job. It must be provided as an int.
* ``disk``: Disk space to be used by the job. It must be provided as a string with the units (e.g. ``2GB``). Keep in
mind that junifer uses a local working directory for each job, and datalad datasets might be cloned in this temporary
directory.
* ``extra_preamble``: Extra lines to be added to the HTCondor submit file. This can be used to add
extra parameters to the job, such as ``requirements``.
* ``collect``: If ``true``, a final job will be added to the DAG to collect the results once all of the individual
element jobs are finished. This is useful if you want to run a ``junifer collect`` job only once all of the
individual element jobs are finished. If not specified, it will default to ``true``.
Example:
.. code-block:: yaml
queue:
jobname: TestHTCondorQueue
kind: HTCondor
env:
kind: conda
name: junifer
mem: 8G
disk: 2GB
collect: true
Once the :ref:`codeless` file is ready, including the ``queue`` section, you can queue the jobs by executing
the ``junifer queue`` command.
The ``queue`` command will create a folder with the name of the job (``jobname``) under the ``junifer_jobs`` directory
in the current working directory.
The ``queue`` command accepts the following arguments:
* ``--help``: Show a help message.
* ``--verbose`` Set the verbosity level. Options are ``warning``, ``info``, ``debug``.
* ``--submit``: Submit the jobs to the queueing system. If not specified, the job submit files will be created but not
submitted.
* ``--overwrite``: Overwrite the job folder if it already exists. If not specified, the command will fail if the job
folder already exists.
* ``--element``: Queue only the specified element(s). If not specified, all elements will be queued.

60
docs/using/running.rst Normal file
View file

@ -0,0 +1,60 @@
.. include:: ../links.inc
synchon commented 2022-11-23 10:30:02 +00:00 (Migrated from github.com)

console instead of bash?

`console` instead of `bash`?
synchon commented 2022-11-23 10:31:41 +00:00 (Migrated from github.com)

console instead of bash?

`console` instead of `bash`?
synchon commented 2022-11-23 10:32:02 +00:00 (Migrated from github.com)

console instead of bash?

`console` instead of `bash`?
synchon commented 2022-11-23 10:32:16 +00:00 (Migrated from github.com)

console instead of bash?

`console` instead of `bash`?
.. _running:
Running jobs
============
Once we have the :ref:`code-less configuration file <codeless>`, we can use the command line interface to extract the
features. This is achieved in a two-step process: ``run`` and ``collect``.
The ``run`` command is used to extract the features from each element in the dataset. However, depending on the
storage interface, this may create one file per subject. The ``collect`` command is then used to collect all of the
individual results into a single file.
Assuming that we have a configuration file named ``config.yaml``, the following commands will extract the features:
.. code-block:: console
junifer run config.yaml
The ``run`` command accepts the following additional arguments:
* ``--help``: Show a help message.
* ``--verbose`` Set the verbosity level. Options are ``warning``, ``info``, ``debug``.
* ``--element``: The *element* to run. If not specified, all elements will be run. This parameter can be specified
synchon commented 2022-11-23 10:30:49 +00:00 (Migrated from github.com)

The rendering for this has some indentation issue.

The rendering for this has some indentation issue.
multiple times to run multiple elements. If the *element* requires several parameters, they can be specified
by separating them with ``,``.
Example on running two elements:
.. code-block:: console
junifer run config.yaml --element sub-01 --element sub-02
Example on elements with multiple parameters and verbose output:
.. code-block:: console
junifer run --verbose info config.yaml --element sub-01,ses-01
.. _collect:
Collecting results
==================
Once the ``run`` command has been executed, the results are stored in the output directory. However, depending on the
storage interface, this may create one file per subject. The ``collect`` command is then used to collect all of the
individual results into a single file.
Assuming that we have a configuration file named ``config.yaml``, the following commands will collect the results:
.. code-block:: console
junifer collect config.yaml
The ``collect`` command accepts the following additional arguments:
* ``--help``: Show a help message.
* ``--verbose`` Set the verbosity level. Options are ``warning``, ``info``, ``debug``.

View file

@ -23,10 +23,10 @@ configure_logging(level="INFO")
# The BIDS datagrabber requires three parameters: the types of data we want, # The BIDS datagrabber requires three parameters: the types of data we want,
# the specific pattern that matches each type, and the variables that will be # the specific pattern that matches each type, and the variables that will be
# replaced int he patterns. # replaced int he patterns.
types = ["T1w", "bold"] types = ["T1w", "BOLD"]
patterns = { patterns = {
"T1w": "{subject}/anat/{subject}_T1w.nii.gz", "T1w": "{subject}/anat/{subject}_T1w.nii.gz",
"bold": "{subject}/func/{subject}_task-rest_bold.nii.gz", "BOLD": "{subject}/func/{subject}_task-rest_bold.nii.gz",
} }
replacements = ["subject"] replacements = ["subject"]
############################################################################### ###############################################################################

View file

@ -52,7 +52,7 @@ with tempfile.TemporaryDirectory() as tmpdir:
# Define the storage interface # Define the storage interface
storage = { storage = {
"kind": "SQLiteFeatureStorage", "kind": "SQLiteFeatureStorage",
"uri": f"{tmpdir}/test.db", "uri": f"{tmpdir}/test.sqlite",
} }
# Run the defined junifer feature extraction pipeline # Run the defined junifer feature extraction pipeline
run( run(

View file

@ -71,7 +71,7 @@ sex = (
# Create a temporary directory for junifer feature extraction: # Create a temporary directory for junifer feature extraction:
with tempfile.TemporaryDirectory() as tmpdir: with tempfile.TemporaryDirectory() as tmpdir:
storage = {"kind": "SQLiteFeatureStorage", "uri": f"{tmpdir}/test.db"} storage = {"kind": "SQLiteFeatureStorage", "uri": f"{tmpdir}/test.sqlite"}
# run the defined junifer feature extraction pipeline # run the defined junifer feature extraction pipeline
run( run(
workdir="/tmp", workdir="/tmp",

View file

@ -43,7 +43,7 @@ storage = {
} }
with tempfile.TemporaryDirectory() as tmpdir: with tempfile.TemporaryDirectory() as tmpdir:
uri = f"{tmpdir}/test.db" uri = f"{tmpdir}/test.sqlite"
storage["uri"] = uri storage["uri"] = uri
run( run(
workdir="/tmp", workdir="/tmp",

View file

@ -21,5 +21,5 @@ markers:
method: std method: std
storage: storage:
kind: SQLiteFeatureStorage kind: SQLiteFeatureStorage
uri: /Users/fraimondo/dev/tbox/junifer/scratch/db/test.db uri: /Users/fraimondo/dev/tbox/junifer/scratch/db/test.sqlite

View file

@ -10,7 +10,7 @@ markers:
method: mean method: mean
storage: storage:
kind: SQLiteFeatureStorage kind: SQLiteFeatureStorage
uri: /data/group/appliedml/fraimondo/junifer_test/test.db uri: /data/group/appliedml/fraimondo/junifer_test/test.sqlite
queue: queue:
jobname: TestHTCondorQueue jobname: TestHTCondorQueue
kind: HTCondor kind: HTCondor

View file

@ -21,5 +21,5 @@ markers:
method: std method: std
storage: storage:
kind: SQLiteFeatureStorage kind: SQLiteFeatureStorage
uri: /data/project/ukb_motor/junifer_test/test.db uri: /data/project/ukb_motor/junifer_test/test.sqlite

View file

@ -33,6 +33,30 @@ def register_datagrabber(klass: Type) -> Type:
return klass return klass
synchon commented 2022-11-23 10:38:07 +00:00 (Migrated from github.com)

datareader

`datareader`
def register_datareader(klass: Type) -> Type:
"""Datareader registration decorator.
Registers the datareader so it can be used by name.
Parameters
----------
klass: class
The class of the datareader to register.
Returns
-------
klass: class
The unmodified input class.
"""
register(
step="datareader",
name=klass.__name__,
klass=klass,
)
return klass
def register_preprocessor(klass: Type) -> Type: def register_preprocessor(klass: Type) -> Type:
"""Preprocessor registration decorator. """Preprocessor registration decorator.

View file

@ -11,5 +11,5 @@ markers:
method: mean method: mean
storage: storage:
kind: SQLiteFeatureStorage kind: SQLiteFeatureStorage
uri: /Users/fraimondo/dev/tbox/junifer/scratch/db/test.db uri: /Users/fraimondo/dev/tbox/junifer/scratch/db/test.sqlite

View file

@ -10,7 +10,7 @@ markers:
method: mean method: mean
storage: storage:
kind: SQLiteFeatureStorage kind: SQLiteFeatureStorage
uri: /Users/fraimondo/dev/tbox/junifer/scratch/db/test.db uri: /Users/fraimondo/dev/tbox/junifer/scratch/db/test.sqlite
queue: queue:
jobname: TestHTCondorQueue jobname: TestHTCondorQueue
kind: HTCondor kind: HTCondor

View file

@ -61,7 +61,7 @@ def test_run_single_element(tmp_path: Path) -> None:
outdir = tmp_path / "out" outdir = tmp_path / "out"
outdir.mkdir() outdir.mkdir()
# Create storage # Create storage
uri = outdir / "test.db" uri = outdir / "test.sqlite"
storage["uri"] = uri # type: ignore storage["uri"] = uri # type: ignore
# Run operations # Run operations
run( run(
@ -72,7 +72,7 @@ def test_run_single_element(tmp_path: Path) -> None:
elements=["sub-01"], elements=["sub-01"],
) )
# Check files # Check files
files = list(outdir.glob("*.db")) files = list(outdir.glob("*.sqlite"))
assert len(files) == 1 assert len(files) == 1
@ -92,7 +92,7 @@ def test_run_multi_element(tmp_path: Path) -> None:
outdir = tmp_path / "out" outdir = tmp_path / "out"
outdir.mkdir() outdir.mkdir()
# Create storage # Create storage
uri = outdir / "test.db" uri = outdir / "test.sqlite"
storage["uri"] = uri # type: ignore storage["uri"] = uri # type: ignore
storage["single_output"] = False # type: ignore storage["single_output"] = False # type: ignore
# Run operations # Run operations
@ -104,7 +104,7 @@ def test_run_multi_element(tmp_path: Path) -> None:
elements=["sub-01", "sub-03"], elements=["sub-01", "sub-03"],
) )
# Check files # Check files
files = list(outdir.glob("*.db")) files = list(outdir.glob("*.sqlite"))
assert len(files) == 2 assert len(files) == 2
@ -124,7 +124,7 @@ def test_run_multi_element_single_output(tmp_path: Path) -> None:
outdir = tmp_path / "out" outdir = tmp_path / "out"
outdir.mkdir() outdir.mkdir()
# Create storage # Create storage
uri = outdir / "test.db" uri = outdir / "test.sqlite"
storage["uri"] = uri # type: ignore storage["uri"] = uri # type: ignore
storage["single_output"] = True # type: ignore storage["single_output"] = True # type: ignore
# Run operations # Run operations
@ -136,9 +136,9 @@ def test_run_multi_element_single_output(tmp_path: Path) -> None:
elements=["sub-01", "sub-03"], elements=["sub-01", "sub-03"],
) )
# Check files # Check files
files = list(outdir.glob("*.db")) files = list(outdir.glob("*.sqlite"))
assert len(files) == 1 assert len(files) == 1
assert files[0].name == "test.db" assert files[0].name == "test.sqlite"
def test_run_and_collect(tmp_path: Path) -> None: def test_run_and_collect(tmp_path: Path) -> None:
@ -157,7 +157,7 @@ def test_run_and_collect(tmp_path: Path) -> None:
outdir = tmp_path / "out" outdir = tmp_path / "out"
outdir.mkdir() outdir.mkdir()
# Create storage # Create storage
uri = outdir / "test.db" uri = outdir / "test.sqlite"
storage["uri"] = uri # type: ignore storage["uri"] = uri # type: ignore
storage["single_output"] = False # type: ignore storage["single_output"] = False # type: ignore
# Run operations # Run operations
@ -173,9 +173,9 @@ def test_run_and_collect(tmp_path: Path) -> None:
) )
elements = dg.get_elements() # type: ignore elements = dg.get_elements() # type: ignore
# This should create 10 files # This should create 10 files
files = list(outdir.glob("*.db")) files = list(outdir.glob("*.sqlite"))
assert len(files) == len(elements) assert len(files) == len(elements)
# But the test.db file should not exist # But the test.sqlite file should not exist
assert not uri.exists() assert not uri.exists()
# Collect in storage # Collect in storage
collect(storage) collect(storage)

View file

@ -12,7 +12,7 @@ from .base import BaseDataGrabber
class MultipleDataGrabber(BaseDataGrabber): class MultipleDataGrabber(BaseDataGrabber):
"""Datagrabber class for data fetching from multiple sources. """Data Grabber class for data fetching from multiple sources.
Defines a Data Grabber which can be used to fetch data from multiple Defines a Data Grabber which can be used to fetch data from multiple
datagrabbers. datagrabbers.

View file

@ -19,7 +19,7 @@ from .utils import validate_patterns, validate_replacements
@register_datagrabber @register_datagrabber
class PatternDataGrabber(BaseDataGrabber): class PatternDataGrabber(BaseDataGrabber):
"""Concrete implementation for data fetching using patterns. """Concrete implementation for data grabbing using patterns.
Implements a Data Grabber that understands patterns to grab data. Implements a Data Grabber that understands patterns to grab data.

View file

@ -19,7 +19,7 @@ def test_validate_types() -> None:
with pytest.raises(TypeError, match="must be a list of strings"): with pytest.raises(TypeError, match="must be a list of strings"):
validate_types([1]) # type: ignore validate_types([1]) # type: ignore
validate_types(["T1w", "bold"]) validate_types(["T1w", "BOLD"])
def test_validate_replacements() -> None: def test_validate_replacements() -> None:
@ -31,7 +31,7 @@ def test_validate_replacements() -> None:
patterns = { patterns = {
"T1w": "{subject}/anat/{subject}_T1w.nii.gz", "T1w": "{subject}/anat/{subject}_T1w.nii.gz",
"bold": "{subject}/func/{subject}_task-rest_bold.nii.gz", "BOLD": "{subject}/func/{subject}_task-rest_bold.nii.gz",
} }
with pytest.raises(TypeError, match="must be a list of strings"): with pytest.raises(TypeError, match="must be a list of strings"):
@ -42,7 +42,7 @@ def test_validate_replacements() -> None:
wrong_patterns = { wrong_patterns = {
"T1w": "{subject}/anat/_T1w.nii.gz", "T1w": "{subject}/anat/_T1w.nii.gz",
"bold": "{session}/func/_task-rest_bold.nii.gz", "BOLD": "{session}/func/_task-rest_bold.nii.gz",
} }
with pytest.raises(ValueError, match="At least one pattern"): with pytest.raises(ValueError, match="At least one pattern"):
@ -53,7 +53,7 @@ def test_validate_replacements() -> None:
def test_validate_patterns() -> None: def test_validate_patterns() -> None:
"""Test validation of patterns.""" """Test validation of patterns."""
types = ["T1w", "bold"] types = ["T1w", "BOLD"]
with pytest.raises(TypeError, match="must be a dict"): with pytest.raises(TypeError, match="must be a dict"):
validate_patterns(types, "wrong") # type: ignore validate_patterns(types, "wrong") # type: ignore
@ -74,12 +74,12 @@ def test_validate_patterns() -> None:
patterns = { patterns = {
"T1w": "{subject}/anat/{subject}_T1w.nii.gz", "T1w": "{subject}/anat/{subject}_T1w.nii.gz",
"bold": "{subject}/func/{subject}_task-rest_bold.nii.gz", "BOLD": "{subject}/func/{subject}_task-rest_bold.nii.gz",
} }
wrongpatterns = { wrongpatterns = {
"T1w": "{subject}/anat/{subject}*.nii", "T1w": "{subject}/anat/{subject}*.nii",
"bold": "{subject}/func/{subject}_task-rest_bold.nii.gz", "BOLD": "{subject}/func/{subject}_task-rest_bold.nii.gz",
} }
with pytest.raises(ValueError, match="following a replacement"): with pytest.raises(ValueError, match="following a replacement"):

View file

@ -28,7 +28,7 @@ def test_multiple() -> None:
"T1w": "{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz", "T1w": "{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz",
} }
pattern2 = { pattern2 = {
"bold": "{subject}/{session}/func/" "BOLD": "{subject}/{session}/func/"
"{subject}_{session}_task-rest_bold.nii.gz", "{subject}_{session}_task-rest_bold.nii.gz",
} }
dg1 = PatternDataladDataGrabber( dg1 = PatternDataladDataGrabber(
@ -42,7 +42,7 @@ def test_multiple() -> None:
dg2 = PatternDataladDataGrabber( dg2 = PatternDataladDataGrabber(
rootdir=rootdir, rootdir=rootdir,
uri=repo_uri, uri=repo_uri,
types=["bold"], types=["BOLD"],
patterns=pattern2, patterns=pattern2,
replacements=replacements, replacements=replacements,
) )
@ -51,7 +51,7 @@ def test_multiple() -> None:
types = dg.get_types() types = dg.get_types()
assert "T1w" in types assert "T1w" in types
assert "bold" in types assert "BOLD" in types
expected_subs = [ expected_subs = [
(f"sub-{i:02d}", f"ses-{j:02d}") (f"sub-{i:02d}", f"ses-{j:02d}")
@ -65,7 +65,7 @@ def test_multiple() -> None:
data = dg[("sub-01", "ses-01")] data = dg[("sub-01", "ses-01")]
assert "T1w" in data assert "T1w" in data
assert "bold" in data assert "BOLD" in data
meta = dg.get_meta() meta = dg.get_meta()
assert "class" in meta assert "class" in meta
@ -86,7 +86,7 @@ def test_multiple_no_intersection() -> None:
"T1w": "{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz", "T1w": "{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz",
} }
pattern2 = { pattern2 = {
"bold": "{subject}/{session}/func/" "BOLD": "{subject}/{session}/func/"
"{subject}_{session}_task-rest_bold.nii.gz", "{subject}_{session}_task-rest_bold.nii.gz",
} }
dg1 = PatternDataladDataGrabber( dg1 = PatternDataladDataGrabber(
@ -100,7 +100,7 @@ def test_multiple_no_intersection() -> None:
dg2 = PatternDataladDataGrabber( dg2 = PatternDataladDataGrabber(
rootdir=rootdir, rootdir=rootdir,
uri=repo_uri2, uri=repo_uri2,
types=["bold"], types=["BOLD"],
patterns=pattern2, patterns=pattern2,
replacements=replacements, replacements=replacements,
) )

View file

@ -46,11 +46,11 @@ def test_bids_PatternDataladDataGrabber(tmp_path: Path) -> None:
""" """
# Define types # Define types
types = ["T1w", "bold"] types = ["T1w", "BOLD"]
# Define patterns # Define patterns
patterns = { patterns = {
"T1w": "{subject}/anat/{subject}_T1w.nii.gz", "T1w": "{subject}/anat/{subject}_T1w.nii.gz",
"bold": "{subject}/func/{subject}_task-rest_bold.nii.gz", "BOLD": "{subject}/func/{subject}_task-rest_bold.nii.gz",
} }
# Define replacements # Define replacements
replacements = ["subject"] replacements = ["subject"]
@ -76,8 +76,8 @@ def test_bids_PatternDataladDataGrabber(tmp_path: Path) -> None:
assert t_sub["T1w"]["path"] == ( assert t_sub["T1w"]["path"] == (
dg.datadir / f"{elem}/anat/{elem}_T1w.nii.gz" dg.datadir / f"{elem}/anat/{elem}_T1w.nii.gz"
) )
assert "path" in t_sub["bold"] assert "path" in t_sub["BOLD"]
assert t_sub["bold"]["path"] == ( assert t_sub["BOLD"]["path"] == (
dg.datadir / f"{elem}/func/{elem}_task-rest_bold.nii.gz" dg.datadir / f"{elem}/func/{elem}_task-rest_bold.nii.gz"
) )
@ -105,11 +105,11 @@ def test_bids_PatternDataladDataGrabber_datadir(tmp_path: Path) -> None:
""" """
# Define types # Define types
types = ["T1w", "bold"] types = ["T1w", "BOLD"]
# Define patterns # Define patterns
patterns = { patterns = {
"T1w": "{subject}/anat/{subject}_T1w.nii.gz", "T1w": "{subject}/anat/{subject}_T1w.nii.gz",
"bold": "{subject}/func/{subject}_task-rest_bold.nii.gz", "BOLD": "{subject}/func/{subject}_task-rest_bold.nii.gz",
} }
# Define replacements # Define replacements
replacements = ["subject"] replacements = ["subject"]
@ -119,7 +119,7 @@ def test_bids_PatternDataladDataGrabber_datadir(tmp_path: Path) -> None:
datadir = "dataset" # use string and not absolute path datadir = "dataset" # use string and not absolute path
patterns = { patterns = {
"T1w": "example_bids/{subject}/anat/{subject}_T*w.nii.gz", "T1w": "example_bids/{subject}/anat/{subject}_T*w.nii.gz",
"bold": "example_bids/{subject}/func/{subject}_task-rest_*.nii.gz", "BOLD": "example_bids/{subject}/func/{subject}_task-rest_*.nii.gz",
} }
with PatternDataladDataGrabber( with PatternDataladDataGrabber(
uri=repo_uri, uri=repo_uri,
@ -135,18 +135,18 @@ def test_bids_PatternDataladDataGrabber_datadir(tmp_path: Path) -> None:
assert t_sub["T1w"]["path"] == ( assert t_sub["T1w"]["path"] == (
dg.datadir / f"{elem}/anat/{elem}_T1w.nii.gz" dg.datadir / f"{elem}/anat/{elem}_T1w.nii.gz"
) )
assert "path" in t_sub["bold"] assert "path" in t_sub["BOLD"]
assert t_sub["bold"]["path"] == ( assert t_sub["BOLD"]["path"] == (
dg.datadir / f"{elem}/func/{elem}_task-rest_bold.nii.gz" dg.datadir / f"{elem}/func/{elem}_task-rest_bold.nii.gz"
) )
def test_bids_PatternDataladDataGrabber_session(): def test_bids_PatternDataladDataGrabber_session():
"""Test a subject and session-based BIDS datalad datagrabber.""" """Test a subject and session-based BIDS datalad datagrabber."""
types = ["T1w", "bold"] types = ["T1w", "BOLD"]
patterns = { patterns = {
"T1w": "{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz", "T1w": "{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz",
"bold": "{subject}/{session}/func/" "BOLD": "{subject}/{session}/func/"
"{subject}_{session}_task-rest_bold.nii.gz", "{subject}_{session}_task-rest_bold.nii.gz",
} }
replacements = ["subject", "session"] replacements = ["subject", "session"]

View file

@ -10,6 +10,7 @@ from typing import Dict, List, Optional
import nibabel as nib import nibabel as nib
import pandas as pd import pandas as pd
from ..api.decorators import register_datareader
from ..pipeline import PipelineStepMixin from ..pipeline import PipelineStepMixin
from ..utils.logging import logger, warn_with_log from ..utils.logging import logger, warn_with_log
@ -28,6 +29,7 @@ _readers["CSV"] = {"func": pd.read_csv, "params": None}
_readers["TSV"] = {"func": pd.read_csv, "params": {"sep": "\t"}} _readers["TSV"] = {"func": pd.read_csv, "params": {"sep": "\t"}}
@register_datareader
class DefaultDataReader(PipelineStepMixin): class DefaultDataReader(PipelineStepMixin):
"""Mixin class for default data reader.""" """Mixin class for default data reader."""

View file

@ -42,7 +42,7 @@ def test_meta() -> None:
nib_data_path = Path(nib_testing.data_path) nib_data_path = Path(nib_testing.data_path)
t_path = nib_data_path / "example4d.nii.gz" t_path = nib_data_path / "example4d.nii.gz"
input = {"bold": {"path": t_path}} input = {"BOLD": {"path": t_path}}
output = reader.fit_transform(input) output = reader.fit_transform(input)
assert "meta" in output assert "meta" in output
assert "datareader" in output["meta"] assert "datareader" in output["meta"]
@ -67,23 +67,23 @@ def test_read_nifti(fname: str) -> None:
t_path = nib_data_path / fname t_path = nib_data_path / fname
input = {"bold": {"path": t_path}} input = {"BOLD": {"path": t_path}}
output = reader.fit_transform(input) output = reader.fit_transform(input)
assert isinstance(output, dict) assert isinstance(output, dict)
assert "bold" in output assert "BOLD" in output
assert isinstance(output["bold"], dict) assert isinstance(output["BOLD"], dict)
assert "path" in output["bold"] assert "path" in output["BOLD"]
assert "data" in output["bold"] assert "data" in output["BOLD"]
read_img = output["bold"]["data"] read_img = output["BOLD"]["data"]
t_read_img = nib.load(t_path) t_read_img = nib.load(t_path)
assert_array_equal(read_img.get_fdata(), t_read_img.get_fdata()) assert_array_equal(read_img.get_fdata(), t_read_img.get_fdata())
input = {"bold": {"path": t_path.as_posix()}} input = {"BOLD": {"path": t_path.as_posix()}}
output2 = reader.fit_transform(input) output2 = reader.fit_transform(input)
assert output["bold"]["path"] == output2["bold"]["path"] assert output["BOLD"]["path"] == output2["BOLD"]["path"]
def test_read_unknown() -> None: def test_read_unknown() -> None:

View file

@ -134,7 +134,7 @@ def test_marker_collection_storage(tmp_path: Path) -> None:
# Test storage # Test storage
dg = OasisVBMTestingDatagrabber() dg = OasisVBMTestingDatagrabber()
uri = tmp_path / "test_marker_collection_storage.db" uri = tmp_path / "test_marker_collection_storage.sqlite"
storage = SQLiteFeatureStorage(uri=uri) storage = SQLiteFeatureStorage(uri=uri)
mc = MarkerCollection( mc = MarkerCollection(
markers=markers, markers=markers,

View file

@ -78,7 +78,7 @@ def test_store(tmp_path: Path) -> None:
parcellation_two=parcellation_TWO, parcellation_two=parcellation_TWO,
correlation_method="spearman", correlation_method="spearman",
) )
uri = tmp_path / "test_crossparcellation.db" uri = tmp_path / "test_crossparcellation.sqlite"
storage = SQLiteFeatureStorage(uri=uri, upsert="ignore") storage = SQLiteFeatureStorage(uri=uri, upsert="ignore")
out = crossparcellation.fit_transform(input_dict, storage=storage) out = crossparcellation.fit_transform(input_dict, storage=storage)

View file

@ -77,6 +77,6 @@ def test_store(tmp_path: Path) -> None:
ets_rss_marker = RSSETSMarker(parcellation=PARCELLATION) ets_rss_marker = RSSETSMarker(parcellation=PARCELLATION)
# Create storage # Create storage
storage = SQLiteFeatureStorage( storage = SQLiteFeatureStorage(
uri=str((tmp_path / "test.db").absolute())) uri=str((tmp_path / "test.sqlite").absolute()))
# Store # Store
ets_rss_marker.fit_transform(input=input_dict, storage=storage) ets_rss_marker.fit_transform(input=input_dict, storage=storage)

View file

@ -78,7 +78,7 @@ def test_FunctionalConnectivityParcels(tmp_path: Path) -> None:
all_out = fc.fit_transform({"BOLD": {"data": fmri_img}}) all_out = fc.fit_transform({"BOLD": {"data": fmri_img}})
uri = tmp_path / "test_fc_parcellation.db" uri = tmp_path / "test_fc_parcellation.sqlite"
# Single storage, must be the uri # Single storage, must be the uri
storage = SQLiteFeatureStorage(uri=uri, upsert="ignore") storage = SQLiteFeatureStorage(uri=uri, upsert="ignore")
meta = { meta = {

View file

@ -62,7 +62,7 @@ def test_FunctionalConnectivitySpheres(tmp_path: Path) -> None:
# check correct output # check correct output
assert fc.get_output_kind(["BOLD"]) == ["matrix"] assert fc.get_output_kind(["BOLD"]) == ["matrix"]
uri = tmp_path / "test_fc_parcel.db" uri = tmp_path / "test_fc_parcel.sqlite"
# Single storage, must be the uri # Single storage, must be the uri
storage = SQLiteFeatureStorage(uri=uri, upsert="ignore") storage = SQLiteFeatureStorage(uri=uri, upsert="ignore")
meta = { meta = {

View file

@ -116,7 +116,7 @@ def test_SphereAggregation_storage(tmp_path: Path) -> None:
oasis_dataset = datasets.fetch_oasis_vbm(n_subjects=1) oasis_dataset = datasets.fetch_oasis_vbm(n_subjects=1)
vbm = oasis_dataset.gray_matter_maps[0] vbm = oasis_dataset.gray_matter_maps[0]
img = nib.load(vbm) img = nib.load(vbm)
uri = tmp_path / "test_sphere_storage_3D.db" uri = tmp_path / "test_sphere_storage_3D.sqlite"
storage = SQLiteFeatureStorage(uri=uri, upsert="ignore") storage = SQLiteFeatureStorage(uri=uri, upsert="ignore")
meta = { meta = {

View file

@ -92,7 +92,7 @@ def test_get_engine_single_output(tmp_path: Path) -> None:
The path to the test directory. The path to the test directory.
""" """
uri = tmp_path / "test_single_output.db" uri = tmp_path / "test_single_output.sqlite"
# Single storage, must be the uri # Single storage, must be the uri
storage = SQLiteFeatureStorage(uri=uri, upsert="ignore") storage = SQLiteFeatureStorage(uri=uri, upsert="ignore")
assert storage.single_output is True assert storage.single_output is True
@ -110,7 +110,7 @@ def test_get_engine_multi_output(tmp_path: Path) -> None:
The path to the test directory. The path to the test directory.
""" """
uri = tmp_path / "test_multi_output.db" uri = tmp_path / "test_multi_output.sqlite"
storage = SQLiteFeatureStorage( storage = SQLiteFeatureStorage(
uri=uri, single_output=False, upsert="ignore" uri=uri, single_output=False, upsert="ignore"
) )
@ -130,7 +130,7 @@ def test_get_engine_single_output_creation(tmp_path: Path) -> None:
tocreate = tmp_path / "tocreate" tocreate = tmp_path / "tocreate"
# Path does not exist yet # Path does not exist yet
assert not tocreate.exists() assert not tocreate.exists()
uri = tocreate.absolute() / "test_single_output.db" uri = tocreate.absolute() / "test_single_output.sqlite"
_ = SQLiteFeatureStorage(uri=uri, upsert="ignore") _ = SQLiteFeatureStorage(uri=uri, upsert="ignore")
# Path exists now # Path exists now
assert tocreate.exists() assert tocreate.exists()
@ -145,7 +145,7 @@ def test_upsert_replace(tmp_path: Path) -> None:
The path to the test directory. The path to the test directory.
""" """
uri = tmp_path / "test_upsert_replace.db" uri = tmp_path / "test_upsert_replace.sqlite"
# Single storage, must be the uri # Single storage, must be the uri
storage = SQLiteFeatureStorage(uri=uri, upsert="ignore") storage = SQLiteFeatureStorage(uri=uri, upsert="ignore")
# Metadata to store # Metadata to store
@ -179,7 +179,7 @@ def test_upsert_ignore(tmp_path: Path) -> None:
The path to the test directory. The path to the test directory.
""" """
uri = tmp_path / "test_upsert_ignore.db" uri = tmp_path / "test_upsert_ignore.sqlite"
# Single storage, must be the uri # Single storage, must be the uri
storage = SQLiteFeatureStorage(uri=uri, upsert="ignore") storage = SQLiteFeatureStorage(uri=uri, upsert="ignore")
# Metadata to store # Metadata to store
@ -217,7 +217,7 @@ def test_upsert_update(tmp_path: Path) -> None:
The path to the test directory. The path to the test directory.
""" """
uri = tmp_path / "test_upsert_delete.db" uri = tmp_path / "test_upsert_delete.sqlite"
storage = SQLiteFeatureStorage(uri=uri) storage = SQLiteFeatureStorage(uri=uri)
# Metadata to store # Metadata to store
meta = {"element": "test", "version": "0.0.1"} meta = {"element": "test", "version": "0.0.1"}
@ -250,7 +250,7 @@ def test_upsert_invalid_option(tmp_path: Path) -> None:
The path to the test directory. The path to the test directory.
""" """
uri = tmp_path / "test_upsert_invalid.db" uri = tmp_path / "test_upsert_invalid.sqlite"
with pytest.raises(ValueError): with pytest.raises(ValueError):
SQLiteFeatureStorage(uri=uri, upsert="wrong") SQLiteFeatureStorage(uri=uri, upsert="wrong")
@ -265,7 +265,7 @@ def test_store_df_and_read_df(tmp_path: Path) -> None:
The path to the test directory. The path to the test directory.
""" """
uri = tmp_path / "test_store_df_and_read_df.db" uri = tmp_path / "test_store_df_and_read_df.sqlite"
storage = SQLiteFeatureStorage(uri=uri, upsert="ignore") storage = SQLiteFeatureStorage(uri=uri, upsert="ignore")
# Metadata to store # Metadata to store
meta = { meta = {
@ -326,7 +326,7 @@ def test_store_metadata(tmp_path: Path) -> None:
The path to the test directory. The path to the test directory.
""" """
uri = tmp_path / "test_metadata_store.db" uri = tmp_path / "test_metadata_store.sqlite"
# Single storage, must be the uri # Single storage, must be the uri
storage = SQLiteFeatureStorage(uri=uri, upsert="ignore") storage = SQLiteFeatureStorage(uri=uri, upsert="ignore")
# Metadata to store # Metadata to store
@ -345,7 +345,7 @@ def test_store_table(tmp_path: Path) -> None:
The path to the test directory. The path to the test directory.
""" """
uri = tmp_path / "test_store_table.db" uri = tmp_path / "test_store_table.sqlite"
storage = SQLiteFeatureStorage(uri=uri) storage = SQLiteFeatureStorage(uri=uri)
# Metadata to store # Metadata to store
meta = {"element": "test", "version": "0.0.1", "marker": {"name": "fc"}} meta = {"element": "test", "version": "0.0.1", "marker": {"name": "fc"}}
@ -404,7 +404,7 @@ def test_store_matrix(tmp_path: Path) -> None:
The path to the test directory. The path to the test directory.
""" """
uri = tmp_path / "test_store_table.db" uri = tmp_path / "test_store_table.sqlite"
storage = SQLiteFeatureStorage(uri=uri) storage = SQLiteFeatureStorage(uri=uri)
# Metadata to store # Metadata to store
meta = {"element": "test", "version": "0.0.1", "marker": {"name": "fc"}} meta = {"element": "test", "version": "0.0.1", "marker": {"name": "fc"}}
@ -435,7 +435,7 @@ def test_store_matrix(tmp_path: Path) -> None:
assert_array_equal(read_df.values[0], data.flatten()) assert_array_equal(read_df.values[0], data.flatten())
assert list(read_df.columns) == stored_names assert list(read_df.columns) == stored_names
# Store without row and column names # Store without row and column names
uri = tmp_path / "test_store_table_nonames.db" uri = tmp_path / "test_store_table_nonames.sqlite"
storage = SQLiteFeatureStorage(uri=uri) storage = SQLiteFeatureStorage(uri=uri)
storage.store_matrix(data=data, meta=meta) storage.store_matrix(data=data, meta=meta)
stored_names = [ stored_names = [
@ -467,7 +467,7 @@ def test_store_matrix(tmp_path: Path) -> None:
data = np.array([[1, 2, 3], [11, 22, 33], [111, 222, 333]]) data = np.array([[1, 2, 3], [11, 22, 33], [111, 222, 333]])
row_names = ["row1", "row2", "row3"] row_names = ["row1", "row2", "row3"]
col_names = ["col1", "col2", "col3"] col_names = ["col1", "col2", "col3"]
uri = tmp_path / "test_store_table_triu.db" uri = tmp_path / "test_store_table_triu.sqlite"
storage = SQLiteFeatureStorage(uri=uri) storage = SQLiteFeatureStorage(uri=uri)
storage.store_matrix( storage.store_matrix(
data=data, data=data,
@ -496,7 +496,7 @@ def test_store_matrix(tmp_path: Path) -> None:
) )
# Store upper triangular matrix without diagonal # Store upper triangular matrix without diagonal
uri = tmp_path / "test_store_table_triu_nodiagonal.db" uri = tmp_path / "test_store_table_triu_nodiagonal.sqlite"
storage = SQLiteFeatureStorage(uri=uri) storage = SQLiteFeatureStorage(uri=uri)
storage.store_matrix( storage.store_matrix(
data=data, data=data,
@ -526,7 +526,7 @@ def test_store_matrix(tmp_path: Path) -> None:
data = np.array([[1, 2, 3], [11, 22, 33], [111, 222, 333]]) data = np.array([[1, 2, 3], [11, 22, 33], [111, 222, 333]])
row_names = ["row1", "row2", "row3"] row_names = ["row1", "row2", "row3"]
col_names = ["col1", "col2", "col3"] col_names = ["col1", "col2", "col3"]
uri = tmp_path / "test_store_table_tril.db" uri = tmp_path / "test_store_table_tril.sqlite"
storage = SQLiteFeatureStorage(uri=uri) storage = SQLiteFeatureStorage(uri=uri)
storage.store_matrix( storage.store_matrix(
data=data, data=data,
@ -555,7 +555,7 @@ def test_store_matrix(tmp_path: Path) -> None:
) )
# Store lower triangular matrix without diagonal # Store lower triangular matrix without diagonal
uri = tmp_path / "test_store_table_tril_nodiagonal.db" uri = tmp_path / "test_store_table_tril_nodiagonal.sqlite"
storage = SQLiteFeatureStorage(uri=uri) storage = SQLiteFeatureStorage(uri=uri)
storage.store_matrix( storage.store_matrix(
data, data,
@ -592,7 +592,7 @@ def test_store_multiple_output(tmp_path: Path):
The path to the test directory. The path to the test directory.
""" """
uri = tmp_path / "test_store_multiple_output.db" uri = tmp_path / "test_store_multiple_output.sqlite"
storage = SQLiteFeatureStorage(uri=uri, single_output=False) storage = SQLiteFeatureStorage(uri=uri, single_output=False)
# Metadata to store # Metadata to store
meta1 = { meta1 = {
@ -691,7 +691,7 @@ def test_collect(tmp_path: Path) -> None:
The path to the test directory. The path to the test directory.
""" """
uri = tmp_path / "test_collect.db" uri = tmp_path / "test_collect.sqlite"
storage = SQLiteFeatureStorage(uri=uri, single_output=False) storage = SQLiteFeatureStorage(uri=uri, single_output=False)
# Metadata for storage # Metadata for storage
meta1 = { meta1 = {