[DOC]: Improve documentation #317

Merged
synchon merged 19 commits from update/improve-docs into main 2024-04-02 07:59:33 +00:00
27 changed files with 842 additions and 338 deletions

View file

@ -25,7 +25,7 @@ Data Grabber
- In Progress - In Progress
- Done - Done
Version added: If the status is "Done", the Junifer version in which the Version added: If the status is "Done", the junifer version in which the
dataset was added. Else, a link to the Github issue or pull request dataset was added. Else, a link to the Github issue or pull request
implementing the dataset. Links to github can be added by using the implementing the dataset. Links to github can be added by using the
following syntax: :gh:`<issue number>` following syntax: :gh:`<issue number>`
@ -122,6 +122,54 @@ Planned
- :gh:`47` - :gh:`47`
Preprocessor
------------
..
Provide a list of the Preprocessors that are implemented or planned.
State: this should indicate the state of the preprocessor. Valid options are
- Planned
- In Progress
- Done
LeSasse commented 2024-03-28 12:56:11 +00:00 (Migrated from github.com)

Does Junifer handle/allow use of multiple preprocessors chained in sequence? how would parametrisation in the YAML look like for that?

Does Junifer handle/allow use of multiple preprocessors chained in sequence? how would parametrisation in the YAML look like for that?
synchon commented 2024-03-28 13:42:22 +00:00 (Migrated from github.com)

Yeah it does. preprocess accepts a list now and the execution follows the sequence you specify in the YAML.

Yeah it does. `preprocess` accepts a list now and the execution follows the sequence you specify in the YAML.
Version added: If the status is "Done", the junifer version in which the
preprocessor was added. Else, a link to the Github issue or pull request
implementing the preprocessor. Links to github can be added by using the
following syntax: :gh:`<issue number>`
Available
~~~~~~~~~
.. list-table::
:widths: auto
:header-rows: 1
* - Class
- Description
- State
- Version Added
* - :class:`.fMRIPrepConfoundRemover`
- Remove confounds from ``fMRIPrep``-ed data
- Done
- 0.0.1
* - :class:`.SpaceWarper`
- | Warp / transform data from one space to another
| (subject-native or other template spaces)
- Done
- 0.0.4
* - ``Smoothing``
- | Apply smoothing to data, particularly useful when dealing with
| ``fMRIPrep``-ed data
- In Progress
- :gh:`161`
..
Planned
~~~~~~~
LeSasse commented 2024-03-28 12:57:00 +00:00 (Migrated from github.com)

Why particularly useful for fMRIPrep'ed data?

Why particularly useful for fMRIPrep'ed data?
synchon commented 2024-03-28 13:44:08 +00:00 (Migrated from github.com)

From your issue description I understand that fMRIPrep doesn't perform smoothing after confound regression.

From your issue description I understand that fMRIPrep doesn't perform smoothing after confound regression.
Marker Marker
------ ------
@ -133,7 +181,7 @@ Marker
- In Progress - In Progress
- Done - Done
Version added: If the status is "Done", the Junifer version in which the Version added: If the status is "Done", the junifer version in which the
marker was added. Else, a link to the Github issue or pull request marker was added. Else, a link to the Github issue or pull request
implementing the marker. Links to github can be added by using the implementing the marker. Links to github can be added by using the
following syntax: :gh:`<issue number>` following syntax: :gh:`<issue number>`
@ -254,10 +302,6 @@ Planned
* - Connectedness * - Connectedness
- Compute connectedness - Compute connectedness
- :gh:`34` - :gh:`34`
* - Permutation entropy, Range entropy, Multiscale entropy and Hurst exponent
- | Calculate Permutation entropy, Range entropy, Multiscale entropy and
| Hurst exponent
- :gh:`61`
Parcellation Parcellation
------------ ------------
@ -265,7 +309,7 @@ Parcellation
.. ..
Provide a list of the Parcellations that are implemented or planned. Provide a list of the Parcellations that are implemented or planned.
Version added: The Junifer version in which the parcellation was added. Version added: The junifer version in which the parcellation was added.
Available Available
~~~~~~~~~ ~~~~~~~~~
@ -277,8 +321,8 @@ Available
* - Name * - Name
LeSasse commented 2024-03-28 12:57:41 +00:00 (Migrated from github.com)

Above you add Version Added with captial A. Similar for Template spaces, should it be Template Spaces?

Above you add Version Added with captial A. Similar for Template spaces, should it be Template Spaces?
synchon commented 2024-03-28 13:45:33 +00:00 (Migrated from github.com)

Yeah missed it, good catch.

Yeah missed it, good catch.
- Options - Options
- Keys - Keys
- Spaces - Template Spaces
- Version added - Version Added
- Publication - Publication
* - Schaefer * - Schaefer
- ``n_rois``, ``yeo_networks`` - ``n_rois``, ``yeo_networks``
@ -439,7 +483,7 @@ Coordinates
.. ..
Provide a list of the Coordinates that are implemented or planned. Provide a list of the Coordinates that are implemented or planned.
Version added: The Junifer version in which the parcellation was added. Version added: The junifer version in which the parcellation was added.
Available Available
~~~~~~~~~ ~~~~~~~~~
@ -450,7 +494,7 @@ Available
* - Name * - Name
- Keys - Keys
- Version added - Version Added
- Publication - Publication
* - Cognitive action control * - Cognitive action control
- ``CogAC`` - ``CogAC``
@ -635,7 +679,7 @@ Mask
.. ..
Provide a list of the masks that are implemented or planned. Provide a list of the masks that are implemented or planned.
Version added: The Junifer version in which the mask was added. Version added: The junifer version in which the mask was added.
Available Available
~~~~~~~~~ ~~~~~~~~~
@ -646,8 +690,8 @@ Available
LeSasse commented 2024-03-28 12:57:54 +00:00 (Migrated from github.com)

see above cmt

see above cmt
* - Name * - Name
- Keys - Keys
- Spaces - Template Space
- Version added - Version Added
- Description - Publication - Description - Publication
* - Vickery-Patil (Gray Matter) * - Vickery-Patil (Gray Matter)
- | ``GM_prob0.2`` - | ``GM_prob0.2``
@ -663,14 +707,14 @@ Available
- | Vickery, Sam, & Patil, Kaustubh. (2022). - | Vickery, Sam, & Patil, Kaustubh. (2022).
LeSasse commented 2024-03-28 12:58:32 +00:00 (Migrated from github.com)

you made a point to fix Junifer to junifer. Should it be nilearn as well rather than Nilearn?

you made a point to fix Junifer to junifer. Should it be nilearn as well rather than Nilearn?
synchon commented 2024-03-28 13:46:36 +00:00 (Migrated from github.com)

That would be better, fair point.

That would be better, fair point.
| Chimpanzee and Human Gray Matter Masks [Data set]. Zenodo. | Chimpanzee and Human Gray Matter Masks [Data set]. Zenodo.
| https://doi.org/10.5281/zenodo.6463123 | https://doi.org/10.5281/zenodo.6463123
* - Nilearn's MNI152 1mm-resolution mask * - ``junifer``'s custom brain mask
- | ``compute_brain_mask`` - | ``compute_brain_mask``
- Adapts to the target data - Adapts to the target data
- 0.0.2 - 0.0.2
- | Compute the whole-brain mask. This mask is calculated using - | Compute the whole-brain, gray-matter or white-matter mask using
| MNI152 1mm-resolution template mask onto the target image. | the template and the resolution from the target image. The
| See :func:`nilearn.masking.compute_brain_mask` | templates are obtained via ``templateflow``.
* - Nilearn's mask computed from FMRI data * - ``nilearn``'s mask computed from fMRI data
- | ``compute_epi_mask`` - | ``compute_epi_mask``
- Adapts to the target data - Adapts to the target data
- 0.0.2 - 0.0.2
@ -678,15 +722,14 @@ Available
| proposed by T.Nichols: find the least dense point of the histogram, | proposed by T.Nichols: find the least dense point of the histogram,
| between fractions ``lower_cutoff`` and ``upper_cutoff`` of the total | between fractions ``lower_cutoff`` and ``upper_cutoff`` of the total
| image histogram. See :func:`nilearn.masking.compute_epi_mask` | image histogram. See :func:`nilearn.masking.compute_epi_mask`
* - Nilearn's background mask * - ``nilearn``'s background mask
- | ``compute_background_mask`` - | ``compute_background_mask``
- Adapts to the target data - Adapts to the target data
- 0.0.2 - 0.0.2
- | Compute a brain mask for the images by guessing the value of the - | Compute a brain mask for the images by guessing the value of the
| background from the border of the image. | background from the border of the image.
| See :func:`nilearn.masking.compute_background_mask` | See :func:`nilearn.masking.compute_background_mask`
* - ``nilearn``'s ICBM152 template gray-matter mask
* - Nilearn's ICBM152 template gray-matter mask
- | ``fetch_icbm152_brain_gm_mask`` - | ``fetch_icbm152_brain_gm_mask``
- ``MNI152NLin2009aAsym`` - ``MNI152NLin2009aAsym``
- 0.0.2 - 0.0.2
@ -695,8 +738,9 @@ Available
| See :func:`nilearn.datasets.fetch_icbm152_brain_gm_mask` | See :func:`nilearn.datasets.fetch_icbm152_brain_gm_mask`
Planned ..
~~~~~~~ Planned
~~~~~~~
.. ..
helpful site for creating tables: https://rest-sphinx-memo.readthedocs.io/en/latest/ReST.html#tables helpful site for creating tables: https://rest-sphinx-memo.readthedocs.io/en/latest/ReST.html#tables

View file

@ -0,0 +1 @@
Improve documentation by adding information about space transformation and writing custom Preprocessors by `Synchon Mandal`_

View file

@ -164,6 +164,8 @@ In case you remove some files or change their filenames, you can run into
errors when using ``make local``. In this situation you can use ``make clean`` errors when using ``make local``. In this situation you can use ``make clean``
to clean up the already build files and then re-run ``make local``. to clean up the already build files and then re-run ``make local``.
Also, we follow British English for the documentation.
Writing Examples Writing Examples
---------------- ----------------

View file

@ -6,19 +6,20 @@ Adding Coordinates
================== ==================
Instead of using whole-brain parcellations to aggregate voxel-wise signals from Instead of using whole-brain parcellations to aggregate voxel-wise signals from
MR images (as for example in the :class:`.ParcelAggregation` marker), junifer MR images (as for example in the :class:`.ParcelAggregation` marker), ``junifer``
allows you to specify a set of coordinates around which to draw spheres to allows you to specify a set of coordinates around which to draw spheres to
aggregate (for example using the :class:`.SphereAggregation` marker) the MR aggregate (for example using the :class:`.SphereAggregation` marker) the MR
signals from individual voxels. Now, before you start specifying your own sets signals from individual voxels. Now, before you start specifying your own sets
of coordinates, check the coordinates that junifer already has of coordinates, check the coordinates that ``junifer`` already has
:ref:`built in <builtin>`. If you simply want to use a well known set of :ref:`built in <builtin>`. If you simply want to use a well known set of
coordinates from the literature, there is a reasonable chance, that junifer coordinates from the literature, there is a reasonable chance, that ``junifer``
provides them already. provides them already.
If you checked the in-built coordinates, and they are not there already (for If you checked the in-built coordinates, and they are not there already (for
example if you came up with your own set of coordinates), then junifer provides example if you came up with your own set of coordinates), then ``junifer``
an easy way for you to register them using the :func:`.register_coordinates` provides an easy way for you to register them using the
function, so you can use your own set of coordinates within a junifer pipeline. :func:`.register_coordinates` function, so you can use your own set of
coordinates within a ``junifer`` pipeline.
From the API reference, we can see that it has 4 positional arguments From the API reference, we can see that it has 4 positional arguments
(``name``, ``coordinates``, ``voi_names`` and ``space``) as well as one (``name``, ``coordinates``, ``voi_names`` and ``space``) as well as one
@ -26,7 +27,7 @@ optional keyword argument (``overwrite``).
The ``name`` argument takes a string indicating the name you want to give to The ``name`` argument takes a string indicating the name you want to give to
this set of coordinates. This ``name`` can be used to obtain and operate on a this set of coordinates. This ``name`` can be used to obtain and operate on a
set of coordinates in junifer. For example, you can obtain your coordinates set of coordinates in ``junifer``. For example, you can obtain your coordinates
after registration by providing ``name`` to :func:`.load_coordinates`. We could after registration by providing ``name`` to :func:`.load_coordinates`. We could
simply call it ``"my_set_of_coordinates"``, but likely you want a more simply call it ``"my_set_of_coordinates"``, but likely you want a more
descriptive and more informative name most of the time. descriptive and more informative name most of the time.
@ -36,9 +37,7 @@ The ``coordinates`` argument takes the actual coordinates as a 2-dimensional
columns (one for each spatial dimension). That is, the first, second, and third columns (one for each spatial dimension). That is, the first, second, and third
columns indicate the x-, y-, and z-coordinates in MNI space respectively. columns indicate the x-, y-, and z-coordinates in MNI space respectively.
The number of rows in the array correspond to the number of coordinates that The number of rows in the array correspond to the number of coordinates that
belong to this set. Note, that junifer (as of yet) only works in MNI space, and belong to this set.
so therefore these coordinates should always be real-world coordinates of the
MNI space.
The ``voi_names`` argument takes a list of strings The ``voi_names`` argument takes a list of strings
indicating the names of each coordinate (i.e. volume-of-interest) in the indicating the names of each coordinate (i.e. volume-of-interest) in the
@ -47,7 +46,7 @@ the number of rows in the coordinates array. Now, we know everything we need to
know to register a set of coordinates. know to register a set of coordinates.
Lastly, we specify the ``space`` that the coordinates are in, for example, Lastly, we specify the ``space`` that the coordinates are in, for example,
``"MNI"`` or ``"Native"`` (scanner-native space). ``"MNI"`` or ``"native"`` (scanner-native space).
Step 1: Prepare code to register a set of coordinates Step 1: Prepare code to register a set of coordinates
----------------------------------------------------- -----------------------------------------------------
@ -62,8 +61,8 @@ packages:
import numpy as np import numpy as np
For the sake of this example, we can create a set of coordinates that belong For the sake of this example, we can create a set of coordinates that belong
to the default mode network (DMN), and register this set of coordinates with to the Default Mode Network (DMN), and register this set of coordinates with
junifer. Note, that junifer already has a ``junifer``. Note, that ``junifer`` already has a
:ref:`set of coordinates built-in <builtin>` ("DMNBuckner") that is associated :ref:`set of coordinates built-in <builtin>` ("DMNBuckner") that is associated
with the DMN. Here, we use the DMN coordinates used in a with the DMN. Here, we use the DMN coordinates used in a
`nilearn example <https://nilearn.github.io/dev/auto_examples/03_connectivity/plot_sphere_based_connectome.html>`_. `nilearn example <https://nilearn.github.io/dev/auto_examples/03_connectivity/plot_sphere_based_connectome.html>`_.
@ -97,7 +96,7 @@ simply use this to register our coordinates:
space="MNI" space="MNI"
) )
Now, when we run this script, junifer registers these coordinates and we can Now, when we run this script, ``junifer`` registers these coordinates and we can
use them in subsequent analyses. Let's now consider how to use coordinate use them in subsequent analyses. Let's now consider how to use coordinate
registration in combination with registration in combination with
:ref:`codeless configuration using a YAML file <codeless>`. :ref:`codeless configuration using a YAML file <codeless>`.
@ -106,7 +105,7 @@ Step 2: Add coordinate registration to the YAML file
---------------------------------------------------- ----------------------------------------------------
In order to register your coordinates for a pipeline configured by a YAML file, In order to register your coordinates for a pipeline configured by a YAML file,
you can use the ``with`` keyword provided by junifer: you can use the ``with`` keyword provided by ``junifer``:
.. code-block:: yaml .. code-block:: yaml

View file

@ -12,8 +12,8 @@ the structure of a dataset and provide two specific functionalities:
element (e.g. the path to the T1 image, the path to the T2 image, etc.) element (e.g. the path to the T1 image, the path to the T2 image, etc.)
#. Provide the list of *elements* available in the dataset. #. Provide the list of *elements* available in the dataset.
In this section, we will see how to create a datagrabber for a dataset. Basic In this section, we will see how to create a DataGrabber for a dataset. Basic
aspects of datagrabbers are covered in the aspects of DataGrabbers are covered in the
:ref:`Understanding Data Grabbers <datagrabber>` section. :ref:`Understanding Data Grabbers <datagrabber>` section.
.. _extending_datagrabbers_think: .. _extending_datagrabbers_think:
@ -22,7 +22,7 @@ Step 1: Think about the element
------------------------------- -------------------------------
Like with any programming-related task, the first step is to think. When Like with any programming-related task, the first step is to think. When
creating a Data Grabber, we need to first define what an *element* is. creating a DataGrabber, we need to first define what an *element* is.
The *element* should be the smallest unit of data that can be processed. That The *element* should be the smallest unit of data that can be processed. That
is, for each element, there should be a set of data that can be processed, but is, for each element, there should be a set of data that can be processed, but
only one of each *data type* (see :ref:`data_types`). only one of each *data type* (see :ref:`data_types`).
@ -42,28 +42,28 @@ then the *element* should be composed of 3 items:
If any of these items were not part of the element, then we will have more than If any of these items were not part of the element, then we will have more than
one ``T1w`` and / or ``BOLD`` image for each subject, which is not allowed. one ``T1w`` and / or ``BOLD`` image for each subject, which is not allowed.
Importantly, nothing prevents that one image is part of two different elements. Importantly, nothing prevents that one image being part of two different
For example, it is usually the case that the ``T1w`` image is not acquired for elements. For example, it is usually the case that the ``T1w`` image is not
each task, but once in the entire session. So in this case, the ``T1w`` image acquired for each task, but once in the entire session. So in this case, the
for the element (``sub001``, ``ses1``, ``rest``) will be the same as the ``T1w`` image for the element (``sub001``, ``ses1``, ``rest``) will be the same
``T1w`` image for the element (``sub001``, ``ses1``, ``stroop``). as the ``T1w`` image for the element (``sub001``, ``ses1``, ``stroop``).
We will now continue this section using as an example, a dataset in BIDS format We will now continue this section using as an example, a dataset in BIDS format
in which 9 subjects (``sub-01`` to ``sub-09``) were scanned each during 3 in which 9 subjects (``sub-01`` to ``sub-09``) were scanned each during 3
sessions (``ses-01``, ``ses-02``, ``ses-03``) and each session included a sessions (``ses-01``, ``ses-02``, ``ses-03``) and each session included a
``T1w`` and a ``BOLD`` image (resting-state), except for ``ses-03`` which was ``T1w`` and a ``BOLD`` image (resting-state), except for ``ses-03`` which was
only anatomical. only anatomical data.
Step 2: Think about the dataset's structure Step 2: Think about the dataset's structure
------------------------------------------- -------------------------------------------
Now that we have our element defined, we need to think about the structure of Now that we have our element defined, we need to think about the structure of
the dataset. Mainly, because the structure of the dataset will determine how the dataset. Mainly, because the structure of the dataset will determine how
the Data Grabber needs to be implemented. the DataGrabber needs to be implemented.
Junifer provides an abstract class to deal with datasets that can be thought in ``junifer`` provides an abstract class to deal with datasets that can be thought
terms of *patterns*. A *pattern* is a string that contains placeholders that are in terms of *patterns*. A *pattern* is a string that contains placeholders that
replaced by the actual values of the element. In our BIDS example, the path are replaced by the actual values of the element. In our BIDS example, the path
to the T1w image of subject ``sub-01`` and session ``ses-01``, relative to the to the T1w image of subject ``sub-01`` and session ``ses-01``, relative to the
dataset location, is ``sub-01/ses-01/anat/sub-01_ses-01_T1w.nii.gz``. By dataset location, is ``sub-01/ses-01/anat/sub-01_ses-01_T1w.nii.gz``. By
replacing ``sub-01`` with ``sub-02``, we can obtain the T1w image of the first replacing ``sub-01`` with ``sub-02``, we can obtain the T1w image of the first
@ -88,7 +88,7 @@ discussion in the `junifer Discussions`_ page. Most probably we can help you
get your dataset in order. get your dataset in order.
If there is no other way, then you can follow :ref:`extending_datagrabbers_base` If there is no other way, then you can follow :ref:`extending_datagrabbers_base`
to create a Data Grabber from scratch. to create a DataGrabber from scratch.
.. _extending_datagrabbers_pattern: .. _extending_datagrabbers_pattern:
@ -101,7 +101,7 @@ Option A: Extending from PatternDataGrabber
The :class:`.PatternDataGrabber` class is an abstract class that has the The :class:`.PatternDataGrabber` class is an abstract class that has the
functionality of understanding patterns embedded in it. functionality of understanding patterns embedded in it.
Before creating the datagrabber, we need to define 3 variables: Before creating the DataGrabber, we need to define 3 variables:
* ``types``: A list with the available :ref:`data_types` in our dataset. * ``types``: A list with the available :ref:`data_types` in our dataset.
* ``patterns``: A dictionary that specifies the pattern for each data type. * ``patterns``: A dictionary that specifies the pattern for each data type.
@ -126,7 +126,7 @@ where the dataset is located. For example, if the dataset is located in
location of the dataset, we can expose the variable in the constructor, as in location of the dataset, we can expose the variable in the constructor, as in
the following example. the following example.
With the variables defined above, we can create our datagrabber and name it With the variables defined above, we can create our DataGrabber and name it
``ExampleBIDSDataGrabber``: ``ExampleBIDSDataGrabber``:
.. code-block:: python .. code-block:: python
@ -152,7 +152,7 @@ With the variables defined above, we can create our datagrabber and name it
replacements=replacements, replacements=replacements,
) )
Our datagrabber is ready to be used by junifer. However, it is still unknown Our DataGrabber is ready to be used by ``junifer``. However, it is still unknown
to the library. We need to register it in the library. To do so, we need to to the library. We need to register it in the library. To do so, we need to
use the :func:`.register_datagrabber` decorator. use the :func:`.register_datagrabber` decorator.
@ -183,9 +183,9 @@ use the :func:`.register_datagrabber` decorator.
) )
Now, we can use our datagrabber in junifer, by setting the ``datagrabber`` kind Now, we can use our DataGrabber in ``junifer``, by setting the ``datagrabber``
in the yaml file to ``ExampleBIDSDataGrabber``. Remember that we still need to kind in the yaml file to ``ExampleBIDSDataGrabber``. Remember that we still need
set the ``datadir``. to set the ``datadir``.
.. code-block:: yaml .. code-block:: yaml
@ -209,7 +209,7 @@ temporary directory. To set the location of the dataset, you can use the
be used to specify the path to the root directory of the dataset after doing be used to specify the path to the root directory of the dataset after doing
``datalad clone``. ``datalad clone``.
In the example, the dataset is hosted in gin In the example, the dataset is hosted in Gin
(``https://gin.g-node.org/juaml/datalad-example-bids``). (``https://gin.g-node.org/juaml/datalad-example-bids``).
When we clone this dataset, we will see the following structure: When we clone this dataset, we will see the following structure:
@ -238,7 +238,7 @@ Now we have our 2 additional variables:
uri = "https://gin.g-node.org/juaml/datalad-example-bids" uri = "https://gin.g-node.org/juaml/datalad-example-bids"
rootdir = "example_bids_ses" rootdir = "example_bids_ses"
And we can create our datagrabber: And we can create our DataGrabber:
.. code-block:: python .. code-block:: python
@ -267,17 +267,34 @@ And we can create our datagrabber:
replacements=replacements, replacements=replacements,
) )
This approach can be used directly from the YAML, like so:
.. code-block:: yaml
datagrabber:
- kind: PatternDataladDataGrabber
types:
- BOLD
- T1w
patterns:
BOLD: "{subject}/{session}/func/{subject}_{session}_task-rest_bold.nii.gz"
T1w: "{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz"
replacements:
- subject
- session
uri: "https://gin.g-node.org/juaml/datalad-example-bids"
rootdir: "example_bids_ses"
.. _extending_datagrabbers_base: .. _extending_datagrabbers_base:
Option B: Extending from BaseDataGrabber Option B: Extending from BaseDataGrabber
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
While we could not think of a use case in which the pattern-based datagrabber While we could not think of a use case in which the pattern-based DataGrabber
would not be suitable, it is still possible to create a datagrabber extending would not be suitable, it is still possible to create a DataGrabber extending
from the :class:`.BaseDataGrabber` class. from the :class:`.BaseDataGrabber` class.
In order to create a datagrabber extending from :class:`.BaseDataGrabber`, we In order to create a DataGrabber extending from :class:`.BaseDataGrabber`, we
need to implement the following methods: need to implement the following methods:
- ``get_item``: to get a single item from the dataset. - ``get_item``: to get a single item from the dataset.
@ -287,7 +304,7 @@ need to implement the following methods:
.. note:: .. note::
The ``__init__`` method could also be implemented, but it is not mandatory. The ``__init__`` method could also be implemented, but it is not mandatory.
This is required if the datagrabber requires any extra parameter. This is required if the DataGrabber requires any extra parameter.
We will now implement our BIDS example with this method. We will now implement our BIDS example with this method.
@ -339,7 +356,7 @@ method, in the same order.
return ["subject", "session"] return ["subject", "session"]
LeSasse commented 2024-03-28 13:55:23 +00:00 (Migrated from github.com)

this is american spelling

this is american spelling
synchon commented 2024-03-28 15:08:44 +00:00 (Migrated from github.com)

Boston tea party reversed.

Boston tea party reversed.
So, to summarize, our datagrabber will look like this: So, to summarise, our DataGrabber will look like this:
.. code-block:: python .. code-block:: python
@ -403,7 +420,7 @@ not standardised.
The ``mappings`` key is only required if the ``format`` is ``adhoc``. If the The ``mappings`` key is only required if the ``format`` is ``adhoc``. If the
``format`` is ``fmriprep``, the ``mappings`` key is not required. ``format`` is ``fmriprep``, the ``mappings`` key is not required.
Currently, junifer provides only one confound remover step Currently, ``junifer`` provides only one confound remover step
(:class:`.fMRIPrepConfoundRemover`), which relies entirely on the ``fmriprep`` (:class:`.fMRIPrepConfoundRemover`), which relies entirely on the ``fmriprep``
confound variable names. Thus, if the confounds are not in ``fmriprep`` format, confound variable names. Thus, if the confounds are not in ``fmriprep`` format,
the user will need to provide the mappings between the *ad-hoc* variable names the user will need to provide the mappings between the *ad-hoc* variable names

View file

@ -0,0 +1,141 @@
.. include:: ../links.inc
LeSasse commented 2024-03-28 13:00:12 +00:00 (Migrated from github.com)

Should dependencies be capitalised?

Should dependencies be capitalised?
LeSasse commented 2024-03-28 13:01:38 +00:00 (Migrated from github.com)

"having two keys" -> "with two keys"

"having two keys" -> "with two keys"
.. _specifying_dependencies:
Specifying Dependencies
=======================
This section describes how you can tackle different situations when writing
your custom Marker and / or Preprocessor and take care of the dependencies for
them.
You might have already come across listing out dependencies for your
:ref:`custom Markers <extending_markers>` and / or
:ref:`custom Preprocessors <extending_preprocessors>`, if not, check them out
first. If you have already gone through them, you are already familiar with using
class attribute ``_DEPENDENCIES`` to keep track of its dependencies. ``junifer``
is a bit more sophisticated about them and we will see here how you can make the
best use of them.
.. _component_dependencies:
Handling dependencies that come as Python packages
--------------------------------------------------
You have already seen this case handled by having a class attribute
``_DEPENDENCIES`` whose value is a set of all the package names that the
component depends on. For example, for :class:`.RSSETSMarker`, we have:
.. code-block:: python
_DEPENDENCIES: ClassVar[Set[str]] = {"nilearn"}
The type annotation is for documentation and static type checking purposes.
Although not required, we highly recommend you use them, your future self
and others who use it will thank you.
.. _component_external_dependencies:
Handling external dependencies from toolboxes
---------------------------------------------
You can also specify dependencies of external toolboxes like AFIN, FSL and ANTs,
by having a class attribute like so:
.. code-block:: python
_EXT_DEPENDENCIES: ClassVar[List[Dict[str, Union[str, List[str]]]]] = [
{
"name": "afni",
"commands": ["3dReHo", "3dAFNItoNIFTI"],
},
]
The above example is taken from the class which computes regional homogeneity
(ReHo) using AFNI. The general pattern is that you need to have the value of
``_EXT_DEPENDENCIES`` as a list of dictionary with two keys:
* ``name`` (str) : lowercased name of the toolbox
* ``commands`` (list of str) : actual names of the commands you need to use
This is simple but powerful as we will see in the following sub-sections.
.. _component_conditional_dependencies:
Handling conditional dependencies
---------------------------------
You might encounter situations where your Marker or Preprocessor needs to have
option for the user to either use a dependency that comes as a package or
use a dependency that relies on external toolboxes. With the foundation we laid
above, it is really simple to solve it while having validation before running
and letting the user know if some dependency is missing.
Let's look at an actual implementation, in this case :class:`.SpaceWarper`, so
that it shows the problem a bit better and how we solve it:
.. code-block:: python
class SpaceWarper(BasePreprocessor):
# docstring
_CONDITIONAL_DEPENDENCIES: ClassVar[List[Dict[str, Union[str, Type]]]] = [
{
"using": "fsl",
"depends_on": FSLWarper,
},
{
"using": "ants",
"depends_on": ANTsWarper,
},
]
def __init__(
self, using: str, reference: str, on: Union[List[str], str]
) -> None:
# validation and setting up
Here, you see a new class attribute ``_CONDITIONAL_DEPENDENCIES`` which is a
list of dictionaries with two keys:
* ``using`` (str) : lowercased name of the toolbox
* ``depends_on`` (object) : a class which implements the particular tool's use
It is mandatory to have the ``using`` positional argument in the constructor in
this case as the validation starts with this and moves further. It is also
mandatory to only allow the value of ``using`` argument to be one of them
specified in the ``using`` key of ``_CONDITIONAL_DEPENDENCIES`` entries.
For brevity, we only show the ``FSLWarper`` here but ``ANTsWarper`` looks very
LeSasse commented 2024-03-28 13:02:31 +00:00 (Migrated from github.com)

``ANTsWarper``` the capitalisation here makes me anxious, but I'll allow it

``ANTsWarper``` the capitalisation here makes me anxious, but I'll allow it
synchon commented 2024-03-28 13:50:02 +00:00 (Migrated from github.com)

It's to follow the convention of the tool name like in other places in the code base.

It's to follow the convention of the tool name like in other places in the code base.
LeSasse commented 2024-03-28 13:51:16 +00:00 (Migrated from github.com)

I understand :)

I understand :)
similar. ``FSLWarper`` looks like this (only the relevant part is shown here):
.. code-block:: python
class FSLWarper:
# docstring
_EXT_DEPENDENCIES: ClassVar[List[Dict[str, Union[str, List[str]]]]] = [
{
"name": "fsl",
"commands": ["flirt", "applywarp"],
},
]
_DEPENDENCIES: ClassVar[Set[str]] = {"numpy", "nibabel"}
def preprocess(
self,
input: Dict[str, Any],
extra_input: Dict[str, Any],
) -> Dict[str, Any]:
# implementation
Here you can see the familiar ``_DEPENDENCIES`` and ``_EXT_DEPENDENCIES`` class
attributes. The validation process starts by looking up the ``using`` value of
the ``_CONDITIONAL_DEPENDENCIES`` entries and then retrieves the object pointed
by ``depends_on``. After that, the ``_DEPENDENCIES`` and ``_EXT_DEPENDENCIES``
class attributes are checked.
This might be a bit too much to get it right away so feel free to check the code
for a better understanding. You can also check ``ALFFBase`` for a Marker
having this pattern.

View file

@ -2,19 +2,20 @@
.. _extending_extension: .. _extending_extension:
Creating a junifer extension Creating a ``junifer`` extension
============================ ================================
Junifer is designed to be easily extensible. Through the use of a registry and ``junifer`` is designed to be easily extensible. Through the use of a registry
decorators, you can easily add new functionality to junifer during runtime. This and decorators, you can easily add new functionality to ``junifer`` during
is done by creating a new Python module and importing it before running junifer. runtime. This is done by creating a new Python module and importing it before
running ``junifer``.
A special consideration has to be made when using the A special consideration has to be made when using the
:ref:`code-less configuration<codeless>`. In this case, the :ref:`code-less configuration<codeless>`. In this case, the
``with`` statement can be used to import a module or run a Python file. ``with`` statement can be used to import a module or run a Python file.
In the following example, we instruct junifer to first import ``my_module`` and In the following example, we instruct ``junifer`` to first import ``my_module``
then run the ``my_file.py`` file. and then run the ``my_file.py`` file:
.. code-block:: yaml .. code-block:: yaml
@ -22,13 +23,13 @@ then run the ``my_file.py`` file.
- my_module - my_module
- my_file.py - my_file.py
Thus, the code from ``my_file.py`` will be executed before running junifer. This Thus, the code from ``my_file.py`` will be executed before running ``junifer``.
is the ideal place to include junifer extensions. This is the ideal place to include ``junifer`` extensions.
.. important:: .. important::
Some junifer commands will not consider files imported from files included Some ``junifer`` commands will not consider files imported from files included
in the ``with`` statement. If ``my_file.py`` imports ``my_other_file.py``, in the ``with`` statement. If ``my_file.py`` imports ``my_other_file.py``,
some of the junifer commands will not consider ``my_other_file.py``. Either some of the ``junifer`` commands will not consider ``my_other_file.py``. Either
place all the code in one file or add multiple files to the ``with`` place all the code in one file or add multiple files to the ``with``
statement. statement.

View file

@ -2,20 +2,20 @@
.. _extending: .. _extending:
Extending junifer Extending ``junifer``
================= =====================
While we aim to provide as many datasets and markers as possible, we are also While we aim to provide as many datasets and markers as possible, we are also
interested in allowing users to extend the functionality with their own interested in allowing users to extend the functionality with their own
datagrabbers, preprocessing, markers, etc., . DataGrabbers, Preprocessors, Markers, etc., .
It's not necessary to have the new functionality included in junifer before It's not necessary to have the new functionality included in ``junifer`` before
the user can use them. The user can simply create a new Python file, code the the user can use them. The user can simply create a new Python file, code the
desired functionality and use it with junifer. This is the first step towards desired functionality and use it with ``junifer``. This is the first step towards
including the new functionality in the junifer pipeline. including the new functionality in the ``junifer`` pipeline.
In this section we will show how to extend junifer, by creating new In this section we will show how to extend ``junifer``, by creating new
datagrabbers, preprocessing, markers, etc., following the *junifer* way. DataGrabbers, Preprocessors, Markers, etc., following the *junifer* way.
.. toctree:: .. toctree::
@ -25,6 +25,8 @@ datagrabbers, preprocessing, markers, etc., following the *junifer* way.
extension extension
datagrabber datagrabber
marker marker
preprocessor
dependencies
parcellations parcellations
coordinates coordinates
masks masks

View file

@ -5,35 +5,35 @@
Creating Markers Creating Markers
================ ================
Computing a marker (a.k.a. *feature*) is the main goal of junifer. While we aim Computing a marker (a.k.a. *feature*) is the main goal of ``junifer``. While we
to provide as many markers as possible, it might be the case that the marker you aim to provide as many Markers as possible, it might be the case that the Marker
are looking for is not available. In this case, you can create your own marker you are looking for is not available. In this case, you can create your own Marker
by following this tutorial. by following this tutorial.
Most of the functionality of a junifer marker has been taken care by the Most of the functionality of a ``junifer`` Marker has been taken care by the
:class:`.BaseMarker` class. Thus, only a few methods are required: :class:`.BaseMarker` class. Thus, only a few methods are required:
#. ``get_valid_inputs``: The method to obtain the list of valid inputs for the #. ``get_valid_inputs``: The method to obtain the list of valid inputs for the
marker. This is used to check that the inputs provided by the user are Marker. This is used to check that the inputs provided by the user are
valid. This method should return a list of strings, representing valid. This method should return a list of strings, representing
:ref:`data types <data_types>`. :ref:`data types <data_types>`.
#. ``get_output_type``: The method to obtain the kind of output of the marker. #. ``get_output_type``: The method to obtain the output type of the Marker.
This is used to check that the output of the marker is compatible with the This is used to check that the output of the Marker is compatible with the
storage. This method should return a string, representing storage. This method should return a string, representing
:ref:`storage types <storage_types>`. :ref:`storage types <storage_types>`.
#. ``compute``: The method that given the data, computes the marker. #. ``compute``: The method that given the data, computes the Marker.
#. ``__init__``: The initialisation method, where the marker is configured. #. ``__init__``: The initialisation method, where the Marker is configured.
As an example, we will develop a ``ParcelMean`` marker, a marker that first As an example, we will develop a ``ParcelMean`` Marker, a Marker that first
applies a parcellation and then computes the mean of the data in each parcel. applies a parcellation and then computes the mean of the data in each parcel.
This is a very simple example, but it will show you how to create a new marker. This is a very simple example, but it will show you how to create a new Marker.
.. _extending_markers_input_output: .. _extending_markers_input_output:
Step 1: Configure input and output Step 1: Configure input and output
---------------------------------- ----------------------------------
This step is quite simple: we need to define the input and output of the marker. This step is quite simple: we need to define the input and output of the Marker.
Based on the current :ref:`data types <data_types>`, we can have ``BOLD``, Based on the current :ref:`data types <data_types>`, we can have ``BOLD``,
``VBM_WM`` and ``VBM_GM`` as valid inputs. ``VBM_WM`` and ``VBM_GM`` as valid inputs.
@ -42,32 +42,32 @@ Based on the current :ref:`data types <data_types>`, we can have ``BOLD``,
def get_valid_inputs(self) -> list[str]: def get_valid_inputs(self) -> list[str]:
return ["BOLD", "VBM_WM", "VBM_GM"] return ["BOLD", "VBM_WM", "VBM_GM"]
The output of the marker depends on the input. For ``BOLD``, it will be The output of the Marker depends on the input. For ``BOLD``, it will be
``timeseries``, while for the rest of the inputs, it will be ``vector``. Thus, ``timeseries``, while for the rest of the inputs, it will be ``vector``. Thus,
we can define the output as: we can define the output as:
.. code-block:: python .. code-block:: python
def get_output_type(self, input_kind: str) -> str: def get_output_type(self, input_type: str) -> str:
if input_kind == "BOLD": if input_type == "BOLD":
return "timeseries" return "timeseries"
else: else:
return "vector" return "vector"
.. _extending_markers_init: .. _extending_markers_init:
Step 2: Initialize the marker Step 2: Initialise the Marker
----------------------------- -----------------------------
In this step we need to define the parameters of the marker the user can provide In this step we need to define the parameters of the Marker the user can provide
to configure how the marker will behave. to configure how the Marker will behave.
The parameters of the marker are defined in the ``__init__`` method. The The parameters of the Marker are defined in the ``__init__`` method. The
:class:`.BaseMarker` class requires two optional parameters: :class:`.BaseMarker` class requires two optional parameters:
1. ``name``: the name of the marker. This is used to identify the marker in the 1. ``name``: the name of the Marker. This is used to identify the Marker in the
configuration file. configuration file.
2. ``on``: a list or string with the data types that the marker will be applied 2. ``on``: a list or string with the data types that the Marker will be applied
to. to.
.. attention:: .. attention::
@ -92,29 +92,29 @@ parcellation to use. Thus, we can define the ``__init__`` method as follows:
.. caution:: .. caution::
Parameters of the marker must be stored as object attributes without using Parameters of the Marker must be stored as object attributes without using
``_`` as prefix. This is because any attribute that starts with ``_`` will ``_`` as prefix. This is because any attribute that starts with ``_`` will
not be considered as a parameter and not stored as part of the metadata of not be considered as a parameter and not stored as part of the metadata of
the marker. the Marker.
.. _extending_markers_compute: .. _extending_markers_compute:
Step 3: Compute the marker Step 3: Compute the Marker
-------------------------- --------------------------
In this step, we will define the method that computes the marker. This method In this step, we will define the method that computes the Marker. This method
will be called by junifer when needed, using the data provided by the will be called by ``junifer`` when needed, using the data provided by the
datagrabber, as configured by the user. The method ``compute`` has two DataGrabber, as configured by the user. The method ``compute`` has two
arguments: arguments:
* ``input``: a dictionary with the data to be used to compute the marker. This * ``input``: a dictionary with the data to be used to compute the Marker. This
will be the corresponding element in the :ref:`Data Object<data_object>` will be the corresponding element in the :ref:`Data Object<data_object>`
already indexed. Thus, the dictionary has at least two keys: ``data`` and already indexed. Thus, the dictionary has at least two keys: ``data`` and
``path``. The first one contains the data, while the second one contains the ``path``. The first one contains the data, while the second one contains the
path to the data. The dictionary can also contain other keys, depending on the path to the data. The dictionary can also contain other keys, depending on the
data type. data type.
* ``extra_input``: the rest of the :ref:`Data Object<data_object>`. This is * ``extra_input``: the rest of the :ref:`Data Object<data_object>`. This is
useful if you want to use other data to compute the marker useful if you want to use other data to compute the Marker
(e.g.: ``BOLD_confounds`` can be used to de-confound the ``BOLD`` data). (e.g.: ``BOLD_confounds`` can be used to de-confound the ``BOLD`` data).
Following the example, we will compute the mean of the data in each parcel using Following the example, we will compute the mean of the data in each parcel using
@ -133,7 +133,7 @@ the ``store`` method.
from typing import Any from typing import Any
from junifer.data import load_parcellation from junifer.data import get_parcellation
from nilearn.maskers import NiftiLabelsMasker from nilearn.maskers import NiftiLabelsMasker
@ -145,13 +145,11 @@ the ``store`` method.
# Get the data # Get the data
data = input["data"] data = input["data"]
# Get the min of the voxels sizes and use it as the resolution # Get the parcellation tailored for the target
resolution = np.min(data.header.get_zooms()[:3]) t_parcellation, t_labels, _ = get_parcellation(
# Load the parcellation
t_parcellation, t_labels, _ = load_parcellation(
name=self.parcellation_name, name=self.parcellation_name,
resolution=resolution, target_data=input,
extra_input=extra_input,
) )
# Create a masker # Create a masker
@ -173,32 +171,33 @@ the ``store`` method.
.. _extending_markers_finalize: .. _extending_markers_finalize:
Step 4: Finalise the marker Step 4: Finalise the Marker
--------------------------- ---------------------------
Once all of the above steps are done, we just need to give our marker a name, Once all of the above steps are done, we just need to give our Marker a name,
state its *dependencies* and register it using the ``@register_marker`` state its *dependencies* and register it using the ``@register_marker``
decorator. decorator.
The *dependencies* are the core packages that are required to compute the marker. The :ref:`dependencies <specifying_dependencies>` are the core packages that are
This will be later used to keep track of the versions of the packages used to required to compute the Marker. This will be later used to keep track of the
compute the marker. To inform junifer about the dependencies of a marker, we need versions of the packages used to compute the Marker. To inform ``junifer``
to define a ``_DEPENDENCIES`` attribute in the class. This attribute must be a about the dependencies of a Marker, we need to define a ``_DEPENDENCIES``
set, with the names of the packages as strings. For example, the ``ParcelMean`` attribute in the class. This attribute must be a set, with the names of the
marker has the following dependencies: packages as strings. For example, the ``ParcelMean`` marker has the
following dependencies:
.. code-block:: python .. code-block:: python
_DEPENDENCIES = {"nilearn", "numpy"} _DEPENDENCIES = {"nilearn", "numpy"}
Finally, we need to register the marker using the ``@register_marker`` decorator. Finally, we need to register the Marker using the ``@register_marker`` decorator.
.. code-block:: python .. code-block:: python
from typing import Any from typing import Any
from junifer.api.decorators import register_marker from junifer.api.decorators import register_marker
from junifer.data import load_parcellation from junifer.data import get_parcellation
from junifer.markers.base import BaseMarker from junifer.markers.base import BaseMarker
from nilearn.maskers import NiftiLabelsMasker from nilearn.maskers import NiftiLabelsMasker
@ -220,8 +219,8 @@ Finally, we need to register the marker using the ``@register_marker`` decorator
def get_valid_inputs(self) -> list[str]: def get_valid_inputs(self) -> list[str]:
return ["BOLD", "VBM_WM", "VBM_GM"] return ["BOLD", "VBM_WM", "VBM_GM"]
def get_output_type(self, input_kind: str) -> str: def get_output_type(self, input_type: str) -> str:
if input_kind == "BOLD": if input_type == "BOLD":
return "timeseries" return "timeseries"
else: else:
return "vector" return "vector"
@ -234,13 +233,11 @@ Finally, we need to register the marker using the ``@register_marker`` decorator
# Get the data # Get the data
data = input["data"] data = input["data"]
# Get the min of the voxels sizes and use it as the resolution # Get the parcellation tailored for the target
resolution = np.min(data.header.get_zooms()[:3]) t_parcellation, t_labels, _ = get_parcellation(
# Load the parcellation
t_parcellation, t_labels, _ = load_parcellation(
name=self.parcellation_name, name=self.parcellation_name,
resolution=resolution, target_data=input,
extra_input=extra_input,
) )
# Create a masker # Create a masker
@ -283,8 +280,8 @@ Template for a custom Marker
valid = [] valid = []
return valid return valid
def get_output_type(self, input_kind): def get_output_type(self, input_type):
# TODO: Return the valid output kind for each input kind # TODO: Return the valid output type for each input type
pass pass
def compute(self, input, extra_input): def compute(self, input, extra_input):

View file

@ -5,7 +5,7 @@
Adding Masks Adding Masks
============ ============
Many processing steps and markers in junifer allow you to specify a binary Many processing steps and Markers in ``junifer`` allow you to specify a binary
mask to select voxels you want to include in the analysis. There are a number mask to select voxels you want to include in the analysis. There are a number
of masks :ref:`in-built in junifer already <builtin>`, so check if any of them of masks :ref:`in-built in junifer already <builtin>`, so check if any of them
suit your needs. Check how to use these masks :ref:`here <using_masks>`. Once suit your needs. Check how to use these masks :ref:`here <using_masks>`. Once
@ -14,23 +14,23 @@ suit your needs, and you have found that they don't, you can come back here to
learn how to use your own masks. learn how to use your own masks.
The principle is fairly simple and quite similar to :ref:`adding_parcellations` The principle is fairly simple and quite similar to :ref:`adding_parcellations`
and :ref:`adding_coordinates`. junifer provides a :func:`.register_mask` and :ref:`adding_coordinates`. ``junifer`` provides a :func:`.register_mask`
function that lets you register your own custom masks. It consists of three function that lets you register your own custom masks. It consists of three
positional arguments (``name``, ``mask_path`` and ``space``) and one optional positional arguments (``name``, ``mask_path`` and ``space``) and one optional
keyword argument (``overwrite``). keyword argument (``overwrite``).
The ``name`` argument is a string indicating the name of the mask. This name The ``name`` argument is a string indicating the name of the mask. This name
is used to refer to that mask in junifer internally in order to obtain the is used to refer to that mask in ``junifer`` internally in order to obtain the
actual mask data and perform operations on it. For example, using the name you actual mask data and perform operations on it. For example, using the name you
can load a mask after registration using the can load a mask after registration using the
:func:`.load_mask` function. :func:`.load_mask` function.
The ``mask_path`` should contain the path to a valid NIfTI image with binary The ``mask_path`` should contain the path to a valid NIfTI image with binary
voxel values (i.e. 0 or 1). This data can then be used by junifer to mask other voxel values (i.e. 0 or 1). This data can then be used by ``junifer`` to mask
MR images. other MR images.
Lastly, we specify the ``space`` that the coordinates are in, for example, Lastly, we specify the ``space`` that the coordinates are in, for example,
``"MNI"`` or ``"Native"`` (scanner-native space). ``"MNI152NLin6Asym"`` or ``"native"`` (scanner-native space).
Step 1: Prepare code to register a mask Step 1: Prepare code to register a mask
--------------------------------------- ---------------------------------------
@ -49,7 +49,7 @@ look as follows:
# on your system: # on your system:
mask_path = Path("..") / ".." / "my_custom_mask.nii.gz" mask_path = Path("..") / ".." / "my_custom_mask.nii.gz"
register_mask(name="my_custom_mask", mask_path=mask_path, space="Native") register_mask(name="my_custom_mask", mask_path=mask_path, space="native")
Simple, right? Now we just have to configure a YAML file to register this mask Simple, right? Now we just have to configure a YAML file to register this mask
so we can use it for :ref:`codeless configuration of junifer <codeless>`. so we can use it for :ref:`codeless configuration of junifer <codeless>`.
@ -57,14 +57,14 @@ so we can use it for :ref:`codeless configuration of junifer <codeless>`.
Step 2: Configure a YAML file for registration of a mask Step 2: Configure a YAML file for registration of a mask
-------------------------------------------------------- --------------------------------------------------------
In order to do this, we can use the ``with`` keyword provided by junifer: In order to do this, we can use the ``with`` keyword provided by ``junifer``:
.. code-block:: yaml .. code-block:: yaml
with: with:
- register_custom_mask.py - register_custom_mask.py
Then we can use this mask for any processing step or marker that takes in a Then we can use this mask for any processing step or Marker that takes in a
mask as an argument. For example: mask as an argument. For example:
.. code-block:: yaml .. code-block:: yaml
@ -76,12 +76,15 @@ mask as an argument. For example:
method: mean method: mean
masks: "my_custom_mask" masks: "my_custom_mask"
Now, you can simply use this YAML file to run your pipeline. One important Now, you can simply use this YAML file to run your pipeline.
point to keep in mind is that if the paths given in ``register_custom_mask.py``
are relative paths, they will be interpreted by junifer as relative to the .. important::
jobs directory (i.e. where junifer will create submit files, logs directory and
so on). For simplicity, you may just want to use absolute paths to avoid It's important to keep in mind that if the paths given in
confusion, yet using relative paths is likely a better way to make your ``register_custom_mask.py`` are relative paths, they will be interpreted
pipeline directory/repository more portable and therefore more reproducible for by junifer as relative to the jobs directory (i.e. where ``junifer`` will
others. Really, once you understand how these paths are interpreted by junifer, create submit files, logs directory and so on). For simplicity, you may just
it is quite easy. want to use absolute paths to avoid confusion, yet using relative paths is
likely a better way to make your pipeline directory / repository more portable
and therefore more reproducible for others. Really, once you understand how
paths are interpreted by ``junifer``, it is quite easy.

View file

@ -5,28 +5,29 @@
Adding Parcellations Adding Parcellations
==================== ====================
Before you start adding your own parcellations, check whether junifer has Before you start adding your own parcellations, check whether ``junifer`` has
the parcellation :ref:`in-built already <builtin>`. Perhaps, what is available the parcellation :ref:`in-built already <builtin>`. Perhaps, what is available
there will suffice to achieve your goals. However, of course junifer will not there will suffice to achieve your goals. However, of course ``junifer`` will not
have every parcellation available that you may want to use, and if so, it will have every parcellation available that you may want to use, and if so, it will
be nice to be able to add it yourself using a format that junifer understands. be nice to be able to add it yourself using a format that ``junifer`` understands.
Similarly, you may even be interested in creating your own custom parcellations Similarly, you may even be interested in creating your own custom parcellations
and then adding them to junifer, so you can use junifer to obtain different and then adding them to ``junifer``, so you can use ``junifer`` to obtain
markers to assess and validate your own parcellation. So, how can you do this? different Markers to assess and validate your own parcellation. So, how can you do
this?
Since both of these use-cases are quite common, and not being able to use your Since both of these use-cases are quite common, and not being able to use your
favourite parcellation is of course quite a buzzkill, junifer actually provides favourite parcellation is of course quite a buzzkill, ``junifer`` actually
the easy-to-use :func:`.register_parcellation` function to do just that. Let's provides the easy-to-use :func:`.register_parcellation` function to do just that.
try to understand the API reference and then use this function to register our Let's try to understand the API reference and then use this function to register our
own parcellation. own parcellation.
From the API reference, we can see that it has 4 positional arguments From the API reference, we can see that it has 4 positional arguments
(``name``, ``parcellation_path``, ``parcels_labels`` and ``space``) as well as (``name``, ``parcellation_path``, ``parcels_labels`` and ``space``) as well as
one optional keyword argument (``overwrite``). one optional keyword argument (``overwrite``).
The ``name`` of the parcellation is up to you and will be the name that junifer The ``name`` of the parcellation is up to you and will be the name that
will use to refer to this particular parcellation. You can think of this as ``junifer`` will use to refer to this particular parcellation. You can think of
being similar to a key in a python dictionary, i.e. a key that is used to this as being similar to a key in a Python dictionary, i.e. a key that is used to
obtain and operate on the actual parcellation data. This ``name`` must always obtain and operate on the actual parcellation data. This ``name`` must always
be a string. For example, we could call our parcellation be a string. For example, we could call our parcellation
``"my_custom_parcellation"`` (Note, that in a real-world use case this is ``"my_custom_parcellation"`` (Note, that in a real-world use case this is
@ -52,14 +53,14 @@ first label in this list corresponds to the first integer label in the
parcellation and so on). parcellation and so on).
LeSasse commented 2024-03-28 13:04:36 +00:00 (Migrated from github.com)

"For example, a simple example could look like this:" -> "A simple example could look like this:"

"For example, a simple example could look like this:" -> "A simple example could look like this:"
Lastly, we specify the ``space`` that the parcellation is in, for example, Lastly, we specify the ``space`` that the parcellation is in, for example,
``"MNI"`` or ``"Native"`` (scanner-native space). ``"MNI152NLin2009cAsym"`` or ``"native"`` (scanner-native space).
Step 1: Prepare code to register a parcellation Step 1: Prepare code to register a parcellation
----------------------------------------------- -----------------------------------------------
Now we know everything that we need to know to make sure junifer can use our Now we know everything that we need to know to make sure ``junifer`` can use our
own parcellation to compute any parcellation-based marker. For example, own parcellation to compute any parcellation-based Marker. A simple example could
a simple example could look like this: look like this:
.. code-block:: python .. code-block:: python
@ -82,20 +83,20 @@ a simple example could look like this:
name="my_custom_parcellation", name="my_custom_parcellation",
parcellation_path=path_to_parcellation, parcellation_path=path_to_parcellation,
parcels_labels=my_labels, parcels_labels=my_labels,
space="MNI" space="MNI152NLin2009cAsym"
) )
We can run this code and it seems to work, however, how can we actually We can run this code and it seems to work, however, how can we actually
include the custom parcellation in a junifer pipeline using a include the custom parcellation in a ``junifer`` pipeline using a
:ref:`code-less YAML configuration <codeless>`? :ref:`code-less YAML configuration <codeless>`?
Step 2: Add parcellation registration to the YAML file Step 2: Add parcellation registration to the YAML file
------------------------------------------------------ ------------------------------------------------------
In order to use the parcellation in a junifer pipeline configured by a YAML In order to use the parcellation in a ``junifer`` pipeline configured by a YAML
file, we can save the above code in a python file, say file, we can save the above code in a Python file, say
``registering_my_parcellation.py``. We can then simply add this file using the ``registering_my_parcellation.py``. We can then simply add this file using the
``with`` keyword provided by junifer: ``with`` keyword provided by ``junifer``:
.. code-block:: yaml .. code-block:: yaml
@ -105,7 +106,7 @@ file, we can save the above code in a python file, say
Afterwards continue configuring the rest of the pipeline in this YAML file, and Afterwards continue configuring the rest of the pipeline in this YAML file, and
you will be able to use this parcellation using the name you gave the you will be able to use this parcellation using the name you gave the
parcellation when registering it. For example, we can add a parcellation when registering it. For example, we can add a
:class:`.ParcelAggregation` marker to demonstrate how this can be done: :class:`.ParcelAggregation` Marker to demonstrate how this can be done:
.. code-block:: yaml .. code-block:: yaml
@ -115,12 +116,15 @@ parcellation when registering it. For example, we can add a
parcellation: my_custom_parcellation parcellation: my_custom_parcellation
method: mean method: mean
Now, you can simply use this YAML file to run your pipeline. One important Now, you can simply use this YAML file to run your pipeline.
point to keep in mind is that if the paths given in
``registering_my_parcellation.py`` are relative paths, they will be interpreted .. important::
by junifer as relative to the jobs directory (i.e. where junifer will create
submit files, logs directory and so on). For simplicity, you may just want to It's important to keep in mind that if the paths given in
use absolute paths to avoid confusion, yet using relative paths is likely a ``registering_my_parcellation.py`` are relative paths, they will be interpreted
better way to make your pipeline directory/repository more portable and by ``junifer`` as relative to the jobs directory (i.e. where ``junifer`` will
therefore more reproducible for others. Really, once you understand how these create submit files, logs directory and so on). For simplicity, you may just
paths are interpreted by junifer, it is quite easy. want to use absolute paths to avoid confusion, yet using relative paths is
likely a better way to make your pipeline directory / repository more portable
and therefore more reproducible for others. Really, once you understand how
these paths are interpreted by ``junifer``, it is quite easy.

View file

@ -0,0 +1,246 @@
.. include:: ../links.inc
.. _extending_preprocessors:
Creating Preprocessors
======================
As already mentioned in the introduction, ``junifer`` does not do traditional
MRI pre-processing but can perform minimal preprocessing of the data that the
DataGrabber provides, for example, smoothing after confound regression or
transforming data to subject-native space before feature extraction. While
there are a few Preprocessors available already and we are constantly adding
new ones, you might need something specific and then you can create your
own Preprocessor.
While implementing your own Preprocessor, you need to always inherit from
:class:`.BasePreprocessor` and implement a few methods:
#. ``get_valid_inputs``: This method should return a list of strings
representing the valid data types that the Preprocessor can work on.
Check :ref:`data types <data_types>` for reference.
#. ``get_output_type``: This method should just return the input as it
is unused as of now.
#. ``preprocess``: The method that given the data, preprocesses the data.
#. ``__init__``: The initialisation method, where the Preprocessor is
configured.
As an example, we will develop a ``NilearnSmoothing`` Preprocessor, which
smoothens the data using :func:`nilearn.image.smooth_img`. This is often
desirable in cases where your data is preprocessed using ``fMRIPrep``, as
``fMRIPrep`` does not perform smoothing.
.. _extending_preprocessors_input_output:
Step 1: Configure input and output
----------------------------------
In this step, we define the input and output data types of the Preprocessor.
For input we can accept ``T1w``, ``T2w`` and ``BOLD``
:ref:`data types <data_types>`.
.. code-block:: python
...
def get_valid_inputs(self) -> list[str]:
return ["T1w", "T2w", "BOLD"]
...
The output definition of the Preprocessor is unused now but is kept for
completeness.
.. code-block:: python
...
def get_output_type(self, input_type: str) -> str:
return input_type
...
.. _extending_preprocessors_init:
Step 2: Initialise the Preprocessor
-----------------------------------
Now we need to define our Preprocessor class' constructor which is also how
you configure it. Our class will have the following arguments:
1. ``fwhm``: The smoothing strength as a full-width at half maximum
(in millimetres). Since we depend on :func:`nilearn.image.smooth_img`, we
pass the value to it.
2. ``on``: The data type we want the Preprocessor to work on. If the user does
not specify, it will work on all the data types given by the
``get_valid_inputs`` function.
.. attention::
Only basic types (*int*, *bool* and *str*), lists, tuples and dictionaries
are allowed as parameters. This is because the parameters are stored in
JSON format, and JSON only supports these types.
.. code-block:: python
from typing import Literal
from numpy.typing import ArrayLike
...
def __init__(
self,
fwhm: int | float | ArrayLike | Literal["fast"] | None,
on: str | list[str] | None = None,
) -> None:
self.fwhm = fwhm
super().__init__(on=on)
...
.. caution::
Parameters of the Preprocessor must be stored as object attributes without
using ``_`` as prefix. This is because any attribute that starts with ``_``
will not be considered as a parameter and not stored as part of the metadata
of the Preprocessor.
.. _extending_preprocessors_preprocess:
Step 3: Preprocess the data
---------------------------
Finally, we will write the actual logic of the Preprocessor. This method will
be called by ``junifer`` when needed, using the data provided by the
DataGrabber, as configured by the user. The method ``preprocess`` has two
arguments:
* ``input``: A dictionary with the data to be used by the Preprocessor. This
will be the corresponding element in the :ref:`Data Object<data_object>`
already indexed. Thus, the dictionary has at least two keys: ``data`` and
``path``. The first one contains the data, while the second one contains the
path to the data. The dictionary can also contain other keys, depending on the
data type.
* ``extra_input``: The rest of the :ref:`Data Object<data_object>`. This is
useful if you want to use other data (e.g., ``Warp`` can be used to provide
the transformation matrix file for transformation to subject-native space).
and it has two return values:
* First is the ``input`` dictionary with necessary data modified. Usually, you
want to replace the ``input["data"]`` with the preprocessed data.
* Second is a dictionary just like ``input`` or ``extra_input`` but with only
specific key-value pairs which you would like to pass down to the Markers.
For example, if your Preprocessor computes some mask with the preprocessed
data, you could pass it through this which would be added and available
in the Marker step with the same key you pass here. Usually, you would
want to pass ``None``.
.. code-block:: python
from typing import Any
from nilearn import image as nimg
...
def preprocess(
self,
input: dict[str, Any],
extra_input: dict[str, Any] | None = None,
) -> tuple[dict[str, Any], dict[str, Any] | None]:
input["data"] = nimg.smooth_img(imgs=input["data"], fwhm=self.fwhm)
return input, None
...
Step 4: Finalise the Preprocessor
LeSasse commented 2024-03-28 13:07:17 +00:00 (Migrated from github.com)

I believe in multiple places I saw the american spelling of things, so this should be Finalize. Can we record in some central place i.e. something like "How to contribute to docs" that we aim to use the american spelling?

I believe in multiple places I saw the american spelling of things, so this should be Finalize. Can we record in some central place i.e. something like "How to contribute to docs" that we aim to use the american spelling?
synchon commented 2024-03-28 13:53:18 +00:00 (Migrated from github.com)

I've made everything follow British English in the docs. Recording it is a good idea.

I've made everything follow British English in the docs. Recording it is a good idea.
LeSasse commented 2024-03-28 13:54:34 +00:00 (Migrated from github.com)

Ok, i will point out the american way when i find them.

Ok, i will point out the american way when i find them.
---------------------------------
Now we just need to combine everything we have above and throw in a couple of
other stuff to get our Preprocessor ready.
First, we specify the :ref:`dependencies <specifying_dependencies>` for our
class, which are basically the packages that are required by the class. This is
used for validation before running to ensure all the packages are installed and
also to keep track of the dependencies and their versions in the metadata. We
define it using a class attribute like so:
.. code-block:: python
_DEPENDENCIES = {"nilearn"}
Then, we just need to register the Preprocessor using ``@register_preprocessor``
decorator and our final code should look like this:
.. code-block:: python
from typing import Any, Literal
from junifer.api.decorators import register_preprocessor
from junifer.preprocess import BasePreprocessor
from nilearn import image as nimg
from numpy.typing import ArrayLike
@register_preprocessor
class NilearnSmoothing(BasePreprocessor):
_DEPENDENCIES = {"nilearn"}
def __init__(
self,
fwhm: int | float | ArrayLike | Literal["fast"] | None,
on: str | list[str] | None = None,
) -> None:
self.fwhm = fwhm
super().__init__(on=on)
def get_valid_inputs(self) -> list[str]:
return ["T1w", "T2w", "BOLD"]
def get_output_type(self, input_type: str) -> str:
return input_type
def preprocess(
self,
input: dict[str, Any],
extra_input: dict[str, Any] | None = None,
) -> tuple[dict[str, Any], dict[str, Any] | None]:
input["data"] = nimg.smooth_img(imgs=input["data"], fwhm=self.fwhm)
return input, None
.. _extending_preprocessors_template:
Template for a custom Preprocessor
----------------------------------
.. code-block:: python
from junifer.api.decorators import register_preprocessor
from junifer.preprocess import BasePreprocessor
@register_preprocessor
class TemplatePreprocessor(BasePreprocessor):
def __init__(self, on=None):
# TODO: add preprocessor-specific parameters
super().__init__(on=on)
def get_valid_inputs(self):
# TODO: Complete with the valid inputs
valid = []
return valid
def get_output_type(self, input_type):
return input_type
def preprocess(self, input, extra_input):
# TODO: add the preprocessor logic
return input, None

View file

@ -1,35 +1,13 @@
.. include:: links.inc .. include:: links.inc
Installing junifer Installing ``junifer``
================== ======================
Requirements
------------
junifer is compatible with `Python`_ >= 3.8 and requires the following packages:
* ``click>=8.1.3,<8.2``
* ``numpy>=1.24,<1.27``
* ``datalad>=0.15.4,<0.20``
* ``pandas>=1.4.0,<2.2``
* ``nibabel>=3.2.0,<5.11``
* ``nilearn>=0.9.0,<=0.11.0``
* ``sqlalchemy>=1.4.27,<= 2.1.0``
* ``ruamel.yaml>=0.17,<0.18``
* ``h5py>=3.8.0,<3.10``
Depending on the installation method, these packages might be installed
automatically.
Installation
------------
Depending on your use-case, ``junifer`` can be installed differently: Depending on your use-case, ``junifer`` can be installed differently:
* Install the :ref:`install_latest_release`. This is the most suitable approach * Install the :ref:`latest stable release <install_latest_release>`. This is the most suitable approach
for end users. for end users.
* Install from :ref:`install_development_git`. This is the most suitable approach * Install from :ref:`latest development release <install_development_git>`. This is the most suitable approach
for developers. for developers.
@ -39,8 +17,8 @@ Either way, we strongly recommend using
.. _install_latest_release: .. _install_latest_release:
Stable release Using a package manager
~~~~~~~~~~~~~~ -----------------------
Use ``pip`` to install ``junifer`` from `PyPI <https://pypi.org>`_, like so: Use ``pip`` to install ``junifer`` from `PyPI <https://pypi.org>`_, like so:
@ -64,8 +42,8 @@ You can also install via ``conda``, like so:
.. _install_development_git: .. _install_development_git:
Local Git repository From the source
~~~~~~~~~~~~~~~~~~~~ ---------------
Follow the `detailed contribution guidelines <contribution.rst>`_. Follow the `detailed contribution guidelines <contribution.rst>`_.
@ -81,11 +59,11 @@ that are required for specific markers.
.. important:: .. important::
The Docker container wrappers add the commands required by junifer. Using The Docker container wrappers add the commands required by ``junifer``. Using
these commands have some limitations, mostly related to handling files and these commands have some limitations, mostly related to handling files and
paths. Junifer knows about this and uses these commands in the proper way. paths. ``junifer`` knows about this and uses these commands in the proper way.
Keep this in mind if you try to use the Docker wrappers outside of Keep this in mind if you try to use the Docker wrappers outside of
junifer. These caveats and limitations are not documented. ``junifer``. These caveats and limitations are not documented.
AFNI AFNI
---- ----

View file

@ -23,6 +23,7 @@
.. _`nilearn`: https://nilearn.github.io .. _`nilearn`: https://nilearn.github.io
.. _`nipype`: https://nipype.readthedocs.io .. _`nipype`: https://nipype.readthedocs.io
.. _`datalad`: https://datalad.org .. _`datalad`: https://datalad.org
.. _`templateflow`: https://www.templateflow.org
.. _`venv`: https://docs.python.org/3/tutorial/venv.html .. _`venv`: https://docs.python.org/3/tutorial/venv.html
.. _`conda env`: https://docs.conda.io/projects/conda/en/latest/user-guide/tasks/manage-environments.html .. _`conda env`: https://docs.conda.io/projects/conda/en/latest/user-guide/tasks/manage-environments.html

View file

@ -2,8 +2,8 @@
.. _starting: .. _starting:
First steps with junifer First steps with ``junifer``
======================== ============================
.. note:: .. note::
@ -13,7 +13,7 @@ First steps with junifer
flowchart TD flowchart TD
start((( Start ))) start((( Start )))
read_understanding(Read Understanding Junifer) read_understanding(Read Understanding junifer)
start --> read_understanding start --> read_understanding
question_features{Can I compute\nthe features I want\nusing junifer?} question_features{Can I compute\nthe features I want\nusing junifer?}
read_understanding --> question_features read_understanding --> question_features
@ -22,7 +22,7 @@ First steps with junifer
question_features -->|No| question_feature_type question_features -->|No| question_feature_type
read_using --> question_features read_using --> question_features
read_using(Read Using Junifer) read_using(Read Using junifer)
question_feature_type{"What am I missing\nfrom junifer"} question_feature_type{"What am I missing\nfrom junifer"}
question_feature_type --> missing_datagrabber question_feature_type --> missing_datagrabber
@ -30,7 +30,7 @@ First steps with junifer
question_feature_type --> missing_marker question_feature_type --> missing_marker
question_feature_type --> missing_other question_feature_type --> missing_other
missing_datagrabber("A Dataset/DataGrabber") missing_datagrabber("A Dataset/DataGrabber")
missing_preprocessing("A Preprocessing") missing_preprocessing("A Preprocessor")
missing_marker("A Marker") missing_marker("A Marker")
missing_other("Something else") missing_other("Something else")
@ -38,7 +38,7 @@ First steps with junifer
question_datagrabber_junifarm{Is the\ndataset/datagrabber\nin juni-farm?} question_datagrabber_junifarm{Is the\ndataset/datagrabber\nin juni-farm?}
question_datagrabber_junifarm -->|Yes| read_using_final question_datagrabber_junifarm -->|Yes| read_using_final
question_datagrabber_junifarm -->|No| read_extending_datagrabber_start question_datagrabber_junifarm -->|No| read_extending_datagrabber_start
read_extending_datagrabber_start(Read Creating a Junifer extension) read_extending_datagrabber_start(Read Creating a junifer extension)
read_extending_datagrabber_start --> read_extending_datagrabber read_extending_datagrabber_start --> read_extending_datagrabber
read_extending_datagrabber(Read Creating Data Grabbers) read_extending_datagrabber(Read Creating Data Grabbers)
read_extending_datagrabber --> question_datagrabber_kind read_extending_datagrabber --> question_datagrabber_kind
@ -55,14 +55,14 @@ First steps with junifer
question_contribute_datagrabber{Do you think\nyour DataGrabber\nis useful for other users?} question_contribute_datagrabber{Do you think\nyour DataGrabber\nis useful for other users?}
question_contribute_datagrabber -->|Yes| contribute_datagrabber question_contribute_datagrabber -->|Yes| contribute_datagrabber
question_contribute_datagrabber -->|No| final_run question_contribute_datagrabber -->|No| final_run
contribute_datagrabber(Create a\nDATASET REQUEST\nissue on Github) contribute_datagrabber(Create a\nDATASET REQUEST\nissue on GitHub)
contribute_datagrabber --> final_run contribute_datagrabber --> final_run
missing_marker --> question_marker_junifarm missing_marker --> question_marker_junifarm
question_marker_junifarm{Is the marker\nin juni-farm?} question_marker_junifarm{Is the marker\nin juni-farm?}
question_marker_junifarm -->|Yes| read_using_final question_marker_junifarm -->|Yes| read_using_final
question_marker_junifarm -->|No| read_extending_marker_start question_marker_junifarm -->|No| read_extending_marker_start
read_extending_marker_start(Read Creating a Junifer extension) read_extending_marker_start(Read Creating a junifer extension)
read_extending_marker_start --> read_extending_marker read_extending_marker_start --> read_extending_marker
read_extending_marker(Read Creating Markers) read_extending_marker(Read Creating Markers)
read_extending_marker --> question_marker_solved read_extending_marker --> question_marker_solved
@ -72,7 +72,7 @@ First steps with junifer
question_contribute_marker{Do you think\nyour Marker\nis useful for other users?} question_contribute_marker{Do you think\nyour Marker\nis useful for other users?}
question_contribute_marker -->|Yes| contribute_marker question_contribute_marker -->|Yes| contribute_marker
question_contribute_marker -->|No| final_run question_contribute_marker -->|No| final_run
contribute_marker(Create a\nMARKER REQUEST\nissue on Github) contribute_marker(Create a\nMARKER REQUEST\nissue on GitHub)
contribute_marker --> final_run contribute_marker --> final_run
missing_preprocessing --> contact_help missing_preprocessing --> contact_help
@ -87,22 +87,22 @@ First steps with junifer
missing_other --> missing_other_other missing_other --> missing_other_other
missing_other_other --> contact_help missing_other_other --> contact_help
contact_help(((Contact the\nJunifer team))) contact_help(((Contact the\njunifer team)))
missing_mask --> read_adding_mask_start missing_mask --> read_adding_mask_start
read_adding_mask_start("Read Creating a Junifer extension") read_adding_mask_start("Read Creating a junifer extension")
read_adding_mask_start --> read_adding_mask read_adding_mask_start --> read_adding_mask
read_adding_mask("Read Adding Masks") read_adding_mask("Read Adding Masks")
read_adding_mask --> missing_other_solved read_adding_mask --> missing_other_solved
missing_parcellation --> read_adding_parcellation_start missing_parcellation --> read_adding_parcellation_start
read_adding_parcellation_start("Read Creating a Junifer extension") read_adding_parcellation_start("Read Creating a junifer extension")
read_adding_parcellation_start --> read_adding_parcellation read_adding_parcellation_start --> read_adding_parcellation
read_adding_parcellation("Read Adding Parcellations") read_adding_parcellation("Read Adding Parcellations")
read_adding_parcellation --> missing_other_solved read_adding_parcellation --> missing_other_solved
missing_coordinates --> read_adding_coordinates_start missing_coordinates --> read_adding_coordinates_start
read_adding_coordinates_start("Read Creating a Junifer extension") read_adding_coordinates_start("Read Creating a junifer extension")
read_adding_coordinates_start --> read_adding_coordinates read_adding_coordinates_start --> read_adding_coordinates
read_adding_coordinates("Read Adding Coordinates") read_adding_coordinates("Read Adding Coordinates")
read_adding_coordinates --> missing_other_solved read_adding_coordinates --> missing_other_solved
@ -110,11 +110,11 @@ First steps with junifer
missing_other_solved{Did you solve your issue?} missing_other_solved{Did you solve your issue?}
missing_other_solved -->|Yes| read_using_final missing_other_solved -->|Yes| read_using_final
missing_other_solved -->|No| missing_other_contact missing_other_solved -->|No| missing_other_contact
missing_other_contact(Contact the\nJunifer team) missing_other_contact(Contact the\njunifer team)
missing_other_contact --> missing_other_issue missing_other_contact --> missing_other_issue
missing_other_issue(((Submit a\nFEATURE REQUEST\nissue in Github))) missing_other_issue(((Submit a\nFEATURE REQUEST\nissue in GitHub)))
read_using_final(Read Using Junifer) read_using_final(Read Using junifer)
read_using_final --> final_yaml read_using_final --> final_yaml
final_yaml(Create/edit the YAML file) final_yaml(Create/edit the YAML file)
final_yaml --> final_run final_yaml --> final_run
@ -125,9 +125,9 @@ First steps with junifer
question_error_run{"Is it an issue\nwith my YAML file?"} question_error_run{"Is it an issue\nwith my YAML file?"}
question_error_run -->|Yes| final_yaml question_error_run -->|Yes| final_yaml
question_error_run -->|No| error_contact question_error_run -->|No| error_contact
error_contact(Contact the\nJunifer team) error_contact(Contact the\njunifer team)
error_contact --> error_issue error_contact --> error_issue
error_issue(((Submit a\nBUG REPORT issue\nin Github))) error_issue(((Submit a\nBUG REPORT issue\nin GitHub)))
question_final_run_worked -->|Yes| final_queue question_final_run_worked -->|Yes| final_queue
final_queue(Use junifer queue to compute your features) final_queue(Use junifer queue to compute your features)
final_queue --> final_magic final_queue --> final_magic

View file

@ -105,7 +105,7 @@ Data Types
- Preprocessed or Raw T1w image - Preprocessed or Raw T1w image
* - ``BOLD`` * - ``BOLD``
- BOLD image (4D) - BOLD image (4D)
- Preprocessed/Denoised BOLD image (fmriprep output) - Preprocessed or Denoised BOLD image (fMRIPrep output)
* - ``BOLD_confounds`` * - ``BOLD_confounds``
- BOLD image confounds (CSV/TSV file) - BOLD image confounds (CSV/TSV file)
- Confounds that can be applied to the BOLD image. - Confounds that can be applied to the BOLD image.

View file

@ -9,9 +9,9 @@ Description
----------- -----------
The ``DataGrabber`` is an object that can provide an interface to datasets you The ``DataGrabber`` is an object that can provide an interface to datasets you
want to work with in junifer. Every concrete implementation of a DataGrabber is want to work with in ``junifer``. Every concrete implementation of a DataGrabber
aware of a particular dataset's structure and thus allows you to fetch specific is aware of a particular dataset's structure and thus allows you to fetch
elements of interest from the dataset. It adds the ``path`` key to each specific elements of interest from the dataset. It adds the ``path`` key to each
:ref:`data type <data_types>` in the :ref:`Data object <data_object>`. :ref:`data type <data_types>` in the :ref:`Data object <data_object>`.
DataGrabbers are intended to be used as context managers. When used within a DataGrabbers are intended to be used as context managers. When used within a
@ -20,18 +20,18 @@ the dataset, for example, downloading and cleaning up. As the interface
is consistent, you always use the same procedure to interact with the DataGrabber. is consistent, you always use the same procedure to interact with the DataGrabber.
For example, a concrete implementation of :class:`.DataladDataGrabber` can For example, a concrete implementation of :class:`.DataladDataGrabber` can
provide junifer with data from a Datalad dataset. Of course, DataGrabbers are not provide ``junifer`` with data from a Datalad dataset. Of course, DataGrabbers are
only meant to work with Datalad datasets but any dataset. not only meant to work with Datalad datasets but any dataset.
If you are interested in using already provided DataGrabbers, please go to If you are interested in using already provided DataGrabbers, please go to
:doc:`../builtin`. And, if you want to implement your own DataGrabber, you need :doc:`../builtin`. And, if you want to implement your own DataGrabber, you need
to provide concrete implementations of base classes already provided. to provide concrete implementations of abstract base classes already provided.
Base Classes Base Classes
------------ ------------
In this section, we showcase different abstract base classes you might want to In this section, we showcase different abstract and concrete base classes you
use to implement your own DataGrabber. might want to use to implement your own DataGrabber.
.. list-table:: .. list-table::
:widths: auto :widths: auto

View file

@ -9,17 +9,17 @@ Description
----------- -----------
The ``DataReader`` is an object that is responsible for actually reading data The ``DataReader`` is an object that is responsible for actually reading data
files in junifer. It reads the value of the key ``path`` for each files in ``junifer``. It reads the value of the key ``path`` for each
:ref:`data type <data_types>` in the :ref:`Data object <data_object>` and loads :ref:`data type <data_types>` in the :ref:`Data object <data_object>` and loads
them into memory. After reading the data into memory, it adds the key ``data`` them into memory. After reading the data into memory, it adds the key ``data``
to the same level as ``path`` and the value is the actual data in the memory. to the same level as ``path`` and the value is the actual data in the memory.
DataReaders are meant to be used inside the datagrabber context but you can DataReaders are meant to be used inside the DataGrabber context but you can
operate on them outside the context as long as the actual data is in the memory operate on them outside the context as long as the actual data is in the memory
and the Python runtime has not garbage-collected it. and the Python runtime has not garbage-collected it.
For data formats not supported by junifer yet, you can either make your own For data formats not supported by ``junifer`` yet, you can either make your own
*Data Reader* or open an issue on `junifer Github`_ and we can help you out. DataReader or open an issue on `junifer Github`_ and we can help you out.
File Formats File Formats
------------ ------------

View file

@ -2,14 +2,14 @@
.. _understanding: .. _understanding:
Understanding junifer Understanding ``junifer``
===================== =========================
Before you start, you should understand how junifer works. Junifer is a Before you start, you should understand how ``junifer`` works. ``junifer`` is a
tool conceived to extract features from neuroimaging data in an easy-to-use tool conceived to extract features from neuroimaging data in an easy-to-use
manner, with minimal coding and minimal user expertise in the internal aspects. manner, with minimal coding and minimal user expertise in the internal aspects.
Unlike other tools like FSL, SPM, AFNI, etc., junifer is not a toolbox to Unlike other tools like FSL, SPM, AFNI, etc., ``junifer`` is not a toolbox to
pre-process data, but a toolbox to extract features from previously pre-process data, but a toolbox to extract features from previously
pre-processed data. pre-processed data.
@ -20,9 +20,9 @@ julearn_).
.. important:: .. important::
Junifer is not a toolbox to create pipelines, but a tool to configure the ``junifer`` is not a toolbox to create pipelines, but a tool to configure the
junifer pipeline, which is intended to be fixed and not to be changed. If you ``junifer`` pipeline, which is intended to be fixed and not to be changed. If
want to create a pipeline, you should use other tools like nipype_. you want to create a pipeline, you should use other tools like nipype_.
.. toctree:: .. toctree::
:maxdepth: 2 :maxdepth: 2

View file

@ -21,12 +21,12 @@ the pipeline.
SPM, AFNI, etc., . For example, one can perform confound removal on loaded SPM, AFNI, etc., . For example, one can perform confound removal on loaded
data and then perform feature extraction. data and then perform feature extraction.
Markers are meant to be used inside the datagrabber context but you can operate Markers are meant to be used inside the DataGrabber context but you can operate
on them outside the context as long as the actual data is in the memory and the on them outside the context as long as the actual data is in the memory and the
Python runtime has not garbage-collected it. Python runtime has not garbage-collected it.
If you are interested in using already provided markers, please go to If you are interested in using already provided Markers, please go to
:doc:`../builtin`. And, if you want to implement your own marker, you need to :doc:`../builtin`. And, if you want to implement your own Marker, you need to
provide concrete implementation of :class:`.BaseMarker`. Specifically, you provide concrete implementation of :class:`.BaseMarker`. Specifically, you
need to override ``get_valid_inputs``, ``get_output_type`` and ``compute`` need to override ``get_valid_inputs``, ``get_output_type`` and ``compute``
methods. methods.

View file

@ -2,8 +2,8 @@
.. _pipeline: .. _pipeline:
The junifer Pipeline The ``junifer`` Pipeline
==================== ========================
The junifer pipeline is the main execution path of junifer. It consists of five The junifer pipeline is the main execution path of junifer. It consists of five
steps: steps:
@ -11,10 +11,10 @@ steps:
1. :ref:`Data Grabber <datagrabber>`: Interpret the dataset and provide a list 1. :ref:`Data Grabber <datagrabber>`: Interpret the dataset and provide a list
of files. of files.
2. :ref:`Data Reader <datareader>`: Read the files. 2. :ref:`Data Reader <datareader>`: Read the files.
3. :ref:`Pre-processing <preprocess>`: Prepare the images for marker 3. :ref:`Preprocess <preprocess>`: Prepare the files' data for marker
computation. computation.
4. :ref:`Marker Computation <marker>`: Compute the marker. 4. :ref:`Marker Computation <marker>`: Compute the marker(s).
5. :ref:`Storage <storage>`: Store the marker values. 5. :ref:`Storage <storage>`: Store the marker(s) values.
The element that is passed across the pipeline is called the The element that is passed across the pipeline is called the
:ref:`Data Object<data_object>`. :ref:`Data Object<data_object>`.
@ -26,7 +26,7 @@ The following is a graphical representation of the pipeline:
flowchart LR flowchart LR
dg[Data Grabber] dg[Data Grabber]
dr[Data Reader] dr[Data Reader]
pp[Pre-processing] pp[Preprocess]
mc[Marker Computation] mc[Marker Computation]
st[Storage] st[Storage]
dg --> dr dg --> dr
@ -45,7 +45,7 @@ on multiple markers:
flowchart LR flowchart LR
dg[Data Grabber] dg[Data Grabber]
dr[Data Reader] dr[Data Reader]
pp[Pre-processing] pp[Preprocess]
mc1[Marker Computation] mc1[Marker Computation]
mc2[Marker Computation] mc2[Marker Computation]
mc3[Marker Computation] mc3[Marker Computation]

View file

@ -22,12 +22,12 @@ want to perform confound removal on ``BOLD`` data before feature extraction.
Confound Removal Confound Removal
---------------- ----------------
The *Confound Removal* step is meant to remove *confounds* from the ``BOLD`` This step is meant to remove *confounds* from the ``BOLD`` data. The confounds
data. The confounds are extracted from the ``BOLD_confounds`` data (must be are extracted from the ``BOLD_confounds`` data (must be provided by the
provided by the :ref:`Data Grabber <datagrabber>`). The confounds are then :ref:`Data Grabber <datagrabber>`). The confounds are then regressed out from
regressed out from the ``BOLD`` data using :func:`nilearn.image.clean_img`. the ``BOLD`` data using :func:`nilearn.image.clean_img`.
Currently, junifer supports only one confound removal class: Currently, ``junifer`` supports only one confound removal class:
:class:`.fMRIPrepConfoundRemover`. This class is meant to remove confounds as :class:`.fMRIPrepConfoundRemover`. This class is meant to remove confounds as
described before, using the output of `fMRIPrep`_ as reference. described before, using the output of `fMRIPrep`_ as reference.
@ -128,3 +128,69 @@ parameters:
| If not, a mask is computed using | If not, a mask is computed using
| :func:`nilearn.masking.compute_brain_mask`. | :func:`nilearn.masking.compute_brain_mask`.
- compute - compute
.. _preprocess_warping:
Warping or Transformation to other spaces
-----------------------------------------
``junifer`` can also warp or transform any supported
:ref:`data type <data_types>` from the template space provided by the dataset
(e.g., ``MNI152NLin6Asym``) to either the subject's
:ref:`native space <preprocess_warping_native>` or to any other
:ref:`template space <preprocess_warping_template>`
(e.g., ``MNI152NLin2009cAsym``). This functionality is provided by
:class:`.SpaceWarper` and depends on external tools like FSL and / or ANTs.
.. _preprocess_warping_native:
Warping to subject's native space
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
To warp to subject's native space, the dataset needs to provide ``T1w`` and
``Warp`` data types and the DataGrabber needs to at least have
``["BOLD", "T1w", "Warp"]`` (if you are warping ``BOLD``) as the ``types``
parameter's value. The :class:`.SpaceWarper`'s ``reference`` parameter needs
to be set to ``T1w``, which means that the ``BOLD`` data will be transformed
using the ``T1w`` as reference (it's resampled internally to match the
resolution of the ``BOLD``). The ``Warp`` data type is new and it's only purpose
is to provide the warp or transformation file (can be linear, non-linear or
linear + non-linear transform) for the purpose. For ``using`` parameter, you can
pass either ``"fsl"`` or ``"ants"`` depending on the warp or transformation file
format.
An example YAML might look like this:
.. code-block:: yaml
preprocess:
- kind: SpaceWarper
using: fsl
reference: T1w
.. _preprocess_warping_template:
Warping to other template space
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
In a situation where your dataset might provide the ``BOLD`` data (or any other
data type that you want to work on) in ``MNI152NLin6Asym`` template space but
you would like to compute features in ``MNI152NLin2009cAsym`` template space,
you can also use the :class:`.SpaceWarper` by setting the ``reference``
parameter to the template space's name, in this case,
``reference="MNI152NLin2009cAsym"``. The ``using`` parameter needs to be set
to ``"ants"`` as we need it to warp the data.
.. note::
We only support template spaces provided by `templateflow`_ and the naming
is similar except that we omit the ``tpl-`` prefix used by ``templateflow``.
For an YAML example:
.. code-block:: yaml
preprocess:
- kind: SpaceWarper
using: ants
reference: MNI152NLin2009cAsym

View file

@ -13,7 +13,7 @@ as computed from :ref:`Marker <marker>` step of the pipeline. If the pipeline is
provided with a ``storage-like`` object, the extracted features are stored via provided with a ``storage-like`` object, the extracted features are stored via
that object else they are kept in memory. that object else they are kept in memory.
Storage is meant to be used inside the datagrabber context but you can operate Storage is meant to be used inside the DataGrabber context but you can operate
on them outside the context as long as the processed data is in the memory and on them outside the context as long as the processed data is in the memory and
the Python runtime has not garbage-collected it. the Python runtime has not garbage-collected it.
@ -25,8 +25,8 @@ storage object in turn declares and provides implementation for specific
``matrix``, ``vector`` and ``timeseries`` via ``store_matrix``, ``store_vector`` ``matrix``, ``vector`` and ``timeseries`` via ``store_matrix``, ``store_vector``
and ``store_timeseries`` methods respectively. and ``store_timeseries`` methods respectively.
For storage interfaces not supported by junifer yet, you can either make your For storage interfaces not supported by ``junifer`` yet, you can either make
own ``Storage`` by providing a concrete implementation of your own ``Storage`` by providing a concrete implementation of
:class:`.BaseFeatureStorage` or open an issue on `junifer Github`_ and we can :class:`.BaseFeatureStorage` or open an issue on `junifer Github`_ and we can
help you out. help you out.

View file

@ -5,7 +5,7 @@
Code-less Configuration Code-less Configuration
fraimondo commented 2024-04-01 15:22:48 +00:00 (Migrated from github.com)

"One of..."

"One of..."
======================= =======================
On of the most important features of junifer is its capacity to run without One of the most important features of ``junifer`` is its capacity to run without
writing a single line of code. This is achieved by using a configuration file writing a single line of code. This is achieved by using a configuration file
that is written in YAML_. In this file, we configure the different steps of that is written in YAML_. In this file, we configure the different steps of
:ref:`pipeline`. :ref:`pipeline`.
@ -17,7 +17,7 @@ As a reminder, this is how the pipeline looks like:
flowchart LR flowchart LR
dg[Data Grabber] dg[Data Grabber]
dr[Data Reader] dr[Data Reader]
pp[Pre-processing] pp[Preprocess]
mc[Marker Computation] mc[Marker Computation]
st[Storage] st[Storage]
dg --> dr dg --> dr
@ -31,21 +31,21 @@ as well as some general parameters.
As an example, we will generate the configuration file for a pipeline that will As an example, we will generate the configuration file for a pipeline that will
extract the mean ``VBM_GM`` values using two different parcellations and one set extract the mean ``VBM_GM`` values using two different parcellations and one set
of coordinates, from the ``Oasis VBM Testing dataset`` included in junifer. of coordinates, from the ``Oasis VBM Testing dataset`` included in ``junifer``.
General Parameters General Parameters
------------------ ------------------
The general parameters are the ones that are not specific to any of the sections The general parameters are the ones that are not specific to any of the sections
of the pipeline, but configure junifer as a whole. These parameters are: of the pipeline, but configure ``junifer`` as a whole. These parameters are:
* ``with``: A section used to specify modules and junifer extensions to use. * ``with``: A section used to specify modules and ``junifer`` extensions to use.
* ``workdir``: The working directory where junifer will store temporary files. * ``workdir``: The working directory where ``junifer`` will store temporary files.
Since the example uses a specific datagrabber for testing, we need to add Since the example uses a specific DataGrabber for testing, we need to add
``junifer.testing.registry`` to the ``with`` section. This will allow junifer ``junifer.testing.registry`` to the ``with`` section. This will allow ``junifer``
to find the datagrabber. We will set the ``workdir`` to ``/tmp``. to find the DataGrabber. We will set the ``workdir`` to ``/tmp``.
.. code-block:: yaml .. code-block:: yaml
@ -66,9 +66,9 @@ In order to configure the pipeline, we need to configure each step:
.. important:: .. important::
The datareader step configuration is optional, as junifer only provides one The ``datareader`` step configuration is optional, as ``junifer`` only
datareader. Nevertheless, it is possible to extend junifer with custom provides one DataReader. Nevertheless, it is possible to extend ``junifer``
datareaders, and thus, it is also possible to configure this step. with custom DataReaders, and thus, it is also possible to configure this step.
Data Grabber Data Grabber
@ -107,9 +107,9 @@ In the ``Oasis VBM Testing dataset`` example, the section will look like this:
Data Reader Data Reader
^^^^^^^^^^^ ^^^^^^^^^^^
As mentioned before, this section is entirely optional, as junifer only provides As mentioned before, this section is entirely optional, as ``junifer`` only
one DataReader (:class:`.DefaultDataReader`), which is the default in case the provides one DataReader (:class:`.DefaultDataReader`), which is the default in
section is not specified. case the section is not specified.
In any case, the syntax of the section is the same as for the ``datagrabber`` In any case, the syntax of the section is the same as for the ``datagrabber``
section, using the ``kind`` key to specify the DataReader to use, and additional section, using the ``kind`` key to specify the DataReader to use, and additional
@ -127,25 +127,26 @@ For the ``Oasis VBM Testing dataset`` example, we will not specify a
Preprocess Preprocess
^^^^^^^^^^ ^^^^^^^^^^
Pre-processing is also an optional step, as it might be the case that no ``preprocess`` is also an optional step, as it might be the case that no
pre-processing is needed. In the case that pre-processing is needed, the section pre-processing is needed. As we can perform multiple preprocessing steps, it's
must be configured using the ``kind`` key to specify the preprocessor to use, passed as a list of Preprocessors. In the case that pre-processing is needed,
and additional keys to pass parameters to the preprocessor. each Preprocessord must be configured using the ``kind`` key to specify the
Preprocessor to use, and additional keys to pass parameters to the Preprocessor.
For example, to use the :class:`.fMRIPrepConfoundRemover` preprocessor, we just For example, to use the :class:`.fMRIPrepConfoundRemover` Preprocessor, we just
need to specify its name as the ``kind`` key, as well as its parameters. need to specify its name as the ``kind`` key, as well as its parameters.
.. code-block:: yaml .. code-block:: yaml
preprocess: preprocess:
kind: fMRIPrepConfoundRemover - kind: fMRIPrepConfoundRemover
strategy: strategy:
motion: full motion: full
wm_csf: full wm_csf: full
global_signal: basic global_signal: basic
spike: 0.2 spike: 0.2
detrend: false detrend: false
standardize: true standardize: true
For the ``Oasis VBM Testing dataset`` example, we will not specify a For the ``Oasis VBM Testing dataset`` example, we will not specify a
@ -155,9 +156,9 @@ preprocessing step.
Marker Marker
^^^^^^ ^^^^^^
The ``markers`` section diverges from the previous ones, as we need to specify The ``markers`` section like the ``preprocess`` section expects a list of
a list of markers. Each marker has a name that we can use to refer to it later, markers. Each Marker has a name that we can use to refer to it later,
and a set of parameters that will be passed to the marker. and a set of parameters that will be passed to the Marker.
For the ``Oasis VBM Testing dataset`` example, we want to compute the mean For the ``Oasis VBM Testing dataset`` example, we want to compute the mean
``VBM_GM`` value for each parcel using the ``Schaefer parcellation (100 parcels, ``VBM_GM`` value for each parcel using the ``Schaefer parcellation (100 parcels,
@ -190,14 +191,14 @@ Finally, we need to define how and where the results will be stored. This is
done using the ``storage`` section, which must be configured using the ``kind`` done using the ``storage`` section, which must be configured using the ``kind``
key to specify the storage to use, and additional keys to pass parameters. key to specify the storage to use, and additional keys to pass parameters.
For example, to use the :class:`.SQLiteFeatureStorage` storage, we just need to For example, to use the :class:`.HDF5FeatureStorage` storage, we just need to
specify where we want to store the results: specify where we want to store the results:
.. code-block:: yaml .. code-block:: yaml
storage: storage:
kind: SQLiteFeatureStorage kind: HDF5FeatureStorage
uri: /data/junifer/example/oasis_vbm_testing.sqlite uri: /data/junifer/example/oasis_vbm_testing.hdf5
Complete Example Complete Example
@ -231,5 +232,5 @@ looks like:
method: mean method: mean
storage: storage:
kind: SQLiteFeatureStorage kind: HDF5FeatureStorage
uri: /data/junifer/example/oasis_vbm_testing.sqlite uri: /data/junifer/example/oasis_vbm_testing.hdf5

View file

@ -2,14 +2,14 @@
.. _using: .. _using:
Using junifer Using ``junifer``
============= =================
In this section, we will cover the main aspects behind using junifer. We will In this section, we will cover the main aspects behind using ``junifer``. We
first explain the basics behind junifer's code-less configuration. Then we will will first explain the basics behind junifer's code-less configuration. Then we
show how to use the command line interface to ``run`` junifer and ``collect`` will show how to use the command line interface to ``run`` the pipeline and
the results. Finally, we will show how to use the ``queue`` command to interact ``collect`` the results. Finally, we will show how to use the ``queue`` command
with HPC and HTC systems. to interact with HPC and HTC systems.
.. toctree:: .. toctree::
:maxdepth: 2 :maxdepth: 2
@ -25,7 +25,7 @@ with HPC and HTC systems.
Using Common Components Using Common Components
----------------------- -----------------------
The following sections explains common components of junifer that can be used The following sections explains common components of ``junifer`` that can be used
across many steps of the pipeline. across many steps of the pipeline.
.. toctree:: .. toctree::

View file

@ -12,7 +12,7 @@ contain a certain ratio of gray matter to white matter / cerebrospinal fluid,
ensuring that the features are not extracted from voxels that contain mostly ensuring that the features are not extracted from voxels that contain mostly
white matter or cerebrospinal fluid, which could add noise to the BOLD signal. white matter or cerebrospinal fluid, which could add noise to the BOLD signal.
Junifer provides a number of built-in masks, which can be listed using ``junifer`` provides a number of built-in masks, which can be listed using
:func:`.list_masks`. Some masks are images, while other masks can be computed :func:`.list_masks`. Some masks are images, while other masks can be computed
using :ref:`nilearn` functions. using :ref:`nilearn` functions.
@ -22,15 +22,14 @@ dictionary in which the **only** key is the built-in mask name and the value is
a dictionary of keyword arguments to pass to the mask function. a dictionary of keyword arguments to pass to the mask function.
For example, the following is a valid mask specification that specified the For example, the following is a valid mask specification that specified the
``GM_prob0.2`` mask. ``GM_prob0.2`` mask:
.. code-block:: yaml .. code-block:: yaml
masks: GM_prob0.2 masks: GM_prob0.2
The following is a valid mask specification that specifies the The following is a valid mask specification that specifies the
``compute_brain_mask`` mask (function from nilearn), with a threshold of ``compute_brain_mask`` mask, with a threshold of ``0.5``.
``0.5``.
.. code-block:: yaml .. code-block:: yaml

View file

@ -5,14 +5,15 @@
Queueing Jobs (HPC, HTC) Queueing Jobs (HPC, HTC)
======================== ========================
Yet another interesting feature of junifer is the ability to queue jobs on Yet another interesting feature of ``junifer`` is the ability to queue jobs on
computational clusters. This is done by adding the ``queue`` section in the computational clusters. This is done by adding the ``queue`` section in the
:ref:`codeless` file and executing the ``junifer queue`` command. :ref:`codeless` file and executing the ``junifer queue`` command.
While junifer is meant to support `HTCondor`_, `SLURM`_ and local queueing While junifer is meant to support `HTCondor`_, `SLURM`_ and local queueing
using `GNU Parallel`_, only HTCondor is currently supported. This will be using `GNU Parallel`_, only HTCondor is currently supported. This will be
implemented in future releases of junifer. If you are in immediate need of any of implemented in future releases of ``junifer``. If you are in immediate need of
these schedulers, please create an issue on the `junifer github`_ repository. any of these schedulers, please create an issue on the `junifer github`_
repository.
The ``queue`` section of the :ref:`codeless` must start by defining the The ``queue`` section of the :ref:`codeless` must start by defining the
following general parameters: following general parameters:
@ -24,7 +25,7 @@ following general parameters:
* ``kind``: The kind of scheduler to be used. Currently, only ``HTCondor`` is * ``kind``: The kind of scheduler to be used. Currently, only ``HTCondor`` is
supported. supported.
Example: Example in YAML:
.. code-block:: yaml .. code-block:: yaml
@ -40,7 +41,7 @@ The rest of the parameters depend on the scheduler you are using.
HTCondor HTCondor
-------- --------
When using HTCondor, junifer will use a DAG to queue one job per element When using HTCondor, ``junifer`` will use a DAG to queue one job per element
(``junifer run``). As an option, the DAG can include a final job (``junifer run``). As an option, the DAG can include a final job
(``junifer collect``) to collect the results once all of the individual element (``junifer collect``) to collect the results once all of the individual element
jobs are finished. jobs are finished.
@ -61,12 +62,13 @@ The following parameters are available for HTCondor:
If relative path is used then it should be relative to the YAML. If relative path is used then it should be relative to the YAML.
* ``mem``: Memory to be used by the job. It must be provided as a string with * ``mem``: Memory to be used by the job. It must be provided as a string with
the units (e.g. ``2GB``). the units (e.g., ``"2GB"``).
* ``cpus``: Number of CPUs to be used by the job. It must be provided as an int. * ``cpus``: Number of CPUs to be used by the job. It must be provided as an
integer (e.g., ``1``).
* ``disk``: Disk space to be used by the job. It must be provided as a string * ``disk``: Disk space to be used by the job. It must be provided as a string
with the units (e.g. ``2GB``). Keep in mind that junifer uses a local working with the units (e.g., ``"2GB"``). Keep in mind that ``junifer`` uses a local
directory for each job, and datalad datasets might be cloned in this temporary working directory for each job, and datalad datasets might be cloned in this
directory. temporary directory.
* ``extra_preamble``: Extra lines to be added to the HTCondor submit file. This * ``extra_preamble``: Extra lines to be added to the HTCondor submit file. This
can be used to add extra parameters to the job, such as ``requirements``. can be used to add extra parameters to the job, such as ``requirements``.
* ``collect``: This parameter allows to include a collect to the DAG to collect * ``collect``: This parameter allows to include a collect to the DAG to collect
@ -81,7 +83,7 @@ The following parameters are available for HTCondor:
* ``no``: Do not include a collect job to the DAG. * ``no``: Do not include a collect job to the DAG.
Example: Example in YAML:
.. code-block:: yaml .. code-block:: yaml