[DOC]: Improve documentation #317

Merged
synchon merged 19 commits from update/improve-docs into main 2024-04-02 07:59:33 +00:00
27 changed files with 842 additions and 338 deletions

View file

@ -25,7 +25,7 @@ Data Grabber
- In Progress
- Done
Version added: If the status is "Done", the Junifer version in which the
Version added: If the status is "Done", the junifer version in which the
dataset was added. Else, a link to the Github issue or pull request
implementing the dataset. Links to github can be added by using the
following syntax: :gh:`<issue number>`
@ -122,6 +122,54 @@ Planned
- :gh:`47`
Preprocessor
------------
..
Provide a list of the Preprocessors that are implemented or planned.
State: this should indicate the state of the preprocessor. Valid options are
- Planned
- In Progress
- Done
LeSasse commented 2024-03-28 12:56:11 +00:00 (Migrated from github.com)

Does Junifer handle/allow use of multiple preprocessors chained in sequence? how would parametrisation in the YAML look like for that?

Does Junifer handle/allow use of multiple preprocessors chained in sequence? how would parametrisation in the YAML look like for that?
synchon commented 2024-03-28 13:42:22 +00:00 (Migrated from github.com)

Yeah it does. preprocess accepts a list now and the execution follows the sequence you specify in the YAML.

Yeah it does. `preprocess` accepts a list now and the execution follows the sequence you specify in the YAML.
Version added: If the status is "Done", the junifer version in which the
preprocessor was added. Else, a link to the Github issue or pull request
implementing the preprocessor. Links to github can be added by using the
following syntax: :gh:`<issue number>`
Available
~~~~~~~~~
.. list-table::
:widths: auto
:header-rows: 1
* - Class
- Description
- State
- Version Added
* - :class:`.fMRIPrepConfoundRemover`
- Remove confounds from ``fMRIPrep``-ed data
- Done
- 0.0.1
* - :class:`.SpaceWarper`
- | Warp / transform data from one space to another
| (subject-native or other template spaces)
- Done
- 0.0.4
* - ``Smoothing``
- | Apply smoothing to data, particularly useful when dealing with
| ``fMRIPrep``-ed data
- In Progress
- :gh:`161`
..
Planned
~~~~~~~
LeSasse commented 2024-03-28 12:57:00 +00:00 (Migrated from github.com)

Why particularly useful for fMRIPrep'ed data?

Why particularly useful for fMRIPrep'ed data?
synchon commented 2024-03-28 13:44:08 +00:00 (Migrated from github.com)

From your issue description I understand that fMRIPrep doesn't perform smoothing after confound regression.

From your issue description I understand that fMRIPrep doesn't perform smoothing after confound regression.
Marker
------
@ -133,7 +181,7 @@ Marker
- In Progress
- Done
Version added: If the status is "Done", the Junifer version in which the
Version added: If the status is "Done", the junifer version in which the
marker was added. Else, a link to the Github issue or pull request
implementing the marker. Links to github can be added by using the
following syntax: :gh:`<issue number>`
@ -254,10 +302,6 @@ Planned
* - Connectedness
- Compute connectedness
- :gh:`34`
* - Permutation entropy, Range entropy, Multiscale entropy and Hurst exponent
- | Calculate Permutation entropy, Range entropy, Multiscale entropy and
| Hurst exponent
- :gh:`61`
Parcellation
------------
@ -265,7 +309,7 @@ Parcellation
..
Provide a list of the Parcellations that are implemented or planned.
Version added: The Junifer version in which the parcellation was added.
Version added: The junifer version in which the parcellation was added.
Available
~~~~~~~~~
@ -277,8 +321,8 @@ Available
* - Name
LeSasse commented 2024-03-28 12:57:41 +00:00 (Migrated from github.com)

Above you add Version Added with captial A. Similar for Template spaces, should it be Template Spaces?

Above you add Version Added with captial A. Similar for Template spaces, should it be Template Spaces?
synchon commented 2024-03-28 13:45:33 +00:00 (Migrated from github.com)

Yeah missed it, good catch.

Yeah missed it, good catch.
- Options
- Keys
- Spaces
- Version added
- Template Spaces
- Version Added
- Publication
* - Schaefer
- ``n_rois``, ``yeo_networks``
@ -439,7 +483,7 @@ Coordinates
..
Provide a list of the Coordinates that are implemented or planned.
Version added: The Junifer version in which the parcellation was added.
Version added: The junifer version in which the parcellation was added.
Available
~~~~~~~~~
@ -450,7 +494,7 @@ Available
* - Name
- Keys
- Version added
- Version Added
- Publication
* - Cognitive action control
- ``CogAC``
@ -635,7 +679,7 @@ Mask
..
Provide a list of the masks that are implemented or planned.
Version added: The Junifer version in which the mask was added.
Version added: The junifer version in which the mask was added.
Available
~~~~~~~~~
@ -646,8 +690,8 @@ Available
LeSasse commented 2024-03-28 12:57:54 +00:00 (Migrated from github.com)

see above cmt

see above cmt
* - Name
- Keys
- Spaces
- Version added
- Template Space
- Version Added
- Description - Publication
* - Vickery-Patil (Gray Matter)
- | ``GM_prob0.2``
@ -663,14 +707,14 @@ Available
- | Vickery, Sam, & Patil, Kaustubh. (2022).
LeSasse commented 2024-03-28 12:58:32 +00:00 (Migrated from github.com)

you made a point to fix Junifer to junifer. Should it be nilearn as well rather than Nilearn?

you made a point to fix Junifer to junifer. Should it be nilearn as well rather than Nilearn?
synchon commented 2024-03-28 13:46:36 +00:00 (Migrated from github.com)

That would be better, fair point.

That would be better, fair point.
| Chimpanzee and Human Gray Matter Masks [Data set]. Zenodo.
| https://doi.org/10.5281/zenodo.6463123
* - Nilearn's MNI152 1mm-resolution mask
* - ``junifer``'s custom brain mask
- | ``compute_brain_mask``
- Adapts to the target data
- 0.0.2
- | Compute the whole-brain mask. This mask is calculated using
| MNI152 1mm-resolution template mask onto the target image.
| See :func:`nilearn.masking.compute_brain_mask`
* - Nilearn's mask computed from FMRI data
- | Compute the whole-brain, gray-matter or white-matter mask using
| the template and the resolution from the target image. The
| templates are obtained via ``templateflow``.
* - ``nilearn``'s mask computed from fMRI data
- | ``compute_epi_mask``
- Adapts to the target data
- 0.0.2
@ -678,15 +722,14 @@ Available
| proposed by T.Nichols: find the least dense point of the histogram,
| between fractions ``lower_cutoff`` and ``upper_cutoff`` of the total
| image histogram. See :func:`nilearn.masking.compute_epi_mask`
* - Nilearn's background mask
* - ``nilearn``'s background mask
- | ``compute_background_mask``
- Adapts to the target data
- 0.0.2
- | Compute a brain mask for the images by guessing the value of the
| background from the border of the image.
| See :func:`nilearn.masking.compute_background_mask`
* - Nilearn's ICBM152 template gray-matter mask
* - ``nilearn``'s ICBM152 template gray-matter mask
- | ``fetch_icbm152_brain_gm_mask``
- ``MNI152NLin2009aAsym``
- 0.0.2
@ -695,8 +738,9 @@ Available
| See :func:`nilearn.datasets.fetch_icbm152_brain_gm_mask`
Planned
~~~~~~~
..
Planned
~~~~~~~
..
helpful site for creating tables: https://rest-sphinx-memo.readthedocs.io/en/latest/ReST.html#tables

View file

@ -0,0 +1 @@
Improve documentation by adding information about space transformation and writing custom Preprocessors by `Synchon Mandal`_

View file

@ -164,6 +164,8 @@ In case you remove some files or change their filenames, you can run into
errors when using ``make local``. In this situation you can use ``make clean``
to clean up the already build files and then re-run ``make local``.
Also, we follow British English for the documentation.
Writing Examples
----------------

View file

@ -6,19 +6,20 @@ Adding Coordinates
==================
Instead of using whole-brain parcellations to aggregate voxel-wise signals from
MR images (as for example in the :class:`.ParcelAggregation` marker), junifer
MR images (as for example in the :class:`.ParcelAggregation` marker), ``junifer``
allows you to specify a set of coordinates around which to draw spheres to
aggregate (for example using the :class:`.SphereAggregation` marker) the MR
signals from individual voxels. Now, before you start specifying your own sets
of coordinates, check the coordinates that junifer already has
of coordinates, check the coordinates that ``junifer`` already has
:ref:`built in <builtin>`. If you simply want to use a well known set of
coordinates from the literature, there is a reasonable chance, that junifer
coordinates from the literature, there is a reasonable chance, that ``junifer``
provides them already.
If you checked the in-built coordinates, and they are not there already (for
example if you came up with your own set of coordinates), then junifer provides
an easy way for you to register them using the :func:`.register_coordinates`
function, so you can use your own set of coordinates within a junifer pipeline.
example if you came up with your own set of coordinates), then ``junifer``
provides an easy way for you to register them using the
:func:`.register_coordinates` function, so you can use your own set of
coordinates within a ``junifer`` pipeline.
From the API reference, we can see that it has 4 positional arguments
(``name``, ``coordinates``, ``voi_names`` and ``space``) as well as one
@ -26,7 +27,7 @@ optional keyword argument (``overwrite``).
The ``name`` argument takes a string indicating the name you want to give to
this set of coordinates. This ``name`` can be used to obtain and operate on a
set of coordinates in junifer. For example, you can obtain your coordinates
set of coordinates in ``junifer``. For example, you can obtain your coordinates
after registration by providing ``name`` to :func:`.load_coordinates`. We could
simply call it ``"my_set_of_coordinates"``, but likely you want a more
descriptive and more informative name most of the time.
@ -36,9 +37,7 @@ The ``coordinates`` argument takes the actual coordinates as a 2-dimensional
columns (one for each spatial dimension). That is, the first, second, and third
columns indicate the x-, y-, and z-coordinates in MNI space respectively.
The number of rows in the array correspond to the number of coordinates that
belong to this set. Note, that junifer (as of yet) only works in MNI space, and
so therefore these coordinates should always be real-world coordinates of the
MNI space.
belong to this set.
The ``voi_names`` argument takes a list of strings
indicating the names of each coordinate (i.e. volume-of-interest) in the
@ -47,7 +46,7 @@ the number of rows in the coordinates array. Now, we know everything we need to
know to register a set of coordinates.
Lastly, we specify the ``space`` that the coordinates are in, for example,
``"MNI"`` or ``"Native"`` (scanner-native space).
``"MNI"`` or ``"native"`` (scanner-native space).
Step 1: Prepare code to register a set of coordinates
-----------------------------------------------------
@ -62,8 +61,8 @@ packages:
import numpy as np
For the sake of this example, we can create a set of coordinates that belong
to the default mode network (DMN), and register this set of coordinates with
junifer. Note, that junifer already has a
to the Default Mode Network (DMN), and register this set of coordinates with
``junifer``. Note, that ``junifer`` already has a
:ref:`set of coordinates built-in <builtin>` ("DMNBuckner") that is associated
with the DMN. Here, we use the DMN coordinates used in a
`nilearn example <https://nilearn.github.io/dev/auto_examples/03_connectivity/plot_sphere_based_connectome.html>`_.
@ -97,7 +96,7 @@ simply use this to register our coordinates:
space="MNI"
)
Now, when we run this script, junifer registers these coordinates and we can
Now, when we run this script, ``junifer`` registers these coordinates and we can
use them in subsequent analyses. Let's now consider how to use coordinate
registration in combination with
:ref:`codeless configuration using a YAML file <codeless>`.
@ -106,7 +105,7 @@ Step 2: Add coordinate registration to the YAML file
----------------------------------------------------
In order to register your coordinates for a pipeline configured by a YAML file,
you can use the ``with`` keyword provided by junifer:
you can use the ``with`` keyword provided by ``junifer``:
.. code-block:: yaml

View file

@ -12,8 +12,8 @@ the structure of a dataset and provide two specific functionalities:
element (e.g. the path to the T1 image, the path to the T2 image, etc.)
#. Provide the list of *elements* available in the dataset.
In this section, we will see how to create a datagrabber for a dataset. Basic
aspects of datagrabbers are covered in the
In this section, we will see how to create a DataGrabber for a dataset. Basic
aspects of DataGrabbers are covered in the
:ref:`Understanding Data Grabbers <datagrabber>` section.
.. _extending_datagrabbers_think:
@ -22,7 +22,7 @@ Step 1: Think about the element
-------------------------------
Like with any programming-related task, the first step is to think. When
creating a Data Grabber, we need to first define what an *element* is.
creating a DataGrabber, we need to first define what an *element* is.
The *element* should be the smallest unit of data that can be processed. That
is, for each element, there should be a set of data that can be processed, but
only one of each *data type* (see :ref:`data_types`).
@ -42,28 +42,28 @@ then the *element* should be composed of 3 items:
If any of these items were not part of the element, then we will have more than
one ``T1w`` and / or ``BOLD`` image for each subject, which is not allowed.
Importantly, nothing prevents that one image is part of two different elements.
For example, it is usually the case that the ``T1w`` image is not acquired for
each task, but once in the entire session. So in this case, the ``T1w`` image
for the element (``sub001``, ``ses1``, ``rest``) will be the same as the
``T1w`` image for the element (``sub001``, ``ses1``, ``stroop``).
Importantly, nothing prevents that one image being part of two different
elements. For example, it is usually the case that the ``T1w`` image is not
acquired for each task, but once in the entire session. So in this case, the
``T1w`` image for the element (``sub001``, ``ses1``, ``rest``) will be the same
as the ``T1w`` image for the element (``sub001``, ``ses1``, ``stroop``).
We will now continue this section using as an example, a dataset in BIDS format
in which 9 subjects (``sub-01`` to ``sub-09``) were scanned each during 3
sessions (``ses-01``, ``ses-02``, ``ses-03``) and each session included a
``T1w`` and a ``BOLD`` image (resting-state), except for ``ses-03`` which was
only anatomical.
only anatomical data.
Step 2: Think about the dataset's structure
-------------------------------------------
Now that we have our element defined, we need to think about the structure of
the dataset. Mainly, because the structure of the dataset will determine how
the Data Grabber needs to be implemented.
the DataGrabber needs to be implemented.
Junifer provides an abstract class to deal with datasets that can be thought in
terms of *patterns*. A *pattern* is a string that contains placeholders that are
replaced by the actual values of the element. In our BIDS example, the path
``junifer`` provides an abstract class to deal with datasets that can be thought
in terms of *patterns*. A *pattern* is a string that contains placeholders that
are replaced by the actual values of the element. In our BIDS example, the path
to the T1w image of subject ``sub-01`` and session ``ses-01``, relative to the
dataset location, is ``sub-01/ses-01/anat/sub-01_ses-01_T1w.nii.gz``. By
replacing ``sub-01`` with ``sub-02``, we can obtain the T1w image of the first
@ -88,7 +88,7 @@ discussion in the `junifer Discussions`_ page. Most probably we can help you
get your dataset in order.
If there is no other way, then you can follow :ref:`extending_datagrabbers_base`
to create a Data Grabber from scratch.
to create a DataGrabber from scratch.
.. _extending_datagrabbers_pattern:
@ -101,7 +101,7 @@ Option A: Extending from PatternDataGrabber
The :class:`.PatternDataGrabber` class is an abstract class that has the
functionality of understanding patterns embedded in it.
Before creating the datagrabber, we need to define 3 variables:
Before creating the DataGrabber, we need to define 3 variables:
* ``types``: A list with the available :ref:`data_types` in our dataset.
* ``patterns``: A dictionary that specifies the pattern for each data type.
@ -126,7 +126,7 @@ where the dataset is located. For example, if the dataset is located in
location of the dataset, we can expose the variable in the constructor, as in
the following example.
With the variables defined above, we can create our datagrabber and name it
With the variables defined above, we can create our DataGrabber and name it
``ExampleBIDSDataGrabber``:
.. code-block:: python
@ -152,7 +152,7 @@ With the variables defined above, we can create our datagrabber and name it
replacements=replacements,
)
Our datagrabber is ready to be used by junifer. However, it is still unknown
Our DataGrabber is ready to be used by ``junifer``. However, it is still unknown
to the library. We need to register it in the library. To do so, we need to
use the :func:`.register_datagrabber` decorator.
@ -183,9 +183,9 @@ use the :func:`.register_datagrabber` decorator.
)
Now, we can use our datagrabber in junifer, by setting the ``datagrabber`` kind
in the yaml file to ``ExampleBIDSDataGrabber``. Remember that we still need to
set the ``datadir``.
Now, we can use our DataGrabber in ``junifer``, by setting the ``datagrabber``
kind in the yaml file to ``ExampleBIDSDataGrabber``. Remember that we still need
to set the ``datadir``.
.. code-block:: yaml
@ -209,7 +209,7 @@ temporary directory. To set the location of the dataset, you can use the
be used to specify the path to the root directory of the dataset after doing
``datalad clone``.
In the example, the dataset is hosted in gin
In the example, the dataset is hosted in Gin
(``https://gin.g-node.org/juaml/datalad-example-bids``).
When we clone this dataset, we will see the following structure:
@ -238,7 +238,7 @@ Now we have our 2 additional variables:
uri = "https://gin.g-node.org/juaml/datalad-example-bids"
rootdir = "example_bids_ses"
And we can create our datagrabber:
And we can create our DataGrabber:
.. code-block:: python
@ -267,17 +267,34 @@ And we can create our datagrabber:
replacements=replacements,
)
This approach can be used directly from the YAML, like so:
.. code-block:: yaml
datagrabber:
- kind: PatternDataladDataGrabber
types:
- BOLD
- T1w
patterns:
BOLD: "{subject}/{session}/func/{subject}_{session}_task-rest_bold.nii.gz"
T1w: "{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz"
replacements:
- subject
- session
uri: "https://gin.g-node.org/juaml/datalad-example-bids"
rootdir: "example_bids_ses"
.. _extending_datagrabbers_base:
Option B: Extending from BaseDataGrabber
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
While we could not think of a use case in which the pattern-based datagrabber
would not be suitable, it is still possible to create a datagrabber extending
While we could not think of a use case in which the pattern-based DataGrabber
would not be suitable, it is still possible to create a DataGrabber extending
from the :class:`.BaseDataGrabber` class.
In order to create a datagrabber extending from :class:`.BaseDataGrabber`, we
In order to create a DataGrabber extending from :class:`.BaseDataGrabber`, we
need to implement the following methods:
- ``get_item``: to get a single item from the dataset.
@ -287,7 +304,7 @@ need to implement the following methods:
.. note::
The ``__init__`` method could also be implemented, but it is not mandatory.
This is required if the datagrabber requires any extra parameter.
This is required if the DataGrabber requires any extra parameter.
We will now implement our BIDS example with this method.
@ -339,7 +356,7 @@ method, in the same order.
return ["subject", "session"]
LeSasse commented 2024-03-28 13:55:23 +00:00 (Migrated from github.com)

this is american spelling

this is american spelling
synchon commented 2024-03-28 15:08:44 +00:00 (Migrated from github.com)

Boston tea party reversed.

Boston tea party reversed.
So, to summarize, our datagrabber will look like this:
So, to summarise, our DataGrabber will look like this:
.. code-block:: python
@ -403,7 +420,7 @@ not standardised.
The ``mappings`` key is only required if the ``format`` is ``adhoc``. If the
``format`` is ``fmriprep``, the ``mappings`` key is not required.
Currently, junifer provides only one confound remover step
Currently, ``junifer`` provides only one confound remover step
(:class:`.fMRIPrepConfoundRemover`), which relies entirely on the ``fmriprep``
confound variable names. Thus, if the confounds are not in ``fmriprep`` format,
the user will need to provide the mappings between the *ad-hoc* variable names

View file

@ -0,0 +1,141 @@
.. include:: ../links.inc
LeSasse commented 2024-03-28 13:00:12 +00:00 (Migrated from github.com)

Should dependencies be capitalised?

Should dependencies be capitalised?
LeSasse commented 2024-03-28 13:01:38 +00:00 (Migrated from github.com)

"having two keys" -> "with two keys"

"having two keys" -> "with two keys"
.. _specifying_dependencies:
Specifying Dependencies
=======================
This section describes how you can tackle different situations when writing
your custom Marker and / or Preprocessor and take care of the dependencies for
them.
You might have already come across listing out dependencies for your
:ref:`custom Markers <extending_markers>` and / or
:ref:`custom Preprocessors <extending_preprocessors>`, if not, check them out
first. If you have already gone through them, you are already familiar with using
class attribute ``_DEPENDENCIES`` to keep track of its dependencies. ``junifer``
is a bit more sophisticated about them and we will see here how you can make the
best use of them.
.. _component_dependencies:
Handling dependencies that come as Python packages
--------------------------------------------------
You have already seen this case handled by having a class attribute
``_DEPENDENCIES`` whose value is a set of all the package names that the
component depends on. For example, for :class:`.RSSETSMarker`, we have:
.. code-block:: python
_DEPENDENCIES: ClassVar[Set[str]] = {"nilearn"}
The type annotation is for documentation and static type checking purposes.
Although not required, we highly recommend you use them, your future self
and others who use it will thank you.
.. _component_external_dependencies:
Handling external dependencies from toolboxes
---------------------------------------------
You can also specify dependencies of external toolboxes like AFIN, FSL and ANTs,
by having a class attribute like so:
.. code-block:: python
_EXT_DEPENDENCIES: ClassVar[List[Dict[str, Union[str, List[str]]]]] = [
{
"name": "afni",
"commands": ["3dReHo", "3dAFNItoNIFTI"],
},
]
The above example is taken from the class which computes regional homogeneity
(ReHo) using AFNI. The general pattern is that you need to have the value of
``_EXT_DEPENDENCIES`` as a list of dictionary with two keys:
* ``name`` (str) : lowercased name of the toolbox
* ``commands`` (list of str) : actual names of the commands you need to use
This is simple but powerful as we will see in the following sub-sections.
.. _component_conditional_dependencies:
Handling conditional dependencies
---------------------------------
You might encounter situations where your Marker or Preprocessor needs to have
option for the user to either use a dependency that comes as a package or
use a dependency that relies on external toolboxes. With the foundation we laid
above, it is really simple to solve it while having validation before running
and letting the user know if some dependency is missing.
Let's look at an actual implementation, in this case :class:`.SpaceWarper`, so
that it shows the problem a bit better and how we solve it:
.. code-block:: python
class SpaceWarper(BasePreprocessor):
# docstring
_CONDITIONAL_DEPENDENCIES: ClassVar[List[Dict[str, Union[str, Type]]]] = [
{
"using": "fsl",
"depends_on": FSLWarper,
},
{
"using": "ants",
"depends_on": ANTsWarper,
},
]
def __init__(
self, using: str, reference: str, on: Union[List[str], str]
) -> None:
# validation and setting up
Here, you see a new class attribute ``_CONDITIONAL_DEPENDENCIES`` which is a
list of dictionaries with two keys:
* ``using`` (str) : lowercased name of the toolbox
* ``depends_on`` (object) : a class which implements the particular tool's use
It is mandatory to have the ``using`` positional argument in the constructor in
this case as the validation starts with this and moves further. It is also
mandatory to only allow the value of ``using`` argument to be one of them
specified in the ``using`` key of ``_CONDITIONAL_DEPENDENCIES`` entries.
For brevity, we only show the ``FSLWarper`` here but ``ANTsWarper`` looks very
LeSasse commented 2024-03-28 13:02:31 +00:00 (Migrated from github.com)

``ANTsWarper``` the capitalisation here makes me anxious, but I'll allow it

``ANTsWarper``` the capitalisation here makes me anxious, but I'll allow it
synchon commented 2024-03-28 13:50:02 +00:00 (Migrated from github.com)

It's to follow the convention of the tool name like in other places in the code base.

It's to follow the convention of the tool name like in other places in the code base.
LeSasse commented 2024-03-28 13:51:16 +00:00 (Migrated from github.com)

I understand :)

I understand :)
similar. ``FSLWarper`` looks like this (only the relevant part is shown here):
.. code-block:: python
class FSLWarper:
# docstring
_EXT_DEPENDENCIES: ClassVar[List[Dict[str, Union[str, List[str]]]]] = [
{
"name": "fsl",
"commands": ["flirt", "applywarp"],
},
]
_DEPENDENCIES: ClassVar[Set[str]] = {"numpy", "nibabel"}
def preprocess(
self,
input: Dict[str, Any],
extra_input: Dict[str, Any],
) -> Dict[str, Any]:
# implementation
Here you can see the familiar ``_DEPENDENCIES`` and ``_EXT_DEPENDENCIES`` class
attributes. The validation process starts by looking up the ``using`` value of
the ``_CONDITIONAL_DEPENDENCIES`` entries and then retrieves the object pointed
by ``depends_on``. After that, the ``_DEPENDENCIES`` and ``_EXT_DEPENDENCIES``
class attributes are checked.
This might be a bit too much to get it right away so feel free to check the code
for a better understanding. You can also check ``ALFFBase`` for a Marker
having this pattern.

View file

@ -2,19 +2,20 @@
.. _extending_extension:
Creating a junifer extension
============================
Creating a ``junifer`` extension
================================
Junifer is designed to be easily extensible. Through the use of a registry and
decorators, you can easily add new functionality to junifer during runtime. This
is done by creating a new Python module and importing it before running junifer.
``junifer`` is designed to be easily extensible. Through the use of a registry
and decorators, you can easily add new functionality to ``junifer`` during
runtime. This is done by creating a new Python module and importing it before
running ``junifer``.
A special consideration has to be made when using the
:ref:`code-less configuration<codeless>`. In this case, the
``with`` statement can be used to import a module or run a Python file.
In the following example, we instruct junifer to first import ``my_module`` and
then run the ``my_file.py`` file.
In the following example, we instruct ``junifer`` to first import ``my_module``
and then run the ``my_file.py`` file:
.. code-block:: yaml
@ -22,13 +23,13 @@ then run the ``my_file.py`` file.
- my_module
- my_file.py
Thus, the code from ``my_file.py`` will be executed before running junifer. This
is the ideal place to include junifer extensions.
Thus, the code from ``my_file.py`` will be executed before running ``junifer``.
This is the ideal place to include ``junifer`` extensions.
.. important::
Some junifer commands will not consider files imported from files included
Some ``junifer`` commands will not consider files imported from files included
in the ``with`` statement. If ``my_file.py`` imports ``my_other_file.py``,
some of the junifer commands will not consider ``my_other_file.py``. Either
some of the ``junifer`` commands will not consider ``my_other_file.py``. Either
place all the code in one file or add multiple files to the ``with``
statement.

View file

@ -2,20 +2,20 @@
.. _extending:
Extending junifer
=================
Extending ``junifer``
=====================
While we aim to provide as many datasets and markers as possible, we are also
interested in allowing users to extend the functionality with their own
datagrabbers, preprocessing, markers, etc., .
DataGrabbers, Preprocessors, Markers, etc., .
It's not necessary to have the new functionality included in junifer before
It's not necessary to have the new functionality included in ``junifer`` before
the user can use them. The user can simply create a new Python file, code the
desired functionality and use it with junifer. This is the first step towards
including the new functionality in the junifer pipeline.
desired functionality and use it with ``junifer``. This is the first step towards
including the new functionality in the ``junifer`` pipeline.
In this section we will show how to extend junifer, by creating new
datagrabbers, preprocessing, markers, etc., following the *junifer* way.
In this section we will show how to extend ``junifer``, by creating new
DataGrabbers, Preprocessors, Markers, etc., following the *junifer* way.
.. toctree::
@ -25,6 +25,8 @@ datagrabbers, preprocessing, markers, etc., following the *junifer* way.
extension
datagrabber
marker
preprocessor
dependencies
parcellations
coordinates
masks

View file

@ -5,35 +5,35 @@
Creating Markers
================
Computing a marker (a.k.a. *feature*) is the main goal of junifer. While we aim
to provide as many markers as possible, it might be the case that the marker you
are looking for is not available. In this case, you can create your own marker
Computing a marker (a.k.a. *feature*) is the main goal of ``junifer``. While we
aim to provide as many Markers as possible, it might be the case that the Marker
you are looking for is not available. In this case, you can create your own Marker
by following this tutorial.
Most of the functionality of a junifer marker has been taken care by the
Most of the functionality of a ``junifer`` Marker has been taken care by the
:class:`.BaseMarker` class. Thus, only a few methods are required:
#. ``get_valid_inputs``: The method to obtain the list of valid inputs for the
marker. This is used to check that the inputs provided by the user are
Marker. This is used to check that the inputs provided by the user are
valid. This method should return a list of strings, representing
:ref:`data types <data_types>`.
#. ``get_output_type``: The method to obtain the kind of output of the marker.
This is used to check that the output of the marker is compatible with the
#. ``get_output_type``: The method to obtain the output type of the Marker.
This is used to check that the output of the Marker is compatible with the
storage. This method should return a string, representing
:ref:`storage types <storage_types>`.
#. ``compute``: The method that given the data, computes the marker.
#. ``__init__``: The initialisation method, where the marker is configured.
#. ``compute``: The method that given the data, computes the Marker.
#. ``__init__``: The initialisation method, where the Marker is configured.
As an example, we will develop a ``ParcelMean`` marker, a marker that first
As an example, we will develop a ``ParcelMean`` Marker, a Marker that first
applies a parcellation and then computes the mean of the data in each parcel.
This is a very simple example, but it will show you how to create a new marker.
This is a very simple example, but it will show you how to create a new Marker.
.. _extending_markers_input_output:
Step 1: Configure input and output
----------------------------------
This step is quite simple: we need to define the input and output of the marker.
This step is quite simple: we need to define the input and output of the Marker.
Based on the current :ref:`data types <data_types>`, we can have ``BOLD``,
``VBM_WM`` and ``VBM_GM`` as valid inputs.
@ -42,32 +42,32 @@ Based on the current :ref:`data types <data_types>`, we can have ``BOLD``,
def get_valid_inputs(self) -> list[str]:
return ["BOLD", "VBM_WM", "VBM_GM"]
The output of the marker depends on the input. For ``BOLD``, it will be
The output of the Marker depends on the input. For ``BOLD``, it will be
``timeseries``, while for the rest of the inputs, it will be ``vector``. Thus,
we can define the output as:
.. code-block:: python
def get_output_type(self, input_kind: str) -> str:
if input_kind == "BOLD":
def get_output_type(self, input_type: str) -> str:
if input_type == "BOLD":
return "timeseries"
else:
return "vector"
.. _extending_markers_init:
Step 2: Initialize the marker
Step 2: Initialise the Marker
-----------------------------
In this step we need to define the parameters of the marker the user can provide
to configure how the marker will behave.
In this step we need to define the parameters of the Marker the user can provide
to configure how the Marker will behave.
The parameters of the marker are defined in the ``__init__`` method. The
The parameters of the Marker are defined in the ``__init__`` method. The
:class:`.BaseMarker` class requires two optional parameters:
1. ``name``: the name of the marker. This is used to identify the marker in the
1. ``name``: the name of the Marker. This is used to identify the Marker in the
configuration file.
2. ``on``: a list or string with the data types that the marker will be applied
2. ``on``: a list or string with the data types that the Marker will be applied
to.
.. attention::
@ -92,29 +92,29 @@ parcellation to use. Thus, we can define the ``__init__`` method as follows:
.. caution::
Parameters of the marker must be stored as object attributes without using
Parameters of the Marker must be stored as object attributes without using
``_`` as prefix. This is because any attribute that starts with ``_`` will
not be considered as a parameter and not stored as part of the metadata of
the marker.
the Marker.
.. _extending_markers_compute:
Step 3: Compute the marker
Step 3: Compute the Marker
--------------------------
In this step, we will define the method that computes the marker. This method
will be called by junifer when needed, using the data provided by the
datagrabber, as configured by the user. The method ``compute`` has two
In this step, we will define the method that computes the Marker. This method
will be called by ``junifer`` when needed, using the data provided by the
DataGrabber, as configured by the user. The method ``compute`` has two
arguments:
* ``input``: a dictionary with the data to be used to compute the marker. This
* ``input``: a dictionary with the data to be used to compute the Marker. This
will be the corresponding element in the :ref:`Data Object<data_object>`
already indexed. Thus, the dictionary has at least two keys: ``data`` and
``path``. The first one contains the data, while the second one contains the
path to the data. The dictionary can also contain other keys, depending on the
data type.
* ``extra_input``: the rest of the :ref:`Data Object<data_object>`. This is
useful if you want to use other data to compute the marker
useful if you want to use other data to compute the Marker
(e.g.: ``BOLD_confounds`` can be used to de-confound the ``BOLD`` data).
Following the example, we will compute the mean of the data in each parcel using
@ -133,7 +133,7 @@ the ``store`` method.
from typing import Any
from junifer.data import load_parcellation
from junifer.data import get_parcellation
from nilearn.maskers import NiftiLabelsMasker
@ -145,13 +145,11 @@ the ``store`` method.
# Get the data
data = input["data"]
# Get the min of the voxels sizes and use it as the resolution
resolution = np.min(data.header.get_zooms()[:3])
# Load the parcellation
t_parcellation, t_labels, _ = load_parcellation(
# Get the parcellation tailored for the target
t_parcellation, t_labels, _ = get_parcellation(
name=self.parcellation_name,
resolution=resolution,
target_data=input,
extra_input=extra_input,
)
# Create a masker
@ -173,32 +171,33 @@ the ``store`` method.
.. _extending_markers_finalize:
Step 4: Finalise the marker
Step 4: Finalise the Marker
---------------------------
Once all of the above steps are done, we just need to give our marker a name,
Once all of the above steps are done, we just need to give our Marker a name,
state its *dependencies* and register it using the ``@register_marker``
decorator.
The *dependencies* are the core packages that are required to compute the marker.
This will be later used to keep track of the versions of the packages used to
compute the marker. To inform junifer about the dependencies of a marker, we need
to define a ``_DEPENDENCIES`` attribute in the class. This attribute must be a
set, with the names of the packages as strings. For example, the ``ParcelMean``
marker has the following dependencies:
The :ref:`dependencies <specifying_dependencies>` are the core packages that are
required to compute the Marker. This will be later used to keep track of the
versions of the packages used to compute the Marker. To inform ``junifer``
about the dependencies of a Marker, we need to define a ``_DEPENDENCIES``
attribute in the class. This attribute must be a set, with the names of the
packages as strings. For example, the ``ParcelMean`` marker has the
following dependencies:
.. code-block:: python
_DEPENDENCIES = {"nilearn", "numpy"}
Finally, we need to register the marker using the ``@register_marker`` decorator.
Finally, we need to register the Marker using the ``@register_marker`` decorator.
.. code-block:: python
from typing import Any
from junifer.api.decorators import register_marker
from junifer.data import load_parcellation
from junifer.data import get_parcellation
from junifer.markers.base import BaseMarker
from nilearn.maskers import NiftiLabelsMasker
@ -220,8 +219,8 @@ Finally, we need to register the marker using the ``@register_marker`` decorator
def get_valid_inputs(self) -> list[str]:
return ["BOLD", "VBM_WM", "VBM_GM"]
def get_output_type(self, input_kind: str) -> str:
if input_kind == "BOLD":
def get_output_type(self, input_type: str) -> str:
if input_type == "BOLD":
return "timeseries"
else:
return "vector"
@ -234,13 +233,11 @@ Finally, we need to register the marker using the ``@register_marker`` decorator
# Get the data
data = input["data"]
# Get the min of the voxels sizes and use it as the resolution
resolution = np.min(data.header.get_zooms()[:3])
# Load the parcellation
t_parcellation, t_labels, _ = load_parcellation(
# Get the parcellation tailored for the target
t_parcellation, t_labels, _ = get_parcellation(
name=self.parcellation_name,
resolution=resolution,
target_data=input,
extra_input=extra_input,
)
# Create a masker
@ -283,8 +280,8 @@ Template for a custom Marker
valid = []
return valid
def get_output_type(self, input_kind):
# TODO: Return the valid output kind for each input kind
def get_output_type(self, input_type):
# TODO: Return the valid output type for each input type
pass
def compute(self, input, extra_input):

View file

@ -5,7 +5,7 @@
Adding Masks
============
Many processing steps and markers in junifer allow you to specify a binary
Many processing steps and Markers in ``junifer`` allow you to specify a binary
mask to select voxels you want to include in the analysis. There are a number
of masks :ref:`in-built in junifer already <builtin>`, so check if any of them
suit your needs. Check how to use these masks :ref:`here <using_masks>`. Once
@ -14,23 +14,23 @@ suit your needs, and you have found that they don't, you can come back here to
learn how to use your own masks.
The principle is fairly simple and quite similar to :ref:`adding_parcellations`
and :ref:`adding_coordinates`. junifer provides a :func:`.register_mask`
and :ref:`adding_coordinates`. ``junifer`` provides a :func:`.register_mask`
function that lets you register your own custom masks. It consists of three
positional arguments (``name``, ``mask_path`` and ``space``) and one optional
keyword argument (``overwrite``).
The ``name`` argument is a string indicating the name of the mask. This name
is used to refer to that mask in junifer internally in order to obtain the
is used to refer to that mask in ``junifer`` internally in order to obtain the
actual mask data and perform operations on it. For example, using the name you
can load a mask after registration using the
:func:`.load_mask` function.
The ``mask_path`` should contain the path to a valid NIfTI image with binary
voxel values (i.e. 0 or 1). This data can then be used by junifer to mask other
MR images.
voxel values (i.e. 0 or 1). This data can then be used by ``junifer`` to mask
other MR images.
Lastly, we specify the ``space`` that the coordinates are in, for example,
``"MNI"`` or ``"Native"`` (scanner-native space).
``"MNI152NLin6Asym"`` or ``"native"`` (scanner-native space).
Step 1: Prepare code to register a mask
---------------------------------------
@ -49,7 +49,7 @@ look as follows:
# on your system:
mask_path = Path("..") / ".." / "my_custom_mask.nii.gz"
register_mask(name="my_custom_mask", mask_path=mask_path, space="Native")
register_mask(name="my_custom_mask", mask_path=mask_path, space="native")
Simple, right? Now we just have to configure a YAML file to register this mask
so we can use it for :ref:`codeless configuration of junifer <codeless>`.
@ -57,14 +57,14 @@ so we can use it for :ref:`codeless configuration of junifer <codeless>`.
Step 2: Configure a YAML file for registration of a mask
--------------------------------------------------------
In order to do this, we can use the ``with`` keyword provided by junifer:
In order to do this, we can use the ``with`` keyword provided by ``junifer``:
.. code-block:: yaml
with:
- register_custom_mask.py
Then we can use this mask for any processing step or marker that takes in a
Then we can use this mask for any processing step or Marker that takes in a
mask as an argument. For example:
.. code-block:: yaml
@ -76,12 +76,15 @@ mask as an argument. For example:
method: mean
masks: "my_custom_mask"
Now, you can simply use this YAML file to run your pipeline. One important
point to keep in mind is that if the paths given in ``register_custom_mask.py``
are relative paths, they will be interpreted by junifer as relative to the
jobs directory (i.e. where junifer will create submit files, logs directory and
so on). For simplicity, you may just want to use absolute paths to avoid
confusion, yet using relative paths is likely a better way to make your
pipeline directory/repository more portable and therefore more reproducible for
others. Really, once you understand how these paths are interpreted by junifer,
it is quite easy.
Now, you can simply use this YAML file to run your pipeline.
.. important::
It's important to keep in mind that if the paths given in
``register_custom_mask.py`` are relative paths, they will be interpreted
by junifer as relative to the jobs directory (i.e. where ``junifer`` will
create submit files, logs directory and so on). For simplicity, you may just
want to use absolute paths to avoid confusion, yet using relative paths is
likely a better way to make your pipeline directory / repository more portable
and therefore more reproducible for others. Really, once you understand how
paths are interpreted by ``junifer``, it is quite easy.

View file

@ -5,28 +5,29 @@
Adding Parcellations
====================
Before you start adding your own parcellations, check whether junifer has
Before you start adding your own parcellations, check whether ``junifer`` has
the parcellation :ref:`in-built already <builtin>`. Perhaps, what is available
there will suffice to achieve your goals. However, of course junifer will not
there will suffice to achieve your goals. However, of course ``junifer`` will not
have every parcellation available that you may want to use, and if so, it will
be nice to be able to add it yourself using a format that junifer understands.
be nice to be able to add it yourself using a format that ``junifer`` understands.
Similarly, you may even be interested in creating your own custom parcellations
and then adding them to junifer, so you can use junifer to obtain different
markers to assess and validate your own parcellation. So, how can you do this?
and then adding them to ``junifer``, so you can use ``junifer`` to obtain
different Markers to assess and validate your own parcellation. So, how can you do
this?
Since both of these use-cases are quite common, and not being able to use your
favourite parcellation is of course quite a buzzkill, junifer actually provides
the easy-to-use :func:`.register_parcellation` function to do just that. Let's
try to understand the API reference and then use this function to register our
favourite parcellation is of course quite a buzzkill, ``junifer`` actually
provides the easy-to-use :func:`.register_parcellation` function to do just that.
Let's try to understand the API reference and then use this function to register our
own parcellation.
From the API reference, we can see that it has 4 positional arguments
(``name``, ``parcellation_path``, ``parcels_labels`` and ``space``) as well as
one optional keyword argument (``overwrite``).
The ``name`` of the parcellation is up to you and will be the name that junifer
will use to refer to this particular parcellation. You can think of this as
being similar to a key in a python dictionary, i.e. a key that is used to
The ``name`` of the parcellation is up to you and will be the name that
``junifer`` will use to refer to this particular parcellation. You can think of
this as being similar to a key in a Python dictionary, i.e. a key that is used to
obtain and operate on the actual parcellation data. This ``name`` must always
be a string. For example, we could call our parcellation
``"my_custom_parcellation"`` (Note, that in a real-world use case this is
@ -52,14 +53,14 @@ first label in this list corresponds to the first integer label in the
parcellation and so on).
LeSasse commented 2024-03-28 13:04:36 +00:00 (Migrated from github.com)

"For example, a simple example could look like this:" -> "A simple example could look like this:"

"For example, a simple example could look like this:" -> "A simple example could look like this:"
Lastly, we specify the ``space`` that the parcellation is in, for example,
``"MNI"`` or ``"Native"`` (scanner-native space).
``"MNI152NLin2009cAsym"`` or ``"native"`` (scanner-native space).
Step 1: Prepare code to register a parcellation
-----------------------------------------------
Now we know everything that we need to know to make sure junifer can use our
own parcellation to compute any parcellation-based marker. For example,
a simple example could look like this:
Now we know everything that we need to know to make sure ``junifer`` can use our
own parcellation to compute any parcellation-based Marker. A simple example could
look like this:
.. code-block:: python
@ -82,20 +83,20 @@ a simple example could look like this:
name="my_custom_parcellation",
parcellation_path=path_to_parcellation,
parcels_labels=my_labels,
space="MNI"
space="MNI152NLin2009cAsym"
)
We can run this code and it seems to work, however, how can we actually
include the custom parcellation in a junifer pipeline using a
include the custom parcellation in a ``junifer`` pipeline using a
:ref:`code-less YAML configuration <codeless>`?
Step 2: Add parcellation registration to the YAML file
------------------------------------------------------
In order to use the parcellation in a junifer pipeline configured by a YAML
file, we can save the above code in a python file, say
In order to use the parcellation in a ``junifer`` pipeline configured by a YAML
file, we can save the above code in a Python file, say
``registering_my_parcellation.py``. We can then simply add this file using the
``with`` keyword provided by junifer:
``with`` keyword provided by ``junifer``:
.. code-block:: yaml
@ -105,7 +106,7 @@ file, we can save the above code in a python file, say
Afterwards continue configuring the rest of the pipeline in this YAML file, and
you will be able to use this parcellation using the name you gave the
parcellation when registering it. For example, we can add a
:class:`.ParcelAggregation` marker to demonstrate how this can be done:
:class:`.ParcelAggregation` Marker to demonstrate how this can be done:
.. code-block:: yaml
@ -115,12 +116,15 @@ parcellation when registering it. For example, we can add a
parcellation: my_custom_parcellation
method: mean
Now, you can simply use this YAML file to run your pipeline. One important
point to keep in mind is that if the paths given in
``registering_my_parcellation.py`` are relative paths, they will be interpreted
by junifer as relative to the jobs directory (i.e. where junifer will create
submit files, logs directory and so on). For simplicity, you may just want to
use absolute paths to avoid confusion, yet using relative paths is likely a
better way to make your pipeline directory/repository more portable and
therefore more reproducible for others. Really, once you understand how these
paths are interpreted by junifer, it is quite easy.
Now, you can simply use this YAML file to run your pipeline.
.. important::
It's important to keep in mind that if the paths given in
``registering_my_parcellation.py`` are relative paths, they will be interpreted
by ``junifer`` as relative to the jobs directory (i.e. where ``junifer`` will
create submit files, logs directory and so on). For simplicity, you may just
want to use absolute paths to avoid confusion, yet using relative paths is
likely a better way to make your pipeline directory / repository more portable
and therefore more reproducible for others. Really, once you understand how
these paths are interpreted by ``junifer``, it is quite easy.

View file

@ -0,0 +1,246 @@
.. include:: ../links.inc
.. _extending_preprocessors:
Creating Preprocessors
======================
As already mentioned in the introduction, ``junifer`` does not do traditional
MRI pre-processing but can perform minimal preprocessing of the data that the
DataGrabber provides, for example, smoothing after confound regression or
transforming data to subject-native space before feature extraction. While
there are a few Preprocessors available already and we are constantly adding
new ones, you might need something specific and then you can create your
own Preprocessor.
While implementing your own Preprocessor, you need to always inherit from
:class:`.BasePreprocessor` and implement a few methods:
#. ``get_valid_inputs``: This method should return a list of strings
representing the valid data types that the Preprocessor can work on.
Check :ref:`data types <data_types>` for reference.
#. ``get_output_type``: This method should just return the input as it
is unused as of now.
#. ``preprocess``: The method that given the data, preprocesses the data.
#. ``__init__``: The initialisation method, where the Preprocessor is
configured.
As an example, we will develop a ``NilearnSmoothing`` Preprocessor, which
smoothens the data using :func:`nilearn.image.smooth_img`. This is often
desirable in cases where your data is preprocessed using ``fMRIPrep``, as
``fMRIPrep`` does not perform smoothing.
.. _extending_preprocessors_input_output:
Step 1: Configure input and output
----------------------------------
In this step, we define the input and output data types of the Preprocessor.
For input we can accept ``T1w``, ``T2w`` and ``BOLD``
:ref:`data types <data_types>`.
.. code-block:: python
...
def get_valid_inputs(self) -> list[str]:
return ["T1w", "T2w", "BOLD"]
...
The output definition of the Preprocessor is unused now but is kept for
completeness.
.. code-block:: python
...
def get_output_type(self, input_type: str) -> str:
return input_type
...
.. _extending_preprocessors_init:
Step 2: Initialise the Preprocessor
-----------------------------------
Now we need to define our Preprocessor class' constructor which is also how
you configure it. Our class will have the following arguments:
1. ``fwhm``: The smoothing strength as a full-width at half maximum
(in millimetres). Since we depend on :func:`nilearn.image.smooth_img`, we
pass the value to it.
2. ``on``: The data type we want the Preprocessor to work on. If the user does
not specify, it will work on all the data types given by the
``get_valid_inputs`` function.
.. attention::
Only basic types (*int*, *bool* and *str*), lists, tuples and dictionaries
are allowed as parameters. This is because the parameters are stored in
JSON format, and JSON only supports these types.
.. code-block:: python
from typing import Literal
from numpy.typing import ArrayLike
...
def __init__(
self,
fwhm: int | float | ArrayLike | Literal["fast"] | None,
on: str | list[str] | None = None,
) -> None:
self.fwhm = fwhm
super().__init__(on=on)
...
.. caution::
Parameters of the Preprocessor must be stored as object attributes without
using ``_`` as prefix. This is because any attribute that starts with ``_``
will not be considered as a parameter and not stored as part of the metadata
of the Preprocessor.
.. _extending_preprocessors_preprocess:
Step 3: Preprocess the data
---------------------------
Finally, we will write the actual logic of the Preprocessor. This method will
be called by ``junifer`` when needed, using the data provided by the
DataGrabber, as configured by the user. The method ``preprocess`` has two
arguments:
* ``input``: A dictionary with the data to be used by the Preprocessor. This
will be the corresponding element in the :ref:`Data Object<data_object>`
already indexed. Thus, the dictionary has at least two keys: ``data`` and
``path``. The first one contains the data, while the second one contains the
path to the data. The dictionary can also contain other keys, depending on the
data type.
* ``extra_input``: The rest of the :ref:`Data Object<data_object>`. This is
useful if you want to use other data (e.g., ``Warp`` can be used to provide
the transformation matrix file for transformation to subject-native space).
and it has two return values:
* First is the ``input`` dictionary with necessary data modified. Usually, you
want to replace the ``input["data"]`` with the preprocessed data.
* Second is a dictionary just like ``input`` or ``extra_input`` but with only
specific key-value pairs which you would like to pass down to the Markers.
For example, if your Preprocessor computes some mask with the preprocessed
data, you could pass it through this which would be added and available
in the Marker step with the same key you pass here. Usually, you would
want to pass ``None``.
.. code-block:: python
from typing import Any
from nilearn import image as nimg
...
def preprocess(
self,
input: dict[str, Any],
extra_input: dict[str, Any] | None = None,
) -> tuple[dict[str, Any], dict[str, Any] | None]:
input["data"] = nimg.smooth_img(imgs=input["data"], fwhm=self.fwhm)
return input, None
...
Step 4: Finalise the Preprocessor
LeSasse commented 2024-03-28 13:07:17 +00:00 (Migrated from github.com)

I believe in multiple places I saw the american spelling of things, so this should be Finalize. Can we record in some central place i.e. something like "How to contribute to docs" that we aim to use the american spelling?

I believe in multiple places I saw the american spelling of things, so this should be Finalize. Can we record in some central place i.e. something like "How to contribute to docs" that we aim to use the american spelling?
synchon commented 2024-03-28 13:53:18 +00:00 (Migrated from github.com)

I've made everything follow British English in the docs. Recording it is a good idea.

I've made everything follow British English in the docs. Recording it is a good idea.
LeSasse commented 2024-03-28 13:54:34 +00:00 (Migrated from github.com)

Ok, i will point out the american way when i find them.

Ok, i will point out the american way when i find them.
---------------------------------
Now we just need to combine everything we have above and throw in a couple of
other stuff to get our Preprocessor ready.
First, we specify the :ref:`dependencies <specifying_dependencies>` for our
class, which are basically the packages that are required by the class. This is
used for validation before running to ensure all the packages are installed and
also to keep track of the dependencies and their versions in the metadata. We
define it using a class attribute like so:
.. code-block:: python
_DEPENDENCIES = {"nilearn"}
Then, we just need to register the Preprocessor using ``@register_preprocessor``
decorator and our final code should look like this:
.. code-block:: python
from typing import Any, Literal
from junifer.api.decorators import register_preprocessor
from junifer.preprocess import BasePreprocessor
from nilearn import image as nimg
from numpy.typing import ArrayLike
@register_preprocessor
class NilearnSmoothing(BasePreprocessor):
_DEPENDENCIES = {"nilearn"}
def __init__(
self,
fwhm: int | float | ArrayLike | Literal["fast"] | None,
on: str | list[str] | None = None,
) -> None:
self.fwhm = fwhm
super().__init__(on=on)
def get_valid_inputs(self) -> list[str]:
return ["T1w", "T2w", "BOLD"]
def get_output_type(self, input_type: str) -> str:
return input_type
def preprocess(
self,
input: dict[str, Any],
extra_input: dict[str, Any] | None = None,
) -> tuple[dict[str, Any], dict[str, Any] | None]:
input["data"] = nimg.smooth_img(imgs=input["data"], fwhm=self.fwhm)
return input, None
.. _extending_preprocessors_template:
Template for a custom Preprocessor
----------------------------------
.. code-block:: python
from junifer.api.decorators import register_preprocessor
from junifer.preprocess import BasePreprocessor
@register_preprocessor
class TemplatePreprocessor(BasePreprocessor):
def __init__(self, on=None):
# TODO: add preprocessor-specific parameters
super().__init__(on=on)
def get_valid_inputs(self):
# TODO: Complete with the valid inputs
valid = []
return valid
def get_output_type(self, input_type):
return input_type
def preprocess(self, input, extra_input):
# TODO: add the preprocessor logic
return input, None

View file

@ -1,35 +1,13 @@
.. include:: links.inc
Installing junifer
==================
Requirements
------------
junifer is compatible with `Python`_ >= 3.8 and requires the following packages:
* ``click>=8.1.3,<8.2``
* ``numpy>=1.24,<1.27``
* ``datalad>=0.15.4,<0.20``
* ``pandas>=1.4.0,<2.2``
* ``nibabel>=3.2.0,<5.11``
* ``nilearn>=0.9.0,<=0.11.0``
* ``sqlalchemy>=1.4.27,<= 2.1.0``
* ``ruamel.yaml>=0.17,<0.18``
* ``h5py>=3.8.0,<3.10``
Depending on the installation method, these packages might be installed
automatically.
Installation
------------
Installing ``junifer``
======================
Depending on your use-case, ``junifer`` can be installed differently:
* Install the :ref:`install_latest_release`. This is the most suitable approach
* Install the :ref:`latest stable release <install_latest_release>`. This is the most suitable approach
for end users.
* Install from :ref:`install_development_git`. This is the most suitable approach
* Install from :ref:`latest development release <install_development_git>`. This is the most suitable approach
for developers.
@ -39,8 +17,8 @@ Either way, we strongly recommend using
.. _install_latest_release:
Stable release
~~~~~~~~~~~~~~
Using a package manager
-----------------------
Use ``pip`` to install ``junifer`` from `PyPI <https://pypi.org>`_, like so:
@ -64,8 +42,8 @@ You can also install via ``conda``, like so:
.. _install_development_git:
Local Git repository
~~~~~~~~~~~~~~~~~~~~
From the source
---------------
Follow the `detailed contribution guidelines <contribution.rst>`_.
@ -81,11 +59,11 @@ that are required for specific markers.
.. important::
The Docker container wrappers add the commands required by junifer. Using
The Docker container wrappers add the commands required by ``junifer``. Using
these commands have some limitations, mostly related to handling files and
paths. Junifer knows about this and uses these commands in the proper way.
paths. ``junifer`` knows about this and uses these commands in the proper way.
Keep this in mind if you try to use the Docker wrappers outside of
junifer. These caveats and limitations are not documented.
``junifer``. These caveats and limitations are not documented.
AFNI
----

View file

@ -23,6 +23,7 @@
.. _`nilearn`: https://nilearn.github.io
.. _`nipype`: https://nipype.readthedocs.io
.. _`datalad`: https://datalad.org
.. _`templateflow`: https://www.templateflow.org
.. _`venv`: https://docs.python.org/3/tutorial/venv.html
.. _`conda env`: https://docs.conda.io/projects/conda/en/latest/user-guide/tasks/manage-environments.html

View file

@ -2,8 +2,8 @@
.. _starting:
First steps with junifer
========================
First steps with ``junifer``
============================
.. note::
@ -13,7 +13,7 @@ First steps with junifer
flowchart TD
start((( Start )))
read_understanding(Read Understanding Junifer)
read_understanding(Read Understanding junifer)
start --> read_understanding
question_features{Can I compute\nthe features I want\nusing junifer?}
read_understanding --> question_features
@ -22,7 +22,7 @@ First steps with junifer
question_features -->|No| question_feature_type
read_using --> question_features
read_using(Read Using Junifer)
read_using(Read Using junifer)
question_feature_type{"What am I missing\nfrom junifer"}
question_feature_type --> missing_datagrabber
@ -30,7 +30,7 @@ First steps with junifer
question_feature_type --> missing_marker
question_feature_type --> missing_other
missing_datagrabber("A Dataset/DataGrabber")
missing_preprocessing("A Preprocessing")
missing_preprocessing("A Preprocessor")
missing_marker("A Marker")
missing_other("Something else")
@ -38,7 +38,7 @@ First steps with junifer
question_datagrabber_junifarm{Is the\ndataset/datagrabber\nin juni-farm?}
question_datagrabber_junifarm -->|Yes| read_using_final
question_datagrabber_junifarm -->|No| read_extending_datagrabber_start
read_extending_datagrabber_start(Read Creating a Junifer extension)
read_extending_datagrabber_start(Read Creating a junifer extension)
read_extending_datagrabber_start --> read_extending_datagrabber
read_extending_datagrabber(Read Creating Data Grabbers)
read_extending_datagrabber --> question_datagrabber_kind
@ -55,14 +55,14 @@ First steps with junifer
question_contribute_datagrabber{Do you think\nyour DataGrabber\nis useful for other users?}
question_contribute_datagrabber -->|Yes| contribute_datagrabber
question_contribute_datagrabber -->|No| final_run
contribute_datagrabber(Create a\nDATASET REQUEST\nissue on Github)
contribute_datagrabber(Create a\nDATASET REQUEST\nissue on GitHub)
contribute_datagrabber --> final_run
missing_marker --> question_marker_junifarm
question_marker_junifarm{Is the marker\nin juni-farm?}
question_marker_junifarm -->|Yes| read_using_final
question_marker_junifarm -->|No| read_extending_marker_start
read_extending_marker_start(Read Creating a Junifer extension)
read_extending_marker_start(Read Creating a junifer extension)
read_extending_marker_start --> read_extending_marker
read_extending_marker(Read Creating Markers)
read_extending_marker --> question_marker_solved
@ -72,7 +72,7 @@ First steps with junifer
question_contribute_marker{Do you think\nyour Marker\nis useful for other users?}
question_contribute_marker -->|Yes| contribute_marker
question_contribute_marker -->|No| final_run
contribute_marker(Create a\nMARKER REQUEST\nissue on Github)
contribute_marker(Create a\nMARKER REQUEST\nissue on GitHub)
contribute_marker --> final_run
missing_preprocessing --> contact_help
@ -87,22 +87,22 @@ First steps with junifer
missing_other --> missing_other_other
missing_other_other --> contact_help
contact_help(((Contact the\nJunifer team)))
contact_help(((Contact the\njunifer team)))
missing_mask --> read_adding_mask_start
read_adding_mask_start("Read Creating a Junifer extension")
read_adding_mask_start("Read Creating a junifer extension")
read_adding_mask_start --> read_adding_mask
read_adding_mask("Read Adding Masks")
read_adding_mask --> missing_other_solved
missing_parcellation --> read_adding_parcellation_start
read_adding_parcellation_start("Read Creating a Junifer extension")
read_adding_parcellation_start("Read Creating a junifer extension")
read_adding_parcellation_start --> read_adding_parcellation
read_adding_parcellation("Read Adding Parcellations")
read_adding_parcellation --> missing_other_solved
missing_coordinates --> read_adding_coordinates_start
read_adding_coordinates_start("Read Creating a Junifer extension")
read_adding_coordinates_start("Read Creating a junifer extension")
read_adding_coordinates_start --> read_adding_coordinates
read_adding_coordinates("Read Adding Coordinates")
read_adding_coordinates --> missing_other_solved
@ -110,11 +110,11 @@ First steps with junifer
missing_other_solved{Did you solve your issue?}
missing_other_solved -->|Yes| read_using_final
missing_other_solved -->|No| missing_other_contact
missing_other_contact(Contact the\nJunifer team)
missing_other_contact(Contact the\njunifer team)
missing_other_contact --> missing_other_issue
missing_other_issue(((Submit a\nFEATURE REQUEST\nissue in Github)))
missing_other_issue(((Submit a\nFEATURE REQUEST\nissue in GitHub)))
read_using_final(Read Using Junifer)
read_using_final(Read Using junifer)
read_using_final --> final_yaml
final_yaml(Create/edit the YAML file)
final_yaml --> final_run
@ -125,9 +125,9 @@ First steps with junifer
question_error_run{"Is it an issue\nwith my YAML file?"}
question_error_run -->|Yes| final_yaml
question_error_run -->|No| error_contact
error_contact(Contact the\nJunifer team)
error_contact(Contact the\njunifer team)
error_contact --> error_issue
error_issue(((Submit a\nBUG REPORT issue\nin Github)))
error_issue(((Submit a\nBUG REPORT issue\nin GitHub)))
question_final_run_worked -->|Yes| final_queue
final_queue(Use junifer queue to compute your features)
final_queue --> final_magic

View file

@ -105,7 +105,7 @@ Data Types
- Preprocessed or Raw T1w image
* - ``BOLD``
- BOLD image (4D)
- Preprocessed/Denoised BOLD image (fmriprep output)
- Preprocessed or Denoised BOLD image (fMRIPrep output)
* - ``BOLD_confounds``
- BOLD image confounds (CSV/TSV file)
- Confounds that can be applied to the BOLD image.

View file

@ -9,9 +9,9 @@ Description
-----------
The ``DataGrabber`` is an object that can provide an interface to datasets you
want to work with in junifer. Every concrete implementation of a DataGrabber is
aware of a particular dataset's structure and thus allows you to fetch specific
elements of interest from the dataset. It adds the ``path`` key to each
want to work with in ``junifer``. Every concrete implementation of a DataGrabber
is aware of a particular dataset's structure and thus allows you to fetch
specific elements of interest from the dataset. It adds the ``path`` key to each
:ref:`data type <data_types>` in the :ref:`Data object <data_object>`.
DataGrabbers are intended to be used as context managers. When used within a
@ -20,18 +20,18 @@ the dataset, for example, downloading and cleaning up. As the interface
is consistent, you always use the same procedure to interact with the DataGrabber.
For example, a concrete implementation of :class:`.DataladDataGrabber` can
provide junifer with data from a Datalad dataset. Of course, DataGrabbers are not
only meant to work with Datalad datasets but any dataset.
provide ``junifer`` with data from a Datalad dataset. Of course, DataGrabbers are
not only meant to work with Datalad datasets but any dataset.
If you are interested in using already provided DataGrabbers, please go to
:doc:`../builtin`. And, if you want to implement your own DataGrabber, you need
to provide concrete implementations of base classes already provided.
to provide concrete implementations of abstract base classes already provided.
Base Classes
------------
In this section, we showcase different abstract base classes you might want to
use to implement your own DataGrabber.
In this section, we showcase different abstract and concrete base classes you
might want to use to implement your own DataGrabber.
.. list-table::
:widths: auto

View file

@ -9,17 +9,17 @@ Description
-----------
The ``DataReader`` is an object that is responsible for actually reading data
files in junifer. It reads the value of the key ``path`` for each
files in ``junifer``. It reads the value of the key ``path`` for each
:ref:`data type <data_types>` in the :ref:`Data object <data_object>` and loads
them into memory. After reading the data into memory, it adds the key ``data``
to the same level as ``path`` and the value is the actual data in the memory.
DataReaders are meant to be used inside the datagrabber context but you can
DataReaders are meant to be used inside the DataGrabber context but you can
operate on them outside the context as long as the actual data is in the memory
and the Python runtime has not garbage-collected it.
For data formats not supported by junifer yet, you can either make your own
*Data Reader* or open an issue on `junifer Github`_ and we can help you out.
For data formats not supported by ``junifer`` yet, you can either make your own
DataReader or open an issue on `junifer Github`_ and we can help you out.
File Formats
------------

View file

@ -2,14 +2,14 @@
.. _understanding:
Understanding junifer
=====================
Understanding ``junifer``
=========================
Before you start, you should understand how junifer works. Junifer is a
Before you start, you should understand how ``junifer`` works. ``junifer`` is a
tool conceived to extract features from neuroimaging data in an easy-to-use
manner, with minimal coding and minimal user expertise in the internal aspects.
Unlike other tools like FSL, SPM, AFNI, etc., junifer is not a toolbox to
Unlike other tools like FSL, SPM, AFNI, etc., ``junifer`` is not a toolbox to
pre-process data, but a toolbox to extract features from previously
pre-processed data.
@ -20,9 +20,9 @@ julearn_).
.. important::
Junifer is not a toolbox to create pipelines, but a tool to configure the
junifer pipeline, which is intended to be fixed and not to be changed. If you
want to create a pipeline, you should use other tools like nipype_.
``junifer`` is not a toolbox to create pipelines, but a tool to configure the
``junifer`` pipeline, which is intended to be fixed and not to be changed. If
you want to create a pipeline, you should use other tools like nipype_.
.. toctree::
:maxdepth: 2

View file

@ -21,12 +21,12 @@ the pipeline.
SPM, AFNI, etc., . For example, one can perform confound removal on loaded
data and then perform feature extraction.
Markers are meant to be used inside the datagrabber context but you can operate
Markers are meant to be used inside the DataGrabber context but you can operate
on them outside the context as long as the actual data is in the memory and the
Python runtime has not garbage-collected it.
If you are interested in using already provided markers, please go to
:doc:`../builtin`. And, if you want to implement your own marker, you need to
If you are interested in using already provided Markers, please go to
:doc:`../builtin`. And, if you want to implement your own Marker, you need to
provide concrete implementation of :class:`.BaseMarker`. Specifically, you
need to override ``get_valid_inputs``, ``get_output_type`` and ``compute``
methods.

View file

@ -2,8 +2,8 @@
.. _pipeline:
The junifer Pipeline
====================
The ``junifer`` Pipeline
========================
The junifer pipeline is the main execution path of junifer. It consists of five
steps:
@ -11,10 +11,10 @@ steps:
1. :ref:`Data Grabber <datagrabber>`: Interpret the dataset and provide a list
of files.
2. :ref:`Data Reader <datareader>`: Read the files.
3. :ref:`Pre-processing <preprocess>`: Prepare the images for marker
3. :ref:`Preprocess <preprocess>`: Prepare the files' data for marker
computation.
4. :ref:`Marker Computation <marker>`: Compute the marker.
5. :ref:`Storage <storage>`: Store the marker values.
4. :ref:`Marker Computation <marker>`: Compute the marker(s).
5. :ref:`Storage <storage>`: Store the marker(s) values.
The element that is passed across the pipeline is called the
:ref:`Data Object<data_object>`.
@ -26,7 +26,7 @@ The following is a graphical representation of the pipeline:
flowchart LR
dg[Data Grabber]
dr[Data Reader]
pp[Pre-processing]
pp[Preprocess]
mc[Marker Computation]
st[Storage]
dg --> dr
@ -45,7 +45,7 @@ on multiple markers:
flowchart LR
dg[Data Grabber]
dr[Data Reader]
pp[Pre-processing]
pp[Preprocess]
mc1[Marker Computation]
mc2[Marker Computation]
mc3[Marker Computation]

View file

@ -22,12 +22,12 @@ want to perform confound removal on ``BOLD`` data before feature extraction.
Confound Removal
----------------
The *Confound Removal* step is meant to remove *confounds* from the ``BOLD``
data. The confounds are extracted from the ``BOLD_confounds`` data (must be
provided by the :ref:`Data Grabber <datagrabber>`). The confounds are then
regressed out from the ``BOLD`` data using :func:`nilearn.image.clean_img`.
This step is meant to remove *confounds* from the ``BOLD`` data. The confounds
are extracted from the ``BOLD_confounds`` data (must be provided by the
:ref:`Data Grabber <datagrabber>`). The confounds are then regressed out from
the ``BOLD`` data using :func:`nilearn.image.clean_img`.
Currently, junifer supports only one confound removal class:
Currently, ``junifer`` supports only one confound removal class:
:class:`.fMRIPrepConfoundRemover`. This class is meant to remove confounds as
described before, using the output of `fMRIPrep`_ as reference.
@ -128,3 +128,69 @@ parameters:
| If not, a mask is computed using
| :func:`nilearn.masking.compute_brain_mask`.
- compute
.. _preprocess_warping:
Warping or Transformation to other spaces
-----------------------------------------
``junifer`` can also warp or transform any supported
:ref:`data type <data_types>` from the template space provided by the dataset
(e.g., ``MNI152NLin6Asym``) to either the subject's
:ref:`native space <preprocess_warping_native>` or to any other
:ref:`template space <preprocess_warping_template>`
(e.g., ``MNI152NLin2009cAsym``). This functionality is provided by
:class:`.SpaceWarper` and depends on external tools like FSL and / or ANTs.
.. _preprocess_warping_native:
Warping to subject's native space
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
To warp to subject's native space, the dataset needs to provide ``T1w`` and
``Warp`` data types and the DataGrabber needs to at least have
``["BOLD", "T1w", "Warp"]`` (if you are warping ``BOLD``) as the ``types``
parameter's value. The :class:`.SpaceWarper`'s ``reference`` parameter needs
to be set to ``T1w``, which means that the ``BOLD`` data will be transformed
using the ``T1w`` as reference (it's resampled internally to match the
resolution of the ``BOLD``). The ``Warp`` data type is new and it's only purpose
is to provide the warp or transformation file (can be linear, non-linear or
linear + non-linear transform) for the purpose. For ``using`` parameter, you can
pass either ``"fsl"`` or ``"ants"`` depending on the warp or transformation file
format.
An example YAML might look like this:
.. code-block:: yaml
preprocess:
- kind: SpaceWarper
using: fsl
reference: T1w
.. _preprocess_warping_template:
Warping to other template space
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
In a situation where your dataset might provide the ``BOLD`` data (or any other
data type that you want to work on) in ``MNI152NLin6Asym`` template space but
you would like to compute features in ``MNI152NLin2009cAsym`` template space,
you can also use the :class:`.SpaceWarper` by setting the ``reference``
parameter to the template space's name, in this case,
``reference="MNI152NLin2009cAsym"``. The ``using`` parameter needs to be set
to ``"ants"`` as we need it to warp the data.
.. note::
We only support template spaces provided by `templateflow`_ and the naming
is similar except that we omit the ``tpl-`` prefix used by ``templateflow``.
For an YAML example:
.. code-block:: yaml
preprocess:
- kind: SpaceWarper
using: ants
reference: MNI152NLin2009cAsym

View file

@ -13,7 +13,7 @@ as computed from :ref:`Marker <marker>` step of the pipeline. If the pipeline is
provided with a ``storage-like`` object, the extracted features are stored via
that object else they are kept in memory.
Storage is meant to be used inside the datagrabber context but you can operate
Storage is meant to be used inside the DataGrabber context but you can operate
on them outside the context as long as the processed data is in the memory and
the Python runtime has not garbage-collected it.
@ -25,8 +25,8 @@ storage object in turn declares and provides implementation for specific
``matrix``, ``vector`` and ``timeseries`` via ``store_matrix``, ``store_vector``
and ``store_timeseries`` methods respectively.
For storage interfaces not supported by junifer yet, you can either make your
own ``Storage`` by providing a concrete implementation of
For storage interfaces not supported by ``junifer`` yet, you can either make
your own ``Storage`` by providing a concrete implementation of
:class:`.BaseFeatureStorage` or open an issue on `junifer Github`_ and we can
help you out.

View file

@ -5,7 +5,7 @@
Code-less Configuration
fraimondo commented 2024-04-01 15:22:48 +00:00 (Migrated from github.com)

"One of..."

"One of..."
=======================
On of the most important features of junifer is its capacity to run without
One of the most important features of ``junifer`` is its capacity to run without
writing a single line of code. This is achieved by using a configuration file
that is written in YAML_. In this file, we configure the different steps of
:ref:`pipeline`.
@ -17,7 +17,7 @@ As a reminder, this is how the pipeline looks like:
flowchart LR
dg[Data Grabber]
dr[Data Reader]
pp[Pre-processing]
pp[Preprocess]
mc[Marker Computation]
st[Storage]
dg --> dr
@ -31,21 +31,21 @@ as well as some general parameters.
As an example, we will generate the configuration file for a pipeline that will
extract the mean ``VBM_GM`` values using two different parcellations and one set
of coordinates, from the ``Oasis VBM Testing dataset`` included in junifer.
of coordinates, from the ``Oasis VBM Testing dataset`` included in ``junifer``.
General Parameters
------------------
The general parameters are the ones that are not specific to any of the sections
of the pipeline, but configure junifer as a whole. These parameters are:
of the pipeline, but configure ``junifer`` as a whole. These parameters are:
* ``with``: A section used to specify modules and junifer extensions to use.
* ``workdir``: The working directory where junifer will store temporary files.
* ``with``: A section used to specify modules and ``junifer`` extensions to use.
* ``workdir``: The working directory where ``junifer`` will store temporary files.
Since the example uses a specific datagrabber for testing, we need to add
``junifer.testing.registry`` to the ``with`` section. This will allow junifer
to find the datagrabber. We will set the ``workdir`` to ``/tmp``.
Since the example uses a specific DataGrabber for testing, we need to add
``junifer.testing.registry`` to the ``with`` section. This will allow ``junifer``
to find the DataGrabber. We will set the ``workdir`` to ``/tmp``.
.. code-block:: yaml
@ -66,9 +66,9 @@ In order to configure the pipeline, we need to configure each step:
.. important::
The datareader step configuration is optional, as junifer only provides one
datareader. Nevertheless, it is possible to extend junifer with custom
datareaders, and thus, it is also possible to configure this step.
The ``datareader`` step configuration is optional, as ``junifer`` only
provides one DataReader. Nevertheless, it is possible to extend ``junifer``
with custom DataReaders, and thus, it is also possible to configure this step.
Data Grabber
@ -107,9 +107,9 @@ In the ``Oasis VBM Testing dataset`` example, the section will look like this:
Data Reader
^^^^^^^^^^^
As mentioned before, this section is entirely optional, as junifer only provides
one DataReader (:class:`.DefaultDataReader`), which is the default in case the
section is not specified.
As mentioned before, this section is entirely optional, as ``junifer`` only
provides one DataReader (:class:`.DefaultDataReader`), which is the default in
case the section is not specified.
In any case, the syntax of the section is the same as for the ``datagrabber``
section, using the ``kind`` key to specify the DataReader to use, and additional
@ -127,18 +127,19 @@ For the ``Oasis VBM Testing dataset`` example, we will not specify a
Preprocess
^^^^^^^^^^
Pre-processing is also an optional step, as it might be the case that no
pre-processing is needed. In the case that pre-processing is needed, the section
must be configured using the ``kind`` key to specify the preprocessor to use,
and additional keys to pass parameters to the preprocessor.
``preprocess`` is also an optional step, as it might be the case that no
pre-processing is needed. As we can perform multiple preprocessing steps, it's
passed as a list of Preprocessors. In the case that pre-processing is needed,
each Preprocessord must be configured using the ``kind`` key to specify the
Preprocessor to use, and additional keys to pass parameters to the Preprocessor.
For example, to use the :class:`.fMRIPrepConfoundRemover` preprocessor, we just
For example, to use the :class:`.fMRIPrepConfoundRemover` Preprocessor, we just
need to specify its name as the ``kind`` key, as well as its parameters.
.. code-block:: yaml
preprocess:
kind: fMRIPrepConfoundRemover
- kind: fMRIPrepConfoundRemover
strategy:
motion: full
wm_csf: full
@ -155,9 +156,9 @@ preprocessing step.
Marker
^^^^^^
The ``markers`` section diverges from the previous ones, as we need to specify
a list of markers. Each marker has a name that we can use to refer to it later,
and a set of parameters that will be passed to the marker.
The ``markers`` section like the ``preprocess`` section expects a list of
markers. Each Marker has a name that we can use to refer to it later,
and a set of parameters that will be passed to the Marker.
For the ``Oasis VBM Testing dataset`` example, we want to compute the mean
``VBM_GM`` value for each parcel using the ``Schaefer parcellation (100 parcels,
@ -190,14 +191,14 @@ Finally, we need to define how and where the results will be stored. This is
done using the ``storage`` section, which must be configured using the ``kind``
key to specify the storage to use, and additional keys to pass parameters.
For example, to use the :class:`.SQLiteFeatureStorage` storage, we just need to
For example, to use the :class:`.HDF5FeatureStorage` storage, we just need to
specify where we want to store the results:
.. code-block:: yaml
storage:
kind: SQLiteFeatureStorage
uri: /data/junifer/example/oasis_vbm_testing.sqlite
kind: HDF5FeatureStorage
uri: /data/junifer/example/oasis_vbm_testing.hdf5
Complete Example
@ -231,5 +232,5 @@ looks like:
method: mean
storage:
kind: SQLiteFeatureStorage
uri: /data/junifer/example/oasis_vbm_testing.sqlite
kind: HDF5FeatureStorage
uri: /data/junifer/example/oasis_vbm_testing.hdf5

View file

@ -2,14 +2,14 @@
.. _using:
Using junifer
=============
Using ``junifer``
=================
In this section, we will cover the main aspects behind using junifer. We will
first explain the basics behind junifer's code-less configuration. Then we will
show how to use the command line interface to ``run`` junifer and ``collect``
the results. Finally, we will show how to use the ``queue`` command to interact
with HPC and HTC systems.
In this section, we will cover the main aspects behind using ``junifer``. We
will first explain the basics behind junifer's code-less configuration. Then we
will show how to use the command line interface to ``run`` the pipeline and
``collect`` the results. Finally, we will show how to use the ``queue`` command
to interact with HPC and HTC systems.
.. toctree::
:maxdepth: 2
@ -25,7 +25,7 @@ with HPC and HTC systems.
Using Common Components
-----------------------
The following sections explains common components of junifer that can be used
The following sections explains common components of ``junifer`` that can be used
across many steps of the pipeline.
.. toctree::

View file

@ -12,7 +12,7 @@ contain a certain ratio of gray matter to white matter / cerebrospinal fluid,
ensuring that the features are not extracted from voxels that contain mostly
white matter or cerebrospinal fluid, which could add noise to the BOLD signal.
Junifer provides a number of built-in masks, which can be listed using
``junifer`` provides a number of built-in masks, which can be listed using
:func:`.list_masks`. Some masks are images, while other masks can be computed
using :ref:`nilearn` functions.
@ -22,15 +22,14 @@ dictionary in which the **only** key is the built-in mask name and the value is
a dictionary of keyword arguments to pass to the mask function.
For example, the following is a valid mask specification that specified the
``GM_prob0.2`` mask.
``GM_prob0.2`` mask:
.. code-block:: yaml
masks: GM_prob0.2
The following is a valid mask specification that specifies the
``compute_brain_mask`` mask (function from nilearn), with a threshold of
``0.5``.
``compute_brain_mask`` mask, with a threshold of ``0.5``.
.. code-block:: yaml

View file

@ -5,14 +5,15 @@
Queueing Jobs (HPC, HTC)
========================
Yet another interesting feature of junifer is the ability to queue jobs on
Yet another interesting feature of ``junifer`` is the ability to queue jobs on
computational clusters. This is done by adding the ``queue`` section in the
:ref:`codeless` file and executing the ``junifer queue`` command.
While junifer is meant to support `HTCondor`_, `SLURM`_ and local queueing
using `GNU Parallel`_, only HTCondor is currently supported. This will be
implemented in future releases of junifer. If you are in immediate need of any of
these schedulers, please create an issue on the `junifer github`_ repository.
implemented in future releases of ``junifer``. If you are in immediate need of
any of these schedulers, please create an issue on the `junifer github`_
repository.
The ``queue`` section of the :ref:`codeless` must start by defining the
following general parameters:
@ -24,7 +25,7 @@ following general parameters:
* ``kind``: The kind of scheduler to be used. Currently, only ``HTCondor`` is
supported.
Example:
Example in YAML:
.. code-block:: yaml
@ -40,7 +41,7 @@ The rest of the parameters depend on the scheduler you are using.
HTCondor
--------
When using HTCondor, junifer will use a DAG to queue one job per element
When using HTCondor, ``junifer`` will use a DAG to queue one job per element
(``junifer run``). As an option, the DAG can include a final job
(``junifer collect``) to collect the results once all of the individual element
jobs are finished.
@ -61,12 +62,13 @@ The following parameters are available for HTCondor:
If relative path is used then it should be relative to the YAML.
* ``mem``: Memory to be used by the job. It must be provided as a string with
the units (e.g. ``2GB``).
* ``cpus``: Number of CPUs to be used by the job. It must be provided as an int.
the units (e.g., ``"2GB"``).
* ``cpus``: Number of CPUs to be used by the job. It must be provided as an
integer (e.g., ``1``).
* ``disk``: Disk space to be used by the job. It must be provided as a string
with the units (e.g. ``2GB``). Keep in mind that junifer uses a local working
directory for each job, and datalad datasets might be cloned in this temporary
directory.
with the units (e.g., ``"2GB"``). Keep in mind that ``junifer`` uses a local
working directory for each job, and datalad datasets might be cloned in this
temporary directory.
* ``extra_preamble``: Extra lines to be added to the HTCondor submit file. This
can be used to add extra parameters to the job, such as ``requirements``.
* ``collect``: This parameter allows to include a collect to the DAG to collect
@ -81,7 +83,7 @@ The following parameters are available for HTCondor:
* ``no``: Do not include a collect job to the DAG.
Example:
Example in YAML:
.. code-block:: yaml