[DOC]: Improve documentation #220

Merged
synchon merged 35 commits from update/docs into main 2023-04-17 11:52:26 +00:00
37 changed files with 934 additions and 642 deletions

View file

@ -4,18 +4,22 @@ channels:
- defaults
dependencies:
- python=3.10
- click>=8.1.3,<8.2
- numpy>=1.22,<1.23
- datalad>=0.15.4,<0.18
- pandas>=1.4.0,<1.5
- click=8.1.*
- numpy>=1.22,<1.24
- pandas>=1.4.0,<1.6
- nibabel>=3.2.0,<4.1
- nilearn>=0.9.0,<1.0
- nilearn>=0.9.0,<=0.10.0
- sqlalchemy>=1.4.27,<= 1.5.0
- pyyaml>=5.1.2,<7.0
- seaborn>=0.11.2,<0.12
- Sphinx>=5.0.2,<5.1
- sphinx-gallery>=0.10.1,<0.11
- numpydoc>=1.4.0,<1.5
- h5py=3.8.*
- seaborn=0.11.*
- Sphinx=5.3.*
- sphinx-gallery=0.11.*
- furo>=2022.9.29,<2023.0.0
- numpydoc=1.5.*
- sphinx-copybutton=0.5.*
- towncrier=22.12.*
- sphinxcontrib-mermaid=0.8.*
- tox
- ipykernel
- isort
@ -28,5 +32,5 @@ dependencies:
- codespell
- pip
- pip:
- sphinx-rtd-theme>=1.0.0,<1.1
- sphinx-multiversion>=0.2.4,<0.3
- datalad>=0.15.4,<0.19
- julearn==0.2.5

View file

@ -2,12 +2,12 @@
.. _builtin:
Built-in Pipeline steps and data
================================
Built-in Pipeline Components
============================
Data Grabbers
-------------
Data Grabber
------------
..
Provide a list of the DataGrabbers that are implemented or planned.
@ -50,13 +50,15 @@ Available
- Done
- 0.0.1
* - :class:`.JuselessDataladUKBVBM`
- UKB VBM dataset preprocessed with CAT. Available for Juseless only.
- | UKB VBM dataset preprocessed with CAT.
| Available for Juseless only.
- Restricted
- ``junifer.configs.juseless``
- Done
- 0.0.1
* - :class:`.JuselessDataladCamCANVBM`
- CamCAN VBM dataset preprocessed with CAT. Available for Juseless only.
- | CamCAN VBM dataset preprocessed with CAT.
| Available for Juseless only.
- Restricted
- ``junifer.configs.juseless``
- Done
@ -80,19 +82,22 @@ Available
- Done
- 0.0.1
* - :class:`.JuselessDataladAOMICID1000VBM`
- AOMIC ID1000 VBM dataset. Available for Juseless only.
- | AOMIC ID1000 VBM dataset.
| Available for Juseless only.
- Restricted
- ``junifer.configs.juseless``
- Done
- 0.0.1
* - :class:`.JuselessDataladIXIVBM`
- `IXI VBM dataset <https://brain-development.org/ixi-dataset/>`_. Available for Juseless only.
- | `IXI VBM dataset <https://brain-development.org/ixi-dataset/>`_.
| Available for Juseless only.
- Restricted
- ``junifer.configs.juseless``
- Done
- 0.0.1
* - :class:`.JuselessUCLA`
- UCLA fMRIPrep dataset. Available for Juseless only.
- | UCLA fMRIPrep dataset.
| Available for Juseless only.
- Restricted
- ``junifer.configs.juseless``
- Done
@ -117,8 +122,8 @@ Planned
- :gh:`47`
Markers
-------
Marker
------
..
Provide a list of the Markers that are implemented or planned.
@ -185,13 +190,15 @@ Available
- Done
- 0.0.1
* - :class:`.EdgeCentricFCParcels`
- Calculate edge-centric functional connectivity over parcellation, as found in
`Jo et al. (2021) <https://doi.org/10.1016/j.neuroimage.2021.118204>`_
- | Calculate edge-centric functional connectivity over parcellation, as
| found in
| `Jo et al. (2021) <https://doi.org/10.1016/j.neuroimage.2021.118204>`_
- Done
- 0.0.2
* - :class:`.EdgeCentricFCSpheres`
- Calculate edge-centric functional connectivity over spheres placed on coordinates,
as found in `Jo et al. (2021) <https://doi.org/10.1016/j.neuroimage.2021.118204>`_
- | Calculate edge-centric functional connectivity over spheres placed on
| coordinates, as found in
| `Jo et al. (2021) <https://doi.org/10.1016/j.neuroimage.2021.118204>`_
- Done
- 0.0.2
* - :class:`.TemporalSNRParcels`
@ -199,7 +206,8 @@ Available
- Done
- 0.0.2
* - :class:`.TemporalSNRSpheres`
- Calculate temporal signal-to-noise ratio using spheres placed on coordinates
- | Calculate temporal signal-to-noise ratio using spheres placed on
| coordinates
- Done
- 0.0.2
@ -217,11 +225,12 @@ Planned
- Compute connectedness
- :gh:`34`
* - Permutation entropy, Range entropy, Multiscale entropy and Hurst exponent
- Calculate Permutation entropy, Range entropy, Multiscale entropy and Hurst exponent
- | Calculate Permutation entropy, Range entropy, Multiscale entropy and
| Hurst exponent
- :gh:`61`
Parcellations
-------------
Parcellation
------------
..
Provide a list of the Parcellations that are implemented or planned.
@ -266,8 +275,10 @@ Available
| ``TianxS2x3TxMNI6thgeneration``, ``TianxS2x7TxMNI6thgeneration``,
| ``TianxS3x3TxMNI6thgeneration``, ``TianxS3x7TxMNI6thgeneration``,
| ``TianxS4x3TxMNI6thgeneration``, ``TianxS4x7TxMNI6thgeneration``,
| ``TianxS1x3TxMNInonlinear2009cAsym``, ``TianxS2x3TxMNInonlinear2009cAsym``,
| ``TianxS3x3TxMNInonlinear2009cAsym``, ``TianxS4x3TxMNInonlinear2009cAsym``
| ``TianxS1x3TxMNInonlinear2009cAsym``,
| ``TianxS2x3TxMNInonlinear2009cAsym``,
| ``TianxS3x3TxMNInonlinear2009cAsym``,
| ``TianxS4x3TxMNInonlinear2009cAsym``
- 0.0.1
- | Tian, Y., Margulies, D.S., Breakspear, M. et al.
| Topographic organization of the human subcortex
@ -309,7 +320,8 @@ Planned
| https://doi.org/10.1016/j.neuroimage.2013.05.081.
* - Mindboggle 101
- | Klein, A., & Tourville, J.
| 101 labeled brain images and a consistent human cortical labeling protocol.
| 101 labeled brain images and a consistent human cortical labeling
| protocol.
| Frontiers in Neuroscience (2012).
| http://doi.org/10.3389/fnins.2012.00171/abstract
* - Destrieux
@ -326,13 +338,13 @@ Planned
| https://doi.org/10.1093/cercor/bhw157
* - Buckner
- | Buckner, R.L., Krienen, F.M., Castellanos, A., Diaz, J.C., Yeo, B.T.T.
| The organization of the human cerebellum estimated by intrinsic functional
| connectivity.
| The organization of the human cerebellum estimated by intrinsic
| functional connectivity.
| Journal of Neurophysiology, Volume 106(5), Pages 2322–2345 (2011).
| https://doi.org/10.1152/jn.00339.2011
| Yeo, B.T.T., Krienen, F.M., Sepulcre, J. et al.
| The organization of the human cerebral cortex estimated by intrinsic functional
| connectivity.
| The organization of the human cerebral cortex estimated by intrinsic
| functional connectivity.
| Journal of Neurophysiology, Volume 106(3), Pages 1125–1165 (2011).
| https://doi.org/10.1152/jn.00338.2011
@ -359,17 +371,18 @@ Available
* - Cognitive action control
- ``CogAC``
- 0.0.1
- | Cieslik, E.C., Mueller, V.I., Eickhoff, C.R., Langner, R., Eickhoff, S.B.
| Three key regions for supervisory attentional control: Evidence from neuroimaging
| meta-analyses.
- | Cieslik, E.C., Mueller, V.I., Eickhoff, C.R., Langner, R.,
| Eickhoff, S.B.
| Three key regions for supervisory attentional control: Evidence from
| neuroimaging meta-analyses.
| Neuroscience & Biobehavioral Reviews, Volume 48, Pages 22-34 (2015).
| https://doi.org/10.1016/j.neubiorev.2014.11.003.
* - Cognitive action regulation
- ``CogAR``
- 0.0.1
- | Langner, R., Leiberg, S., Hoffstaedter, F., Eickhoff, S.B.
| Towards a human self-regulation system: Common and distinct neural signatures
| of emotional and behavioural control.
| Towards a human self-regulation system: Common and distinct neural
| signatures of emotional and behavioural control.
| Neuroscience & Biobehavioral Reviews, Volume 90, Pages 400-410 (2018).
| https://doi.org/10.1016/j.neubiorev.2018.04.022.
* - Default mode network
@ -381,8 +394,10 @@ Available
| Journal of neurophysiology, Volume 103(1), Pages 297-321 (2010).
| https://doi.org/10.1152/jn.00783.2009
| Buckner, R.L., Andrews‐Hanna, J.R., & Schacter, D.L.
| The brain's default network: anatomy, function, and relevance to disease.
| Annals of the New York Academy of Sciences, Volume 1124(1), Pages 1-38 (2008).
| The brain's default network: anatomy, function, and relevance to
| disease.
| Annals of the New York Academy of Sciences, Volume 1124(1), Pages 1-38
| (2008).
| https://doi.org/10.1196/annals.1440.011
* - Missing formal name
- ``eMDN``
@ -392,15 +407,16 @@ Available
- ``Empathy``
- 0.0.1
- | Bzdok, D., Schilbach, L., Vogeley, K. et al.
| Parsing the neural correlates of moral cognition: ALE meta-analysis on morality,
| theory of mind, and empathy.
| Parsing the neural correlates of moral cognition: ALE meta-analysis on
| morality, theory of mind, and empathy.
| Brain Structure and Function, Volume 217(4), Pages 783-796 (2012).
| https://doi.org/10.1007/s00429-012-0380-y
* - Extended social-affective default
- ``eSAD``
- 0.0.1
- | Amft, M., Bzdok, D., Laird, A.R. et al.
| Definition and characterization of an extended social-affective default network.
| Definition and characterization of an extended social-affective default
| network.
| Brain structure & function, Volume 220, Pages 1031–1049 (2015).
| https://doi.org/10.1007/s00429-013-0698-0
* - Extended multiple-demand network
@ -422,25 +438,28 @@ Available
- ``MultiTask``
- 0.0.1
- | Worringer, B., Langner, R., Koch, I. et al.
| Common and distinct neural correlates of dual-tasking and task-switching:
| a meta-analytic review and a neuro-cognitive processing model of human multitasking.
| Common and distinct neural correlates of dual-tasking and
| task-switching: a meta-analytic review and a neuro-cognitive processing
| model of human multitasking.
| Brain structure & function, Volume 224(5), Pages 1845–1869 (2019).
| https://doi.org/10.1007/s00429-019-01870-4
* - Physiological stress
- ``PhysioStress``
- 0.0.1
- | Kogler, L., Müller, V.I., Chang, A. et al.
| Psychosocial versus physiological stress — Meta-analyses on deactivations and
| activations of the neural correlates of stress reactions.
| Psychosocial versus physiological stress — Meta-analyses on
| deactivations and activations of the neural correlates of stress
| reactions.
| NeuroImage, Volume 119, Pages 235-251 (2015).
| https://doi.org/10.1016/j.neuroimage.2015.06.059.
* - Reward-related decision making
- ``Rew``
- 0.0.1
- | Liu, X., Hairston, J., Schrier, M., Fan, J.
| Common and distinct networks underlying reward valence and processing stages:
| A meta-analysis of functional neuroimaging studies.
| Neuroscience & Biobehavioral Reviews, Volume 35(5), Pages 1219-1236 (2011).
| Common and distinct networks underlying reward valence and processing
| stages: A meta-analysis of functional neuroimaging studies.
| Neuroscience & Biobehavioral Reviews, Volume 35(5), Pages 1219-1236
| (2011).
| https://doi.org/10.1016/j.neubiorev.2010.12.012.
* - Missing formal name
- ``Somatosensory``
@ -450,23 +469,24 @@ Available
- ``ToM``
- 0.0.1
- | Bzdok, D., Schilbach, L., Vogeley, K. et al.
| Parsing the neural correlates of moral cognition: ALE meta-analysis on morality,
| theory of mind, and empathy.
| Parsing the neural correlates of moral cognition: ALE meta-analysis on
| morality, theory of mind, and empathy.
| Brain Structure and Function, Volume 217(4), Pages 783-796 (2012).
| https://doi.org/10.1007/s00429-012-0380-y
* - Vigilant attention
- ``VigAtt``
- 0.0.1
- | Langner, R., & Eickhoff, S.B.
| Sustaining attention to simple tasks: a meta-analytic review of the neural
| mechanisms of vigilant attention.
| Sustaining attention to simple tasks: a meta-analytic review of the
| neural mechanisms of vigilant attention.
| Psychological bulletin, Volume 139 4, Pages 870-900 (2013).
| https://doi.org/10.1037/a0030694
* - Working memory
- ``WM``
- 0.0.1
- | Rottschy, C., Langner, R., Dogan, I. et al.
| Modelling neural correlates of working memory: A coordinate-based meta-analysis.
| Modelling neural correlates of working memory: A coordinate-based
| meta-analysis.
| NeuroImage, Volume 60, Pages 830-846 (2012).
| https://doi.org/10.1016/j.neuroimage.2011.11.050.
* - Areal functional network from Power et al. (2011)
@ -496,19 +516,20 @@ Planned
- Publication
* - Emotional scene and face processing (EmoSF)
- | Sabatinelli, D., Fortune, E.E., Li, Q. et al.
| Emotional perception: Meta-analyses of face and natural scene processing.
| Emotional perception: Meta-analyses of face and natural scene
| processing.
| NeuroImage, Volume 54(3), Pages 2524-2533 (2011).
| https://doi.org/10.1016/j.neuroimage.2010.10.011.
* - Perceptuo-motor network
- | Heckner, M.K., Cieslik, E.C., Eickhoff, S.B. et al.
| The Aging Brain and Executive Functions Revisited: Implications from Meta-analytic
| and Functional-Connectivity Evidence.
| The Aging Brain and Executive Functions Revisited: Implications from
| Meta-analytic and Functional-Connectivity Evidence.
| Journal of Cognitive Neuroscience, Volume 33(9), Pages 1716–1752 (2021).
| https://doi.org/10.1162/jocn_a_01616
Masks
-----
Mask
----
..
Provide a list of the masks that are implemented or planned.
@ -541,30 +562,33 @@ Available
* - Nilearn's MNI152 1mm-resolution mask
- | ``compute_brain_mask``
- 0.0.2
- | Compute the whole-brain mask. This mask is calculated using MNI152 1mm-resolution template mask onto the
| target image. See :func:`nilearn.masking.compute_brain_mask`
- | Compute the whole-brain mask. This mask is calculated using
| MNI152 1mm-resolution template mask onto the target image.
| See :func:`nilearn.masking.compute_brain_mask`
* - Nilearn's mask computed from FMRI data
- | ``compute_epi_mask``
- 0.0.2
- | Compute a brain mask from fMRI data. This is based on an heuristic proposed by T.Nichols: find the least
| dense point of the histogram, between fractions ``lower_cutoff`` and ``upper_cutoff`` of the total image
| histogram. See :func:`nilearn.masking.compute_epi_mask`
- | Compute a brain mask from fMRI data. This is based on an heuristic
| proposed by T.Nichols: find the least dense point of the histogram,
| between fractions ``lower_cutoff`` and ``upper_cutoff`` of the total
| image histogram. See :func:`nilearn.masking.compute_epi_mask`
* - Nilearn's background mask
- | ``compute_background_mask``
- 0.0.2
- | Compute a brain mask for the images by guessing the value of the background from the border of the image.
- | Compute a brain mask for the images by guessing the value of the
| background from the border of the image.
| See :func:`nilearn.masking.compute_background_mask`
* - Nilearn's ICBM152 template gray-matter mask
- | ``fetch_icbm152_brain_gm_mask``
- 0.0.2
- | Compute a gray-matter mask from the asymmetrical ICBM152 2009 template, release a.
- | Compute a gray-matter mask from the asymmetrical ICBM152 2009 template,
| release a.
| See :func:`nilearn.datasets.fetch_icbm152_brain_gm_mask`
Planned
~~~~~~~
..
helpful site for creating tables: https://rest-sphinx-memo.readthedocs.io/en/latest/ReST.html#tables

View file

@ -0,0 +1 @@
Improve general prose, formatting and code blocks in docs and set line length for ``.rst`` files to 80 by `Synchon Mandal`_

View file

@ -86,11 +86,12 @@ Before you submit a pull request, check that it meets these guidelines:
updated. Consider creating a Python file that demonstrates the usage in
``examples/`` directory.
#. Make sure to create a Draft Pull Request. If you are not sure how to do it,
check `here <https://github.blog/2019-02-14-introducing-draft-pull-requests/>`_.
#. Note the pull request ID assigned after completing the previous step and create
a short one-liner file of your contribution named as ``<pull-request-ID>.<type>``
in ``docs/changes/newsfragments/``, ``<type>`` being as per the following
convention:
check
`here <https://github.blog/2019-02-14-introducing-draft-pull-requests/>`_.
#. Note the pull request ID assigned after completing the previous step and
create a short one-liner file of your contribution named as
``<pull-request-ID>.<type>`` in ``docs/changes/newsfragments/``, ``<type>``
being as per the following convention:
* API change : ``change``
* Bug fix : ``bugfix``
@ -155,9 +156,10 @@ Writing Examples
----------------
The format used for text is reST. Check the `sphinx reST reference`_ for more
details. The examples are run and displayed in HTML format using `sphinx gallery`_. To add an
example, just create a ``.py`` file that starts either with ``plot_`` or ``run_``,
dependending on whether the example generates a figure or not.
details. The examples are run and displayed in HTML format using
`sphinx gallery`_. To add an example, just create a ``.py`` file that starts
either with ``plot_`` or ``run_``, dependending on whether the example generates
a figure or not.
The first lines of the example should be a Python block comment with a title,
a description of the example, authors and license name.

View file

@ -6,19 +6,19 @@ Adding Coordinates
==================
Instead of using whole-brain parcellations to aggregate voxel-wise signals from
MR images (as for example in the :class:`.ParcelAggregation` marker), Junifer
MR images (as for example in the :class:`.ParcelAggregation` marker), junifer
allows you to specify a set of coordinates around which to draw spheres to
aggregate (for example using the :class:`.SphereAggregation` marker) the MR
signals from individual voxels. Now, before you start specifying your own sets
of coordinates, check the coordinates that Junifer already has
of coordinates, check the coordinates that junifer already has
:ref:`built in <builtin>`. If you simply want to use a well known set of
coordinates from the literature, there is a reasonable chance, that Junifer
coordinates from the literature, there is a reasonable chance, that junifer
provides them already.
If you checked the in-built coordinates, and they are not there already (for
example if you came up with your own set of coordinates), then Junifer provides
example if you came up with your own set of coordinates), then junifer provides
an easy way for you to register them using the :func:`.register_coordinates`
function, so you can use your own set of coordinates within a Junifer pipeline.
function, so you can use your own set of coordinates within a junifer pipeline.
From the API reference, we can see that it has 3 positional arguments
(``name``, ``coordinates``, and ``voi_names``) as well as one
@ -26,7 +26,7 @@ optional keyword argument (``overwrite``).
The ``name`` argument takes a string indicating the name you want to give to
this set of coordinates. This ``name`` can be used to obtain and operate on a
set of coordinates in Junifer. For example, you can obtain your coordinates
set of coordinates in junifer. For example, you can obtain your coordinates
after registration by providing ``name`` to :func:`.load_coordinates`. We could
simply call it ``"my_set_of_coordinates"``, but likely you want a more
descriptive and more informative name most of the time.
@ -36,7 +36,7 @@ The ``coordinates`` argument takes the actual coordinates as a 2-dimensional
columns (one for each spatial dimension). That is, the first, second, and third
columns indicate the x-, y-, and z-coordinates in MNI space respectively.
The number of rows in the array correspond to the number of coordinates that
belong to this set. Note, that Junifer (as of yet) only works in MNI space, and
belong to this set. Note, that junifer (as of yet) only works in MNI space, and
so therefore these coordinates should always be real-world coordinates of the
MNI space.
@ -60,8 +60,8 @@ packages:
For the sake of this example, we can create a set of coordinates that belong
to the default mode network (DMN), and register this set of coordinates with
Junifer. Note, that Junifer already has a
:ref:`set of coordinates built-in<builtin>` ("DMNBuckner") that is associated
junifer. Note, that junifer already has a
:ref:`set of coordinates built-in <builtin>` ("DMNBuckner") that is associated
with the DMN. Here, we use the DMN coordinates used in a
`nilearn example <https://nilearn.github.io/dev/auto_examples/03_connectivity/plot_sphere_based_connectome.html>`_.
@ -93,16 +93,16 @@ simply use this to register our coordinates:
voi_names=voi_names
)
Now, when we run this script, Junifer registers these coordinates and we can
Now, when we run this script, junifer registers these coordinates and we can
use them in subsequent analyses. Let's now consider how to use coordinate
registration in combination with
:ref:`codeless configuration using a YAML file<codeless>`.
:ref:`codeless configuration using a YAML file <codeless>`.
Step 2: Add coordinate registration to the YAML file
----------------------------------------------------
In order to register your coordinates for a pipeline configured by a YAML file,
you can use the ``with`` keyword provided by Junifer:
you can use the ``with`` keyword provided by junifer:
.. code-block:: yaml

View file

@ -8,9 +8,9 @@ Creating Data Grabbers
Data Grabbers are the first step of the pipeline. Its purpose is to interpret
the structure of a dataset and provide two specific functionalities:
1) Given an *element*, provide the path to each kind of data available for this
#. Given an *element*, provide the path to each kind of data available for this
element (e.g. the path to the T1 image, the path to the T2 image, etc.)
2) Provide the list of *elements* available in the dataset.
#. Provide the list of *elements* available in the dataset.
In this section, we will see how to create a datagrabber for a dataset. Basic
aspects of datagrabbers are covered in the
@ -27,20 +27,20 @@ The *element* should be the smallest unit of data that can be processed. That
is, for each element, there should be a set of data that can be processed, but
only one of each *data type* (see :ref:`data_types`).
For example, if we have a dataset from an fMRI study in which:
For example, if we have a dataset from a fMRI study in which:
a) both T1w and fMRI was acquired
b) 20 subjects went through an experiment twice
c) the experiment included resting-stage fMRI and a task named *stroop*
a. both T1w and fMRI was acquired
b. 20 subjects went through an experiment twice
c. the experiment included resting-stage fMRI and a task named *stroop*
then the *element* should be composed of 3 items:
* ``subject``: The subject IDs, e.g. `sub001`, `sub002`, ... `sub020`
* ``session``: The sesion number, e.g. `ses1`, `ses2`
* ``task``: The task performed, e.g. `rest`, `stroop`
* ``subject``: The subject IDs, e.g. `sub001`, `sub002`, ... `sub020`
* ``session``: The sesion number, e.g. `ses1`, `ses2`
* ``task``: The task performed, e.g. `rest`, `stroop`
If any of these items were not part of the element, then we will have more than
one ``T1w`` and/or ``BOLD`` image for each subject, which is not allowed.
one ``T1w`` and / or ``BOLD`` image for each subject, which is not allowed.
Importantly, nothing prevents that one image is part of two different elements.
For example, it is usually the case that the ``T1w`` image is not acquired for
@ -74,12 +74,11 @@ expressed as a pattern:
where ``{subject}`` is the replacement for the subject id and ``{session}``
is the replacement for the session id.
Since it is a BIDS dataaset, the same happens with the BOLD images. The path to
Since it is a BIDS dataset, the same happens with the BOLD images. The path to
the BOLD images can be expressed as a pattern:
``{subject}/{session}/func/{subject}_{session}_task-rest_bold.nii.gz``
This will be the norm in most of the datasets. If your dataset can be expressed
in terms of patterns, then follow :ref:`extending_datagrabbers_pattern`.
Otherwise, we recommend that you take time to re-think about your dataset
@ -90,7 +89,6 @@ get your dataset in order.
If there is no other way, then you can follow :ref:`extending_datagrabbers_base`
to create a Data Grabber from scratch.
.. _extending_datagrabbers_pattern:
Step 3: Create a Data Grabber
@ -99,13 +97,12 @@ Step 3: Create a Data Grabber
Option A: Extending from PatternDataGrabber
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
The :py:class:`~junifer.datagrabber.PatternDataGrabber` class is an
abstract class that has the functionality of understanding patterns embeded
in it.
The :class:`.PatternDataGrabber` class is an abstract class that has the
functionality of understanding patterns embeded in it.
Before creating the datagrabber, we need to define 3 variables:
* ``types``: A list with the available :ref:`data_types` in our dataset
* ``types``: A list with the available :ref:`data_types` in our dataset.
* ``patterns``: A dictionary that specifies the pattern for each data type.
* ``replacements``: A list indicating which of the elements in the patterns
should be replaced by the values of the element.
@ -126,19 +123,21 @@ where the dataset is located. For example, if the dataset is located in
``/data/project/test/data``, then ``datadir`` should be
``/data/project/test/data``. Or, if we want to allow the user to specify the
location of the dataset, we can expose the variable in the constructor, as in
this example
the following example.
With this defined, we can now create our datagrabber, we will name it
With the variables defined above, we can create our datagrabber and name it
``ExampleBIDSDataGrabber``:
.. code-block:: python
from junifer.datagrabber.pattern import PatternDataGrabber
from pathlib import Path
from junifer.datagrabber import PatternDataGrabber
class ExampleBIDSDataGrabber(PatternDataGrabber):
def __init__(self, datadir):
def __init__(self, datadir: str | Path) -> None:
types = ["T1w", "BOLD"]
patterns = {
"T1w": "{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz",
@ -154,19 +153,21 @@ With this defined, we can now create our datagrabber, we will name it
Our datagrabber is ready to be used by junifer. However, it is still unknown
to the library. We need to register it in the library. To do so, we need to
use the :py:func:`~junifer.api.decorators.register_datagrabber` decorator.
use the :func:`.register_datagrabber` decorator.
.. code-block:: python
from junifer.datagrabber.pattern import PatternDataGrabber
from pathlib import Path
from junifer.api.decorators import register_datagrabber
from junifer.datagrabber import PatternDataGrabber
@register_datagrabber
class ExampleBIDSDataGrabber(PatternDataGrabber):
def __init__(self, datadir):
def __init__(self, datadir: str | Path) -> None:
types = ["T1w", "BOLD"]
patterns = {
"T1w": "{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz",
@ -193,11 +194,11 @@ set the ``datadir``.
Optional: Using datalad
"""""""""""""""""""""""
~~~~~~~~~~~~~~~~~~~~~~~
If you are using `datalad`_, you can use the :class:`.PatternDataladDataGrabber`
instead of the :class:`.PatternDataGrabber`. This class will not only
interpret patterns, but also use `datalad`_ to `clone` and `get` the data.
interpret patterns, but also use `datalad`_ to ``clone`` and ``get`` the data.
The main difference between the two is that the ``datadir`` is not the actual
location of the dataset, but the location where the dataset will be cloned. It
@ -238,17 +239,16 @@ Now we have our 2 additional variables:
And we can create our datagrabber:
.. code-block:: python
from junifer.datagrabber.pattern import PatternDataladDataGrabber
from junifer.api.decorators import register_datagrabber
from junifer.datagrabber import PatternDataladDataGrabber
@register_datagrabber
class ExampleBIDSDataGrabber(PatternDataladDataGrabber):
def __init__(self):
def __init__(self) -> None:
types = ["T1w", "BOLD"]
patterns = {
"T1w": "{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz",
@ -272,29 +272,31 @@ And we can create our datagrabber:
Option B: Extending from BaseDataGrabber
fraimondo commented 2023-04-17 08:41:02 +00:00 (Migrated from github.com)

same, see above

same, see above
fraimondo commented 2023-04-17 08:42:06 +00:00 (Migrated from github.com)

I removed the typing from the examples on purpose.

While it's better for the actual code, it just makes the example more complicated to read with this kind of things.

I removed the typing from the examples on purpose. While it's better for the actual code, it just makes the example more complicated to read with this kind of things.
synchon commented 2023-04-17 09:24:09 +00:00 (Migrated from github.com)

I'll simplify it.

I'll simplify it.
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
While we could not think of a use case in which the pattern-based data grabber would not be suitable, it is still
possible to create a datagrabber extending from the :py:class:`~junifer.datagrabber.base.BaseDataGrabber` class.
While we could not think of a use case in which the pattern-based datagrabber
would not be suitable, it is still possible to create a datagrabber extending
from the :class:`.BaseDataGrabber` class.
In order to create a datagrabber extending from :py:class:`~junifer.datagrabber.base.BaseDataGrabber`, we need to
implement the following methods:
In order to create a datagrabber extending from :class:`.BaseDataGrabber`, we
need to implement the following methods:
- ``get_item``: to get a single item from the dataset.
- ``get_elements``: to get the list of all elements present in the dataset
- ``get_element_keys``: to get the keys of the elements in the dataset.
.. note::
The ``__init__`` method could also be implemented, but it is not mandatory. This is required if the datagrabber
requires any parameter.
The ``__init__`` method could also be implemented, but it is not mandatory.
This is required if the datagrabber requires any extra parameter.
We will now implement our BIDS example with this method.
The first method, ``get_item``, needs to obtain a single
item from the dataset. Since this dataset requires two variables, ``subject`` and ``session``, we will use them
as parameters of ``get_item``:
item from the dataset. Since this dataset requires two variables, ``subject``
and ``session``, we will use them as parameters of ``get_item``:
.. code-block:: python
def get_item(self, subject, session):
def get_item(self, subject: str, session: str) -> dict[str, str]:
out = {
"T1w": f"{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz",
"BOLD": f"{subject}/{session}/func/{subject}_{session}_task-rest_bold.nii.gz",
@ -302,13 +304,17 @@ as parameters of ``get_item``:
return out
The second method, ``get_elements``, needs to return a list of all the elements in the dataset. In this case, we
know that the dataset contains 3 subjects and 3 sessions, so we can create a list of all the possible combinations.
However, we need to remember that for session *ses-03* there is no BOLD data.
The second method, ``get_elements``, needs to return a list of all the elements
in the dataset. In this case, we know that the dataset contains 3 subjects and 3
sessions, so we can create a list of all the possible combinations. However, we
need to remember that for session *ses-03* there is no BOLD data.
.. code-block:: python
def get_elements(self):
from itertools import product
def get_elements(self) -> list[str]:
subjects = ["sub-01", "sub-02", "sub-03"]
sessions = ["ses-01", "ses-02"]
@ -316,19 +322,19 @@ However, we need to remember that for session *ses-03* there is no BOLD data.
if "BOLD" not in self.types:
sessions.append("ses-03")
elements = []
for subject in subjects:
for session in sessions:
elements.append({"subject": subject, "session": session})
for subject, element in product(subjects, sessions):
elements.append({"subject": subject, "session": session})
return elements
And finally, we can implement the ``get_element_keys`` method. This method needs to return a list of the keys that
represent each of the items in the element tuple. As a rule of thumb, they should be the parameters of the
``get_item`` method, in the same order.
And finally, we can implement the ``get_element_keys`` method. This method needs
to return a list of the keys that represent each of the items in the element
tuple. As a rule of thumb, they should be the parameters of the ``get_item``
method, in the same order.
.. code-block:: python
def get_element_keys(self):
def get_element_keys(self) -> list[str]:
return ["subject", "session"]
@ -336,20 +342,21 @@ So, to summarize, our datagrabber will look like this:
.. code-block:: python
from junifer.datagrabber.base import BaseDataGrabber
from junifer.api.decorators import register_datagrabber
from junifer.datagrabber import BaseDataGrabber
@register_datagrabber
class ExampleBIDSDataGrabber(BaseDataGrabber):
def get_item(self, subject, session):
def get_item(self, subject: str, session: str) -> dict[str, str]:
out = {
"T1w": f"{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz",
"BOLD": f"{subject}/{session}/func/{subject}_{session}_task-rest_bold.nii.gz",
}
return out
def get_elements(self):
def get_elements(self) -> list[str]:
subjects = ["sub-01", "sub-02", "sub-03"]
sessions = ["ses-01", "ses-02"]
@ -362,62 +369,74 @@ So, to summarize, our datagrabber will look like this:
elements.append({"subject": subject, "session": session})
return elements
def get_element_keys(self):
def get_element_keys(self) -> list[str]:
return ["subject", "session"]
Optional: Using datalad
"""""""""""""""""""""""
~~~~~~~~~~~~~~~~~~~~~~~
If this dataset is in a datalad dataset, we can extend from
:class:`.DataladDataGrabber` instead of :class:`.BaseDataGrabber`. This will
allow us to use the datalad API to obtain the data.
Step 4: Optional: Adding *BOLD confounds*
-----------------------------------------
For some analyses, it is useful to have the confounds associated with the BOLD data. This corresponds to the
``BOLD_confounds`` item in the :ref:`Data Object <data_object>` (see :ref:`data_types`). However, the ``BOLD_confounds``
element does not only consists of a ``path``, but it requries more information about the format of the confounds file.
Thus, the ``BOLD_confounds`` element is a dictionary with the following keys:
For some analyses, it is useful to have the confounds associated with the BOLD
data. This corresponds to the ``BOLD_confounds`` item in the
:ref:`Data Object <data_object>` (see :ref:`data_types`). However, the
``BOLD_confounds`` element does not only consists of a ``path``, but it requries
more information about the format of the confounds file. Thus, the
``BOLD_confounds`` element is a dictionary with the following keys:
- ``path``: the path to the confounds file.
- ``format``: the format of the confounds file. Currently, this can be either ``fmriprep`` or ``adhoc``.
- ``format``: the format of the confounds file. Currently, this can be either
``fmriprep`` or ``adhoc``.
The ``fmriprep`` format corresponds to the format of the confounds files generated by `fMRIPrep`_. The
``adhoc`` format corresponds to a format that is not standardized.
The ``fmriprep`` format corresponds to the format of the confounds files
generated by `fMRIPrep`_. The ``adhoc`` format corresponds to a format that is
not standardised.
.. note::
The ``mappings`` key is only required if the ``format`` is ``adhoc``. If the ``format`` is ``fmriprep``, the
``mappings`` key is not required.
The ``mappings`` key is only required if the ``format`` is ``adhoc``. If the
``format`` is ``fmriprep``, the ``mappings`` key is not required.
Currently, Junifer provides only one confound remover step
(:class:`.fMRIPrepConfoundRemover`), which relies entirely on the ``fmriprep`` confound
variable names. Thus, if the confounds are not in ``fmriprep`` format, the user will need to provide the mappings
between the *ad-hoc* variable names and the ``fmriprep`` variable names.
This is done by specifying the ``adhoc`` format and providing the mappings as a dictionary in the ``mappings`` key.
Currently, junifer provides only one confound remover step
(:class:`.fMRIPrepConfoundRemover`), which relies entirely on the ``fmriprep``
confound variable names. Thus, if the confounds are not in ``fmriprep`` format,
the user will need to provide the mappings between the *ad-hoc* variable names
and the ``fmriprep`` variable names. This is done by specifying the ``adhoc``
format and providing the mappings as a dictionary in the ``mappings`` key.
In the following example, the confounds file has 3 variables that are not in the ``fmriprep`` format. Thus, we will
provide the mappings for these variables to the ``fmriprep`` format.
In the following example, the confounds file has 3 variables that are not in the
``fmriprep`` format. Thus, we will provide the mappings for these variables to
the ``fmriprep`` format. For example, the ``get_item`` method could look like
this:
.. code-block:: python
out["BOLD_confounds"]: {
"path": f"{subject}/{session}/func/{subject}_{session}_confounds.tsv",
"format": "adhoc",
"mappings": {
"fmriprep": {
"variable1": "rot_x",
"variable2": "rot_z",
"variable3": "rot_y",
}
},
}
def get_item(
self, subject: str, session: str
) -> dict:
out = {
"BOLD": f"{subject}/{session}/func/{subject}_{session}_task-rest_bold.nii.gz",
"BOLD_confounds": {
"path": f"{subject}/{session}/func/{subject}_{session}_confounds.tsv",
"format": "adhoc",
"mappings": {
"fmriprep": {
"variable1": "rot_x",
"variable2": "rot_z",
"variable3": "rot_y",
}
},
}
.. note::
Not all of the mappings need to be provided. For the moment, this is used only by the
:class:`.fMRIPrepConfoundRemover` step, which requires variables based on the
strategy selected. However, it is recommended to provide all the mappings, as this will allow the user to
choose different strategies with the same dataset.
Not all of the mappings need to be provided. For the moment, this is used
only by the :class:`.fMRIPrepConfoundRemover` step, which requires variables
based on the strategy selected. However, it is recommended to provide all the
mappings, as this will allow the user to choose different strategies with the
same dataset.

View file

@ -2,17 +2,19 @@
.. _extending_extension:
Creating a Junifer extension
Creating a junifer extension
============================
Junifer is designed to be easily extensible. Through the use of a registry and decorators, we can easily add new
functionality to junifer on runtime. This is done by creating a new python module and importing it before running
junifer.
Junifer is designed to be easily extensible. Through the use of a registry and
decorators, you can easily add new functionality to junifer during runtime. This
is done by creating a new Python module and importing it before running junifer.
A special consideration has to be made when using the :ref:`code-less configuration<codeless>`. In this case, the
``with`` statement can be used to import a module or run a python ``.py`` file.
A special consideration has to be made when using the
:ref:`code-less configuration<codeless>`. In this case, the
``with`` statement can be used to import a module or run a Python file.
In the following example, we instruct junifer to first import ``my_module`` and then run the ``my_file.py`` file.
In the following example, we instruct junifer to first import ``my_module`` and
then run the ``my_file.py`` file.
.. code-block:: yaml
@ -20,9 +22,13 @@ In the following example, we instruct junifer to first import ``my_module`` and
- my_module
- my_file.py
Thus, the code from ``my_file.py`` will be executed before running junifer. This is the ideal place to create junifer
extensions.
Thus, the code from ``my_file.py`` will be executed before running junifer. This
is the ideal place to include junifer extensions.
.. important:: Some junifer commands will not consider files imported from files included in the ``with`` statement.
That is, if ``my_file.py`` imports ``my_other_file.py``, some of the junifer commands will not consider
``my_other_file.py``. Either place all the code in one file or add multiple files to the ``with`` statement.
.. important::
Some junifer commands will not consider files imported from files included
in the ``with`` statement. If ``my_file.py`` imports ``my_other_file.py``,
some of the junifer commands will not consider ``my_other_file.py``. Either
place all the code in one file or add multiple files to the ``with``
statement.

View file

@ -7,16 +7,15 @@ Extending junifer
While we aim to provide as many datasets and markers as possible, we are also
interested in allowing users to extend the functionality with their own
datagrabbers, preprocessing, markers, etc.
datagrabbers, preprocessing, markers, etc., .
This does not mean that the new functionality will have to be included in
junifer before the user can use them. Instead, the user can simply
create a new python file, code the desired functionality and use it with
junifer. This is the first step towards including the new functionality in
the junifer package.
It's not necessary to have the new functionality included in junifer before
the user can use them. The user can simply create a new Python file, code the
desired functionality and use it with junifer. This is the first step towards
including the new functionality in the junifer pipeline.
In this section we will show how to extend junifer, by creating new
datagrabbers, preprocessing and markers, following the *junifer* way.
datagrabbers, preprocessing, markers, etc., following the *junifer* way.
.. toctree::
@ -28,4 +27,4 @@ datagrabbers, preprocessing and markers, following the *junifer* way.
marker
parcellations
coordinates
masks
masks

View file

@ -5,111 +5,143 @@
Creating Markers
================
Computing a marker (a.k.a. *feature*) is the main goal of junifer. While we aim to provide as many markers as possible,
it might be the case that the marker you are looking for is not available. In this case, you can create your own marker
Computing a marker (a.k.a. *feature*) is the main goal of junifer. While we aim
to provide as many markers as possible, it might be the case that the marker you
are looking for is not available. In this case, you can create your own marker
by following this tutorial.
Most of the functionality of a junifer marker has been taken care by the :class:`.BaseMarker` class.
Thus, only a few methods are required:
Most of the functionality of a junifer marker has been taken care by the
:class:`.BaseMarker` class. Thus, only a few methods are required:
1. ``get_valid_inputs``: a method to obtain the list of valid inputs for the marker. This is used to check that the
inputs provided by the user are valid. This method should return a list of strings, representing
:ref:`data types <data_types>`
2. ``get_output_type``: a method to obtain the kind of output of the marker. This is used to check that the output
of the marker is compatible with the storage. This method should return a string, representing
:ref:`storage types <storage_types>`
3. ``compute``: the method that given the data, computes the marker.
4. ``__init__``: the initialization method, where the marker is configured.
#. ``get_valid_inputs``: The method to obtain the list of valid inputs for the
marker. This is used to check that the inputs provided by the user are
valid. This method should return a list of strings, representing
:ref:`data types <data_types>`.
#. ``get_output_type``: The method to obtain the kind of output of the marker.
This is used to check that the output of the marker is compatible with the
storage. This method should return a string, representing
:ref:`storage types <storage_types>`.
#. ``compute``: The method that given the data, computes the marker.
#. ``__init__``: The initialisation method, where the marker is configured.
As an example, we will develop a Parcel Mean marker, that is, a marker that first applies a parcellation and
then computes the mean of the data in each parcel. This is a very simple example, but it will show you how to create
a new marker.
As an example, we will develop a ``ParcelMean`` marker, a marker that first
applies a parcellation and then computes the mean of the data in each parcel.
This is a very simple example, but it will show you how to create a new marker.
.. _extending_markers_input_output:
Step 1: Configure input and output
----------------------------------
This step is quite simple: we need to define the input and output of the marker. Based on the current
:ref:`data types <data_types>`, we can define as valid inputs ``BOLD``, ``VBM_WM`` and ``VBM_GM``.
This step is quite simple: we need to define the input and output of the marker.
Based on the current :ref:`data types <data_types>`, we can have ``BOLD``,
``VBM_WM`` and ``VBM_GM`` as valid inputs.
.. code-block:: python
def get_valid_inputs(self):
return ['BOLD', 'VBM_WM', 'VBM_GM']
def get_valid_inputs(self) -> list[str]:
return ["BOLD", "VBM_WM", "VBM_GM"]
The output of the marker depends on the input. For ``BOLD``, it will be ``timeseries``, while for the rest of the inputs,
it will be ``table``. Thus, we can define the output as:
The output of the marker depends on the input. For ``BOLD``, it will be
``timeseries``, while for the rest of the inputs, it will be ``vector``. Thus,
we can define the output as:
.. code-block:: python
def get_output_type(self, input_kind):
if input_kind == 'BOLD':
return 'timeseries'
def get_output_type(self, input_kind: str) -> str:
if input_kind == "BOLD":
return "timeseries"
else:
return 'table'
return "vector"
.. _extending_markers_init:
Step 2: Initialize the marker
Step 2: Initialise the marker
-----------------------------
In this step we need to define the parameters of the marker. That is, all the parameters that the user can provide
In this step we need to define the parameters of the marker the user can provide
to configure how the marker will behave.
The parameters of the marker are defined in the ``__init__`` method. The :class:`.BaseMarker` class
requires two optional parameters:
The parameters of the marker are defined in the ``__init__`` method. The
:class:`.BaseMarker` class requires two optional parameters:
1. ``name``: the name of the marker. This is used to identify the marker in the configuration file.
2. ``on``: a list or string with the data types that the marker will be applied to.
1. ``name``: the name of the marker. This is used to identify the marker in the
configuration file.
2. ``on``: a list or string with the data types that the marker will be applied
to.
.. attention:: Only basic types (*int*, *bool* and *str*) as well as Lists, Tuples and Dictionaries are allowed as
parameters. This is because the parameters are stored in a JSON file, and JSON only supports these types.
.. attention::
Only basic types (*int*, *bool* and *str*), lists, tuples and dictionaries
are allowed as parameters. This is because the parameters are stored in
JSON format, and JSON only supports these types.
In this example, the is only paramater required for the computation is the name of the parcellation to use. Thus, we can
define the ``__init__`` method as follows:
In this example, only paramater required for the computation is the name of the
parcellation to use. Thus, we can define the ``__init__`` method as follows:
.. code-block:: python
def __init__(self, parcellation_name, on=None, name=None):
self.parcellation_name = parcellation_name
def __init__(
self,
parcellation: str,
on: str | list[str] | None = None,
name: str | None = None,
) -> None:
self.parcellation = parcellation
super().__init__(on=on, name=name)
.. caution:: Parameters of the marker must be stored as object attributes without using ``_`` as prefix. This is
because any attribute that starts with ``_`` will not be considered as a parameter and not stored as
part of the metadata of the marker.
.. caution::
Parameters of the marker must be stored as object attributes without using
``_`` as prefix. This is because any attribute that starts with ``_`` will
not be considered as a parameter and not stored as part of the metadata of
the marker.
.. _extending_markers_compute:
Step 3: Compute the marker
--------------------------
In this step, we will define the method that computes the marker. This method will be called by junifer when needed,
using the data provided by the datagrabber, as configured by the user. The function ``compute`` has two arguments:
In this step, we will define the method that computes the marker. This method
will be called by junifer when needed, using the data provided by the
datagrabber, as configured by the user. The method ``compute`` has two
arguments:
* ``input``: a dictionary with the data to be used to compute the marker. This will be the corresponding element in the
:ref:`Data Object<data_object>` alredy indexing. Thus, the dictionary has at least two keys: ``data`` and ``path``.
The first one contains the data, while the second one contains the path to the data. The dictionary can also contain
other keys, depending on the data type.
* ``extra_input``: the rest of the :ref:`Data Object<data_object>`. This is useful if you want to use other data to
compute the marker (e.g.: ``BOLD_confounds`` can be used to de-confound the ``BOLD`` data).
* ``input``: a dictionary with the data to be used to compute the marker. This
will be the corresponding element in the :ref:`Data Object<data_object>`
alredy indexed. Thus, the dictionary has at least two keys: ``data`` and
``path``. The first one contains the data, while the second one contains the
path to the data. The dictionary can also contain other keys, depending on the
data type.
* ``extra_input``: the rest of the :ref:`Data Object<data_object>`. This is
useful if you want to use other data to compute the marker
(e.g.: ``BOLD_confounds`` can be used to de-confound the ``BOLD`` data).
Following the example, we will compute the mean of the data in each parcel using
:class:`nilearn.maskers.NiftiLabelsMasker`. Importantly, the output of the compute function must be a dictionary.
This dictionary will later be passed onto the ``store`` method.
:class:`nilearn.maskers.NiftiLabelsMasker`. Importantly, the output of the
compute function must be a dictionary. This dictionary will later be passed onto
the ``store`` method.
.. hint:: To simplify the ``store`` method, define keys of the dictionary based on the corresponding store functions
in the :ref:`storage types <storage_types>`. For example, if the output is a ``table``, the keys of the
dictionary should be ``data`` and ``columns``.
.. hint::
To simplify the ``store`` method, define keys of the dictionary based on the
corresponding store functions in the :ref:`storage types <storage_types>`.
For example, if the output is a ``vector``, the keys of the dictionary should
be ``data`` and ``col_names``.
.. code-block:: python
from nilearn.maskers import NiftiLabelsMasker
from junifer.data import load_parcellation
from typing import Any
def compute(self, input, extra_input):
from junifer.data import load_parcellation
from nilearn.maskers import NiftiLabelsMasker
def compute(
self,
input: dict[str, Any],
extra_input: dict[str, Any] | None = None,
) -> dict[str, Any]:
# Get the data
data = input["data"]
@ -126,7 +158,7 @@ This dictionary will later be passed onto the ``store`` method.
masker = NiftiLabelsMasker(
labels_img=t_parcellation,
standardize=True,
memory='nilearn_cache',
memory="nilearn_cache",
verbose=5,
)
@ -134,59 +166,71 @@ This dictionary will later be passed onto the ``store`` method.
out_values = masker.fit_transform([data])
# Create the output dictionary
out = {"data": out_values, "columns": t_labels}
out = {"data": out_values, "col_names": t_labels}
# If its 3D (BOLD), name each row as "scan"
if out_values.shape[0] > 1:
out["row_names"] = "scan"
return out
.. _extending_markers_finalize:
Step 4: Finalize the marker
Step 4: Finalise the marker
---------------------------
Once all of the above steps are done, we just need to give our marker a name, state its *dependencies* and register it
using the ``@register_marker`` decorator.
Once all of the above steps are done, we just need to give our marker a name,
state its *dependencies* and register it using the ``@register_marker``
decorator.
The *dependencies* are the core packages that are required to compute the marker. This will be later used to keep track
of the versions of the packages used to compute the marker. To inform junifer about the dependencies of a marker,
we need to define a ``_DEPENDENCIES`` attribute in the class. This attribute must be a set, with the names of the
packages as strings. For example, the ``ParcelMean`` marker has the following dependencies:
The *dependencies* are the core packages that are required to compute the marker.
This will be later used to keep track of the versions of the packages used to
compute the marker. To inform junifer about the dependencies of a marker, we need
to define a ``_DEPENDENCIES`` attribute in the class. This attribute must be a
set, with the names of the packages as strings. For example, the ``ParcelMean``
marker has the following dependencies:
.. code-block:: python
_DEPENDENCIES = {"nilearn"}
_DEPENDENCIES = {"nilearn", "numpy"}
Finally, we need to register the marker using the ``@register_marker`` decorator. This decorator takes the name of the
Finally, we need to register the marker using the ``@register_marker`` decorator.
.. code-block:: python
from nilearn.maskers import NiftiLabelsMasker
from junifer.data import load_parcellation
from typing import Any
from junifer.api.decorators import register_marker
from junifer.data import load_parcellation
from junifer.markers.base import BaseMarker
from nilearn.maskers import NiftiLabelsMasker
@register_marker
class ParcelMean(BaseMarker):
_DEPENDENCIES = {"nilearn", "numpy"}
_DEPENDENCIES = {"nilearn", "numpy"}
def __init__(self, parcellation_name, on=None, name=None):
self.parcellation_name = parcellation_name
def __init__(
self,
parcellation: str,
on: str | list[str] | None = None,
name: str | None = None,
) -> None:
self.parcellation = parcellation
super().__init__(on=on, name=name)
def get_valid_inputs(self):
return ['BOLD', 'VBM_WM', 'VBM_GM']
def get_valid_inputs(self) -> list[str]:
return ["BOLD", "VBM_WM", "VBM_GM"]
def get_output_type(self, input_kind):
if input_kind == 'BOLD':
return 'timeseries'
def get_output_type(self, input_kind: str) -> str:
if input_kind == "BOLD":
return "timeseries"
else:
return 'table'
return "vector"
def compute(self, input, extra_input):
def compute(
self,
input: dict[str, Any],
extra_input: dict[str, Any] | None = None,
) -> dict[str, Any]:
# Get the data
data = input["data"]
@ -203,7 +247,7 @@ Finally, we need to register the marker using the ``@register_marker`` decorator
masker = NiftiLabelsMasker(
labels_img=t_parcellation,
standardize=True,
memory='nilearn_cache',
memory="nilearn_cache",
verbose=5,
)
@ -211,11 +255,8 @@ Finally, we need to register the marker using the ``@register_marker`` decorator
out_values = masker.fit_transform([data])
# Create the output dictionary
out = {"data": out_values, "columns": t_labels}
out = {"data": out_values, "col_names": t_labels}
# If its 3D (BOLD), name each row as "scan"
if out_values.shape[0] > 1:
out["row_names"] = "scan"
return out
@ -227,7 +268,8 @@ Template for a custom Marker
.. code-block:: python
from junifer.api.decorators import register_marker
from junifer.markers.base import BaseMarker
from junifer.markers import BaseMarker
@register_marker
class TemplateMarker(BaseMarker):
@ -249,5 +291,5 @@ Template for a custom Marker
# TODO: compute the marker and create the output dictionary
# Create the output dictionary
out = {"data": None, "columns": None}
out = {"data": None, "col_names": None}
return out

View file

@ -5,28 +5,28 @@
Adding Masks
============
Many processing steps and markers in Junifer allow you to specify a binary
Many processing steps and markers in junifer allow you to specify a binary
mask to select voxels you want to include in the analysis. There are a number
of masks :ref:`in-built in Junifer already<builtin>`, so check if any of them
suit your needs. Check how to use these masks :ref:`here<using_masks>`. Once
of masks :ref:`in-built in junifer already <builtin>`, so check if any of them
suit your needs. Check how to use these masks :ref:`here <using_masks>`. Once
you know how to use these masks, and you checked whether the in-built masks
suit your needs, and you have found that they don't, you can come back here to
learn how to use your own masks.
The principle is fairly simple and quite similar to :ref:`adding_parcellations`
and :ref:`adding_coordinates`. Junifer provides a :func:`.register_mask`
and :ref:`adding_coordinates`. junifer provides a :func:`.register_mask`
function that lets you register your own custom masks. It consists of two
positional arguments (``name`` and ``mask_path``) and one optional keyword
argument (``overwrite``).
The ``name`` argument is a string indicating the name of the mask. This name
is used to refer to that mask in Junifer internally in order to obtain the
is used to refer to that mask in junifer internally in order to obtain the
actual mask data and perform operations on it. For example, using the name you
can load a mask after registration using the
:func:`.load_mask` function.
The ``mask_path`` should contain the path to a valid NIfTI image with binary
voxel values (i.e. 0 or 1). This data can then be used by Junifer to mask other
voxel values (i.e. 0 or 1). This data can then be used by junifer to mask other
MR images.
Step 1: Prepare code to register a mask
@ -37,9 +37,11 @@ look as follows:
.. code-block:: python
from junifer.data import register_custom_mask
from pathlib import Path
from junifer.data import register_custom_mask
# this path is only an example, of course use the correct path
# on your system:
mask_path = Path("..") / ".." / "my_custom_mask.nii.gz"
@ -47,12 +49,12 @@ look as follows:
register_mask(name="my_custom_mask", mask_path=mask_path)
Simple, right? Now we just have to configure a YAML file to register this mask
so we can use it for :ref:`codeless configuration of junifer<codeless>`.
so we can use it for :ref:`codeless configuration of junifer <codeless>`.
Step 2: Configure a YAML file for registration of a mask
--------------------------------------------------------
In order to do this, we can use the ``with`` keyword provided by Junifer:
In order to do this, we can use the ``with`` keyword provided by junifer:
.. code-block:: yaml
@ -73,10 +75,10 @@ mask as an argument. For example:
Now, you can simply use this YAML file to run your pipeline. One important
point to keep in mind is that if the paths given in ``register_custom_mask.py``
are relative paths, they will be interpreted by Junifer as relative to the
jobs directory (i.e. where Junifer will create submit files, logs directory and
are relative paths, they will be interpreted by junifer as relative to the
jobs directory (i.e. where junifer will create submit files, logs directory and
so on). For simplicity, you may just want to use absolute paths to avoid
confusion, yet using relative paths is likely a better way to make your
pipeline directory/repository more portable and therefore more reproducible for
others. Really, once you understand how these paths are interpreted by Junifer,
others. Really, once you understand how these paths are interpreted by junifer,
it is quite easy.

View file

@ -5,17 +5,17 @@
Adding Parcellations
====================
Before you start adding your own parcellations, check whether Junifer has
the parcellation :ref:`in-built already<builtin>`. Perhaps, what is available
there will suffice to achieve your goals. However, of course Junifer will not
Before you start adding your own parcellations, check whether junifer has
the parcellation :ref:`in-built already <builtin>`. Perhaps, what is available
there will suffice to achieve your goals. However, of course junifer will not
have every parcellation available that you may want to use, and if so, it will
be nice to be able to add it yourself using a format that Junifer understands.
be nice to be able to add it yourself using a format that junifer understands.
Similarly, you may even be interested in creating your own custom parcellations
and then adding them to Junifer, so you can use Junifer to obtain different
and then adding them to junifer, so you can use junifer to obtain different
markers to assess and validate your own parcellation. So, how can you do this?
Since both of these use-cases are quite common, and not being able to use your
favourite parcellation is of course quite a buzzkill, Junifer actually provides
favourite parcellation is of course quite a buzzkill, junifer actually provides
the easy-to-use :func:`.register_parcellation` function to do just that. Let's
try to understand the API reference and then use this function to register our
own parcellation.
@ -24,7 +24,7 @@ From the API reference, we can see that it has 3 positional arguments
(``name``, ``parcellation_path``, and ``parcels_labels``) as well as one
optional keyword argument (``overwrite``).
The ``name`` of the parcellation is up to you and will be the name that Junifer
The ``name`` of the parcellation is up to you and will be the name that junifer
will use to refer to this particular parcellation. You can think of this as
being similar to a key in a python dictionary, i.e. a key that is used to
obtain and operate on the actual parcellation data. This ``name`` must always
@ -46,7 +46,7 @@ features that parcellation-based markers produce in an unambiguous way, such
that a user can easily identify which ROIs were used to produce a specific
feature (multiple ROIs, because some features consist of information from two
or more ROIs, as for example in functional connectivity). Therefore, we provide
Junifer with a list of strings, that contains the names for each ROI. In this
junifer with a list of strings, that contains the names for each ROI. In this
list, the label at the i-th position indicates the i-th integer label (i.e. the
first label in this list corresponds to the first integer label in the
parcellation and so on).
@ -54,7 +54,7 @@ parcellation and so on).
Step 1: Prepare code to register a parcellation
-----------------------------------------------
Now we know everything that we need to know to make sure Junifer can use our
Now we know everything that we need to know to make sure junifer can use our
own parcellation to compute any parcellation-based marker. For example,
a simple example could look like this:
@ -82,16 +82,16 @@ a simple example could look like this:
)
We can run this code and it seems to work, however, how can we actually
include the custom parcellation in a Junifer pipeline using a
:ref:`code-less YAML configuration<codeless>`?
include the custom parcellation in a junifer pipeline using a
:ref:`code-less YAML configuration <codeless>`?
Step 2: Add parcellation registration to the YAML file
------------------------------------------------------
In order to use the parcellation in a Junifer pipeline configured by a YAML
In order to use the parcellation in a junifer pipeline configured by a YAML
file, we can save the above code in a python file, say
``registering_my_parcellation.py``. We can then simply add this file using the
``with`` keyword provided by Junifer:
``with`` keyword provided by junifer:
.. code-block:: yaml
@ -114,9 +114,9 @@ parcellation when registering it. For example, we can add a
Now, you can simply use this YAML file to run your pipeline. One important
point to keep in mind is that if the paths given in
``registering_my_parcellation.py`` are relative paths, they will be interpreted
by Junifer as relative to the jobs directory (i.e. where Junifer will create
by junifer as relative to the jobs directory (i.e. where junifer will create
submit files, logs directory and so on). For simplicity, you may just want to
use absolute paths to avoid confusion, yet using relative paths is likely a
better way to make your pipeline directory/repository more portable and
therefore more reproducible for others. Really, once you understand how these
paths are interpreted by Junifer, it is quite easy.
paths are interpreted by junifer, it is quite easy.

View file

@ -23,7 +23,7 @@ The following steps are specific to VSCode and you can choose to go with it:
.. code-block:: bash
conda env create -n <your-environment-name> -f conda-env.yml python=3.9
conda env create -n <your-environment-name> -f conda-env.yml python=3.10
conda activate <your-environment-name>
The ``conda-env.yml`` can be found at the root of the repository.

View file

@ -12,14 +12,16 @@ Getting Help
-- The Beatles, Help!
This fragment of the song originally appeared in the 1965 film Help! and was written by Lennon and McCartney, just a
few years after the development of Mark 1, the first commercial computer. Maybe just a random coincidence, or maybe
This fragment of the song originally appeared in the 1965 film Help! and was
written by Lennon and McCartney, just a few years after the development of
Mark 1, the first commercial computer. Maybe just a random coincidence, or maybe
Lennon was just trying to write an email to McCartney.
While the song might have been written with another meaning in mind, it is a good way to describe the situation of many
researchers who are presented with a new toolbox. Indeed, the situation of many researchers is that the projects they
are working on are becoming more and more complex in terms of methods and data. Thus, we *open up the doors* to new
possibilites:
While the song might have been written with another meaning in mind, it is a good
way to describe the situation of many researchers who are presented with a new
toolbox. Indeed, the situation of many researchers is that the projects they are
working on are becoming more and more complex in terms of methods and data. Thus,
we *open up the doors* to new possibilites:
| When I was younger, so much younger than today
| I never needed anybody's help in any way
@ -28,10 +30,12 @@ possibilites:
-- The Beatles, Help!
The setback with modern research is that current methods are often more complex and require more computing, which
means that we need to learn concepts from computer science, mathematics, statistics, etc. This is a good thing, but it
also means that we need to learn new tools and new ways of thinking, which can be a bit overwhelming, to the point that
we start relying more and more on other researchers. In the end, we might feel like we lose our independence:
The setback with modern research is that current methods are often more complex
and require more computing, which means that we need to learn concepts from
computer science, mathematics, statistics, etc. This is a good thing, but it
also means that we need to learn new tools and new ways of thinking, which can be
a bit overwhelming, to the point that we start relying more and more on other
researchers. In the end, we might feel like we lose our independence:
| And now my life has changed in oh so many ways
| My independence seems to vanish in the haze
@ -40,34 +44,44 @@ we start relying more and more on other researchers. In the end, we might feel l
-- The Beatles, Help!
We can continue with the song, but we think you get the point. The point is that we will need help, and we will need
to ask for it. We will need to ask for help from our colleagues, from our supervisors, from our friends.
We can continue with the song, but we think you get the point. The point is that
we will need help, and we will need to ask for it. We will need to ask for help
from our colleagues, from our supervisors, from our friends.
We are a small team of researchers and developers, and we are not experts in *everything at once*. Each one of us has
a specific expertise, and we are trying to use this expertise to create Junifer. When we conceived Junifer, we thought
of researchers' problems and tried to come up with the best way to help them, by building a tool that is easy to
understand, learn and use. Most importantly, we made it to help. We are here to help you and your research.
We are a small team of researchers and developers, and we are not experts in
*everything at once*. Each one of us has a specific expertise, and we are trying
to use this expertise to create Junifer. When we conceived Junifer, we thought
of researchers' problems and tried to come up with the best way to help them, by
building a tool that is easy to understand, learn and use. Most importantly, we
made it to help. We are here to help you and your research.
If you have any questions, problems and / or suggestions, please do not hesitate to contact us. We will be happy to help you
and we will be happy to hear from you.
If you have any questions, problems and / or suggestions, please do not hesitate
to contact us. We will be happy to help you and we will be happy to hear from
you.
Seems nice, no? But we have one condition: **help us help you**.
Communication is the key for you to help us and in turn help you solve your problems. We cannot know what you are
trying to do, unless you tell us. **The more detailed explanation you give us, the faster we can help you**.
We have opened several communication channels so that you can contact us in the way that is most convenient for you.
Communication is the key for you to help us and in turn help you solve your
problems. We cannot know what you are trying to do, unless you tell us.
**The more detailed explanation you give us, the faster we can help you**.
We have opened several communication channels so that you can contact us in
the way that is most convenient for you.
Some people prefer to **write, in detail, with code and figures**. If you are one of those, use the
`junifer Discussions`_. site on GitHub. This is a place where you can ask questions, and where you can discuss topics
such as potential new features, or potential new methods.
Some people prefer to **write, in detail, with code and figures**. If you are
one of those, use the `junifer Discussions`_. site on GitHub. This is a place
where you can ask questions, and where you can discuss topics such as
potential new features, or potential new methods.
Some people do **not have a clear idea of what they want**, but they know that they need help. This is not a
problem, but it is a bit more involved. Since it will require more frequent interactions to try to understand
what you are trying to do, we have the `junifer matrix channel`_ in which you can chat with us and other junifer users.
Some people do **not have a clear idea of what they want**, but they know that
they need help. This is not a problem, but it is a bit more involved. Since it
will require more frequent interactions to try to understand what you are trying
to do, we have the `junifer matrix channel`_ in which you can chat with us and
other junifer users.
Finally, some people **prefer to communicate verbally**. If you are one of those, you might want to join our
*office hours*. Given that our agenda might vary, office hours will be announced on the
`junifer matrix channel`_ chat. Feel free to join and just *listen in* if you are too shy to write.
Finally, some people **prefer to communicate verbally**. If you are one of those,
you might want to join our *office hours*. Given that our agenda might vary,
office hours will be announced on the `junifer matrix channel`_ chat. Feel free
to join and just *listen in* if you are too shy to write.
In short, these are the 3 communication channels to get help:
@ -93,7 +107,8 @@ In short, these are the 3 communication channels to get help:
Cons:
* Real-time depends on the availability of the other users.
* It might be difficult to follow if several conversation happens at the same time.
* It might be difficult to follow if several conversation happens at the same
time.
#. Video Calls (*office hours*)

View file

@ -49,6 +49,8 @@ Funding
=======
We thank the `Helmholtz Imaging Platform <https://helmholtz-imaging.de/>`_,
`SMHB <https://www.fz-juelich.de/en/smhb>`_ and `eBRAIN Health <https://www.ebrain-health.eu/>`_
(HORIZON-INFRA-2021-TECH-01) for supporting development of Junifer.
(The funding sources had no role in the design, implementation and evaluation of the pipeline.)
`SMHB <https://www.fz-juelich.de/en/smhb>`_ and
`eBRAIN Health <https://www.ebrain-health.eu/>`_ (HORIZON-INFRA-2021-TECH-01)
for supporting development of Junifer.
(The funding sources had no role in the design, implementation and evaluation of
the pipeline.)

View file

@ -19,7 +19,8 @@ junifer is compatible with `Python`_ >= 3.8 and requires the following packages:
* ``pyyaml>=5.1.2,<7.0``
* ``h5py>=3.8.0,<3.9``
Depending on the installation method, these packages might be installed automatically.
Depending on the installation method, these packages might be installed
automatically.
Installation
------------
@ -32,7 +33,8 @@ Depending on your use-case, junifer can be installed differently:
for developers.
Either way, we strongly recommend using `virtual environments <https://realpython.com/python-virtual-environments-a-primer>`_.
Either way, we strongly recommend using
`virtual environments <https://realpython.com/python-virtual-environments-a-primer>`_.
.. _install_latest_release:
@ -60,40 +62,47 @@ Follow the `detailed contribution guidelines <contribution.rst>`_.
Installing external dependencies
================================
Some markers will require optional external dependencies to be installed. In this section you will
find a list of all external dependencies that are required for specific markers.
Some markers will require optional external dependencies to be installed. In
this section you will find a list of all external dependencies that are required
for specific markers.
AFNI
----
To install AFNI, you can always follow the `AFNI official instructions
<https://afni.nimh.nih.gov/pub/dist/doc/htmldoc/background_install/main_toc.html>`_. Additionally, you can also follow
the following steps to install and configure the AFNI Docker container in your local system.
<https://afni.nimh.nih.gov/pub/dist/doc/htmldoc/background_install/main_toc.html>`_.
Additionally, you can also follow the following steps to install and configure
the AFNI Docker container in your local system.
.. important::
The AFNI Docker container wrappers add the commands required by junifer. Using these commands have
some limitations, mostly related to handling files and paths. Junifer knows about this and uses these
commands in the proper way. Keep this in mind if you try to use the AFNI Docker wrappers outside of junifer.
These caveats and limitations are not documented.
The AFNI Docker container wrappers add the commands required by junifer. Using
these commands have some limitations, mostly related to handling files and
paths. Junifer knows about this and uses these commands in the proper way.
Keep this in mind if you try to use the AFNI Docker wrappers outside of
junifer. These caveats and limitations are not documented.
1. Install Docker. You can follow the `Docker official instructions <https://docs.docker.com/get-docker/>`_.
2. Pull the AFNI Docker image from `Docker Hub <https://hub.docker.com/r/afni/afni>`_:
1. Install Docker. You can follow the
`Docker official instructions <https://docs.docker.com/get-docker/>`_.
2. Pull the AFNI Docker image from
`Docker Hub <https://hub.docker.com/r/afni/afni>`_:
.. code-block:: shell
.. code-block:: bash
docker pull afni/afni_make_build
3. Add the Junifer AFNI scripts to your PATH environmental variable. Run the following command:
3. Add the Junifer AFNI scripts to your PATH environmental variable. Run the
following command:
.. code-block:: shell
.. code-block:: bash
junifer setup afni-docker
Take the last line and copy it to your ``.bashrc`` or ``.zshrc`` file.
Or, alternatively, you can exceute this command which will update the ``~/.bashrc`` for you:
Or, alternatively, you can exceute this command which will update the
``~/.bashrc`` for you:
.. code-block:: shell
.. code-block:: bash
junifer setup afni-docker | grep "PATH=" | xargs | >> ~/.bashrc

View file

@ -78,7 +78,8 @@ to generate the proper changelog that should be reflected in
git push origin --follow-tags
#. Optional: bump the *MAJOR* or *MINOR* segment of next release (replace ``D.E.0`` with the proper version).
#. Optional: bump the *MAJOR* or *MINOR* segment of next release (replace
``D.E.0`` with the proper version).
.. code-block:: bash

View file

@ -2,12 +2,12 @@
.. _starting:
First steps with Junifer
First steps with junifer
========================
.. note::
.. note::
To scroll the graph left and right, click on the graph and use the arrows on your keyboard.
To scroll the graph left and right, click on the graph and use the arrows on your keyboard.
.. mermaid::
@ -74,7 +74,6 @@ First steps with Junifer
question_contribute_marker -->|No| final_run
contribute_marker(Create a\nMARKER REQUEST\nissue on Github)
contribute_marker --> final_run
missing_preprocessing --> contact_help
@ -101,13 +100,13 @@ First steps with Junifer
read_adding_parcellation_start --> read_adding_parcellation
read_adding_parcellation("Read Adding Parcellations")
read_adding_parcellation --> missing_other_solved
missing_coordinates --> read_adding_coordinates_start
read_adding_coordinates_start("Read Creating a Junifer extension")
read_adding_coordinates_start --> read_adding_coordinates
read_adding_coordinates("Read Adding Coordinates")
read_adding_coordinates --> missing_other_solved
missing_other_solved{Did you solve your issue?}
missing_other_solved -->|Yes| read_using_final
missing_other_solved -->|No| missing_other_contact
@ -133,7 +132,7 @@ First steps with Junifer
final_queue(Use junifer queue to compute your features)
final_queue --> final_magic
final_magic(((Let junifer do its magic!)))
click read_understanding href "https://juaml.github.io/junifer/main/understanding/index.html"
click read_using href "https://juaml.github.io/junifer/main/using/index.html"
click read_using_final href "https://juaml.github.io/junifer/main/using/index.html"
@ -155,4 +154,4 @@ First steps with Junifer
click error_contact href "https://juaml.github.io/junifer/main/help.html"
click contact_help href "https://juaml.github.io/junifer/main/help.html"
click missing_other_contact href "https://juaml.github.io/junifer/main/help.html"
click missing_other_contact href "https://juaml.github.io/junifer/main/help.html"

View file

@ -10,23 +10,87 @@ Description
This is the *object* that traverses the steps of the pipeline. It is indeed a
dictionary of dictionaries. The first level of keys are the :ref:`data types <data_types>`
and a special key named ``meta`` that contains all the information on the data
object including source and previous transformation steps.
and the values are the corresponding information as dictionaries.
The second level of keys are the actual data. So far, there are two keys used:
.. code-block:: python
- ``path``: path to the file containing the data.
- ``data``: the data loaded in memory.
{'BOLD': {...}, 'T1w': {...}}
The :ref:`Data Grabber <datagrabber>` step will only fill the ``path`` value.
The ``data`` value will be filled by the :ref:`DataReader <datareader>` step, if it is one of the possible file types
that the datareader can read.
The second level of keys are the actual data. A special second-level key named
``meta`` is present in each step, that contains all the information on the
data type including source and previous transformation steps.
A point to note is that you never directly interact with the *data object* but it's important to know where and how the object is being manipulated to reason about your pipeline.
The :ref:`Data Grabber <datagrabber>` step adds the ``path`` second-level key
which gives the path to the file containing the data. The ``meta`` key in this
step only contains information about the datagrabber used.
.. code-block:: python
{'BOLD': {'meta': {'datagrabber': {'class': 'SPMAuditoryTestingDatagrabber',
'types': ['BOLD', 'T1w']},
'dependencies': set(),
'element': {'subject': 'sub001'}},
'path': PosixPath('/var/folders/dv/2lbr8f8j0q12zrx3mz3ll5m40000gp/T/tmpgxcyjfo1/sub001_bold.nii.gz')},
'T1w': {'meta': {'datagrabber': {'class': 'SPMAuditoryTestingDatagrabber',
'types': ['BOLD', 'T1w']},
'dependencies': set(),
'element': {'subject': 'sub001'}},
'path': PosixPath('/var/folders/dv/2lbr8f8j0q12zrx3mz3ll5m40000gp/T/tmpgxcyjfo1/sub001_T1w.nii.gz')}}
The :ref:`Data Reader <datareader>` step adds the ``data`` second-level key
which is the actual data loaded into memory. The ``meta`` key in this step
adds information about the datareader used to read the data.
.. code-block:: python
{'BOLD': {'data': <nibabel.nifti1.Nifti1Image object at 0x16b5d8910>,
'meta': {'datagrabber': {'class': 'SPMAuditoryTestingDatagrabber',
'types': ['BOLD', 'T1w']},
'datareader': {'class': 'DefaultDataReader'},
'dependencies': {'nilearn'},
'element': {'subject': 'sub001'}},
'path': PosixPath('/var/folders/dv/2lbr8f8j0q12zrx3mz3ll5m40000gp/T/tmpe49321ce/sub001_bold.nii.gz')},
'T1w': {'data': <nibabel.nifti1.Nifti1Image object at 0x16b5d78d0>,
'meta': {'datagrabber': {'class': 'SPMAuditoryTestingDatagrabber',
'types': ['BOLD', 'T1w']},
'datareader': {'class': 'DefaultDataReader'},
'dependencies': set(),
'element': {'subject': 'sub001'}},
'path': PosixPath('/var/folders/dv/2lbr8f8j0q12zrx3mz3ll5m40000gp/T/tmpe49321ce/sub001_T1w.nii.gz')}}
The :ref:`Preprocess <preprocess>` step, if used, modifies the ``data``
second-level key's value and appends the ``meta`` key with information about
the preprocessor.
The :ref:`Marker <marker>` step removes the ``path`` second-level key,
replaces the ``data`` second-level key's value with the marker's computed value
and adds further keys needed for the storage, for example, ``col_names``.
.. code-block:: python
{'BOLD': {'col_names': ['root_sum_of_squares_ets'],
'data': ...,
'meta': {'datagrabber': {'class': 'SPMAuditoryTestingDatagrabber',
'types': ['BOLD', 'T1w']},
'datareader': {'class': 'DefaultDataReader'},
'dependencies': {'nilearn'},
'element': {'subject': 'sub001'},
'marker': {'agg_method': 'mean',
'agg_method_params': None,
'class': 'RSSETSMarker',
'masks': None,
'name': 'RSSETSMarker',
'parcellation': 'Schaefer100x17'},
'type': 'BOLD'}}}
.. note::
You never directly interact with the *data object* but it's important to know
where and how the object is being manipulated to reason about your pipeline.
.. _data_types:
Data types
Data Types
----------
.. list-table::
@ -59,4 +123,4 @@ Data types
- GCOR computed with CONN toolbox
* - ``LCOR``
- Local Correlation image (3D)
- LCOR computed with CONN toolbox
- LCOR computed with CONN toolbox

View file

@ -8,27 +8,30 @@ Data Grabber
Description
-----------
The *Data Grabber* is an object that can provide an interface to datasets you want to work with in junifer.
Every concrete implementation of a datagrabber is aware of a particular dataset's structure and thus allows
you to fetch specific elements of interest from the dataset. It adds the ``path`` key to each :ref:`data type <data_types>`
in the :ref:`Data object <data_object>`.
The ``DataGrabber`` is an object that can provide an interface to datasets you
want to work with in junifer. Every concrete implementation of a datagrabber is
aware of a particular dataset's structure and thus allows you to fetch specific
elements of interest from the dataset. It adds the ``path`` key to each
:ref:`data type <data_types>` in the :ref:`Data object <data_object>`.
Datagrabbers are intended to be used as context managers. When used within a context, a datagrabber takes care
of any pre and post steps for interacting with the dataset, for example, downloading and cleaning up. As the interface
Datagrabbers are intended to be used as context managers. When used within a
context, a datagrabber takes care of any pre and post steps for interacting with
the dataset, for example, downloading and cleaning up. As the interface
is consistent, you always use the same procedure to interact with the datagrabber.
For example, a concrete implementation of :class:`.DataladDataGrabber` can provide junifer
with data from a Datalad dataset. Of course, datagrabbers are not only meant to work with Datalad datasets but
any dataset.
For example, a concrete implementation of :class:`.DataladDataGrabber` can
provide junifer with data from a Datalad dataset. Of course, datagrabbers are not
only meant to work with Datalad datasets but any dataset.
If you are interested in using already provided datagrabbers, please go to :doc:`../builtin`. And, if you want
to implement your own datagrabber, you need to provide concrete implementations of base classes already
provided.
If you are interested in using already provided datagrabbers, please go to
:doc:`../builtin`. And, if you want to implement your own datagrabber, you need
to provide concrete implementations of base classes already provided.
Base classes
Base Classes
------------
In this section, we showcase different abstract base classes you might want to use to implement your own datagrabber.
In this section, we showcase different abstract base classes you might want to
use to implement your own datagrabber.
.. list-table::
:widths: auto
@ -37,18 +40,26 @@ In this section, we showcase different abstract base classes you might want to u
* - Name
- Description
* - :class:`.BaseDataGrabber`
- | The abstract base class providing you an interface to implement your own datagrabber.
| You should try to avoid using this directly and instead use
| :class:`.PatternDataGrabber` or :class:`.DataladDataGrabber`.
| To build your own custom *low-level* datagrabber, you need to at least implement the ``get_elements`` method,
| but most of the time you should also override other existing methods like ``__enter__`` and ``__exit__``.
- | The abstract base class providing you an interface to implement your
| own datagrabber. You should try to avoid using this directly and
| instead use :class:`.PatternDataGrabber` or
| :class:`.DataladDataGrabber`. To build your own custom *low-level*
| datagrabber, you need to override the ``get_elements_keys``,
| ``get_elements`` and ``get_item`` methods, and most of the time you
| should also override other existing methods like ``__enter__`` and
| ``__exit__``.
* - :class:`.PatternDataGrabber`
- | It implements functionality to help you define the pattern of the dataset you want to get. For example,
| you know that T1 images are found in a directory following this pattern ``{subject}/anat/{subject}_T1w.nii.gz``
| inside of the dataset. Now you can provide this to the **PatternDataGrabber** and it will be able to get the file.
- | It implements functionality to help you define the pattern of the
| dataset you want to get. For example, you know that T1 images are
| found in a directory following the pattern:
| ``{subject}/anat/{subject}_T1w.nii.gz`` inside of the dataset. Now you
| can provide this to the :class:`.PatternDataGrabber` and it will be
| able to get the file.
* - :class:`.DataladDataGrabber`
- | It implements functionality to deal with Datalad datasets. Specifically, the ``__enter__`` and ``__exit__`` methods
| take care of cloning and removing the Datalad dataset.
- | It implements functionality to deal with Datalad datasets. Specifically,
| the ``__enter__`` and ``__exit__`` methods take care of cloning and
| removing the Datalad dataset.
* - :class:`.PatternDataladDataGrabber`
- | It is a combination of :class:`.PatternDataladDataGrabber` and
| :class:`.DataladDataGrabber`. This is probably the class you are looking for when using Datalad.
- | It is a combination of :class:`.PatternDataGrabber` and
| :class:`.DataladDataGrabber`. This is probably the class you are looking
| for when using Datalad datasets.

View file

@ -8,22 +8,24 @@ Data Reader
Description
-----------
The *Data Reader* is an object that is responsible for actually reading data files in junifer.
It reads the value of the key ``path`` for each :ref:`data type <data_types>` in the :ref:`Data object <data_object>`
and loads them to memory. After reading the data into memory, it adds the key ``data`` to the same level as ``path``
and the value is the actual data in the memory.
The ``DataReader`` is an object that is responsible for actually reading data
files in junifer. It reads the value of the key ``path`` for each
:ref:`data type <data_types>` in the :ref:`Data object <data_object>` and loads
them into memory. After reading the data into memory, it adds the key ``data``
to the same level as ``path`` and the value is the actual data in the memory.
Datareaders are meant to be used inside the datagrabber context but you can operate on them outside the context as long
as the actual data is in the memory and the Python runtime has not garbage-collected it.
Datareaders are meant to be used inside the datagrabber context but you can
operate on them outside the context as long as the actual data is in the memory
and the Python runtime has not garbage-collected it.
For data formats not supported by junifer yet, you can either make your own *Data Reader* or open an issue on
`junifer Github`_ and we can help you out.
For data formats not supported by junifer yet, you can either make your own
*Data Reader* or open an issue on `junifer Github`_ and we can help you out.
Currently supported file-formats
--------------------------------
File Formats
------------
We already provide a concrete implementation :class:`.DefaultDataReader` which knows how to
read the following file formats:
We already provide a concrete implementation :class:`.DefaultDataReader` which
knows how to read the following file formats:
.. list-table::
:widths: auto

View file

@ -10,17 +10,19 @@ tool conceived to extract features from neuroimaging data in an easy-to-use
manner, with minimal coding and minimal user expertise in the internal aspects.
Unlike other tools like FSL, SPM, AFNI, etc., junifer is not a toolbox to
pre-process data, but a toolbox to extract features from previously pre-processed
data.
pre-process data, but a toolbox to extract features from previously
pre-processed data.
The main idea is that you have a set of images (e.g. a set of functional MRI,
structural MRI, diffusion MRI, etc.) and you want to extract features to
later use in statistical analyses or machine learning (for example, using
julearn_).
.. important:: Junifer is not a toolbox to create pipelines, but a tool to configure the junifer pipeline, which is
intended to be fixed and not to be changed. If you want to create a pipeline, you should use
other tools like nipype_.
.. important::
Junifer is not a toolbox to create pipelines, but a tool to configure the
junifer pipeline, which is intended to be fixed and not to be changed. If you
want to create a pipeline, you should use other tools like nipype_.
.. toctree::
:maxdepth: 2

View file

@ -8,16 +8,25 @@ Marker
Description
-----------
The ``Marker`` is an object that is responsible for feature extraction. It primarily operates on data loaded
in memory by :ref:`Data Reader <datareader>` and stored in the ``data`` key of each :ref:`data type <data_types>`
in the :ref:`Data object <data_object>`. In some cases, it can also operate on pre-processed data as obtained
from the :ref:`Preprocess <preprocess>` step of the pipeline. It is important to note that this pre-process is
not similar to pre-processing done by tools like FSL, SPM, AFNI, etc. . For example, one can perform confound
removal on loaded data and then perform feature extraction.
The ``Marker`` is an object that is responsible for feature extraction. It
primarily operates on data loaded into memory by :ref:`Data Reader <datareader>`
and stored in the ``data`` key of each :ref:`data type <data_types>` in the
:ref:`Data object <data_object>`. In some cases, it can also operate on
pre-processed data as obtained from the :ref:`Preprocess <preprocess>` step of
the pipeline.
Markers are meant to be used inside the datagrabber context but you can operate on them outside the context as long
as the actual data is in the memory and the Python runtime has not garbage-collected it.
.. important::
If you are interested in using already provided markers, please go to :doc:`../builtin`. And, if you want to implement
your own marker, you need to provide concrete implementation of :class:`.BaseMarker`. Specifically, you
need to override ``get_output_type``, ``store`` and ``compute`` methods.
This pre-process is not similar to pre-processing done by tools like FSL,
SPM, AFNI, etc., . For example, one can perform confound removal on loaded
data and then perform feature extraction.
Markers are meant to be used inside the datagrabber context but you can operate
on them outside the context as long as the actual data is in the memory and the
Python runtime has not garbage-collected it.
If you are interested in using already provided markers, please go to
:doc:`../builtin`. And, if you want to implement your own marker, you need to
provide concrete implementation of :class:`.BaseMarker`. Specifically, you
need to override ``get_valid_inputs``, ``get_output_type`` and ``compute``
methods.

View file

@ -2,22 +2,26 @@
.. _pipeline:
The Junifer Pipeline
The junifer Pipeline
====================
The junifer pipeline is the main execution path of junifer. It consists of five steps:
The junifer pipeline is the main execution path of junifer. It consists of five
steps:
1. :ref:`Data Grabber <datagrabber>`: Interpret the dataset and provide a list of files.
1. :ref:`Data Grabber <datagrabber>`: Interpret the dataset and provide a list
of files.
2. :ref:`Data Reader <datareader>`: Read the files.
3. :ref:`Pre-processing <preprocess>`: Prepare the images for marker computation.
3. :ref:`Pre-processing <preprocess>`: Prepare the images for marker
computation.
4. :ref:`Marker Computation <marker>`: Compute the marker.
5. :ref:`Storage <storage>`: Store the marker values.
The element that is passed accross the pipeline is called the :ref:`Data Object<data_object>`.
The element that is passed accross the pipeline is called the
:ref:`Data Object<data_object>`.
The following is a graphical representation of the pipeline:
.. mermaid::
.. mermaid::
flowchart LR
dg[Data Grabber]
@ -31,11 +35,12 @@ The following is a graphical representation of the pipeline:
mc --> st
However, it is usually the case that several markers are computed for the same data. Thus, the *markers* step
of the pipeline is defined as a list of markers. The following is a graphical representation of the pipeline execution
However, it is usually the case that several markers are computed for the same
data. Thus, the ``Marker Computation`` step of the pipeline is defined as a list
of markers. The following is a graphical representation of the pipeline execution
on multiple markers:
.. mermaid::
.. mermaid::
flowchart LR
dg[Data Grabber]
@ -64,5 +69,8 @@ on multiple markers:
mc4 --> st4
mc5 --> st5
.. note:: To avoid keeping in memory all of the computed marker, the storage step is called after each marker
computation, releasing the memory used to compute each marker.
.. note::
To avoid keeping in memory all of the computed marker, the storage step is
called after each marker computation, releasing the memory used to compute
each marker.

View file

@ -8,32 +8,38 @@ Preprocess
Description
-----------
The ``Preprocess`` step of the pipeline is meant for pre-processing before or after :ref:`Marker <marker>` step
depending on the use-case. For example, you might want to perform confound removal on ``BOLD`` data before
feature extraction.
The ``Preprocess`` is an object meant for pre-processing before or after
:ref:`Marker <marker>` step depending on the use-case. For example, you might
want to perform confound removal on ``BOLD`` data before feature extraction.
This step is still under development and is an optional one for the pipeline to work.
.. note::
This step is optional for the pipeline to work.
.. _preprocess_confounds:
Confound Removal
----------------
The *Confound Removal* step is meant to remove *confounds* from the ``BOLD`` data. The confounds are
extracted from the ``BOLD_confounds`` data (must be provided by the :ref:`Data Grabber <datagrabber>`).
The confounds are then regressed out from the ``BOLD`` data using :func:`nilearn.image.clean_img`.
The *Confound Removal* step is meant to remove *confounds* from the ``BOLD``
data. The confounds are extracted from the ``BOLD_confounds`` data (must be
provided by the :ref:`Data Grabber <datagrabber>`). The confounds are then
regressed out from the ``BOLD`` data using :func:`nilearn.image.clean_img`.
Currently, junifer supports only one confound removal class:
:class:`.fMRIPrepConfoundRemover`. This class is meant to remove confounds as described
before, using the output of `fMRIPrep`_ as reference.
:class:`.fMRIPrepConfoundRemover`. This class is meant to remove confounds as
described before, using the output of `fMRIPrep`_ as reference.
Strategy
~~~~~~~~
This confound remover uses the `nilearn`_ API from :func:`nilearn.interfaces.fmriprep.load_confounds`. That is,
define a *strategy* to extract the confounds from the ``BOLD_confounds`` data. The *strategy* is defined
by choosing the *noise components* to be used and the *confounds* to be extracted from each noise components.
The *noise components* currently supported are:
This confound remover uses the `nilearn`_ API from
:func:`nilearn.interfaces.fmriprep.load_confounds`. That is, define a *strategy*
to extract the confounds from the ``BOLD_confounds`` data. The *strategy* is
defined by choosing the *noise components* to be used and the *confounds* to be
extracted from each noise components. The *noise components* currently supported
are:
* ``motion``
* ``wm_csf``
@ -41,20 +47,24 @@ The *noise components* currently supported are:
The confounds options for each *noise component* are:
* ``basic``: the basic confounds for each *noise component*. For example, for ``motion``, the basic confounds
are the 6 motion parameters (3 translations and 3 rotations). For ``wm_csf``, the basic
confounds are the mean signal of the white matter and CSF regions. For ``global_signal``, the basic confound
* ``basic``: the basic confounds for each *noise component*. For example, for
``motion``, the basic confounds are the 6 motion parameters (3 translations
and 3 rotations). For ``wm_csf``, the basic confounds are the mean signal of
the white matter and CSF regions. For ``global_signal``, the basic confound
is the mean signal of the whole brain.
* ``power2``: the basic confounds plus the square of each basic confound.
* ``derivatives``: the basic confounds plus the derivative of each basic confound.
* ``full``: the basic confounds, the derivative of each basic confound, the square of each basic confound and the
square of each derivative of each basic confound.
* ``derivatives``: the basic confounds plus the derivative of each basic
confound.
* ``full``: the basic confounds, the derivative of each basic confound, the
square of each basic confound and the square of each derivative of each basic
confound.
The *strategy* is defined as a dictionary, with the *noise components* as keys and the *confounds* as values.
The *strategy* is defined as a dictionary, with the *noise components* as keys
and the *confounds* as values.
Example in python format:
.. code-block::
.. code-block:: python
strategy = {
"motion": "basic",
@ -64,7 +74,7 @@ Example in python format:
or in YAML format:
.. code-block::
.. code-block:: yaml
strategy:
motion: basic
@ -73,7 +83,7 @@ or in YAML format:
The default value is to use all the *noise components* with the ``full`` *confounds*:
.. code-block::
.. code-block:: python
strategy = {
"motion": "full",
@ -81,20 +91,22 @@ The default value is to use all the *noise components* with the ``full`` *confou
"global_signal": "full"
}
Other parameters
Other Parameters
~~~~~~~~~~~~~~~~
Additionaly, the :class:`.fMRIPrepConfoundRemover` supports the following parameters:
Additionaly, the :class:`.fMRIPrepConfoundRemover` supports the following
parameters:
.. list-table::
:widths: 10, 30, 5
:widths: auto
:header-rows: 1
* - Parameter
- Description
- Default
* - ``spike``
- Add a spike regressor in the timepoints when the framewise displacement exceeds this threshold.
- | Add a spike regressor in the timepoints when the framewise
| displacement exceeds this threshold.
- deactivated
* - ``detrend``
- Apply detrending on timeseries, before confound removal.
@ -112,5 +124,7 @@ Additionaly, the :class:`.fMRIPrepConfoundRemover` supports the following parame
- Repetition time, in second (sampling period).
- from nifti header
* - ``mask``
- If provided, signal is only cleaned from voxels inside the mask. If not, a mask is computed using :func:`nilearn.masking.compute_brain_mask`.
- | If provided, signal is only cleaned from voxels inside the mask.
| If not, a mask is computed using
| :func:`nilearn.masking.compute_brain_mask`.
- compute

View file

@ -8,27 +8,32 @@ Storage
Description
-----------
The ``Storage`` is an object that is responsible for storing extracted features as computed from :ref:`Marker <marker>`
step of the pipeline. If the pipeline is provided with a ``storage-like`` object, the extracted features are stored via
The ``Storage`` is an object that is responsible for storing extracted features
as computed from :ref:`Marker <marker>` step of the pipeline. If the pipeline is
provided with a ``storage-like`` object, the extracted features are stored via
that object else they are kept in memory.
Storage is meant to be used inside the datagrabber context but you can operate on them outside the context as long
as the processed data is in the memory and the Python runtime has not garbage-collected it.
Storage is meant to be used inside the datagrabber context but you can operate
on them outside the context as long as the processed data is in the memory and
the Python runtime has not garbage-collected it.
The :ref:`Markers <marker>` are responsible for defining what *storage kind* (``matrix``, ``vector``, ``timeseries``)
they support for which :ref:`data type <data_types>` by overriding its ``store`` method. The storage object in turn
declares and provides implementation for specific *storage kind*. For example, :class:`.SQLiteFeatureStorage`
supports saving ``matrix``, ``vector`` and ``timeseries`` via ``store_matrix``, ``store_vector`` and ``store_timeseries``
methods respectively.
For storage interfaces not supported by junifer yet, you can either make your own ``Storage`` by providing a concrete
implementation of :class:`.BaseFeatureStorage` or open an issue on `junifer Github`_ and we can help you out.
The :ref:`Markers <marker>` are responsible for defining what *storage kind*
(``matrix``, ``vector``, ``timeseries``) they support for which
:ref:`data type <data_types>` by overriding its ``get_output_type`` method. The
storage object in turn declares and provides implementation for specific
*storage kind*. For example, :class:`.SQLiteFeatureStorage` supports saving
``matrix``, ``vector`` and ``timeseries`` via ``store_matrix``, ``store_vector``
and ``store_timeseries`` methods respectively.
For storage interfaces not supported by junifer yet, you can either make your
own ``Storage`` by providing a concrete implementation of
:class:`.BaseFeatureStorage` or open an issue on `junifer Github`_ and we can
help you out.
.. _storage_types:
Currently supported storage types
---------------------------------
Storage Types
-------------
.. list-table::
:widths: auto
@ -43,18 +48,18 @@ Currently supported storage types
- ``col_names``, ``row_names``, ``matrix_kind``, ``diagonal``
- :meth:`.BaseFeatureStorage.store_matrix`
* - ``vector``
- A vector of values with column names
- ``columns``, ``row_names``
- A 1D row vector of values with column names
- ``col_names``
- :meth:`.BaseFeatureStorage.store_vector`
* - ``timeseries``
- A 2D matrix of values with column names
- ``columns``, ``row_names``
- ``col_names``
- :meth:`.BaseFeatureStorage.store_timeseries`
.. _storage_interfaces:
Currently supported storage interfaces
--------------------------------------
Storage Interfaces
------------------
.. list-table::
:widths: auto

View file

@ -2,11 +2,12 @@
.. _codeless:
Code-less configuration
Code-less Configuration
=======================
On of the most important features of junifer is its capacity to run without writing a single line of code. This is
achieved by using a configuration file that is written in YAML_. In this file, we configure the different steps of
On of the most important features of junifer is its capacity to run without
writing a single line of code. This is achieved by using a configuration file
that is written in YAML_. In this file, we configure the different steps of
:ref:`pipeline`.
As a reminder, this is how the pipeline looks like:
@ -25,31 +26,34 @@ As a reminder, this is how the pipeline looks like:
mc --> st
Thus, the configuration file must configure each of the sections of the pipeline, as well as some general parameters.
Thus, the configuration file must configure each of the sections of the pipeline,
as well as some general parameters.
As an example, we will generate the configuration file for a pipeline that will extract the mean ``VBM_GM`` values
using two different parcellations and one set of coordinates, from the *Oasis VBM Testing dataset* included in
junifer.
As an example, we will generate the configuration file for a pipeline that will
extract the mean ``VBM_GM`` values using two different parcellations and one set
of coordinates, from the ``Oasis VBM Testing dataset`` included in junifer.
General Parameters
------------------
The general parameters are the ones that are not specific to any of the sections of the pipeline, but configure
junifer as a whole. These parameters are:
The general parameters are the ones that are not specific to any of the sections
of the pipeline, but configure junifer as a whole. These parameters are:
* ``with``: A section used to specify modules and junifer extensions to use.
* ``workdir``: The working directory where junifer will store temporary files.
Since the example uses a specific datagrabber for testing, we need to add ``junifer.testing.registry`` to the
``with`` section. This will allow junifer to find the datagrabber. We will set the ``workdir`` to ``/tmp``.
Since the example uses a specific datagrabber for testing, we need to add
``junifer.testing.registry`` to the ``with`` section. This will allow junifer
to find the datagrabber. We will set the ``workdir`` to ``/tmp``.
.. code-block:: yaml
with: junifer.testing.registry
workdir: /tmp
Step-by-step configuration
Step-by-step Configuration
--------------------------
In order to configure the pipeline, we need to configure each step:
@ -60,15 +64,19 @@ In order to configure the pipeline, we need to configure each step:
* ``markers``
* ``storage``
.. important:: The datareader step configuration is optional, as junifer only provides one datareader. Nevertheless,
it is possible to extend junifer with custom datareaders, and thus, it is also possible to configure this step.
.. important::
The datareader step configuration is optional, as junifer only provides one
datareader. Nevertheless, it is possible to extend junifer with custom
datareaders, and thus, it is also possible to configure this step.
Data Grabber
^^^^^^^^^^^^
The ``datagrabber`` section must be configured using the ``kind`` key to specify the datagrabber to use. Additional
keys correspond to the parameters of the datagrabber.
The ``datagrabber`` section must be configured using the ``kind`` key to specify
the datagrabber to use. Additional keys correspond to the parameters of the
datagrabber.
For example, to use the :class:`.DataladAOMICPIOP1` datagrabber, we just need to
specify its name as the ``kind`` key.
@ -78,8 +86,8 @@ specify its name as the ``kind`` key.
datagrabber:
kind: DataladAOMICPIOP1
However, it is also possible to pass parameters to the datagrabber. In this case, we can restrict the datagrabber to
fetch only the ``restingstate`` task.
However, it is also possible to pass parameters to the datagrabber. In this case,
we can restrict the datagrabber to fetch only the ``restingstate`` task.
.. code-block:: yaml
@ -87,22 +95,24 @@ fetch only the ``restingstate`` task.
kind: DataladAOMICPIOP1
tasks: restingstate
In the *Oasis VBM Testing dataset* example, the section will look like this:
In the ``Oasis VBM Testing dataset`` example, the section will look like this:
.. code-block:: yaml
datagrabber:
kind: OasisVBMTesting
kind: OasisVBMTestingDatagrabber
Data Reader
^^^^^^^^^^^
As mentioned before, this section is entirely optional, as junifer only provides one data reader
(:class:`.DefaultDataReader`), which is the default in case the section is not specified.
As mentioned before, this section is entirely optional, as junifer only provides
one data reader (:class:`.DefaultDataReader`), which is the default in case the
section is not specified.
In any case, the syntax of the section is the same as for the ``datagrabber`` section, using the ``kind`` key to
specify the data reader to use, and additional keys to pass parameters to the data reader:
In any case, the syntax of the section is the same as for the ``datagrabber``
section, using the ``kind`` key to specify the datareader to use, and additional
keys to pass parameters to the datareader:
.. code-block:: yaml
@ -110,18 +120,19 @@ specify the data reader to use, and additional keys to pass parameters to the da
kind: DefaultDataReader
For the *Oasis VBM Testing dataset* example, we will not specify a ``datareader`` step.
For the ``Oasis VBM Testing dataset`` example, we will not specify a
``datareader`` step.
Preprocessing
^^^^^^^^^^^^^
Preprocess
^^^^^^^^^^
Preprocessing is also an optional step, as it might be the case that no pre-processing is needed. In the case that
preprocessing is needed, the section must be configured using the ``kind`` key to specify the preprocessor to use,
Pre-processing is also an optional step, as it might be the case that no
pre-processing is needed. In the case that pre-processing is needed, the section
must be configured using the ``kind`` key to specify the preprocessor to use,
and additional keys to pass parameters to the preprocessor.
For example, to use the :class:`.fMRIPrepConfoundRemover` preprocessor, we just need to specify its
name as the ``kind`` key, as well as its parameters.
For example, to use the :class:`.fMRIPrepConfoundRemover` preprocessor, we just
need to specify its name as the ``kind`` key, as well as its parameters.
.. code-block:: yaml
@ -136,18 +147,22 @@ name as the ``kind`` key, as well as its parameters.
standardize: true
For the *Oasis VBM Testing dataset* example, we will not specify a preprocessing step.
For the ``Oasis VBM Testing dataset`` example, we will not specify a
preprocessing step.
Markers
^^^^^^^
Marker
^^^^^^
The ``markers`` section diverges from the previous ones, as we need to specify a list of markers. Each marker has a
name that we can use to refer to it later, and a set of parameters that will be passed to the marker.
The ``markers`` section diverges from the previous ones, as we need to specify
a list of markers. Each marker has a name that we can use to refer to it later,
and a set of parameters that will be passed to the marker.
For the *Oasis VBM Testing dataset* example, we want to compute the mean ``VBM_GM`` value for each parcel using the
Schaefer parcellation (100 parcels, 7 networks), Schaefer parcellation (200 parcels, 7 networks), and the *DMNBuckner*
network, using 5mm spheres. Thus, we will configure the ``markers`` section as follows:
For the ``Oasis VBM Testing dataset`` example, we want to compute the mean
``VBM_GM`` value for each parcel using the ``Schaefer parcellation (100 parcels,
7 networks)``, ``Schaefer parcellation (200 parcels, 7 networks)``, and the
``DMNBuckner`` network, using ``5mm`` spheres. Thus, we will configure the
``markers`` section as follows:
.. code-block:: yaml
@ -170,11 +185,12 @@ network, using 5mm spheres. Thus, we will configure the ``markers`` section as f
Storage
^^^^^^^
Finally, we need to define how and where the results will be stored. This is done using the ``storage`` section,
which must be configured using the ``kind`` key to specify the storage to use, and additional keys to pass parameters.
Finally, we need to define how and where the results will be stored. This is
done using the ``storage`` section, which must be configured using the ``kind``
key to specify the storage to use, and additional keys to pass parameters.
For example, to use the :class:`.SQLiteFeatureStorage` storage, we just need to specify where we want
to store the results:
For example, to use the :class:`.SQLiteFeatureStorage` storage, we just need to
specify where we want to store the results:
.. code-block:: yaml
@ -183,18 +199,20 @@ to store the results:
uri: /data/junifer/example/oasis_vbm_testing.sqlite
The full example
Complete Example
----------------
This is how the full *Oasis VBM Testing dataset* example configuration file looks like:
This is how the full ``Oasis VBM Testing dataset`` example configuration file
looks like:
.. code-block:: yaml
with: junifer.testing.registry
workdir: /tmp
datagrabber:
kind: OasisVBMTesting
kind: OasisVBMTestingDatagrabber
markers:
- name: Schaefer100x7_mean

View file

@ -5,9 +5,11 @@
Using junifer
=============
In this section, we will cover the main aspects behind using junifer. We will first explain the basics behind junifer's
code-less configuration. Then we will show how to use the command line interface to ``run`` junifer and ``collect``
the results. Finally, we will show how to use the ``queue`` command to interact with HPC and HTC systems.
In this section, we will cover the main aspects behind using junifer. We will
first explain the basics behind junifer's code-less configuration. Then we will
show how to use the command line interface to ``run`` junifer and ``collect``
the results. Finally, we will show how to use the ``queue`` command to interact
with HPC and HTC systems.
.. toctree::
:maxdepth: 2
@ -20,13 +22,14 @@ the results. Finally, we will show how to use the ``queue`` command to interact
.. _using_components:
Using junifer common components
-------------------------------
Using Common Components
-----------------------
The following sections explains common components of junifer that can be used across many steps of the pipeline.
The following sections explains common components of junifer that can be used
across many steps of the pipeline.
.. toctree::
:maxdepth: 2
:caption: Contents:
masks
masks

View file

@ -5,27 +5,32 @@
Masks
=====
Masks are essentially boolean arrays that are used to constrain the extraction of features to voxels that are
meaningful. For example, in an fMRI imaging study, a mask can be used to constrain the extraction of features to
voxels that contain a certain ratio of gray matter to white matter / cerebrospinal fluid, ensuring that the features
are not extracted from voxels that contain mostly white matter or cerebrospinal fluid, which could add noise to the
BOLD signal.
Masks are essentially boolean arrays that are used to constrain the extraction
of features to voxels that are meaningful. For example, in an fMRI imaging
study, a mask can be used to constrain the extraction of features to voxels that
contain a certain ratio of gray matter to white matter / cerebrospinal fluid,
ensuring that the features are not extracted from voxels that contain mostly
white matter or cerebrospinal fluid, which could add noise to the BOLD signal.
Junifer provides a number of built-in masks, which can be listed using the :func:`.list_masks`. Some
masks are images, while other masks can be computed using :ref:`nilearn` functions.
Junifer provides a number of built-in masks, which can be listed using
:func:`.list_masks`. Some masks are images, while other masks can be computed
using :ref:`nilearn` functions.
For markers and steps that accept ``masks`` as an argument, the mask can be specified as a string, which will be the
name of a built-in mask, or as a dictionary in which the **only** key is the built-in mask name and the value is a
dictionary of keyword arguments to pass to the mask function.
For markers and steps that accept ``masks`` as an argument, the mask can be
specified as a string, which will be the name of a built-in mask, or as a
dictionary in which the **only** key is the built-in mask name and the value is
a dictionary of keyword arguments to pass to the mask function.
For example, the following is a valid mask specification that specified the ``GM_prob0.2`` mask.
For example, the following is a valid mask specification that specified the
``GM_prob0.2`` mask.
.. code-block:: yaml
masks: GM_prob0.2
The following is a valid mask specification that specifies the ``compute_brain_mask`` mask (function from nilearn),
with a threshold of 0.5.
The following is a valid mask specification that specifies the
``compute_brain_mask`` mask (function from nilearn), with a threshold of
``0.5``.
.. code-block:: yaml
@ -33,9 +38,11 @@ with a threshold of 0.5.
compute_brain_mask:
threshold: 0.5
Furthermore, junifer allows you to combine several masks using :func:`nilearn.masking.intersect_masks`. This is done by
specifying a list of masks, where each mask is a string or dictionary as described above. For example, the following
is a valid mask specification that specifies the intersection of the ``GM_prob0.2`` and ``compute_brain_mask`` masks.
Furthermore, junifer allows you to combine several masks using
:func:`nilearn.masking.intersect_masks`. This is done by specifying a list of
masks, where each mask is a string or dictionary as described above. For example,
the following is a valid mask specification that specifies the intersection of
the ``GM_prob0.2`` and ``compute_brain_mask`` masks.
.. code-block:: yaml
@ -44,9 +51,9 @@ is a valid mask specification that specifies the intersection of the ``GM_prob0.
- compute_brain_mask:
threshold: 0.5
We can also specify the arguments of :func:`nilearn.masking.intersect_masks` (``threshold`` and ``connected``). The
following example combines the same masks as the previous one, but computing the full intersection.
We can also specify the arguments of :func:`nilearn.masking.intersect_masks`
(``threshold`` and ``connected``). The following example combines the same masks
as the previous one, but computing the full intersection.
.. code-block:: yaml
@ -56,7 +63,8 @@ following example combines the same masks as the previous one, but computing the
threshold: 0.5
- threshold: 1 # intersection
Alternatively, we can also compute the union, even if the voxels do not form a connected component:
Alternatively, we can also compute the union, even if the voxels do not form a
connected component:
.. code-block:: yaml

View file

@ -2,22 +2,27 @@
.. _queueing:
Queueing jobs (HPC, HTC)
Queueing Jobs (HPC, HTC)
========================
Yet another interesting feature of junifer is the ability to queue jobs on computational clusters. This is done by
adding the ``queue`` section in the :ref:`codeless` file and executing the ``junifer queue`` command.
Yet another interesting feature of junifer is the ability to queue jobs on
computational clusters. This is done by adding the ``queue`` section in the
:ref:`codeless` file and executing the ``junifer queue`` command.
While junifer is meant to support `HTCondor`_, `SLURM`_ and local queueing using `GNU Parallel`_, only HTCondor is
currently supported. This will be implemented in future relases of junifer. If you are in immediate need of any of these
schedulers, please create an issue on the `junifer github`_ repository.
While junifer is meant to support `HTCondor`_, `SLURM`_ and local queueing
using `GNU Parallel`_, only HTCondor is currently supported. This will be
implemented in future relases of junifer. If you are in immediate need of any of
these schedulers, please create an issue on the `junifer github`_ repository.
The ``queue`` section of the :ref:`codeless` must start by defining the following general parameters:
The ``queue`` section of the :ref:`codeless` must start by defining the
following general parameters:
* ``jobname``: name of the job to be queued. This will be used to name the folder where the job files will be created,
as well as any relevant file. Depending on the scheduler, it will also be listed in the queueing system with this
name.
* ``kind``: The kind of scheduler to be used. Currently, only ``HTCondor`` is supported.
* ``jobname``: Name of the job to be queued. This will be used to name the
folder where the job files will be created, as well as any relevant file.
Depending on the scheduler, it will also be listed in the queueing system
with this name.
* ``kind``: The kind of scheduler to be used. Currently, only ``HTCondor`` is
supported.
Example:
@ -35,34 +40,45 @@ The rest of the parameters depend on the scheduler you are using.
HTCondor
--------
When using HTCondor, junifer will use a DAG to queue one job per element (``junifer run``). As an option, the DAG can
include a final job (``junifer collect``) to collect the results once all of the individual element jobs are finished.
When using HTCondor, junifer will use a DAG to queue one job per element
(``junifer run``). As an option, the DAG can include a final job
(``junifer collect``) to collect the results once all of the individual element
jobs are finished.
The following parameters are avilable for HTCondor:
* ``env``: Definition of the Python enviroment. It must provide two variables: ``kind`` and ``name``. The ``kind``
corresponds to the kind of virtual environment to use: ``conda``, ``virtualenv`` (not yet supported) or
``local`` (no virtual enviroment). The ``name`` is the name of the enviroment to use in case a virtual environment
is used.
* ``mem``: Memory to be used by the job. It must be provided as a string with the units (e.g. ``2GB``).
* ``env``: Definition of the Python enviroment. It must provide two variables:
* ``kind``: This is the kind of virtual environment to use:
* ``conda``
* ``virtualenv`` (not yet supported)
* ``local`` (no virtual enviroment)
* ``name``: This is the name of the enviroment to use in case a virtual
environment is used.
* ``mem``: Memory to be used by the job. It must be provided as a string with
the units (e.g. ``2GB``).
* ``cpus``: Number of CPUs to be used by the job. It must be provided as an int.
* ``disk``: Disk space to be used by the job. It must be provided as a string with the units (e.g. ``2GB``). Keep in
mind that junifer uses a local working directory for each job, and datalad datasets might be cloned in this temporary
* ``disk``: Disk space to be used by the job. It must be provided as a string
with the units (e.g. ``2GB``). Keep in mind that junifer uses a local working
directory for each job, and datalad datasets might be cloned in this temporary
directory.
* ``extra_preamble``: Extra lines to be added to the HTCondor submit file. This can be used to add
extra parameters to the job, such as ``requirements``.
* ``collect``: This parameter allows to include a collect to the DAG to collect the results once all of the individual
element jobs are finished. This is useful if you want to run a ``junifer collect`` job only once all of the
* ``extra_preamble``: Extra lines to be added to the HTCondor submit file. This
can be used to add extra parameters to the job, such as ``requirements``.
* ``collect``: This parameter allows to include a collect to the DAG to collect
the results once all of the individual element jobs are finished. This is
useful if you want to run a ``junifer collect`` job only once all of the
individual element jobs are finished. Valid options are:
* ``yes``: Include a collect job to the DAG that will be executed even if some of the individual element
jobs fail.
* ``on_success_only``: Include a collect job to the DAG, but will only run if all of the individual element jobs are
successful.
* ``yes``: Include a collect job in the DAG that will be executed even if some
of the individual element jobs fail.
* ``on_success_only``: Include a collect job to the DAG, but will only run if
all of the individual element jobs are successful.
* ``no``: Do not include a collect job to the DAG.
Example:
.. code-block:: yaml
@ -74,23 +90,23 @@ Example:
kind: conda
name: junifer
mem: 8G
disk: 2GB
collect: true
disk: 2G
collect: "yes" # wrap it in string to avoid boolean
Once the :ref:`codeless` file is ready, including the ``queue`` section, you can
queue the jobs by executing the ``junifer queue`` command.
Once the :ref:`codeless` file is ready, including the ``queue`` section, you can queue the jobs by executing
the ``junifer queue`` command.
The ``queue`` command will create a folder with the name of the job (``jobname``) under the ``junifer_jobs`` directory
in the current working directory.
The ``queue`` command will create a folder with the name of the job (``jobname``)
under the ``junifer_jobs`` directory in the current working directory.
The ``queue`` command accepts the following arguments:
* ``--help``: Show a help message.
* ``--verbose`` Set the verbosity level. Options are ``warning``, ``info``, ``debug``.
* ``--submit``: Submit the jobs to the queueing system. If not specified, the job submit files will be created but not
submitted.
* ``--overwrite``: Overwrite the job folder if it already exists. If not specified, the command will fail if the job
folder already exists.
* ``--element``: Queue only the specified element(s). If not specified, all elements will be queued.
* ``--verbose``: Set the verbosity level. Options are ``warning``, ``info``,
``debug``.
* ``--submit``: Submit the jobs to the queueing system. If not specified, the
job submit files will be created but not submitted.
* ``--overwrite``: Overwrite the job folder if it already exists. If not
specified, the command will fail if the job folder already exists.
* ``--element``: Queue only the specified element(s). If not specified, all
elements will be queued.

View file

@ -2,17 +2,20 @@
.. _running:
Running jobs
Running Jobs
============
Once we have the :ref:`code-less configuration file <codeless>`, we can use the command line interface to extract the
features. This is achieved in a two-step process: ``run`` and ``collect``.
Once we have the :ref:`code-less configuration file <codeless>`, we can use the
command line interface to extract the features. This is achieved in a two-step
process: ``run`` and ``collect``.
The ``run`` command is used to extract the features from each element in the dataset. However, depending on the
storage interface, this may create one file per subject. The ``collect`` command is then used to collect all of the
The ``run`` command is used to extract the features from each element in the
dataset. However, depending on the storage interface, this may create one file
per subject. The ``collect`` command is then used to collect all of the
individual results into a single file.
Assuming that we have a configuration file named ``config.yaml``, the following commands will extract the features:
Assuming that we have a configuration file named ``config.yaml``, the following
commands will extract the features:
.. code-block:: bash
@ -21,19 +24,20 @@ Assuming that we have a configuration file named ``config.yaml``, the following
The ``run`` command accepts the following additional arguments:
* ``--help``: Show a help message.
* ``--verbose`` Set the verbosity level. Options are ``warning``, ``info``, ``debug``.
* ``--element``: The *element* to run. If not specified, all elements will be run. This parameter can be specified
multiple times to run multiple elements. If the *element* requires several parameters, they can be specified
by separating them with ``,``.
* ``--verbose``: Set the verbosity level. Options are ``warning``, ``info``,
``debug``.
* ``--element``: The *element* to run. If not specified, all elements will be
run. This parameter can be specified multiple times to run multiple elements.
If the *element* requires several parameters, they can be specified by
separating them with ``,``.
Example on running two elements:
Example of running two elements:
.. code-block:: bash
junifer run config.yaml --element sub-01 --element sub-02
Example on elements with multiple parameters and verbose output:
Example of elements with multiple parameters and verbose output:
.. code-block:: bash
@ -41,14 +45,16 @@ Example on elements with multiple parameters and verbose output:
.. _collect:
Collecting results
Collecting Results
==================
Once the ``run`` command has been executed, the results are stored in the output directory. However, depending on the
storage interface, this may create one file per subject. The ``collect`` command is then used to collect all of the
Once the ``run`` command has been executed, the results are stored in the output
directory. However, depending on the storage interface, this may create one file
per subject. The ``collect`` command is then used to collect all of the
individual results into a single file.
Assuming that we have a configuration file named ``config.yaml``, the following commands will collect the results:
Assuming that we have a configuration file named ``config.yaml``, the following
commands will collect the results:
.. code-block:: bash
@ -57,4 +63,5 @@ Assuming that we have a configuration file named ``config.yaml``, the following
The ``collect`` command accepts the following additional arguments:
* ``--help``: Show a help message.
* ``--verbose`` Set the verbosity level. Options are ``warning``, ``info``, ``debug``.
* ``--verbose``: Set the verbosity level. Options are ``warning``, ``info``,
``debug``.

View file

@ -147,5 +147,4 @@ def test_piop1_invalid_tasks():
"the AOMIC PIOP1 dataset!"
),
):
DataladAOMICPIOP1(tasks="thisisnotarealtask")

View file

@ -20,7 +20,6 @@ def test_aomic_piop2_datagrabber() -> None:
task_params = [None, "restingstate"]
for task_param in task_params:
dg = DataladAOMICPIOP2(tasks=task_param)
# change uri here to use fake data instead of real dataset
@ -142,5 +141,4 @@ def test_piop2_invalid_tasks():
"the AOMIC PIOP2 dataset!"
),
):
DataladAOMICPIOP2(tasks="thisisnotarealtask")

View file

@ -20,6 +20,7 @@ def test_BaseDataGrabber_abstractness() -> None:
def test_BaseDataGrabber() -> None:
"""Test BaseDataGrabber."""
# Create concrete class.
class MyDataGrabber(BaseDataGrabber):
def get_item(self, subject):

View file

@ -365,8 +365,7 @@ def test_hcp1200_datagrabber_elements(
],
)
def test_hcp1200_datagrabber_incorrect_access_icafix(
tasks: Optional[str],
ica_fix: bool
tasks: Optional[str], ica_fix: bool
) -> None:
"""Test HCP1200 datagrabber incorrect access for icafix.

View file

@ -17,6 +17,7 @@ def test_base_marker_abstractness() -> None:
def test_base_marker_subclassing() -> None:
"""Test proper subclassing of BaseMarker."""
# Create concrete class
class MyBaseMarker(BaseMarker):
def __init__(self, on, name=None) -> None:

View file

@ -17,6 +17,7 @@ def test_base_preprocessor_abstractness() -> None:
def test_base_preprocessor_subclassing() -> None:
"""Test proper subclassing of BasePreprocessor."""
# Create concrete class
class MyBasePreprocessor(BasePreprocessor):
def __init__(self, on):

View file

@ -19,6 +19,7 @@ def test_BaseFeatureStorage_abstractness() -> None:
def test_BaseFeatureStorage() -> None:
"""Test proper subclassing of BaseFeatureStorage."""
# Create concrete class
class MyFeatureStorage(BaseFeatureStorage):
"""Implement concrete class."""