[DOC]: Provide a thorough description on how junifer works #39

Merged
synchon merged 18 commits from docs/junifer-working into main 2022-10-26 07:19:24 +00:00
13 changed files with 202 additions and 64 deletions

View file

@ -1,3 +1,5 @@
![Junifer logo](docs/images/junifer_logo.png)
# junifer - JUelich NeuroImaging FEature extractoR # junifer - JUelich NeuroImaging FEature extractoR
![PyPI](https://img.shields.io/pypi/v/junifer?style=flat-square) ![PyPI](https://img.shields.io/pypi/v/junifer?style=flat-square)

View file

@ -14,7 +14,7 @@ Data Grabbers
- Open with registration - Open with registration
- Restricted - Restricted
Type/config: this should mention weather the class is built-in in the Type/config: this should mention whether the class is built-in in the
core of junifer or needs to be imported from a specific configuration in core of junifer or needs to be imported from a specific configuration in
the `junifer.configs` module. the `junifer.configs` module.
@ -47,7 +47,7 @@ Available
- Built-in - Built-in
- In Progress - In Progress
- :gh:`4` - :gh:`4`
* - :class:`junifer.configs.juseless.JuselessDataladUKBVBM` * - :class:`junifer.configs.juseless.datagrabbers.JuselessDataladUKBVBM`
- UKB VBM dataset preprocessed with CAT. Available for Juseless only. - UKB VBM dataset preprocessed with CAT. Available for Juseless only.
- Restricted - Restricted
- ``junifer.configs.juseless`` - ``junifer.configs.juseless``
@ -56,7 +56,7 @@ Available
* - :class:`junifer.configs.juseless.datagrabbers.JuselessDataladCamCANVBM` * - :class:`junifer.configs.juseless.datagrabbers.JuselessDataladCamCANVBM`
- CamCAN VBM dataset preprocessed with CAT. Available for Juseless only. - CamCAN VBM dataset preprocessed with CAT. Available for Juseless only.
- Restricted - Restricted
- ``junifer.configs.juseless.datagrabbers`` - ``junifer.configs.juseless``
- Done - Done
- 0.0.1 - 0.0.1
* - :class:`junifer.datagrabber.DataladAOMICID1000` * - :class:`junifer.datagrabber.DataladAOMICID1000`
@ -77,7 +77,7 @@ Available
- Built-in - Built-in
- Done - Done
- 0.0.1 - 0.0.1
* - :class:`junifer.configs.juseless.JuselessDataladAOMICVBM` * - :class:`junifer.configs.juseless.datagrabbers.JuselessDataladAOMICVBM`
- AOMIC VBM dataset. Available for Juseless only. - AOMIC VBM dataset. Available for Juseless only.
- Restricted - Restricted
- ``junifer.configs.juseless`` - ``junifer.configs.juseless``
@ -86,7 +86,7 @@ Available
* - :class:`junifer.configs.juseless.datagrabbers.JuselessDataladIXIVBM` * - :class:`junifer.configs.juseless.datagrabbers.JuselessDataladIXIVBM`
- `IXI VBM dataset <https://brain-development.org/ixi-dataset/>`_. Available for Juseless only. - `IXI VBM dataset <https://brain-development.org/ixi-dataset/>`_. Available for Juseless only.
- Restricted - Restricted
- ``junifer.configs.juseless.datagrabbers`` - ``junifer.configs.juseless``
- Done - Done
- 0.0.1 - 0.0.1
@ -154,6 +154,10 @@ Available
- Spherical aggregation using mean - Spherical aggregation using mean
- Done - Done
- 0.0.1 - 0.0.1
* - :class:`junifer.markers.FunctionalConnectivitySpheres`
- Perform spherical aggregation and compute functional connectivity
- Done
- 0.0.1
* - :class:`junifer.markers.RSSETSMarker` * - :class:`junifer.markers.RSSETSMarker`
- Compute root sum of squares of edgewise timeseries - Compute root sum of squares of edgewise timeseries
- Done - Done

View file

@ -64,9 +64,19 @@ exclude_patterns = ["_build", "Thumbs.db", ".DS_Store"]
# a list of builtin themes. # a list of builtin themes.
# #
html_theme = "sphinx_rtd_theme" html_theme = "sphinx_rtd_theme"
html_theme_options = {
html_sidebars = {"**": ["globaltoc.html", "sourcelink.html", "searchbox.html"]} "display_version": True,
"style_external_links": True,
"logo_only": True,
}
html_sidebars = {
"**": [
"globaltoc.html",
"sourcelink.html",
"searchbox.html",
]
}
html_logo = "./images/junifer_logo.png"
# Add any paths that contain custom static files (such as style sheets) here, # Add any paths that contain custom static files (such as style sheets) here,
# relative to this directory. They are copied after the builtin static files, # relative to this directory. They are copied after the builtin static files,

Binary file not shown.

After

Width:  |  Height:  |  Size: 61 KiB

View file

@ -1,5 +1,9 @@
.. include:: links.inc .. include:: links.inc
.. image:: images/junifer_logo.png
:width: 300px
:alt: junifer logo
Welcome to junifer's documentation! Welcome to junifer's documentation!
=================================== ===================================

View file

@ -9,14 +9,14 @@ Requirements
junifer is compatible with `Python`_ >= 3.8 and requires the following packages: junifer is compatible with `Python`_ >= 3.8 and requires the following packages:
* click>=8.1.3,<8.2 * ``click>=8.1.3,<8.2``
* numpy>=1.22,<1.23 * ``numpy>=1.22,<1.23``
* datalad>=0.15.4,<0.18 * ``datalad>=0.15.4,<0.18``
* pandas>=1.4.0,<1.5 * ``pandas>=1.4.0,<1.5``
* nibabel>=3.2.0,<4.1 * ``nibabel>=3.2.0,<4.1``
* nilearn>=0.9.0,<1.0 * ``nilearn>=0.9.0,<1.0``
* sqlalchemy>=1.4.27,<= 1.5.0 * ``sqlalchemy>=1.4.27,<= 1.5.0``
* pyyaml>=5.1.2,<7.0 * ``pyyaml>=5.1.2,<7.0``
Depending on the installation method, these packages might be installed automatically. Depending on the installation method, these packages might be installed automatically.

View file

@ -1,5 +1,7 @@
.. include:: ../links.inc .. include:: ../links.inc
.. _data_object:
The Data Object The Data Object
=============== ===============
@ -7,7 +9,7 @@ Description
^^^^^^^^^^^ ^^^^^^^^^^^
This is the *object* that traverses the steps of the pipeline. It is indeed a This is the *object* that traverses the steps of the pipeline. It is indeed a
dictionary of dictionaries. The first level of keys are the :ref:`data_types` dictionary of dictionaries. The first level of keys are the :ref:`data types <data_types>`
and a special key named ``meta`` that contains all the information on the data and a special key named ``meta`` that contains all the information on the data
object including source and previous transformation steps. object including source and previous transformation steps.
@ -16,17 +18,19 @@ The second level of keys are the actual data. So far, there are two keys used:
- ``path``: path to the file containing the data. - ``path``: path to the file containing the data.
- ``data``: the data loaded in memory. - ``data``: the data loaded in memory.
The :ref:`datagrabber` step will only fill the ``path`` value. The :ref:`DataGrabber <datagrabber>` step will only fill the ``path`` value.
The ``data`` value will be filled by the :ref:`datareader` step, if it is one of the possible file types The ``data`` value will be filled by the :ref:`DataReader <datareader>` step, if it is one of the possible file types
that the datareader can read. that the datareader can read.
A point to note is that you never directly interact with the *data object* but it's important to know where and how the object is being manipulated to reason about your pipeine.
.. _data_types: .. _data_types:
Data types Data types
^^^^^^^^^^ ^^^^^^^^^^
.. list-table:: Built-in data types .. list-table::
:widths: 30 80 40 :widths: auto
:header-rows: 1 :header-rows: 1
* - Name * - Name

View file

@ -2,50 +2,52 @@
.. _datagrabber: .. _datagrabber:
Data Grabber DataGrabber
============ ===========
Description Description
^^^^^^^^^^^ ^^^^^^^^^^^
A datagrabber is an object that can provide datasets you want to junifer.
For example, a DataladDataGrabber can provide data from a Datalad dataset to junifer.
Of course, datagrabbers are not only possible for Datalad but any origin of a dataset.
It is intended to use them as context managers.
If you are interested in just using already provided datagrabbers please go to :doc:`../builtin`.
If you want to implement your own Data Grabbers you need to inherit from different types of
Data Grabbers we already provide.
Typical Data Grabbers The ``DataGrabber`` is an object that can provide an interface to datasets you want to work with in junifer.
^^^^^^^^^^^^^^^^^^^^^ Every concrete implementation of a datagrabber is aware of a particular dataset's structure and thus allows
In this section we will showcase different types of datagrabber classes you might want to use you to fetch specific elements of interest from the dataset. It adds the ``path`` key to each :ref:`data type <data_types>`
to implement your own datagrabbers for your own data. in the :ref:`Data object <data_object>`.
.. list-table:: Data Grabbers Type Datagrabbers are intended to be used as context managers. When used within a context, a datagrabber takes care
:widths: 25 35 of any pre and post steps for interacting with the dataset, for example, downloading and cleaning up. As the interface
is consistent, you always use the same procedure to interact with the datagrabber.
For example, a concrete implementation of :class:`junifer.datagrabber.DataladDataGrabber` can provide junifer
with data from a Datalad dataset. Of course, datagrabbers are not only meant to work with Datalad datasets but
any dataset.
If you are interested in using already provided datagrabbers, please go to :doc:`../builtin`. And, if you want
to implement your own datagrabber, you need to provide concrete implementations of base classes already
provided.
Base classes
^^^^^^^^^^^^
In this section, we showcase different abstract base classes you might want to use to implement your own datagrabber.
.. list-table::
:widths: auto
:header-rows: 1 :header-rows: 1
* - Name * - Name
- Description - Description
* - :py:class:`~junifer.datagrabber.base.BaseDataGrabber` * - :class:`junifer.datagrabber.BaseDataGrabber`
- | An abstract class providing you an interface to implement for you own datagrabber. - | The abstract base class providing you an interface to implement your own datagrabber.
| This is not intedent to be used in general. | You should try to avoid using this directly and instead use
| Instead you should use the DataladDataGrabber or PatternDataGrabber if possible. | :class:`junifer.datagrabber.PatternDataGrabber` or :class:`junifer.datagrabber.DataladDataGrabber`.
| You have to at least implement the ``get_elements`` method, but most of the time | To build your own custom *low-level* datagrabber, you need to at least implement the ``get_elements`` method,
| you should also overwrite other existing methods like ``__enter__`` and ``__exit__``. | but most of the time you should also override other existing methods like ``__enter__`` and ``__exit__``.
* - :py:class:`~junifer.datagrabber.pattern.PatternDataGrabber` * - :class:`junifer.datagrabber.PatternDataGrabber`
- | Implements some functionality to help you to define the pattern of the dataset you want to get. - | It implements functionality to help you define the pattern of the dataset you want to get. For example,
| E.g. you know that T1 images are found in a directory following this pattern | you know that T1 images are found in a directory following this pattern ``{subject}/anat/{subject}_T1w.nii.gz``
| inside of the dataset "{subject}/anat/{subject}_T1w.nii.gz". | inside of the dataset. Now you can provide this to the **PatternDataGrabber** and it will be able to get the file.
| Now you can provide this to the PatternDataGrabber * - :class:`junifer.datagrabber.DataladDataGrabber`
| and it will be able to get the image to junifer. - | It implements functionality to deal with Datalad datasets. Specifically, the ``__enter__`` and ``__exit__`` methods
* - :py:class:`~junifer.datagrabber.datalad_base.DataladDataGrabber` | take care of cloning and removing the Datalad dataset.
- | Implements some functionality specific to basic usage of Datalad datasets. * - :class:`junifer.datagrabber.PatternDataladDataGrabber`
| This mostly takes care of the ``__enter__`` and ``__exit__`` methods to clone your Datalad dataset - | It is a combination of :class:`junifer.datagrabber.PatternPatternDataGrabber` and
| and remove it on ``__exit__``. | :class:`junifer.datagrabber.DataladDataGrabber`. This is probably the class you are looking for when using Datalad.
| This means that even if your code fails inside of the context of the
| datagrabber the DataladDataGrabber cleans up the created Datalad directories.
* - :py:class:`~junifer.datagrabber.pattern_datalad.PatternDataladDataGrabber`
- | Combination of PatternDataGrabber and DataladDataGrabber.
| Probably the class you are looking for when using Datalad.

View file

@ -2,5 +2,45 @@
.. _datareader: .. _datareader:
Data Reader DataReader
=========== ==========
Description
^^^^^^^^^^^
The ``DataReader`` is an object that is responsible for actually reading data files in junifer.
It reads the value of the key ``path`` for each :ref:`data type <data_types>` in the :ref:`Data object <data_object>`
and loads them to memory. After reading the data into memory, it adds the key ``data`` to the same level as ``path``
and the value is the actual data in the memory.
Datareaders are meant to be used inside the datagrabber context but you can operate on them outside the context as long
as the actual data is in the memory and the Python runtime has not garbage-collected it.
For data formats not supported by junifer yet, you can either make your own ``DataReader`` or open an issue on
`junifer Github`_ and we can help you out.
Currently supported file-formats
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
We already provide a concrete implementation :class:`junifer.datareader.DefaultDataReader` which knows how to
read the following file formats:
.. list-table::
:widths: auto
:header-rows: 1
* - File extension
- File type
- Description
* - ``.nii``
- NIfTI (uncompressed)
- Uncompressed NIfTI
* - ``.nii.gz``
- NIfTI (compressed)
- Compressed NIfTI
* - ``.csv``
- CSV
- Comma-separated values file
* - ``.tsv``
- TSV
- Tab-separated values file

View file

@ -4,7 +4,7 @@ Understanding junifer
===================== =====================
Before you start, you should understand how junifer works. Junifer is a Before you start, you should understand how junifer works. Junifer is a
tool conceived to extract features from neuroimaging data in a easy-to-use tool conceived to extract features from neuroimaging data in an easy-to-use
manner, with minimal coding and minimal user expertise in the internal aspects. manner, with minimal coding and minimal user expertise in the internal aspects.
Unlike other tools like FSL, SPM, AFNI, etc., junifer is not a toolbox to Unlike other tools like FSL, SPM, AFNI, etc., junifer is not a toolbox to
@ -24,5 +24,6 @@ julearn_).
data data
datagrabber datagrabber
datareader datareader
preprocess
marker marker
storage storage

View file

@ -1,4 +1,23 @@
.. include:: ../links.inc .. include:: ../links.inc
.. _marker:
Marker Marker
====== ======
Description
^^^^^^^^^^^
The ``Marker`` is an object that is responsible for feature extraction. It primarily operates on data loaded
in memory by :ref:`datareader <DataReader>` and stored in the ``data`` key of each :ref:`data type <data_types>`
in the :ref:`Data object <data_object>`. In some cases, it can also operate on pre-processed data as obtained
from the :ref:`Preprocess <preprocess>` step of the pipeline. It is important to note that this pre-process is
not similar to pre-processing done by tools like FSL, SPM, AFNI, etc. . For example, one can perform confound
removal on loaded data and then perform feature extraction.
Markers are meant to be used inside the datagrabber context but you can operate on them outside the context as long
as the actual data is in the memory and the Python runtime has not garbage-collected it.
If you are interested in using already provided markers, please go to :doc:`../builtin`. And, if you want to implement
your own marker, you need to provide concrete implementation of :class:`junifer.markers.BaseMarker`. Specifically, you
need to override ``get_output_kind``, ``store`` and ``compute`` methods.

View file

@ -0,0 +1,15 @@
.. include:: ../links.inc
.. _preprocess:
Preprocess
==========
Description
^^^^^^^^^^^
The ``Preprocess`` step of the pipeline is meant for pre-processing before or after :ref:`Marker <marker>` step
depending on the use-case. For example, you might want to perform confound removal on ``BOLD`` data before
feature extraction.
This step is still under development and is an optional one for the pipeline to work.

View file

@ -1,4 +1,41 @@
.. include:: ../links.inc .. include:: ../links.inc
.. _storage:
Storage Storage
======= =======
Description
^^^^^^^^^^^
The ``Storage`` is an object that is responsible for storing extracted features as computed from :ref:`Marker <marker>`
step of the pipeline. If the pipeline is provided with a ``storage-like`` object, the extracted features are stored via
that object else they are kept in memory.
Storage is meant to be used inside the datagrabber context but you can operate on them outside the context as long
as the processed data is in the memory and the Python runtime has not garbage-collected it.
The :ref:`Markers <marker>` are responsible for defining what *storage kind* (``matrix``, ``table``, ``timeseries``)
they support for which :ref:`data type <data_types>` by overriding its ``store`` method. The storage object in turn
declares and provides implementation for specific *storage kind*. For example, :class:`junifer.storage.SQLiteFeatureStorage`
supports saving ``matrix``, ``table`` and ``timeseries`` via ``store_matrix``, ``store_table`` and ``store_timeseries``
methods respectively.
For storage interfaces not supported by junifer yet, you can either make your own ``Storage`` by providing a concrete
implementation of :class:`junifer.storage.BaseFeatureStorage` or open an issue on `junifer Github`_ and we can help you out.
Currently supported storage interfaces
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
.. list-table::
:widths: auto
:header-rows: 1
* - Storage class
- File extension
- File type
- Storage kinds
* - :class:`junifer.storage.SQLiteFeatureStorage`
- ``.db``
- SQLite
- ``matrix``, ``table``, ``timeseries``