[DOC] Section on extending junifer (datagrabbers and markers) #124
|
|
@ -1,10 +1,17 @@
|
|||
API Functions
|
||||
=============
|
||||
|
||||
Main API functions
|
||||
------------------
|
||||
|
||||
.. automodule:: junifer.api
|
||||
:members:
|
||||
:imported-members:
|
||||
|
||||
|
||||
Decorators
|
||||
----------
|
||||
|
||||
.. automodule:: junifer.api.decorators
|
||||
:members:
|
||||
:imported-members:
|
||||
|
|
|
|||
|
|
@ -1,5 +1,5 @@
|
|||
DataGrabbers
|
||||
============
|
||||
Data Grabbers
|
||||
=============
|
||||
|
||||
.. automodule:: junifer.datagrabber
|
||||
:members:
|
||||
|
|
|
|||
|
|
@ -4,8 +4,8 @@ Built-in Pipeline steps and data
|
|||
================================
|
||||
|
||||
|
||||
DataGrabbers
|
||||
------------
|
||||
Data Grabbers
|
||||
-------------
|
||||
|
||||
..
|
||||
Provide a list of the DataGrabbers that are implemented or planned.
|
||||
|
|
|
|||
|
|
@ -173,11 +173,11 @@ texts.
|
|||
###############################################################################
|
||||
# The BIDS datagrabber requires three parameters: the types of data we want,
|
||||
|
|
||||
# the specific pattern that matches each type, and the variables that will be
|
||||
# replaced int he patterns.
|
||||
types = ["T1w", "bold"]
|
||||
# replaced in the patterns.
|
||||
types = ["T1w", "BOLD"]
|
||||
patterns = {
|
||||
"T1w": "{subject}/anat/{subject}_T1w.nii.gz",
|
||||
"bold": "{subject}/func/{subject}_task-rest_bold.nii.gz",
|
||||
"BOLD": "{subject}/func/{subject}_task-rest_bold.nii.gz",
|
||||
}
|
||||
replacements = ["subject"]
|
||||
###############################################################################
|
||||
|
|
|
|||
424
docs/extending/datagrabber.rst
Normal file
|
|
@ -0,0 +1,424 @@
|
|||
.. include:: ../links.inc
|
||||
|
`Its`
`covered`
`an`
`these`
`... were scanned during 3 sessions ...`
```... ``{session}`` ...```
```
...,
replacements=replacements,
)
```
```
...,
replacements=replacements,
)
```
Can we have a hyperlink reference for datalad? Can we have a hyperlink reference for datalad?
```This class will not only interpret patterns but also use datalad to `clone` and `get` the data.```
`... .`
```
...,
replacements=replacements,
)
```
`of`
`... BOLD.`
`can`
`represent each of the items ...`
One One `the` should be removed.
|
||||
|
||||
.. _extending_datagrabbers:
|
||||
|
||||
Creating Data Grabbers
|
||||
======================
|
||||
|
||||
Data Grabbers are the first step of the pipeline. Its purpose is to interpret
|
||||
the structure of a dataset and provide two specific functionalities:
|
||||
|
||||
1) Given an *element*, provide the path to each kind of data available for this
|
||||
element (e.g. the path to the T1 image, the path to the T2 image, etc.)
|
||||
2) Provide the list of *elements* available in the dataset.
|
||||
|
||||
In this section, we will see how to create a datagrabber for a dataset. Basic
|
||||
aspects of datagrabbers are covered in the
|
||||
:ref:`Understanding Data Grabbers <datagrabber>` section.
|
||||
|
||||
.. _extending_datagrabbers_think:
|
||||
|
||||
Step 1: Think about the element
|
||||
-------------------------------
|
||||
|
||||
Like with any programming-related task, the first step is to think. When
|
||||
creating a Data Grabber, we need to first define what an *element* is.
|
||||
The *element* should be the smallest unit of data that can be processed. That
|
||||
is, for each element, there should be a set of data that can be processed, but
|
||||
only one of each *data type* (see :ref:`data_types`).
|
||||
|
||||
For example, if we have a dataset from an fMRI study in which:
|
||||
|
||||
a) both T1w and fMRI was acquired
|
||||
b) 20 subjects went through an experiment twice
|
||||
c) the experiment included resting-stage fMRI and a task named *stroop*
|
||||
|
||||
then the *element* should be composed of 3 items:
|
||||
|
||||
* ``subject``: The subject IDs, e.g. `sub001`, `sub002`, ... `sub020`
|
||||
* ``session``: The sesion number, e.g. `ses1`, `ses2`
|
||||
* ``task``: The task performed, e.g. `rest`, `stroop`
|
||||
|
||||
If any of these items were not part of the element, then we will have more than
|
||||
one ``T1w`` and/or ``BOLD`` image for each subject, which is not allowed.
|
||||
|
||||
Importantly, nothing prevents that one image is part of two different elements.
|
||||
For example, it is usually the case that the ``T1w`` image is not acquired for
|
||||
each task, but once in the entire session. So in this case, the ``T1w`` image
|
||||
for the element (``sub001``, ``ses1``, ``rest``) will be the same as the
|
||||
|
Maybe: Maybe: ``("sub001", "ses1", "rest")``?
not as strings. I like it like that. not as strings. I like it like that.
With strings, you directly link it to the code which IMO is simpler. With strings, you directly link it to the code which IMO is simpler.
Also, the rendering is not very pretty. Also, the rendering is not very pretty.
What I actually meant was putting the whole thing as monospace. What I actually meant was putting the whole thing as monospace.
|
||||
``T1w`` image for the element (``sub001``, ``ses1``, ``stroop``).
|
||||
|
And here maybe: And here maybe: ``("sub001", "ses1", "stroop")``?
|
||||
|
||||
We will now continue this section using as an example, a dataset in BIDS format
|
||||
in which 9 subjects (`sub-01` to `sub-09`) were scanned each during 3
|
||||
sessions (`ses-01`, `ses-02`, `ses-03`) and each session included a `T1w` and
|
||||
a `BOLD` image (resting-state), except for `ses-03` which was only anatomical.
|
||||
|
||||
Step 2: Think about the dataset's structure
|
||||
-------------------------------------------
|
||||
|
||||
Now that we have our element defined, we need to think about the structure of
|
||||
the dataset. Mainly, because the structure of the dataset will determine how
|
||||
the Data Grabber needs to be implemented.
|
||||
|
||||
Junifer provides an abstract class to deal with datasets that can be thought in
|
||||
terms of *patterns*. A *pattern* is a string that contains placeholders that are
|
||||
replaced by the actual values of the element. In our BIDS example, the path
|
||||
to the T1w image of subject `sub-01` and session `ses-01`, relative to the
|
||||
dataset location, is ``sub-01/ses-01/anat/sub-01_ses-01_T1w.nii.gz``. By
|
||||
replacing ``sub-01`` with ``sub-02``, we can obtain the T1w image of the first
|
||||
session of the second subject. Indeed, the path to the T1w images can be
|
||||
expressed as a pattern:
|
||||
|
||||
``{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz``
|
||||
|
||||
where ``{subject}`` is the replacement for the subject id and ``{session}``
|
||||
is the replacement for the session id.
|
||||
|
||||
Since it is a BIDS dataaset, the same happens with the BOLD images. The path to
|
||||
the BOLD images can be expressed as a pattern:
|
||||
|
||||
``{subject}/{session}/func/{subject}_{session}_task-rest_bold.nii.gz``
|
||||
|
||||
|
||||
This will be the norm in most of the datasets. If your dataset can be expressed
|
||||
in terms of patterns, then follow :ref:`extending_datagrabbers_pattern`.
|
||||
Otherwise, we recommend that you take time to re-think about your dataset
|
||||
structure and why it does not have clear *patterns*. Feel free to open a
|
||||
discussion in the `junifer Discussions`_ page. Most probably we can help you
|
||||
get your dataset in order.
|
||||
|
||||
If there is no other way, then you can follow :ref:`extending_datagrabbers_base`
|
||||
to create a Data Grabber from scratch.
|
||||
|
||||
|
||||
.. _extending_datagrabbers_pattern:
|
||||
|
||||
Step 3: Create a Data Grabber
|
||||
-----------------------------
|
||||
|
||||
Option A: Extending from PatternDataGrabber
|
||||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
|
||||
The :py:class:`~junifer.datagrabber.PatternDataGrabber` class is an
|
||||
abstract class that has the functionality of understanding patterns embeded
|
||||
in it.
|
||||
|
||||
Before creating the datagrabber, we need to define 3 variables:
|
||||
|
||||
* ``types``: A list with the available :ref:`data_types` in our dataset
|
||||
* ``patterns``: A dictionary that specifies the pattern for each data type.
|
||||
* ``replacements``: A list indicating which of the elements in the patterns
|
||||
should be replaced by the values of the element.
|
||||
|
||||
For example, in our BIDS example, the variables will be:
|
||||
|
||||
.. code-block:: python
|
||||
|
||||
types = ["T1w", "BOLD"]
|
||||
patterns = {
|
||||
"T1w": "{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz",
|
||||
"BOLD": "{subject}/{session}/func/{subject}_{session}_task-rest_bold.nii.gz",
|
||||
}
|
||||
replacements = ["subject", "session"]
|
||||
|
||||
An additional fourth variable is the ``datadir``, which should be the path to
|
||||
where the dataset is located. For example, if the dataset is located in
|
||||
``/data/project/test/data``, then ``datadir`` should be
|
||||
``/data/project/test/data``. Or, if we want to allow the user to specify the
|
||||
location of the dataset, we can expose the variable in the constructor, as in
|
||||
this example
|
||||
|
||||
With this defined, we can now create our datagrabber, we will name it
|
||||
``ExampleBIDSDataGrabber``:
|
||||
|
||||
|
||||
.. code-block:: python
|
||||
|
||||
from junifer.datagrabber.pattern import PatternDataGrabber
|
||||
|
||||
class ExampleBIDSDataGrabber(PatternDataGrabber):
|
||||
|
||||
def __init__(self, datadir):
|
||||
types = ["T1w", "BOLD"]
|
||||
patterns = {
|
||||
"T1w": "{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz",
|
||||
"BOLD": "{subject}/{session}/func/{subject}_{session}_task-rest_bold.nii.gz",
|
||||
}
|
||||
replacements = ["subject", "session"]
|
||||
super().__init__(
|
||||
datadir=datadir,
|
||||
types=types,
|
||||
patterns=patterns,
|
||||
replacements=replacements,
|
||||
)
|
||||
|
||||
Our datagrabber is ready to be used by junifer. However, it is still unknown
|
||||
to the library. We need to register it in the library. To do so, we need to
|
||||
use the :py:func:`~junifer.api.decorators.register_datagrabber` decorator.
|
||||
|
||||
|
||||
.. code-block:: python
|
||||
|
||||
from junifer.datagrabber.pattern import PatternDataGrabber
|
||||
from junifer.api.decorators import register_datagrabber
|
||||
|
||||
|
||||
@register_datagrabber
|
||||
class ExampleBIDSDataGrabber(PatternDataGrabber):
|
||||
|
||||
def __init__(self, datadir):
|
||||
types = ["T1w", "BOLD"]
|
||||
patterns = {
|
||||
"T1w": "{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz",
|
||||
"BOLD": "{subject}/{session}/func/{subject}_{session}_task-rest_bold.nii.gz",
|
||||
}
|
||||
replacements = ["subject", "session"]
|
||||
super().__init__(
|
||||
datadir=datadir,
|
||||
types=types,
|
||||
patterns=patterns,
|
||||
replacements=replacements,
|
||||
)
|
||||
|
||||
|
||||
Now, we can use our datagrabber in junifer, by setting the ``datagrabber`` kind
|
||||
in the yaml file to ``ExampleBIDSDataGrabber``. Remember that we still need to
|
||||
set the ``datadir``.
|
||||
|
||||
.. code-block:: yaml
|
||||
|
||||
datagrabber:
|
||||
kind: ExampleBIDSDataGrabber
|
||||
datadir: /data/project/test/data
|
||||
|
||||
|
||||
Optional: Using datalad
|
||||
|
`..., but also ...`
|
||||
"""""""""""""""""""""""
|
||||
|
||||
If you are using `datalad`_, you can use the
|
||||
:py:class:`~junifer.datagrabber.PatternDataladDataGrabber` instead of the
|
||||
:py:class:`~junifer.datagrabber.PatternDataGrabber`. This class will not only
|
||||
interpret patterns, but also use `datalad`_ to `clone` and `get` the data.
|
||||
|
||||
The main difference between the two is that the ``datadir`` is not the actual
|
||||
location of the dataset, but the location where the dataset will be cloned. It
|
||||
can now be ``None``, which means that the data will be downloaded to a
|
||||
temporary directory. To set the location of the dataset, you can use the
|
||||
``uri`` argument in the constructor. Additionally, a ``rootdir`` argument can
|
||||
be used to specify the path to the root directory of the dataset after doing
|
||||
``datalad clone``.
|
||||
|
||||
In the example, the dataset is hosted in gin
|
||||
(``https://gin.g-node.org/juaml/datalad-example-bids``).
|
||||
|
||||
When we clone this dataset, we will see the following structure:
|
||||
|
||||
.. code-block::
|
||||
|
||||
.
|
||||
└── example_bids_ses
|
||||
├── sub-01
|
||||
│ ├── ses-01
|
||||
│ ├── ses-02
|
||||
│ └── ses-03
|
||||
├── sub-02
|
||||
│ ├── ses-01
|
||||
│ ├── ses-02
|
||||
│ └── ses-03
|
||||
├── sub-03
|
||||
...
|
||||
|
||||
So the patterns will start after ``example_bids_ses``. This is our ``rootdir``.
|
||||
|
||||
Now we have our 2 additional variables:
|
||||
|
||||
.. code-block:: python
|
||||
|
||||
uri = "https://gin.g-node.org/juaml/datalad-example-bids"
|
||||
rootdir = "example_bids_ses"
|
||||
|
||||
And we can create our datagrabber:
|
||||
|
||||
|
||||
.. code-block:: python
|
||||
|
||||
from junifer.datagrabber.pattern import PatternDataladDataGrabber
|
||||
from junifer.api.decorators import register_datagrabber
|
||||
|
||||
|
||||
@register_datagrabber
|
||||
class ExampleBIDSDataGrabber(PatternDataladDataGrabber):
|
||||
|
||||
def __init__(self):
|
||||
types = ["T1w", "BOLD"]
|
||||
patterns = {
|
||||
"T1w": "{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz",
|
||||
"BOLD": "{subject}/{session}/func/{subject}_{session}_task-rest_bold.nii.gz",
|
||||
}
|
||||
replacements = ["subject", "session"]
|
||||
uri = "https://gin.g-node.org/juaml/datalad-example-bids"
|
||||
rootdir = "example_bids_ses"
|
||||
super().__init__(
|
||||
datadir=None,
|
||||
uri=uri,
|
||||
rootdir=rootdir,
|
||||
types=types,
|
||||
patterns=patterns,
|
||||
replacements=replacements,
|
||||
)
|
||||
|
||||
|
||||
.. _extending_datagrabbers_base:
|
||||
|
||||
Option B: Extending from BaseDataGrabber
|
||||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
|
||||
While we could not think of a use case in which the pattern-based data grabber would not be suitable, it is still
|
||||
possible to create a datagrabber extending from the :py:class:`~junifer.datagrabber.base.BaseDataGrabber` class.
|
||||
|
||||
In order to create a datagrabber extending from :py:class:`~junifer.datagrabber.base.BaseDataGrabber`, we need to
|
||||
implement the following methods:
|
||||
|
||||
- ``get_item``: to get a single item from the dataset.
|
||||
- ``get_elements``: to get the list of all elements present in the dataset
|
||||
- ``get_element_keys``: to get the keys of the elements in the dataset.
|
||||
|
||||
.. note::
|
||||
The ``__init__`` method could also be implemented, but it is not mandatory. This is required if the datagrabber
|
||||
requires any parameter.
|
||||
|
||||
We will now implement our BIDS example with this method.
|
||||
|
||||
The first method, ``get_item``, needs to obtain a single
|
||||
item from the dataset. Since this dataset requires two variables, ``subject`` and ``session``, we will use them
|
||||
as parameters of ``get_item``:
|
||||
|
||||
.. code-block:: python
|
||||
|
||||
def get_item(self, subject, session):
|
||||
out = {
|
||||
"T1w": f"{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz",
|
||||
"BOLD": f"{subject}/{session}/func/{subject}_{session}_task-rest_bold.nii.gz",
|
||||
}
|
||||
return out
|
||||
|
||||
|
||||
The second method, ``get_elements``, needs to return a list of all the elements in the dataset. In this case, we
|
||||
know that the dataset contains 3 subjects and 3 sessions, so we can create a list of all the possible combinations.
|
||||
However, we need to remember that for session *ses-03* there is no BOLD data.
|
||||
|
||||
.. code-block:: python
|
||||
|
||||
def get_elements(self):
|
||||
subjects = ["sub-01", "sub-02", "sub-03"]
|
||||
sessions = ["ses-01", "ses-02"]
|
||||
|
||||
# If we are not working on BOLD data, we can add "ses-03"
|
||||
if "BOLD" not in self.types:
|
||||
sessions.append("ses-03")
|
||||
elements = []
|
||||
for subject in subjects:
|
||||
for session in sessions:
|
||||
elements.append({"subject": subject, "session": session})
|
||||
return elements
|
||||
|
||||
|
||||
And finally, we can implement the ``get_element_keys`` method. This method needs to return a list of the keys that
|
||||
represent each of the items in the element tuple. As a rule of thumb, they should be the parameters of the
|
||||
``get_item`` method, in the same order.
|
||||
|
||||
.. code-block:: python
|
||||
|
||||
def get_element_keys(self):
|
||||
return ["subject", "session"]
|
||||
|
||||
|
||||
So, to summarize, our datagrabber will look like this:
|
||||
|
||||
.. code-block:: python
|
||||
|
||||
from junifer.datagrabber.base import BaseDataGrabber
|
||||
from junifer.api.decorators import register_datagrabber
|
||||
|
||||
@register_datagrabber
|
||||
class ExampleBIDSDataGrabber(BaseDataGrabber):
|
||||
|
||||
def get_item(self, subject, session):
|
||||
out = {
|
||||
"T1w": f"{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz",
|
||||
"BOLD": f"{subject}/{session}/func/{subject}_{session}_task-rest_bold.nii.gz",
|
||||
}
|
||||
return out
|
||||
|
||||
def get_elements(self):
|
||||
subjects = ["sub-01", "sub-02", "sub-03"]
|
||||
sessions = ["ses-01", "ses-02"]
|
||||
|
||||
# If we are not working on BOLD data, we can add "ses-03"
|
||||
if "BOLD" not in self.types:
|
||||
sessions.append("ses-03")
|
||||
elements = []
|
||||
for subject in subjects:
|
||||
for session in sessions:
|
||||
elements.append({"subject": subject, "session": session})
|
||||
return elements
|
||||
|
||||
def get_element_keys(self):
|
||||
return ["subject", "session"]
|
||||
|
||||
Optional: Using datalad
|
||||
"""""""""""""""""""""""
|
||||
|
||||
If this dataset is in a datalad dataset, we can extend from :class:`junifer.datagrabber.DataladDataGrabber` instead of
|
||||
:class:`junifer.datagrabber.BaseDataGrabber`. This will allow us to use the datalad API to obtain the data.
|
||||
|
||||
|
||||
Step 4: Optional: Adding *BOLD confounds*
|
||||
-----------------------------------------
|
||||
|
||||
For some analyses, it is useful to have the confounds associated with the BOLD data. This corresponds to the
|
||||
``BOLD_confounds`` item in the :ref:`Data Object <data_object>` (see :ref:`data_types`). However, the ``BOLD_confounds``
|
||||
element does not only consists of a ``path``, but it requries more information about the format of the confounds file.
|
||||
Thus, the ``BOLD_confounds`` element is a dictionary with the following keys:
|
||||
|
||||
- ``path``: the path to the confounds file.
|
||||
- ``format``: the format of the confounds file. Currently, this can be either ``fmriprep`` or ``adhoc``.
|
||||
|
||||
The ``fmriprep`` format corresponds to the format of the confounds files generated by `fMRIPrep`_. The
|
||||
``adhoc`` format corresponds to a format that is not standardized.
|
||||
|
||||
.. note::
|
||||
The ``mappings`` key is only required if the ``format`` is ``adhoc``. If the ``format`` is ``fmriprep``, the
|
||||
``mappings`` key is not required.
|
||||
|
||||
|
||||
Currently, Junifer provides only one confound remover step
|
||||
(:class:`junifer.preprocess.fMRIPrepConfoundRemover`), which relies entirely on the ``fmriprep`` confound
|
||||
variable names. Thus, if the confounds are not in ``fmriprep`` format, the user will need to provide the mappings
|
||||
between the *ad-hoc* variable names and the ``fmriprep`` variable names.
|
||||
This is done by specifying the ``adhoc`` format and providing the mappings as a dictionary in the ``mappings`` key.
|
||||
|
||||
In the following example, the confounds file has 3 variables that are not in the ``fmriprep`` format. Thus, we will
|
||||
provide the mappings for these variables to the ``fmriprep`` format.
|
||||
|
||||
.. code-block:: python
|
||||
|
||||
out["BOLD_confounds"]: {
|
||||
"path": f"{subject}/{session}/func/{subject}_{session}_confounds.tsv",
|
||||
"format": "adhoc",
|
||||
"mappings": {
|
||||
"fmriprep": {
|
||||
"variable1": "rot_x",
|
||||
"variable2": "rot_z",
|
||||
"variable3": "rot_y",
|
||||
}
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
.. note::
|
||||
Not all of the mappings need to be provided. For the moment, this is used only by the
|
||||
:class:`junifer.preprocess.fMRIPrepConfoundRemover` step, which requires variables based on the
|
||||
strategy selected. However, it is recommended to provide all the mappings, as this will allow the user to
|
||||
choose different strategies with the same dataset.
|
||||
|
||||
28
docs/extending/extension.rst
Normal file
|
|
@ -0,0 +1,28 @@
|
|||
.. include:: ../links.inc
|
||||
|
`functionality to junifer at runtime.`
|
||||
|
||||
.. _extending_extension:
|
||||
|
||||
Creating a Junifer extension
|
||||
============================
|
||||
|
||||
Junifer is designed to be easily extensible. Through the use of a registry and decorators, we can easily add new
|
||||
functionality to junifer on runtime. This is done by creating a new python module and importing it before running
|
||||
junifer.
|
||||
|
||||
A special consideration has to be made when using the :ref:`code-less configuration<codeless>`. In this case, the
|
||||
``with`` statement can be used to import a module or run a python ``.py`` file.
|
||||
|
||||
In the following example, we instruct junifer to first import ``my_module`` and then run the ``my_file.py`` file.
|
||||
|
||||
.. code-block:: yaml
|
||||
|
||||
with:
|
||||
- my_module
|
||||
- my_file.py
|
||||
|
||||
Thus, the code from ``my_file.py`` will be executed before running junifer. This is the ideal place to create junifer
|
||||
extensions.
|
||||
|
||||
.. important:: Some junifer commands will not consider files imported from files included in the ``with`` statement.
|
||||
That is, if ``my_file.py`` imports ``my_other_file.py``, some of the junifer commands will not consider
|
||||
``my_other_file.py``. Either place all the code in one file or add multiple files to the ``with`` statement.
|
||||
28
docs/extending/index.rst
Normal file
|
|
@ -0,0 +1,28 @@
|
|||
.. include:: ../links.inc
|
||||
|
`functionality`
|
||||
|
||||
.. _extending:
|
||||
|
||||
Extending junifer
|
||||
=================
|
||||
|
||||
While we aim to provide as many datasets and markers as possible, we are also
|
||||
interested in allowing users to extend the functionality with their own
|
||||
datagrabbers, preprocessing, markers, etc.
|
||||
|
||||
This does not mean that the new functionality will have to be included in
|
||||
junifer before the user can use them. Instead, the user can simply
|
||||
create a new python file, code the desired functionality and use it with
|
||||
junifer. This is the first step towards including the new functionality in
|
||||
the junifer package.
|
||||
|
||||
In this section we will show how to extend junifer, by creating new
|
||||
datagrabbers, preprocessing and markers, following the *junifer* way.
|
||||
|
||||
|
||||
.. toctree::
|
||||
:maxdepth: 2
|
||||
:caption: Contents:
|
||||
|
||||
extension
|
||||
datagrabber
|
||||
marker
|
||||
276
docs/extending/marker.rst
Normal file
|
|
@ -0,0 +1,276 @@
|
|||
.. include:: ../links.inc
|
||||
|
`... with the data types that the marker ... `
`In this example, the only parameter required for computation is the name of the parcellation to use.`
Rendering for this as well is weird. Rendering for this as well is weird.
`useful`
```
:ref:`data type <data_types>`
```
```The method ``store`` ...```
`... simply ...`
`Once all of the above steps are done, we ...`
`parcellation`
`parcellation_name`
`self.parcellation_name = parcellation_name`
`parcellation_name`
`self.parcellation_name = parcellation_name`
|
||||
|
||||
.. _extending_markers:
|
||||
|
||||
Creating Markers
|
||||
================
|
||||
|
||||
Computing a marker (a.k.a. *feature*) is the main goal of junifer. While we aim to provide as many markers as possible,
|
||||
it might be the case that the marker you are looking for is not available. In this case, you can create your own marker
|
||||
by following this tutorial.
|
||||
|
||||
Most of the functionality of a junifer marker has been taken care by the :class:`junifer.markers.BaseMarker` class.
|
||||
Thus, only a few methods are required:
|
||||
|
||||
1. ``get_valid_inputs``: a method to obtain the list of valid inputs for the marker. This is used to check that the
|
||||
inputs provided by the user are valid. This method should return a list of strings, representing
|
||||
:ref:`data types <data_types>`
|
||||
2. ``get_output_kind``: a method to obtain the kind of output of the marker. This is used to check that the output
|
||||
of the marker is compatible with the storage. This method should return a string, representing
|
||||
:ref:`storage types <storage_types>`
|
||||
3. ``compute``: the method that given the data, computes the marker.
|
||||
4. ``store``: the method that stores the computed marker.
|
||||
5. ``__init__``: the initialization method, where the marker is configured.
|
||||
|
||||
As an example, we will develop a Parcel Mean marker, that is, a marker that first applies a parcellation and
|
||||
then computes the mean of the data in each parcel. This is a very simple example, but it will show you how to create
|
||||
a new marker.
|
||||
|
||||
.. _extending_markers_input_output:
|
||||
|
||||
Step 1: Configure input and output
|
||||
----------------------------------
|
||||
|
||||
This step is quite simple: we need to define the input and output of the marker. Based on the current
|
||||
:ref:`data types <data_types>`, we can define as valid inputs ``BOLD``, ``VBM_WM`` and ``VBM_GM``.
|
||||
|
||||
.. code-block:: python
|
||||
|
||||
def get_valid_inputs(self):
|
||||
return ['BOLD', 'VBM_WM', 'VBM_GM']
|
||||
|
||||
The output of the marker depends on the input. For ``BOLD``, it will be ``timeseries``, while for the rest of the inputs,
|
||||
it will be ``table``. Thus, we can define the output as:
|
||||
|
||||
.. code-block:: python
|
||||
|
||||
def get_output_kind(self, input_kind):
|
||||
if input_kind == 'BOLD':
|
||||
return 'timeseries'
|
||||
else:
|
||||
return 'table'
|
||||
|
||||
.. _extending_markers_init:
|
||||
|
||||
Step 2: Initialize the marker
|
||||
-----------------------------
|
||||
|
||||
In this step we need to define the parameters of the marker. That is, all the parameters that the user can provide
|
||||
to configure how the marker will behave.
|
||||
|
||||
The parameters of the marker are defined in the ``__init__`` method. The :class:`junifer.markers.BaseMarker` class
|
||||
requires two optional parameters:
|
||||
|
||||
1. ``name``: the name of the marker. This is used to identify the marker in the configuration file.
|
||||
2. ``on``: a list or string with the data types that the marker will be applied to.
|
||||
|
||||
.. attention:: Only basic types (*int*, *bool* and *str*) as well as Lists, Tuples and Dictionaries are allowed as
|
||||
parameters. This is because the parameters are stored in a JSON file, and JSON only supports these types.
|
||||
|
||||
|
||||
In this example, the is only paramater required for the computation is the name of the parcellation to use. Thus, we can
|
||||
define the ``__init__`` method as follows:
|
||||
|
||||
.. code-block:: python
|
||||
|
||||
def __init__(self, parcellation_name, on=None, name=None):
|
||||
self.parcellation_name = parcellation_name
|
||||
super().__init__(on=on, name=name)
|
||||
|
||||
.. caution:: Parameters of the marker must be stored as object attributes without using ``_`` as prefix. This is
|
||||
because any attribute that starts with ``_`` will not be considered as a parameter and not stored as
|
||||
part of the metadata of the marker.
|
||||
|
||||
|
||||
.. _extending_markers_compute:
|
||||
|
||||
Step 3: Compute the marker
|
||||
--------------------------
|
||||
|
||||
In this step, we will define the method that computes the marker. This method will be called by junifer when needed,
|
||||
using the data provided by the datagrabber, as configured by the user. The function ``compute`` has two arguments:
|
||||
|
||||
* ``input``: a dictionary with the data to be used to compute the marker. This will be the corresponding element in the
|
||||
|
The indentation for this is a bit weird when rendered. The indentation for this is a bit weird when rendered.
|
||||
:ref:`Data Object<data_object>` alredy indexing. Thus, the dictionary has at least two keys: ``data`` and ``path``.
|
||||
The first one contains the data, while the second one contains the path to the data. The dictionary can also contain
|
||||
other keys, depending on the data type.
|
||||
* ``extra_input``: the rest of the :ref:`Data Object<data_object>`. This is useful if you want to use other data to
|
||||
compute the marker (e.g.: ``BOLD_confounds`` can be used to de-confound the ``BOLD`` data).
|
||||
|
||||
Following the example, we will compute the mean of the data in each parcel using
|
||||
:class:`nilearn.maskers.NiftiLabelsMasker`. Importantly, the output of the compute function must be a dictionary.
|
||||
This dictionary will later be passed onto the ``store`` method.
|
||||
|
||||
.. hint:: To simplify the ``store`` method, define keys of the dictionary based on the corresponding store functions
|
||||
in the :ref:`storage types <storage_types>`. For example, if the output is a ``table``, the keys of the
|
||||
dictionary should be ``data`` and ``columns``.
|
||||
|
||||
.. code-block:: python
|
||||
|
||||
from nilearn.maskers import NiftiLabelsMasker
|
||||
from junifer.data import load_parcellation
|
||||
|
||||
def compute(self, input, extra_input):
|
||||
# Get the data
|
||||
data = input["data"]
|
||||
|
||||
# Get the min of the voxels sizes and use it as the resolution
|
||||
resolution = np.min(data.header.get_zooms()[:3])
|
||||
|
||||
# Load the parcellation
|
||||
t_parcellation, t_labels, _ = load_parcellation(
|
||||
name=self.parcellation_name,
|
||||
resolution=resolution,
|
||||
)
|
||||
|
||||
# Create a masker
|
||||
masker = NiftiLabelsMasker(
|
||||
```
masker = NiftiLabelsMasker(
labels_img=t_parcellation,
standardize=True,
memory="nilearn_cache",
verbose=5,
)
```
|
||||
labels_img=t_parcellation,
|
||||
standardize=True,
|
||||
memory='nilearn_cache',
|
||||
verbose=5,
|
||||
)
|
||||
|
||||
# mask the data
|
||||
out_values = masker.fit_transform([data])
|
||||
|
||||
# Create the output dictionary
|
||||
out = {"data": out_values, "columns": t_labels}
|
||||
|
||||
# If its 3D (BOLD), name each row as "scan"
|
||||
if out_values.shape[0] > 1:
|
||||
out["row_names"] = "scan"
|
||||
return out
|
||||
|
||||
|
||||
.. _extending_markers_store:
|
||||
|
||||
Step 4: Store the marker
|
||||
------------------------
|
||||
|
||||
In this step, we will define the method that stores the marker. This method will be called by junifer when needed,
|
||||
using the data provided by the ``compute`` method. The method ``store`` has three arguments:
|
||||
|
||||
* ``kind``: A string indicating the :ref:`data type <data_types>` that was used to compute the marker.
|
||||
* ``out``: The output of the ``compute`` method.
|
||||
* ``storage``: The storage object, that will be used to store the marker.
|
||||
|
||||
.. code-block:: python
|
||||
|
||||
def store(self, kind, out, storage):
|
||||
if kind in ["VBM_GM", "VBM_WM"]:
|
||||
storage.store(kind="table", **out)
|
||||
elif kind in ["BOLD"]:
|
||||
storage.store(kind="timeseries", **out)
|
||||
|
||||
.. hint:: Check the hint on :ref:`extending_markers_compute`. If the output of the ``compute`` method is a dictionary
|
||||
with keys based on the :ref:`storage types <storage_types>`, the ``store`` method can simply call the right
|
||||
storage function, based on the ``kind`` parameter, with ``**out``.
|
||||
|
||||
|
||||
.. _extending_markers_finalize:
|
||||
|
||||
Step 5: Finalize the marker
|
||||
---------------------------
|
||||
|
||||
Once all of the above steps are done, we just need to give our marker a name an register it using the
|
||||
``@register_marker`` decorator:
|
||||
|
||||
.. code-block:: python
|
||||
|
||||
from nilearn.maskers import NiftiLabelsMasker
|
||||
from junifer.data import load_parcellation
|
||||
from junifer.api.decorators import register_marker
|
||||
from junifer.markers.base import BaseMarker
|
||||
|
||||
@register_marker
|
||||
class ParcelMean(BaseMarker):
|
||||
|
||||
def __init__(self, parcellation_name, on=None, name=None):
|
||||
self.parcellation_name = parcellation_name
|
||||
super().__init__(on=on, name=name)
|
||||
|
||||
def get_valid_inputs(self):
|
||||
return ['BOLD', 'VBM_WM', 'VBM_GM']
|
||||
|
||||
def get_output_kind(self, input_kind):
|
||||
if input_kind == 'BOLD':
|
||||
return 'timeseries'
|
||||
else:
|
||||
return 'table'
|
||||
|
||||
def compute(self, input, extra_input):
|
||||
# Get the data
|
||||
data = input["data"]
|
||||
|
||||
# Get the min of the voxels sizes and use it as the resolution
|
||||
resolution = np.min(data.header.get_zooms()[:3])
|
||||
|
||||
# Load the parcellation
|
||||
t_parcellation, t_labels, _ = load_parcellation(
|
||||
name=self.parcellation_name,
|
||||
resolution=resolution,
|
||||
)
|
||||
|
||||
# Create a masker
|
||||
masker = NiftiLabelsMasker(
|
||||
```
masker = NiftiLabelsMasker(
labels_img=t_parcellation,
standardize=True,
memory="nilearn_cache",
verbose=5,
)
```
|
||||
labels_img=t_parcellation,
|
||||
standardize=True,
|
||||
memory='nilearn_cache',
|
||||
verbose=5,
|
||||
)
|
||||
|
||||
# mask the data
|
||||
out_values = masker.fit_transform([data])
|
||||
|
||||
# Create the output dictionary
|
||||
out = {"data": out_values, "columns": t_labels}
|
||||
|
||||
# If its 3D (BOLD), name each row as "scan"
|
||||
if out_values.shape[0] > 1:
|
||||
out["row_names"] = "scan"
|
||||
return out
|
||||
|
||||
def store(self, kind, out, storage):
|
||||
if kind in ["VBM_GM", "VBM_WM"]:
|
||||
storage.store(kind="table", **out)
|
||||
elif kind in ["BOLD"]:
|
||||
storage.store(kind="timeseries", **out)
|
||||
|
||||
|
||||
.. _extending_markers_template:
|
||||
|
||||
Template for a custom Marker
|
||||
----------------------------
|
||||
|
||||
.. code-block:: python
|
||||
|
||||
from junifer.api.decorators import register_marker
|
||||
from junifer.markers.base import BaseMarker
|
||||
|
||||
@register_marker
|
||||
class TemplateMarker(BaseMarker):
|
||||
|
||||
def __init__(self, on=None, name=None):
|
||||
# TODO: add marker-specific parameters
|
||||
super().__init__(on=on, name=name)
|
||||
|
||||
def get_valid_inputs(self):
|
||||
# TODO: Complete with the valid inputs
|
||||
valid = []
|
||||
return valid
|
||||
|
||||
def get_output_kind(self, input_kind):
|
||||
# TODO: Return the valid output kind for each input kind
|
||||
pass
|
||||
|
||||
def compute(self, input, extra_input):
|
||||
# TODO: compute the marker and create the output dictionary
|
||||
|
||||
# Create the output dictionary
|
||||
out = {"data": None, "columns": None}
|
||||
return out
|
||||
|
||||
def store(self, kind, out, storage):
|
||||
# TODO: store out using the storage object, based on the kind of data
|
||||
pass
|
||||
BIN
docs/images/pipeline/pipeline.001.png
Normal file
|
After Width: | Height: | Size: 28 KiB |
BIN
docs/images/pipeline/pipeline.002.png
Normal file
|
After Width: | Height: | Size: 103 KiB |
|
|
@ -24,7 +24,9 @@ enabling others to extend it easily.
|
|||
|
||||
installation
|
||||
understanding/index.rst
|
||||
using/index.rst
|
||||
builtin
|
||||
extending/index.rst
|
||||
auto_examples/index.rst
|
||||
api/index.rst
|
||||
contribution
|
||||
|
|
|
|||
|
|
@ -21,16 +21,24 @@
|
|||
.. _`matplotlib`: https://matplotlib.org
|
||||
.. _`fMRIPrep`: https://fmriprep.org
|
||||
.. _`nilearn`: https://nilearn.github.io
|
||||
.. _`nipype`: https://nipype.readthedocs.io
|
||||
.. _`datalad`: https://datalad.org
|
||||
|
||||
.. _`venv`: https://docs.python.org/3/tutorial/venv.html
|
||||
.. _`conda env`: https://docs.conda.io/projects/conda/en/latest/user-guide/tasks/manage-environments.html
|
||||
|
||||
.. _`Github`: https://github.com/
|
||||
.. _`junifer Github`: https://github.com/juaml/junifer
|
||||
.. _`junifer Discussions`: https://github.com/juaml/junifer/discussions
|
||||
|
||||
.. _`Wikipedia`: https://wikipedia.org
|
||||
.. _`YAML`: https://yaml.org
|
||||
|
||||
.. _`setuptools_scm`: https://github.com/pypa/setuptools_scm/
|
||||
|
||||
.. _`sphinx gallery`: https://sphinx-gallery.github.io/stable/index.html
|
||||
.. _`sphinx reST reference`: https://www.sphinx-doc.org/en/master/usage/restructuredtext/basics.html#inline-markup
|
||||
|
||||
.. _`HTCondor`: https://research.cs.wisc.edu/htcondor/
|
||||
.. _`SLURM`: https://slurm.schedmd.com
|
||||
.. _`GNU Parallel`: https://www.gnu.org/software/parallel/
|
||||
|
|
@ -18,7 +18,7 @@ The second level of keys are the actual data. So far, there are two keys used:
|
|||
- ``path``: path to the file containing the data.
|
||||
- ``data``: the data loaded in memory.
|
||||
|
||||
The :ref:`DataGrabber <datagrabber>` step will only fill the ``path`` value.
|
||||
The :ref:`Data Grabber <datagrabber>` step will only fill the ``path`` value.
|
||||
The ``data`` value will be filled by the :ref:`DataReader <datareader>` step, if it is one of the possible file types
|
||||
that the datareader can read.
|
||||
|
||||
|
|
@ -42,9 +42,21 @@ Data types
|
|||
* - ``BOLD``
|
||||
|
`CONN toolbox`?
`CONN toolbox`?
|
||||
- BOLD image (4D)
|
||||
- Preprocessed/Denoised BOLD image (fmriprep output)
|
||||
* - ``BOLD_confounds``
|
||||
- BOLD image confounds (CSV/TSV file)
|
||||
- Confounds that can be applied to the BOLD image.
|
||||
* - ``VBM_GM``
|
||||
- VBM Gray Matter segmentation (3D)
|
||||
- CAT output (`m0wp1` images)
|
||||
* - ``VBM_WM``
|
||||
- VBM White Matter segmentation (3D)
|
||||
- CAT output (`m0wp2` images)
|
||||
* - ``fALFF``
|
||||
- Voxel-wise fALFF image (3D)
|
||||
- fALFF computed with CONN toolbox
|
||||
* - ``GCOR``
|
||||
- Global Correlation image (3D)
|
||||
- GCOR computed with CONN toolbox
|
||||
* - ``LCOR``
|
||||
- Local Correlation image (3D)
|
||||
- LCOR computed with CONN toolbox
|
||||
|
|
@ -2,13 +2,13 @@
|
|||
|
||||
.. _datagrabber:
|
||||
|
||||
DataGrabber
|
||||
===========
|
||||
Data Grabber
|
||||
============
|
||||
|
||||
Description
|
||||
-----------
|
||||
|
||||
The ``DataGrabber`` is an object that can provide an interface to datasets you want to work with in junifer.
|
||||
The *Data Grabber* is an object that can provide an interface to datasets you want to work with in junifer.
|
||||
Every concrete implementation of a datagrabber is aware of a particular dataset's structure and thus allows
|
||||
you to fetch specific elements of interest from the dataset. It adds the ``path`` key to each :ref:`data type <data_types>`
|
||||
in the :ref:`Data object <data_object>`.
|
||||
|
|
|
|||
|
|
@ -2,13 +2,13 @@
|
|||
|
||||
.. _datareader:
|
||||
|
||||
DataReader
|
||||
==========
|
||||
Data Reader
|
||||
===========
|
||||
|
||||
Description
|
||||
-----------
|
||||
|
||||
The ``DataReader`` is an object that is responsible for actually reading data files in junifer.
|
||||
The *Data Reader* is an object that is responsible for actually reading data files in junifer.
|
||||
It reads the value of the key ``path`` for each :ref:`data type <data_types>` in the :ref:`Data object <data_object>`
|
||||
and loads them to memory. After reading the data into memory, it adds the key ``data`` to the same level as ``path``
|
||||
and the value is the actual data in the memory.
|
||||
|
|
@ -16,7 +16,7 @@ and the value is the actual data in the memory.
|
|||
Datareaders are meant to be used inside the datagrabber context but you can operate on them outside the context as long
|
||||
as the actual data is in the memory and the Python runtime has not garbage-collected it.
|
||||
|
||||
For data formats not supported by junifer yet, you can either make your own ``DataReader`` or open an issue on
|
||||
For data formats not supported by junifer yet, you can either make your own *Data Reader* or open an issue on
|
||||
`junifer Github`_ and we can help you out.
|
||||
|
||||
Currently supported file-formats
|
||||
|
|
|
|||
|
|
@ -1,5 +1,7 @@
|
|||
.. include:: ../links.inc
|
||||
|
||||
.. _understanding:
|
||||
|
||||
Understanding junifer
|
||||
=====================
|
||||
|
||||
|
|
@ -16,11 +18,15 @@ structural MRI, diffusion MRI, etc.) and you want to extract features to
|
|||
later use in statistical analyses or machine learning (for example, using
|
||||
julearn_).
|
||||
|
||||
.. important:: Junifer is not a toolbox to create pipelines, but a tool to configure the junifer pipeline, which is
|
||||
intended to be fixed and not to be changed. If you want to create a pipeline, you should use
|
||||
other tools like nipype_.
|
||||
|
||||
.. toctree::
|
||||
:maxdepth: 2
|
||||
:caption: Contents:
|
||||
|
||||
pipeline
|
||||
data
|
||||
datagrabber
|
||||
datareader
|
||||
|
|
|
|||
|
|
@ -9,7 +9,7 @@ Description
|
|||
-----------
|
||||
|
||||
The ``Marker`` is an object that is responsible for feature extraction. It primarily operates on data loaded
|
||||
in memory by :ref:`datareader <DataReader>` and stored in the ``data`` key of each :ref:`data type <data_types>`
|
||||
in memory by :ref:`Data Reader <datareader>` and stored in the ``data`` key of each :ref:`data type <data_types>`
|
||||
in the :ref:`Data object <data_object>`. In some cases, it can also operate on pre-processed data as obtained
|
||||
from the :ref:`Preprocess <preprocess>` step of the pipeline. It is important to note that this pre-process is
|
||||
not similar to pre-processing done by tools like FSL, SPM, AFNI, etc. . For example, one can perform confound
|
||||
|
|
|
|||
30
docs/understanding/pipeline.rst
Normal file
|
|
@ -0,0 +1,30 @@
|
|||
.. include:: ../links.inc
|
||||
|
`DataGrabber`
`dataset` instead of `database`?
It can be either. Given that we are coupled with datalad, I would keep it as dataset. It can be either. Given that we are coupled with datalad, I would keep it as dataset.
I stil prefer to use the two words to describe the concept and not the Class name. I stil prefer to use the two words to describe the concept and not the Class name.
In the Understanding section, we have In the Understanding section, we have `DataGrabber` for the concept as well. Again my reasoning is that if we keep it as one name throughout, users are not confused. And also, one can mentally link better to the name of the step being the class category's name.
|
||||
|
||||
.. _pipeline:
|
||||
|
||||
The Junifer Pipeline
|
||||
====================
|
||||
|
||||
The junifer pipeline is the main execution path of junifer. It consists of five steps:
|
||||
|
||||
1. :ref:`Data Grabber <datagrabber>`: Interpret the dataset and provide a list of files.
|
||||
2. :ref:`Data Reader <datareader>`: Read the files.
|
||||
|
`DataReader`
same as before same as before
|
||||
3. :ref:`Pre-processing <preprocess>`: Prepare the images for marker computation.
|
||||
|
`Preprocess`
same same
|
||||
4. :ref:`Marker Computation <marker>`: Compute the marker.
|
||||
|
`Marker`
same same
|
||||
5. :ref:`Storage <storage>`: Store the marker values.
|
||||
|
||||
The element that is passed accross the pipeline is called the :ref:`Data Object<data_object>`.
|
||||
|
||||
The following is a graphical representation of the pipeline:
|
||||
|
||||
.. image:: ../images/pipeline/pipeline.001.png
|
||||
|
||||
However, it is usually the case that several markers are computed for the same data. Thus, the *markers* step
|
||||
of the pipeline is defined as a list of markers. The following is a graphical representation of the pipeline execution
|
||||
on multiple markers:
|
||||
|
||||
.. image:: ../images/pipeline/pipeline.002.png
|
||||
|
||||
|
||||
.. note:: To avoid keeping in memory all of the computed marker, the storage step is called after each marker
|
||||
computation, releasing the memory used to compute each marker.
|
||||
|
|
@ -20,7 +20,7 @@ Confound Removal
|
|||
----------------
|
||||
|
||||
The *Confound Removal* step is meant to remove *confounds* from the ``BOLD`` data. The confounds are
|
||||
extracted from the ``BOLD_confounds`` data (must be provided by the :ref:`DataGrabber <datagrabber>`).
|
||||
extracted from the ``BOLD_confounds`` data (must be provided by the :ref:`Data Grabber <datagrabber>`).
|
||||
The confounds are then regressed out from the ``BOLD`` data using :func:`nilearn.image.clean_img`.
|
||||
|
||||
Currently, junifer supports only one confound removal class:
|
||||
|
|
|
|||
|
|
@ -24,6 +24,35 @@ methods respectively.
|
|||
For storage interfaces not supported by junifer yet, you can either make your own ``Storage`` by providing a concrete
|
||||
implementation of :class:`junifer.storage.BaseFeatureStorage` or open an issue on `junifer Github`_ and we can help you out.
|
||||
|
||||
|
||||
.. _storage_types:
|
||||
|
||||
Currently supported storage types
|
||||
---------------------------------
|
||||
|
||||
.. list-table::
|
||||
:widths: auto
|
||||
:header-rows: 1
|
||||
|
||||
* - Storage Type
|
||||
- Description
|
||||
- Options
|
||||
- Reference
|
||||
* - ``matrix``
|
||||
- A 2D matrix with row and column names
|
||||
- ``col_names``, ``row_names``, ``matrix_kind``, ``diagonal``
|
||||
- :meth:`junifer.storage.BaseFeatureStorage.store_matrix`
|
||||
* - ``table``
|
||||
- A vector of values with column names
|
||||
- ``columns``, ``row_names``
|
||||
- :meth:`junifer.storage.BaseFeatureStorage.store_table`
|
||||
* - ``timeseries``
|
||||
- A 2D matrix of values with column names
|
||||
- ``columns``, ``row_names``
|
||||
- :meth:`junifer.storage.BaseFeatureStorage.store_timeseries`
|
||||
|
||||
.. _storage_interfaces:
|
||||
|
||||
Currently supported storage interfaces
|
||||
--------------------------------------
|
||||
|
||||
|
|
@ -36,6 +65,6 @@ Currently supported storage interfaces
|
|||
- File type
|
||||
- Storage kinds
|
||||
* - :class:`junifer.storage.SQLiteFeatureStorage`
|
||||
- ``.db``
|
||||
- ``.sqlite``
|
||||
- SQLite
|
||||
- ``matrix``, ``table``, ``timeseries``
|
||||
|
|
|
|||
204
docs/using/codeless.rst
Normal file
|
|
@ -0,0 +1,204 @@
|
|||
.. include:: ../links.inc
|
||||
```
``preprocess``
```
`... datareader ...`
I believe the YAML syntax is I believe the YAML syntax is `false` and when we load YAML, it should automatically convert it to `False`.
Same as above but for Same as above but for `true` and `True`.
`compute`
was not sure about this. was not sure about this.
I think it does, can you please check it? I think it does, can you please check it?
|
||||
|
||||
.. _codeless:
|
||||
|
||||
Code-less configuration
|
||||
=======================
|
||||
|
||||
On of the most important features of junifer is its capacity to run without writing a single line of code. This is
|
||||
achieved by using a configuration file that is written in YAML_. In this file, we configure the different steps of
|
||||
:ref:`pipeline`.
|
||||
|
||||
As a reminder, this is how the pipeline looks like:
|
||||
|
||||
.. image:: ../images/pipeline/pipeline.001.png
|
||||
|
||||
Thus, the configuration file must configure each of the sections of the pipeline, as well as some general parameters.
|
||||
|
||||
As an example, we will generate the configuration file for a pipeline that will extract the mean ``VBM_GM`` values
|
||||
using two different parcellations and one set of coordinates, from the *Oasis VBM Testing dataset* included in
|
||||
junifer.
|
||||
|
||||
|
||||
General Parameters
|
||||
------------------
|
||||
|
||||
The general parameters are the ones that are not specific to any of the sections of the pipeline, but configure
|
||||
junifer as a whole. These parameters are:
|
||||
|
||||
* ``with``: A section used to specify modules and junifer extensions to use.
|
||||
* ``workdir``: The working directory where junifer will store temporary files.
|
||||
|
||||
Since the example uses a specific datagrabber for testing, we need to add ``junifer.testing.registry`` to the
|
||||
``with`` section. This will allow junifer to find the datagrabber. We will set the ``workdir`` to ``/tmp``.
|
||||
|
||||
.. code-block:: yaml
|
||||
|
||||
with: junifer.testing.registry
|
||||
workdir: /tmp
|
||||
|
||||
Step-by-step configuration
|
||||
--------------------------
|
||||
|
||||
In order to configure the pipeline, we need to configure each step:
|
||||
|
||||
* ``datagrabber``
|
||||
* ``datareader``
|
||||
* ``preprocess``
|
||||
* ``markers``
|
||||
* ``storage``
|
||||
|
||||
.. important:: The datareader step configuration is optional, as junifer only provides one datareader. Nevertheless,
|
||||
it is possible to extend junifer with custom datareaders, and thus, it is also possible to configure this step.
|
||||
|
||||
|
||||
Data Grabber
|
||||
|
`DataGrabber`
will keep it as concepts and not class names will keep it as concepts and not class names
|
||||
^^^^^^^^^^^^
|
||||
|
||||
The ``datagrabber`` section must be configured using the ``kind`` key to specify the datagrabber to use. Additional
|
||||
keys correspond to the parameters of the datagrabber.
|
||||
|
||||
For example, to use the :class:`junifer.datagrabber.DataladAOMICPIOP1` datagrabber, we just need to
|
||||
specify its name as the ``kind`` key.
|
||||
|
||||
.. code-block:: yaml
|
||||
|
||||
datagrabber:
|
||||
kind: DataladAOMICPIOP1
|
||||
|
||||
However, it is also possible to pass parameters to the datagrabber. In this case, we can restrict the datagrabber to
|
||||
fetch only the ``restingstate`` task.
|
||||
|
||||
.. code-block:: yaml
|
||||
|
||||
datagrabber:
|
||||
kind: DataladAOMICPIOP1
|
||||
tasks: restingstate
|
||||
|
||||
In the *Oasis VBM Testing dataset* example, the section will look like this:
|
||||
|
||||
.. code-block:: yaml
|
||||
|
||||
datagrabber:
|
||||
kind: OasisVBMTesting
|
||||
|
||||
|
||||
Data Reader
|
||||
|
`DataReader`
|
||||
^^^^^^^^^^^
|
||||
|
||||
As mentioned before, this section is entirely optional, as junifer only provides one data reader
|
||||
(:class:`junifer.datareader.DefaultDataReader`), which is the default in case the section is not specified.
|
||||
|
||||
In any case, the syntax of the section is the same as for the ``datagrabber`` section, using the ``kind`` key to
|
||||
specify the data reader to use, and additional keys to pass parameters to the data reader:
|
||||
|
||||
.. code-block:: yaml
|
||||
|
||||
datareader:
|
||||
kind: DefaultDataReader
|
||||
|
||||
|
||||
For the *Oasis VBM Testing dataset* example, we will not specify a ``datareader`` step.
|
||||
|
||||
Preprocessing
|
||||
^^^^^^^^^^^^^
|
||||
|
||||
Preprocessing is also an optional step, as it might be the case that no pre-processing is needed. In the case that
|
||||
preprocessing is needed, the section must be configured using the ``kind`` key to specify the preprocessor to use,
|
||||
and additional keys to pass parameters to the preprocessor.
|
||||
|
||||
For example, to use the :class:`junifer.preprocess.fMRIPrepConfoundRemover` preprocessor, we just need to specify its
|
||||
name as the ``kind`` key, as well as its parameters.
|
||||
|
||||
|
||||
.. code-block:: yaml
|
||||
|
||||
preproces:
|
||||
kind: fMRIPrepConfoundRemover
|
||||
strategy:
|
||||
motion: full
|
||||
wm_csf: full
|
||||
global_signal: basic
|
||||
spike: 0.2
|
||||
detrend: false
|
||||
standardize: true
|
||||
|
||||
|
||||
For the *Oasis VBM Testing dataset* example, we will not specify a preprocessing step.
|
||||
|
||||
|
||||
Markers
|
||||
^^^^^^^
|
||||
|
||||
The ``markers`` section diverges from the previous ones, as we need to specify a list of markers. Each marker has a
|
||||
name that we can use to refer to it later, and a set of parameters that will be passed to the marker.
|
||||
|
||||
For the *Oasis VBM Testing dataset* example, we want to compute the mean ``VBM_GM`` value for each parcel using the
|
||||
Schaefer parcellation (100 parcels, 7 networks), Schaefer parcellation (200 parcels, 7 networks), and the *DMNBuckner*
|
||||
network, using 5mm spheres. Thus, we will configure the ``markers`` section as follows:
|
||||
|
||||
.. code-block:: yaml
|
||||
|
||||
markers:
|
||||
- name: Schaefer100x7_mean
|
||||
kind: ParcelAggregation
|
||||
parcellation: Schaefer100x7
|
||||
method: mean
|
||||
- name: Schaefer200x7_mean
|
||||
kind: ParcelAggregation
|
||||
parcellation: Schaefer200x7
|
||||
method: mean
|
||||
- name: DMNBuckner_5mm_mean
|
||||
kind: SphereAggregation
|
||||
coords: DMNBuckner
|
||||
radius: 5
|
||||
method: mean
|
||||
|
||||
|
||||
Storage
|
||||
^^^^^^^
|
||||
|
||||
Finally, we need to define how and where the results will be stored. This is done using the ``storage`` section,
|
||||
which must be configured using the ``kind`` key to specify the storage to use, and additional keys to pass parameters.
|
||||
|
||||
For example, to use the :class:`junifer.storage.SQLiteFeatureStorage` storage, we just need to specify where we want
|
||||
to store the results:
|
||||
|
||||
.. code-block:: yaml
|
||||
|
||||
storage:
|
||||
kind: SQLiteFeatureStorage
|
||||
uri: /data/junifer/example/oasis_vbm_testing.sqlite
|
||||
|
In the storage types, we have In the storage types, we have `.db` extension for SQLite. Just to be consistent, maybe we use it here?
it does not matter, the user sets the name and extension. It should be sqlite. it does not matter, the user sets the name and extension. It should be sqlite.
It indeed does not matter, my argument is just for the sake of consistency. I can imagine it be confusing for users who are not familiar with SQLite in that detail. It indeed does not matter, my argument is just for the sake of consistency. I can imagine it be confusing for users who are not familiar with SQLite in that detail.
|
||||
|
||||
|
||||
The full example
|
||||
----------------
|
||||
|
||||
This is how the full *Oasis VBM Testing dataset* example configuration file looks like:
|
||||
|
||||
.. code-block:: yaml
|
||||
|
||||
with: junifer.testing.registry
|
||||
workdir: /tmp
|
||||
|
||||
datagrabber:
|
||||
kind: OasisVBMTesting
|
||||
|
||||
markers:
|
||||
- name: Schaefer100x7_mean
|
||||
kind: ParcelAggregation
|
||||
parcellation: Schaefer100x7
|
||||
method: mean
|
||||
- name: Schaefer200x7_mean
|
||||
kind: ParcelAggregation
|
||||
parcellation: Schaefer200x7
|
||||
method: mean
|
||||
- name: DMNBuckner_5mm_mean
|
||||
kind: SphereAggregation
|
||||
coords: DMNBuckner
|
||||
radius: 5
|
||||
method: mean
|
||||
|
||||
storage:
|
||||
kind: SQLiteFeatureStorage
|
||||
uri: /data/junifer/example/oasis_vbm_testing.sqlite
|
||||
18
docs/using/index.rst
Normal file
|
|
@ -0,0 +1,18 @@
|
|||
.. include:: ../links.inc
|
||||
|
||||
.. _using:
|
||||
|
||||
Using junifer
|
||||
=============
|
||||
|
||||
In this section, we will cover the main aspects behind using junifer. We will first explain the basics behind junifer's
|
||||
code-less configuration. Then we will show how to use the command line interface to ``run`` junifer and ``collect``
|
||||
the results. Finally, we will show how to use the ``queue`` command to interact with HPC and HTC systems.
|
||||
|
||||
.. toctree::
|
||||
:maxdepth: 2
|
||||
:caption: Contents:
|
||||
|
||||
codeless
|
||||
running
|
||||
queueing
|
||||
89
docs/using/queueing.rst
Normal file
|
|
@ -0,0 +1,89 @@
|
|||
.. include:: ../links.inc
|
||||
|
`If you are in immediate need of any of these ...`
`scheduler`
```
``HTCondor``
```
`Python`
Please check my argument for having Please check my argument for having ``true``.
|
||||
|
||||
.. _queueing:
|
||||
|
||||
Queueing jobs (HPC, HTC)
|
||||
========================
|
||||
|
||||
Yet another interesting feature of junifer is the ability to queue jobs on computational clusters. This is done by
|
||||
adding the ``queue`` section in the :ref:`codeless` file and executing the ``junifer queue`` command.
|
||||
|
||||
While junifer is meant to support `HTCondor`_, `SLURM`_ and local queueing using `GNU Parallel`_, only HTCondor is
|
||||
currently supported. This will be implemented in future relases of junifer. If you are in immediate need of any of these
|
||||
schedulers, please create an issue on the `junifer github`_ repository.
|
||||
|
||||
The ``queue`` section of the :ref:`codeless` must start by defining the following general parameters:
|
||||
|
||||
* ``jobname``: name of the job to be queued. This will be used to name the folder where the job files will be created,
|
||||
as well as any relevant file. Depending on the scheduler, it will also be listed in the queueing system with this
|
||||
name.
|
||||
* ``kind``: The kind of scheduler to be used. Currently, only ``HTCondor`` is supported.
|
||||
|
||||
Example:
|
||||
|
||||
.. code-block:: yaml
|
||||
|
||||
queue:
|
||||
jobname: TestHTCondorQueue
|
||||
kind: HTCondor
|
||||
|
||||
|
||||
The rest of the parameters depend on the scheduler you are using.
|
||||
|
||||
.. _queueing_condor:
|
||||
|
||||
HTCondor
|
||||
--------
|
||||
|
||||
When using HTCondor, junifer will use a DAG to queue one job per element (``junifer run``). As an option, the DAG can
|
||||
include a final job (``junifer collect``) to collect the results once all of the individual element jobs are finished.
|
||||
|
||||
The following parameters are avilable for HTCondor:
|
||||
|
||||
* ``env``: Definition of the Python enviroment. It must provide two variables: ``kind`` and ``name``. The ``kind``
|
||||
corresponds to the kind of virtual environment to use: ``conda``, ``virtualenv`` (not yet supported) or
|
||||
``local`` (no virtual enviroment). The ``name`` is the name of the enviroment to use in case a virtual environment
|
||||
is used.
|
||||
* ``mem``: Memory to be used by the job. It must be provided as a string with the units (e.g. ``2GB``).
|
||||
* ``cpus``: Number of CPUs to be used by the job. It must be provided as an int.
|
||||
* ``disk``: Disk space to be used by the job. It must be provided as a string with the units (e.g. ``2GB``). Keep in
|
||||
mind that junifer uses a local working directory for each job, and datalad datasets might be cloned in this temporary
|
||||
directory.
|
||||
* ``extra_preamble``: Extra lines to be added to the HTCondor submit file. This can be used to add
|
||||
extra parameters to the job, such as ``requirements``.
|
||||
* ``collect``: If ``true``, a final job will be added to the DAG to collect the results once all of the individual
|
||||
element jobs are finished. This is useful if you want to run a ``junifer collect`` job only once all of the
|
||||
individual element jobs are finished. If not specified, it will default to ``true``.
|
||||
|
||||
|
||||
Example:
|
||||
|
||||
.. code-block:: yaml
|
||||
|
||||
queue:
|
||||
jobname: TestHTCondorQueue
|
||||
kind: HTCondor
|
||||
env:
|
||||
kind: conda
|
||||
name: junifer
|
||||
mem: 8G
|
||||
disk: 2GB
|
||||
collect: true
|
||||
|
||||
|
||||
Once the :ref:`codeless` file is ready, including the ``queue`` section, you can queue the jobs by executing
|
||||
the ``junifer queue`` command.
|
||||
|
||||
The ``queue`` command will create a folder with the name of the job (``jobname``) under the ``junifer_jobs`` directory
|
||||
in the current working directory.
|
||||
|
||||
The ``queue`` command accepts the following arguments:
|
||||
|
||||
* ``--help``: Show a help message.
|
||||
* ``--verbose`` Set the verbosity level. Options are ``warning``, ``info``, ``debug``.
|
||||
* ``--submit``: Submit the jobs to the queueing system. If not specified, the job submit files will be created but not
|
||||
submitted.
|
||||
* ``--overwrite``: Overwrite the job folder if it already exists. If not specified, the command will fail if the job
|
||||
folder already exists.
|
||||
* ``--element``: Queue only the specified element(s). If not specified, all elements will be queued.
|
||||
|
||||
60
docs/using/running.rst
Normal file
|
|
@ -0,0 +1,60 @@
|
|||
.. include:: ../links.inc
|
||||
|
`console` instead of `bash`?
`console` instead of `bash`?
`console` instead of `bash`?
`console` instead of `bash`?
|
||||
|
||||
.. _running:
|
||||
|
||||
Running jobs
|
||||
============
|
||||
|
||||
Once we have the :ref:`code-less configuration file <codeless>`, we can use the command line interface to extract the
|
||||
features. This is achieved in a two-step process: ``run`` and ``collect``.
|
||||
|
||||
The ``run`` command is used to extract the features from each element in the dataset. However, depending on the
|
||||
storage interface, this may create one file per subject. The ``collect`` command is then used to collect all of the
|
||||
individual results into a single file.
|
||||
|
||||
Assuming that we have a configuration file named ``config.yaml``, the following commands will extract the features:
|
||||
|
||||
.. code-block:: console
|
||||
|
||||
junifer run config.yaml
|
||||
|
||||
The ``run`` command accepts the following additional arguments:
|
||||
|
||||
* ``--help``: Show a help message.
|
||||
* ``--verbose`` Set the verbosity level. Options are ``warning``, ``info``, ``debug``.
|
||||
* ``--element``: The *element* to run. If not specified, all elements will be run. This parameter can be specified
|
||||
|
The rendering for this has some indentation issue. The rendering for this has some indentation issue.
|
||||
multiple times to run multiple elements. If the *element* requires several parameters, they can be specified
|
||||
by separating them with ``,``.
|
||||
|
||||
|
||||
Example on running two elements:
|
||||
|
||||
.. code-block:: console
|
||||
|
||||
junifer run config.yaml --element sub-01 --element sub-02
|
||||
|
||||
Example on elements with multiple parameters and verbose output:
|
||||
|
||||
.. code-block:: console
|
||||
|
||||
junifer run --verbose info config.yaml --element sub-01,ses-01
|
||||
|
||||
.. _collect:
|
||||
|
||||
Collecting results
|
||||
==================
|
||||
|
||||
Once the ``run`` command has been executed, the results are stored in the output directory. However, depending on the
|
||||
storage interface, this may create one file per subject. The ``collect`` command is then used to collect all of the
|
||||
individual results into a single file.
|
||||
|
||||
Assuming that we have a configuration file named ``config.yaml``, the following commands will collect the results:
|
||||
|
||||
.. code-block:: console
|
||||
|
||||
junifer collect config.yaml
|
||||
|
||||
The ``collect`` command accepts the following additional arguments:
|
||||
|
||||
* ``--help``: Show a help message.
|
||||
* ``--verbose`` Set the verbosity level. Options are ``warning``, ``info``, ``debug``.
|
||||
|
|
@ -23,10 +23,10 @@ configure_logging(level="INFO")
|
|||
# The BIDS datagrabber requires three parameters: the types of data we want,
|
||||
# the specific pattern that matches each type, and the variables that will be
|
||||
# replaced int he patterns.
|
||||
types = ["T1w", "bold"]
|
||||
types = ["T1w", "BOLD"]
|
||||
patterns = {
|
||||
"T1w": "{subject}/anat/{subject}_T1w.nii.gz",
|
||||
"bold": "{subject}/func/{subject}_task-rest_bold.nii.gz",
|
||||
"BOLD": "{subject}/func/{subject}_task-rest_bold.nii.gz",
|
||||
}
|
||||
replacements = ["subject"]
|
||||
###############################################################################
|
||||
|
|
|
|||
|
|
@ -52,7 +52,7 @@ with tempfile.TemporaryDirectory() as tmpdir:
|
|||
# Define the storage interface
|
||||
storage = {
|
||||
"kind": "SQLiteFeatureStorage",
|
||||
"uri": f"{tmpdir}/test.db",
|
||||
"uri": f"{tmpdir}/test.sqlite",
|
||||
}
|
||||
# Run the defined junifer feature extraction pipeline
|
||||
run(
|
||||
|
|
|
|||
|
|
@ -71,7 +71,7 @@ sex = (
|
|||
# Create a temporary directory for junifer feature extraction:
|
||||
with tempfile.TemporaryDirectory() as tmpdir:
|
||||
|
||||
storage = {"kind": "SQLiteFeatureStorage", "uri": f"{tmpdir}/test.db"}
|
||||
storage = {"kind": "SQLiteFeatureStorage", "uri": f"{tmpdir}/test.sqlite"}
|
||||
# run the defined junifer feature extraction pipeline
|
||||
run(
|
||||
workdir="/tmp",
|
||||
|
|
|
|||
|
|
@ -43,7 +43,7 @@ storage = {
|
|||
}
|
||||
|
||||
with tempfile.TemporaryDirectory() as tmpdir:
|
||||
uri = f"{tmpdir}/test.db"
|
||||
uri = f"{tmpdir}/test.sqlite"
|
||||
storage["uri"] = uri
|
||||
run(
|
||||
workdir="/tmp",
|
||||
|
|
|
|||
|
|
@ -21,5 +21,5 @@ markers:
|
|||
method: std
|
||||
storage:
|
||||
kind: SQLiteFeatureStorage
|
||||
uri: /Users/fraimondo/dev/tbox/junifer/scratch/db/test.db
|
||||
uri: /Users/fraimondo/dev/tbox/junifer/scratch/db/test.sqlite
|
||||
|
||||
|
|
|
|||
|
|
@ -10,7 +10,7 @@ markers:
|
|||
method: mean
|
||||
storage:
|
||||
kind: SQLiteFeatureStorage
|
||||
uri: /data/group/appliedml/fraimondo/junifer_test/test.db
|
||||
uri: /data/group/appliedml/fraimondo/junifer_test/test.sqlite
|
||||
queue:
|
||||
jobname: TestHTCondorQueue
|
||||
kind: HTCondor
|
||||
|
|
|
|||
|
|
@ -21,5 +21,5 @@ markers:
|
|||
method: std
|
||||
storage:
|
||||
kind: SQLiteFeatureStorage
|
||||
uri: /data/project/ukb_motor/junifer_test/test.db
|
||||
uri: /data/project/ukb_motor/junifer_test/test.sqlite
|
||||
|
||||
|
|
|
|||
|
|
@ -33,6 +33,30 @@ def register_datagrabber(klass: Type) -> Type:
|
|||
return klass
|
||||
|
`datareader`
|
||||
|
||||
|
||||
def register_datareader(klass: Type) -> Type:
|
||||
"""Datareader registration decorator.
|
||||
|
||||
Registers the datareader so it can be used by name.
|
||||
|
||||
Parameters
|
||||
----------
|
||||
klass: class
|
||||
The class of the datareader to register.
|
||||
|
||||
Returns
|
||||
-------
|
||||
klass: class
|
||||
The unmodified input class.
|
||||
|
||||
"""
|
||||
register(
|
||||
step="datareader",
|
||||
name=klass.__name__,
|
||||
klass=klass,
|
||||
)
|
||||
return klass
|
||||
|
||||
|
||||
def register_preprocessor(klass: Type) -> Type:
|
||||
"""Preprocessor registration decorator.
|
||||
|
||||
|
|
|
|||
|
|
@ -11,5 +11,5 @@ markers:
|
|||
method: mean
|
||||
storage:
|
||||
kind: SQLiteFeatureStorage
|
||||
uri: /Users/fraimondo/dev/tbox/junifer/scratch/db/test.db
|
||||
uri: /Users/fraimondo/dev/tbox/junifer/scratch/db/test.sqlite
|
||||
|
||||
|
|
|
|||
|
|
@ -10,7 +10,7 @@ markers:
|
|||
method: mean
|
||||
storage:
|
||||
kind: SQLiteFeatureStorage
|
||||
uri: /Users/fraimondo/dev/tbox/junifer/scratch/db/test.db
|
||||
uri: /Users/fraimondo/dev/tbox/junifer/scratch/db/test.sqlite
|
||||
queue:
|
||||
jobname: TestHTCondorQueue
|
||||
kind: HTCondor
|
||||
|
|
|
|||
|
|
@ -61,7 +61,7 @@ def test_run_single_element(tmp_path: Path) -> None:
|
|||
outdir = tmp_path / "out"
|
||||
outdir.mkdir()
|
||||
# Create storage
|
||||
uri = outdir / "test.db"
|
||||
uri = outdir / "test.sqlite"
|
||||
storage["uri"] = uri # type: ignore
|
||||
# Run operations
|
||||
run(
|
||||
|
|
@ -72,7 +72,7 @@ def test_run_single_element(tmp_path: Path) -> None:
|
|||
elements=["sub-01"],
|
||||
)
|
||||
# Check files
|
||||
files = list(outdir.glob("*.db"))
|
||||
files = list(outdir.glob("*.sqlite"))
|
||||
assert len(files) == 1
|
||||
|
||||
|
||||
|
|
@ -92,7 +92,7 @@ def test_run_multi_element(tmp_path: Path) -> None:
|
|||
outdir = tmp_path / "out"
|
||||
outdir.mkdir()
|
||||
# Create storage
|
||||
uri = outdir / "test.db"
|
||||
uri = outdir / "test.sqlite"
|
||||
storage["uri"] = uri # type: ignore
|
||||
storage["single_output"] = False # type: ignore
|
||||
# Run operations
|
||||
|
|
@ -104,7 +104,7 @@ def test_run_multi_element(tmp_path: Path) -> None:
|
|||
elements=["sub-01", "sub-03"],
|
||||
)
|
||||
# Check files
|
||||
files = list(outdir.glob("*.db"))
|
||||
files = list(outdir.glob("*.sqlite"))
|
||||
assert len(files) == 2
|
||||
|
||||
|
||||
|
|
@ -124,7 +124,7 @@ def test_run_multi_element_single_output(tmp_path: Path) -> None:
|
|||
outdir = tmp_path / "out"
|
||||
outdir.mkdir()
|
||||
# Create storage
|
||||
uri = outdir / "test.db"
|
||||
uri = outdir / "test.sqlite"
|
||||
storage["uri"] = uri # type: ignore
|
||||
storage["single_output"] = True # type: ignore
|
||||
# Run operations
|
||||
|
|
@ -136,9 +136,9 @@ def test_run_multi_element_single_output(tmp_path: Path) -> None:
|
|||
elements=["sub-01", "sub-03"],
|
||||
)
|
||||
# Check files
|
||||
files = list(outdir.glob("*.db"))
|
||||
files = list(outdir.glob("*.sqlite"))
|
||||
assert len(files) == 1
|
||||
assert files[0].name == "test.db"
|
||||
assert files[0].name == "test.sqlite"
|
||||
|
||||
|
||||
def test_run_and_collect(tmp_path: Path) -> None:
|
||||
|
|
@ -157,7 +157,7 @@ def test_run_and_collect(tmp_path: Path) -> None:
|
|||
outdir = tmp_path / "out"
|
||||
outdir.mkdir()
|
||||
# Create storage
|
||||
uri = outdir / "test.db"
|
||||
uri = outdir / "test.sqlite"
|
||||
storage["uri"] = uri # type: ignore
|
||||
storage["single_output"] = False # type: ignore
|
||||
# Run operations
|
||||
|
|
@ -173,9 +173,9 @@ def test_run_and_collect(tmp_path: Path) -> None:
|
|||
)
|
||||
elements = dg.get_elements() # type: ignore
|
||||
# This should create 10 files
|
||||
files = list(outdir.glob("*.db"))
|
||||
files = list(outdir.glob("*.sqlite"))
|
||||
assert len(files) == len(elements)
|
||||
# But the test.db file should not exist
|
||||
# But the test.sqlite file should not exist
|
||||
assert not uri.exists()
|
||||
# Collect in storage
|
||||
collect(storage)
|
||||
|
|
|
|||
|
|
@ -13,9 +13,9 @@ from ....datagrabber import PatternDataladDataGrabber
|
|||
|
||||
@register_datagrabber
|
||||
class JuselessDataladAOMICID1000VBM(PatternDataladDataGrabber):
|
||||
"""Juseless AOMICID1000 VBM DataGrabber class.
|
||||
"""Juseless AOMICID1000 VBM Data Grabber class.
|
||||
|
||||
Implements a DataGrabber to access the AOMICID1000 VBM data in Juseless.
|
||||
Implements a Data Grabber to access the AOMICID1000 VBM data in Juseless.
|
||||
|
||||
Parameters
|
||||
----------
|
||||
|
|
|
|||
|
|
@ -14,9 +14,9 @@ from ....datagrabber import PatternDataladDataGrabber
|
|||
|
||||
@register_datagrabber
|
||||
class JuselessDataladCamCANVBM(PatternDataladDataGrabber):
|
||||
"""Juseless CamCAN VBM DataGrabber class.
|
||||
"""Juseless CamCAN VBM Data Grabber class.
|
||||
|
||||
Implements a DataGrabber to access the CamCAN VBM data in Juseless.
|
||||
Implements a Data Grabber to access the CamCAN VBM data in Juseless.
|
||||
|
||||
Parameters
|
||||
----------
|
||||
|
|
|
|||
|
|
@ -15,9 +15,9 @@ from ....utils import raise_error
|
|||
|
||||
@register_datagrabber
|
||||
class JuselessDataladIXIVBM(PatternDataladDataGrabber):
|
||||
"""Juseless IXI VBM DataGrabber class.
|
||||
"""Juseless IXI VBM Data Grabber class.
|
||||
|
||||
Implements a DataGrabber to access the IXI VBM data in Juseless.
|
||||
Implements a Data Grabber to access the IXI VBM data in Juseless.
|
||||
|
||||
Parameters
|
||||
----------
|
||||
|
|
|
|||
|
|
@ -14,9 +14,9 @@ from ....datagrabber import PatternDataladDataGrabber
|
|||
|
||||
@register_datagrabber
|
||||
class JuselessDataladUKBVBM(PatternDataladDataGrabber):
|
||||
"""Juseless UKB VBM DataGrabber class.
|
||||
"""Juseless UKB VBM Data Grabber class.
|
||||
|
||||
Implements a DataGrabber to access the UKB VBM data in Juseless.
|
||||
Implements a Data Grabber to access the UKB VBM data in Juseless.
|
||||
|
||||
Parameters
|
||||
----------
|
||||
|
|
|
|||
|
|
@ -21,7 +21,7 @@ from .base import BaseDataGrabber
|
|||
class DataladDataGrabber(BaseDataGrabber):
|
||||
"""Abstract base class for data fetching via Datalad.
|
||||
|
||||
Defines a DataGrabber that gets data from a datalad sibling.
|
||||
Defines a Data Grabber that gets data from a datalad sibling.
|
||||
|
||||
Parameters
|
||||
----------
|
||||
|
|
|
|||
|
|
@ -12,9 +12,9 @@ from .base import BaseDataGrabber
|
|||
|
||||
|
||||
class MultipleDataGrabber(BaseDataGrabber):
|
||||
"""Datagrabber class for data fetching from multiple sources.
|
||||
"""Data Grabber class for data fetching from multiple sources.
|
||||
|
||||
Defines a DataGrabber which can be used to fetch data from multiple
|
||||
Defines a Data Grabber which can be used to fetch data from multiple
|
||||
datagrabbers.
|
||||
|
||||
Parameters
|
||||
|
|
|
|||
|
|
@ -19,9 +19,9 @@ from .utils import validate_patterns, validate_replacements
|
|||
|
||||
@register_datagrabber
|
||||
class PatternDataGrabber(BaseDataGrabber):
|
||||
"""Concrete implementation for data fetching using patterns.
|
||||
"""Concrete implementation for data grabbing using patterns.
|
||||
|
||||
Implements a DataGrabber that understands patterns to grab data.
|
||||
Implements a Data Grabber that understands patterns to grab data.
|
||||
|
||||
Parameters
|
||||
----------
|
||||
|
|
|
|||
|
|
@ -19,7 +19,7 @@ def test_validate_types() -> None:
|
|||
with pytest.raises(TypeError, match="must be a list of strings"):
|
||||
validate_types([1]) # type: ignore
|
||||
|
||||
validate_types(["T1w", "bold"])
|
||||
validate_types(["T1w", "BOLD"])
|
||||
|
||||
|
||||
def test_validate_replacements() -> None:
|
||||
|
|
@ -31,7 +31,7 @@ def test_validate_replacements() -> None:
|
|||
|
||||
patterns = {
|
||||
"T1w": "{subject}/anat/{subject}_T1w.nii.gz",
|
||||
"bold": "{subject}/func/{subject}_task-rest_bold.nii.gz",
|
||||
"BOLD": "{subject}/func/{subject}_task-rest_bold.nii.gz",
|
||||
}
|
||||
|
||||
with pytest.raises(TypeError, match="must be a list of strings"):
|
||||
|
|
@ -42,7 +42,7 @@ def test_validate_replacements() -> None:
|
|||
|
||||
wrong_patterns = {
|
||||
"T1w": "{subject}/anat/_T1w.nii.gz",
|
||||
"bold": "{session}/func/_task-rest_bold.nii.gz",
|
||||
"BOLD": "{session}/func/_task-rest_bold.nii.gz",
|
||||
}
|
||||
|
||||
with pytest.raises(ValueError, match="At least one pattern"):
|
||||
|
|
@ -53,7 +53,7 @@ def test_validate_replacements() -> None:
|
|||
|
||||
def test_validate_patterns() -> None:
|
||||
"""Test validation of patterns."""
|
||||
types = ["T1w", "bold"]
|
||||
types = ["T1w", "BOLD"]
|
||||
with pytest.raises(TypeError, match="must be a dict"):
|
||||
validate_patterns(types, "wrong") # type: ignore
|
||||
|
||||
|
|
@ -74,12 +74,12 @@ def test_validate_patterns() -> None:
|
|||
|
||||
patterns = {
|
||||
"T1w": "{subject}/anat/{subject}_T1w.nii.gz",
|
||||
"bold": "{subject}/func/{subject}_task-rest_bold.nii.gz",
|
||||
"BOLD": "{subject}/func/{subject}_task-rest_bold.nii.gz",
|
||||
}
|
||||
|
||||
wrongpatterns = {
|
||||
"T1w": "{subject}/anat/{subject}*.nii",
|
||||
"bold": "{subject}/func/{subject}_task-rest_bold.nii.gz",
|
||||
"BOLD": "{subject}/func/{subject}_task-rest_bold.nii.gz",
|
||||
}
|
||||
|
||||
with pytest.raises(ValueError, match="following a replacement"):
|
||||
|
|
|
|||
|
|
@ -28,7 +28,7 @@ def test_multiple() -> None:
|
|||
"T1w": "{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz",
|
||||
}
|
||||
pattern2 = {
|
||||
"bold": "{subject}/{session}/func/"
|
||||
"BOLD": "{subject}/{session}/func/"
|
||||
"{subject}_{session}_task-rest_bold.nii.gz",
|
||||
}
|
||||
dg1 = PatternDataladDataGrabber(
|
||||
|
|
@ -42,7 +42,7 @@ def test_multiple() -> None:
|
|||
dg2 = PatternDataladDataGrabber(
|
||||
rootdir=rootdir,
|
||||
uri=repo_uri,
|
||||
types=["bold"],
|
||||
types=["BOLD"],
|
||||
patterns=pattern2,
|
||||
replacements=replacements,
|
||||
)
|
||||
|
|
@ -51,7 +51,7 @@ def test_multiple() -> None:
|
|||
|
||||
types = dg.get_types()
|
||||
assert "T1w" in types
|
||||
assert "bold" in types
|
||||
assert "BOLD" in types
|
||||
|
||||
expected_subs = [
|
||||
(f"sub-{i:02d}", f"ses-{j:02d}")
|
||||
|
|
@ -65,7 +65,7 @@ def test_multiple() -> None:
|
|||
|
||||
data = dg[("sub-01", "ses-01")]
|
||||
assert "T1w" in data
|
||||
assert "bold" in data
|
||||
assert "BOLD" in data
|
||||
|
||||
meta = dg.get_meta()
|
||||
assert "class" in meta
|
||||
|
|
@ -86,7 +86,7 @@ def test_multiple_no_intersection() -> None:
|
|||
"T1w": "{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz",
|
||||
}
|
||||
pattern2 = {
|
||||
"bold": "{subject}/{session}/func/"
|
||||
"BOLD": "{subject}/{session}/func/"
|
||||
"{subject}_{session}_task-rest_bold.nii.gz",
|
||||
}
|
||||
dg1 = PatternDataladDataGrabber(
|
||||
|
|
@ -100,7 +100,7 @@ def test_multiple_no_intersection() -> None:
|
|||
dg2 = PatternDataladDataGrabber(
|
||||
rootdir=rootdir,
|
||||
uri=repo_uri2,
|
||||
types=["bold"],
|
||||
types=["BOLD"],
|
||||
patterns=pattern2,
|
||||
replacements=replacements,
|
||||
)
|
||||
|
|
|
|||
|
|
@ -46,11 +46,11 @@ def test_bids_PatternDataladDataGrabber(tmp_path: Path) -> None:
|
|||
|
||||
"""
|
||||
# Define types
|
||||
types = ["T1w", "bold"]
|
||||
types = ["T1w", "BOLD"]
|
||||
# Define patterns
|
||||
patterns = {
|
||||
"T1w": "{subject}/anat/{subject}_T1w.nii.gz",
|
||||
"bold": "{subject}/func/{subject}_task-rest_bold.nii.gz",
|
||||
"BOLD": "{subject}/func/{subject}_task-rest_bold.nii.gz",
|
||||
}
|
||||
# Define replacements
|
||||
replacements = ["subject"]
|
||||
|
|
@ -76,8 +76,8 @@ def test_bids_PatternDataladDataGrabber(tmp_path: Path) -> None:
|
|||
assert t_sub["T1w"]["path"] == (
|
||||
dg.datadir / f"{elem}/anat/{elem}_T1w.nii.gz"
|
||||
)
|
||||
assert "path" in t_sub["bold"]
|
||||
assert t_sub["bold"]["path"] == (
|
||||
assert "path" in t_sub["BOLD"]
|
||||
assert t_sub["BOLD"]["path"] == (
|
||||
dg.datadir / f"{elem}/func/{elem}_task-rest_bold.nii.gz"
|
||||
)
|
||||
|
||||
|
|
@ -105,11 +105,11 @@ def test_bids_PatternDataladDataGrabber_datadir(tmp_path: Path) -> None:
|
|||
|
||||
"""
|
||||
# Define types
|
||||
types = ["T1w", "bold"]
|
||||
types = ["T1w", "BOLD"]
|
||||
# Define patterns
|
||||
patterns = {
|
||||
"T1w": "{subject}/anat/{subject}_T1w.nii.gz",
|
||||
"bold": "{subject}/func/{subject}_task-rest_bold.nii.gz",
|
||||
"BOLD": "{subject}/func/{subject}_task-rest_bold.nii.gz",
|
||||
}
|
||||
# Define replacements
|
||||
replacements = ["subject"]
|
||||
|
|
@ -119,7 +119,7 @@ def test_bids_PatternDataladDataGrabber_datadir(tmp_path: Path) -> None:
|
|||
datadir = "dataset" # use string and not absolute path
|
||||
patterns = {
|
||||
"T1w": "example_bids/{subject}/anat/{subject}_T*w.nii.gz",
|
||||
"bold": "example_bids/{subject}/func/{subject}_task-rest_*.nii.gz",
|
||||
"BOLD": "example_bids/{subject}/func/{subject}_task-rest_*.nii.gz",
|
||||
}
|
||||
with PatternDataladDataGrabber(
|
||||
uri=repo_uri,
|
||||
|
|
@ -135,18 +135,18 @@ def test_bids_PatternDataladDataGrabber_datadir(tmp_path: Path) -> None:
|
|||
assert t_sub["T1w"]["path"] == (
|
||||
dg.datadir / f"{elem}/anat/{elem}_T1w.nii.gz"
|
||||
)
|
||||
assert "path" in t_sub["bold"]
|
||||
assert t_sub["bold"]["path"] == (
|
||||
assert "path" in t_sub["BOLD"]
|
||||
assert t_sub["BOLD"]["path"] == (
|
||||
dg.datadir / f"{elem}/func/{elem}_task-rest_bold.nii.gz"
|
||||
)
|
||||
|
||||
|
||||
def test_bids_PatternDataladDataGrabber_session():
|
||||
"""Test a subject and session-based BIDS datalad datagrabber."""
|
||||
types = ["T1w", "bold"]
|
||||
types = ["T1w", "BOLD"]
|
||||
patterns = {
|
||||
"T1w": "{subject}/{session}/anat/{subject}_{session}_T1w.nii.gz",
|
||||
"bold": "{subject}/{session}/func/"
|
||||
"BOLD": "{subject}/{session}/func/"
|
||||
"{subject}_{session}_task-rest_bold.nii.gz",
|
||||
}
|
||||
replacements = ["subject", "session"]
|
||||
|
|
|
|||
|
|
@ -10,6 +10,7 @@ from typing import Dict, List, Optional
|
|||
import nibabel as nib
|
||||
import pandas as pd
|
||||
|
||||
from ..api.decorators import register_datareader
|
||||
from ..pipeline import PipelineStepMixin
|
||||
from ..utils.logging import logger, warn_with_log
|
||||
|
||||
|
|
@ -28,6 +29,7 @@ _readers["CSV"] = {"func": pd.read_csv, "params": None}
|
|||
_readers["TSV"] = {"func": pd.read_csv, "params": {"sep": "\t"}}
|
||||
|
||||
|
||||
@register_datareader
|
||||
class DefaultDataReader(PipelineStepMixin):
|
||||
"""Mixin class for default data reader."""
|
||||
|
||||
|
|
|
|||
|
|
@ -42,7 +42,7 @@ def test_meta() -> None:
|
|||
|
||||
nib_data_path = Path(nib_testing.data_path)
|
||||
t_path = nib_data_path / "example4d.nii.gz"
|
||||
input = {"bold": {"path": t_path}}
|
||||
input = {"BOLD": {"path": t_path}}
|
||||
output = reader.fit_transform(input)
|
||||
assert "meta" in output
|
||||
assert "datareader" in output["meta"]
|
||||
|
|
@ -67,23 +67,23 @@ def test_read_nifti(fname: str) -> None:
|
|||
|
||||
t_path = nib_data_path / fname
|
||||
|
||||
input = {"bold": {"path": t_path}}
|
||||
input = {"BOLD": {"path": t_path}}
|
||||
output = reader.fit_transform(input)
|
||||
|
||||
assert isinstance(output, dict)
|
||||
assert "bold" in output
|
||||
assert isinstance(output["bold"], dict)
|
||||
assert "path" in output["bold"]
|
||||
assert "data" in output["bold"]
|
||||
assert "BOLD" in output
|
||||
assert isinstance(output["BOLD"], dict)
|
||||
assert "path" in output["BOLD"]
|
||||
assert "data" in output["BOLD"]
|
||||
|
||||
read_img = output["bold"]["data"]
|
||||
read_img = output["BOLD"]["data"]
|
||||
|
||||
t_read_img = nib.load(t_path)
|
||||
assert_array_equal(read_img.get_fdata(), t_read_img.get_fdata())
|
||||
|
||||
input = {"bold": {"path": t_path.as_posix()}}
|
||||
input = {"BOLD": {"path": t_path.as_posix()}}
|
||||
output2 = reader.fit_transform(input)
|
||||
assert output["bold"]["path"] == output2["bold"]["path"]
|
||||
assert output["BOLD"]["path"] == output2["BOLD"]["path"]
|
||||
|
||||
|
||||
def test_read_unknown() -> None:
|
||||
|
|
|
|||
|
|
@ -62,7 +62,7 @@ class MarkerCollection:
|
|||
----------
|
||||
|
`Data Grabber`
|
||||
input : dict
|
||||
The input data to fit the pipeline on. Should be the output of
|
||||
indexing the DataGrabber with one element.
|
||||
indexing the Data Grabber with one element.
|
||||
|
||||
Returns
|
||||
-------
|
||||
|
|
|
|||
|
|
@ -134,7 +134,7 @@ def test_marker_collection_storage(tmp_path: Path) -> None:
|
|||
# Test storage
|
||||
dg = OasisVBMTestingDatagrabber()
|
||||
|
||||
uri = tmp_path / "test_marker_collection_storage.db"
|
||||
uri = tmp_path / "test_marker_collection_storage.sqlite"
|
||||
storage = SQLiteFeatureStorage(uri=uri)
|
||||
mc = MarkerCollection(
|
||||
markers=markers,
|
||||
|
|
|
|||
|
|
@ -78,7 +78,7 @@ def test_store(tmp_path: Path) -> None:
|
|||
parcellation_two=parcellation_TWO,
|
||||
correlation_method="spearman",
|
||||
)
|
||||
uri = tmp_path / "test_crossparcellation.db"
|
||||
uri = tmp_path / "test_crossparcellation.sqlite"
|
||||
storage = SQLiteFeatureStorage(uri=uri, upsert="ignore")
|
||||
out = crossparcellation.fit_transform(input_dict, storage=storage)
|
||||
|
||||
|
|
|
|||
|
|
@ -77,6 +77,6 @@ def test_store(tmp_path: Path) -> None:
|
|||
ets_rss_marker = RSSETSMarker(parcellation=PARCELLATION)
|
||||
# Create storage
|
||||
storage = SQLiteFeatureStorage(
|
||||
uri=str((tmp_path / "test.db").absolute()))
|
||||
uri=str((tmp_path / "test.sqlite").absolute()))
|
||||
# Store
|
||||
ets_rss_marker.fit_transform(input=input_dict, storage=storage)
|
||||
|
|
|
|||
|
|
@ -78,7 +78,7 @@ def test_FunctionalConnectivityParcels(tmp_path: Path) -> None:
|
|||
|
||||
all_out = fc.fit_transform({"BOLD": {"data": fmri_img}})
|
||||
|
||||
uri = tmp_path / "test_fc_parcellation.db"
|
||||
uri = tmp_path / "test_fc_parcellation.sqlite"
|
||||
# Single storage, must be the uri
|
||||
storage = SQLiteFeatureStorage(uri=uri, upsert="ignore")
|
||||
meta = {
|
||||
|
|
|
|||
|
|
@ -62,7 +62,7 @@ def test_FunctionalConnectivitySpheres(tmp_path: Path) -> None:
|
|||
# check correct output
|
||||
assert fc.get_output_kind(["BOLD"]) == ["matrix"]
|
||||
|
||||
uri = tmp_path / "test_fc_parcel.db"
|
||||
uri = tmp_path / "test_fc_parcel.sqlite"
|
||||
# Single storage, must be the uri
|
||||
storage = SQLiteFeatureStorage(uri=uri, upsert="ignore")
|
||||
meta = {
|
||||
|
|
|
|||
|
|
@ -116,7 +116,7 @@ def test_SphereAggregation_storage(tmp_path: Path) -> None:
|
|||
oasis_dataset = datasets.fetch_oasis_vbm(n_subjects=1)
|
||||
vbm = oasis_dataset.gray_matter_maps[0]
|
||||
img = nib.load(vbm)
|
||||
uri = tmp_path / "test_sphere_storage_3D.db"
|
||||
uri = tmp_path / "test_sphere_storage_3D.sqlite"
|
||||
|
||||
storage = SQLiteFeatureStorage(uri=uri, upsert="ignore")
|
||||
meta = {
|
||||
|
|
|
|||
|
|
@ -92,7 +92,7 @@ def test_get_engine_single_output(tmp_path: Path) -> None:
|
|||
The path to the test directory.
|
||||
|
||||
"""
|
||||
uri = tmp_path / "test_single_output.db"
|
||||
uri = tmp_path / "test_single_output.sqlite"
|
||||
# Single storage, must be the uri
|
||||
storage = SQLiteFeatureStorage(uri=uri, upsert="ignore")
|
||||
assert storage.single_output is True
|
||||
|
|
@ -110,7 +110,7 @@ def test_get_engine_multi_output(tmp_path: Path) -> None:
|
|||
The path to the test directory.
|
||||
|
||||
"""
|
||||
uri = tmp_path / "test_multi_output.db"
|
||||
uri = tmp_path / "test_multi_output.sqlite"
|
||||
storage = SQLiteFeatureStorage(
|
||||
uri=uri, single_output=False, upsert="ignore"
|
||||
)
|
||||
|
|
@ -130,7 +130,7 @@ def test_get_engine_single_output_creation(tmp_path: Path) -> None:
|
|||
tocreate = tmp_path / "tocreate"
|
||||
# Path does not exist yet
|
||||
assert not tocreate.exists()
|
||||
uri = tocreate.absolute() / "test_single_output.db"
|
||||
uri = tocreate.absolute() / "test_single_output.sqlite"
|
||||
_ = SQLiteFeatureStorage(uri=uri, upsert="ignore")
|
||||
# Path exists now
|
||||
assert tocreate.exists()
|
||||
|
|
@ -145,7 +145,7 @@ def test_upsert_replace(tmp_path: Path) -> None:
|
|||
The path to the test directory.
|
||||
|
||||
"""
|
||||
uri = tmp_path / "test_upsert_replace.db"
|
||||
uri = tmp_path / "test_upsert_replace.sqlite"
|
||||
# Single storage, must be the uri
|
||||
storage = SQLiteFeatureStorage(uri=uri, upsert="ignore")
|
||||
# Metadata to store
|
||||
|
|
@ -179,7 +179,7 @@ def test_upsert_ignore(tmp_path: Path) -> None:
|
|||
The path to the test directory.
|
||||
|
||||
"""
|
||||
uri = tmp_path / "test_upsert_ignore.db"
|
||||
uri = tmp_path / "test_upsert_ignore.sqlite"
|
||||
# Single storage, must be the uri
|
||||
storage = SQLiteFeatureStorage(uri=uri, upsert="ignore")
|
||||
# Metadata to store
|
||||
|
|
@ -217,7 +217,7 @@ def test_upsert_update(tmp_path: Path) -> None:
|
|||
The path to the test directory.
|
||||
|
||||
"""
|
||||
uri = tmp_path / "test_upsert_delete.db"
|
||||
uri = tmp_path / "test_upsert_delete.sqlite"
|
||||
storage = SQLiteFeatureStorage(uri=uri)
|
||||
# Metadata to store
|
||||
meta = {"element": "test", "version": "0.0.1"}
|
||||
|
|
@ -250,7 +250,7 @@ def test_upsert_invalid_option(tmp_path: Path) -> None:
|
|||
The path to the test directory.
|
||||
|
||||
"""
|
||||
uri = tmp_path / "test_upsert_invalid.db"
|
||||
uri = tmp_path / "test_upsert_invalid.sqlite"
|
||||
with pytest.raises(ValueError):
|
||||
SQLiteFeatureStorage(uri=uri, upsert="wrong")
|
||||
|
||||
|
|
@ -265,7 +265,7 @@ def test_store_df_and_read_df(tmp_path: Path) -> None:
|
|||
The path to the test directory.
|
||||
|
||||
"""
|
||||
uri = tmp_path / "test_store_df_and_read_df.db"
|
||||
uri = tmp_path / "test_store_df_and_read_df.sqlite"
|
||||
storage = SQLiteFeatureStorage(uri=uri, upsert="ignore")
|
||||
# Metadata to store
|
||||
meta = {
|
||||
|
|
@ -326,7 +326,7 @@ def test_store_metadata(tmp_path: Path) -> None:
|
|||
The path to the test directory.
|
||||
|
||||
"""
|
||||
uri = tmp_path / "test_metadata_store.db"
|
||||
uri = tmp_path / "test_metadata_store.sqlite"
|
||||
# Single storage, must be the uri
|
||||
storage = SQLiteFeatureStorage(uri=uri, upsert="ignore")
|
||||
# Metadata to store
|
||||
|
|
@ -345,7 +345,7 @@ def test_store_table(tmp_path: Path) -> None:
|
|||
The path to the test directory.
|
||||
|
||||
"""
|
||||
uri = tmp_path / "test_store_table.db"
|
||||
uri = tmp_path / "test_store_table.sqlite"
|
||||
storage = SQLiteFeatureStorage(uri=uri)
|
||||
# Metadata to store
|
||||
meta = {"element": "test", "version": "0.0.1", "marker": {"name": "fc"}}
|
||||
|
|
@ -404,7 +404,7 @@ def test_store_matrix(tmp_path: Path) -> None:
|
|||
The path to the test directory.
|
||||
|
||||
"""
|
||||
uri = tmp_path / "test_store_table.db"
|
||||
uri = tmp_path / "test_store_table.sqlite"
|
||||
storage = SQLiteFeatureStorage(uri=uri)
|
||||
# Metadata to store
|
||||
meta = {"element": "test", "version": "0.0.1", "marker": {"name": "fc"}}
|
||||
|
|
@ -435,7 +435,7 @@ def test_store_matrix(tmp_path: Path) -> None:
|
|||
assert_array_equal(read_df.values[0], data.flatten())
|
||||
assert list(read_df.columns) == stored_names
|
||||
# Store without row and column names
|
||||
uri = tmp_path / "test_store_table_nonames.db"
|
||||
uri = tmp_path / "test_store_table_nonames.sqlite"
|
||||
storage = SQLiteFeatureStorage(uri=uri)
|
||||
storage.store_matrix(data=data, meta=meta)
|
||||
stored_names = [
|
||||
|
|
@ -467,7 +467,7 @@ def test_store_matrix(tmp_path: Path) -> None:
|
|||
data = np.array([[1, 2, 3], [11, 22, 33], [111, 222, 333]])
|
||||
row_names = ["row1", "row2", "row3"]
|
||||
col_names = ["col1", "col2", "col3"]
|
||||
uri = tmp_path / "test_store_table_triu.db"
|
||||
uri = tmp_path / "test_store_table_triu.sqlite"
|
||||
storage = SQLiteFeatureStorage(uri=uri)
|
||||
storage.store_matrix(
|
||||
data=data,
|
||||
|
|
@ -496,7 +496,7 @@ def test_store_matrix(tmp_path: Path) -> None:
|
|||
)
|
||||
|
||||
# Store upper triangular matrix without diagonal
|
||||
uri = tmp_path / "test_store_table_triu_nodiagonal.db"
|
||||
uri = tmp_path / "test_store_table_triu_nodiagonal.sqlite"
|
||||
storage = SQLiteFeatureStorage(uri=uri)
|
||||
storage.store_matrix(
|
||||
data=data,
|
||||
|
|
@ -526,7 +526,7 @@ def test_store_matrix(tmp_path: Path) -> None:
|
|||
data = np.array([[1, 2, 3], [11, 22, 33], [111, 222, 333]])
|
||||
row_names = ["row1", "row2", "row3"]
|
||||
col_names = ["col1", "col2", "col3"]
|
||||
uri = tmp_path / "test_store_table_tril.db"
|
||||
uri = tmp_path / "test_store_table_tril.sqlite"
|
||||
storage = SQLiteFeatureStorage(uri=uri)
|
||||
storage.store_matrix(
|
||||
data=data,
|
||||
|
|
@ -555,7 +555,7 @@ def test_store_matrix(tmp_path: Path) -> None:
|
|||
)
|
||||
|
||||
# Store lower triangular matrix without diagonal
|
||||
uri = tmp_path / "test_store_table_tril_nodiagonal.db"
|
||||
uri = tmp_path / "test_store_table_tril_nodiagonal.sqlite"
|
||||
storage = SQLiteFeatureStorage(uri=uri)
|
||||
storage.store_matrix(
|
||||
data,
|
||||
|
|
@ -592,7 +592,7 @@ def test_store_multiple_output(tmp_path: Path):
|
|||
The path to the test directory.
|
||||
|
||||
"""
|
||||
uri = tmp_path / "test_store_multiple_output.db"
|
||||
uri = tmp_path / "test_store_multiple_output.sqlite"
|
||||
storage = SQLiteFeatureStorage(uri=uri, single_output=False)
|
||||
# Metadata to store
|
||||
meta1 = {
|
||||
|
|
@ -691,7 +691,7 @@ def test_collect(tmp_path: Path) -> None:
|
|||
The path to the test directory.
|
||||
|
||||
"""
|
||||
uri = tmp_path / "test_collect.db"
|
||||
uri = tmp_path / "test_collect.sqlite"
|
||||
storage = SQLiteFeatureStorage(uri=uri, single_output=False)
|
||||
# Metadata for storage
|
||||
meta1 = {
|
||||
|
|
|
|||
|
|
@ -15,7 +15,7 @@ from ..datagrabber.base import BaseDataGrabber
|
|||
|
||||
|
||||
class OasisVBMTestingDatagrabber(BaseDataGrabber):
|
||||
"""DataGrabber for Oasis VBM testing data."""
|
||||
"""Data Grabber for Oasis VBM testing data."""
|
||||
|
||||
def __init__(self) -> None:
|
||||
# Create temporary directory
|
||||
|
|
@ -79,7 +79,7 @@ class OasisVBMTestingDatagrabber(BaseDataGrabber):
|
|||
|
||||
|
||||
class SPMAuditoryTestingDatagrabber(BaseDataGrabber):
|
||||
"""DataGrabber for SPM Auditory dataset.
|
||||
"""Data Grabber for SPM Auditory dataset.
|
||||
|
||||
Wrapper for :func:`nilearn.datasets.fetch_spm_auditory`.
|
||||
|
||||
|
|
@ -143,7 +143,7 @@ class SPMAuditoryTestingDatagrabber(BaseDataGrabber):
|
|||
|
||||
|
||||
class PartlyCloudyTestingDataGrabber(BaseDataGrabber):
|
||||
"""DataGrabber for Partly Cloudy dataset.
|
||||
"""Data Grabber for Partly Cloudy dataset.
|
||||
|
||||
Wrapper for :func:`nilearn.datasets.fetch_development_fmri`
|
||||
|
||||
|
|
|
|||
# replaced in the patterns.