[DOC]: Add doc for analysing storage-like objects #359

Merged
synchon merged 2 commits from docs/storage-analysis into main 2024-08-12 11:36:04 +00:00
2 changed files with 171 additions and 0 deletions

View file

@ -0,0 +1 @@
Add documentation on analysing extracted features by `Synchon Mandal`_

View file

@ -113,3 +113,173 @@ The ``collect`` command accepts the following additional arguments:
* ``--help``: Show a help message.
* ``--verbose``: Set the verbosity level. Options are ``warning``, ``info``,
``debug``.
.. _analysing_extracted_features:
Analysing Results
=================
After ``collect``-ing the results into a single file, we can analyse them as we
wish. We would need to do this programmatically so feel free to choose your
Python interpreter of choice or use it via a Python script.
First we load the storage like so:
.. code-block:: python
from junifer.storage import HDF5FeatureStorage
# You need to import and use SQLiteFeatureStorage if you chose that
# for storage while extracting features
storage = HDF5FeatureStorage("<path/to/your/collected/file>")
The best way to start analysing would be to list all the extracted features
like so:
.. code-block:: python
# This would output a dictionary with MD5 checksum of the features as keys
# and metadata of the features as values
storage.list_features()
.. code-block::
{'eb85b61eefba61f13d712d425264697b': {'datagrabber': {'class': 'SPMAuditoryTestingDataGrabber',
'types': ['BOLD', 'T1w']},
'dependencies': {'scikit-learn': '1.3.0', 'nilearn': '0.9.2'},
'datareader': {'class': 'DefaultDataReader'},
'type': 'BOLD',
'marker': {'class': 'FunctionalConnectivityParcels',
'parcellation': 'Schaefer100x7',
'agg_method': 'mean',
'agg_method_params': None,
'cor_method': 'covariance',
'cor_method_params': {'empirical': False},
'masks': None,
'name': 'schaefer_100x7_fc_parcels'},
'_element_keys': ['subject'],
'name': 'BOLD_schaefer_100x7_fc_parcels'}}
Once we have this, we can retrieve a single feature like so:
.. code-block:: python
feature_dict = storage.read("<name-key-from-feature-metadata>")
to get the stored dictionary or,
.. code-block:: python
feature_df = storage.read_df("<name-key-from-feature-metadata>")
to get it as a :class:`pandas.DataFrame`.
If there are features with duplicate ``name`` s, then we would need to use
the MD5 checksum and pass it like so:
.. code-block:: python
feature_df = storage.read_df(feature_md5="<md5-hash-of-feature>")
We can now manipulate the dictionary or the :class:`pandas.DataFrame` as we
wish or use the DataFrame directly with `julearn`_ if desired.
On-the-fly transforms
---------------------
``junifer`` supports performing some computationally cheap operations
(like computing brain connectivity or BrainPrint post-analysis) directly
on the storage objects. To make sure everything works, install ``junifer``
like so:
.. code-block:: bash
pip install junifer[onthefly]
or if installed via ``conda``, everything should be there and no further action
is required.
Computing brain connectivity via `bctpy <https://github.com/aestrivex/bctpy>`_
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
You can compute for example, the degree of each node considering an undirected
graph via `bctpy`_ like so:
.. code-block:: python
from junifer.onthefly import read_transform
transformed_df = read_transform(
storage,
# md5 hash can also be passed via `feature_md5`
feature_name="<name-key-from-feature-metadata>",
# format is `package_function`
transform="bctpy_degrees_und",
)
Post-analysis of BrainPrint results
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
To perform post-analysis on BrainPrint results, start by importing:
.. code-block:: python
from junifer.onthefly import brainprint
Normalising:
~~~~~~~~~~~~
Surface normalisation can be done by:
.. code-block:: python
surface_normalised_df = brainprint.normalize(
storage,
{
"areas": {
"feature_name": "<name-key-from-areas-feature>",
# if md5 hash is passed, the above should be None
"feature_md5": None,
},
"eigenvalues": {
"feature_name": "<name-key-from-eigenvalues-feature>",
"feature_md5": None,
},
},
kind="surface",
)
and volume normalisation can be done by:
.. code-block:: python
volume_normalised_df = brainprint.normalize(
storage,
{
"volumes": {
"feature_name": "<name-key-from-volumes-feature>",
"feature_md5": None,
},
"eigenvalues": {
"feature_name": "<name-key-from-eigenvalues-feature>",
"feature_md5": None,
},
},
kind="volume",
)
Re-weighting:
~~~~~~~~~~~~~
To perform re-weighting, run like so:
.. code-block:: python
reweighted_df = brainprint.reweight(
storage,
# md5 hash can also be passed via `feature_md5`
feature_name="<name-key-from-eigenvalues-feature>",
)