[DOC]: venv example #327

Merged
synchon merged 2 commits from docs/improve-queueing into main 2024-06-19 10:43:09 +00:00
2 changed files with 81 additions and 7 deletions

View file

@ -0,0 +1 @@
Update "Queueing Jobs (HPC, HTC)" section by `Synchon Mandal`_

View file

@ -10,10 +10,10 @@ computational clusters. This is done by adding the ``queue`` section in the
:ref:`codeless` file and executing the ``junifer queue`` command. :ref:`codeless` file and executing the ``junifer queue`` command.
While junifer is meant to support `HTCondor`_, `SLURM`_ and local queueing While junifer is meant to support `HTCondor`_, `SLURM`_ and local queueing
using `GNU Parallel`_, only HTCondor is currently supported. This will be using `GNU Parallel`_, only ``HTCondor`` and ``GNU Parallel`` are currently
implemented in future releases of ``junifer``. If you are in immediate need of supported. This will be implemented in future releases of ``junifer``. If you are
any of these schedulers, please create an issue on the `junifer github`_ in immediate need of any of these schedulers, please create an issue on the
repository. `junifer github`_ repository.
The ``queue`` section of the :ref:`codeless` must start by defining the The ``queue`` section of the :ref:`codeless` must start by defining the
following general parameters: following general parameters:
@ -22,8 +22,8 @@ following general parameters:
folder where the job files will be created, as well as any relevant file. folder where the job files will be created, as well as any relevant file.
Depending on the scheduler, it will also be listed in the queueing system Depending on the scheduler, it will also be listed in the queueing system
with this name. with this name.
* ``kind``: The kind of scheduler to be used. Currently, only ``HTCondor`` is * ``kind``: The kind of scheduler to be used. Currently, only ``HTCondor`` and
supported. ``GNUParallelLocal`` are supported.
Example in YAML: Example in YAML:
@ -48,7 +48,9 @@ jobs are finished.
The following parameters are available for HTCondor: The following parameters are available for HTCondor:
* ``env``: Definition of the Python environment. It must provide two variables: * ``pre_run``: Extra shell commands to run before ``junifer run``.
* ``pre_collect``: Extra shell commands to run before ``junifer collect``.
* ``env``: Definition of the Python environment. It has the following parameters:
* ``kind``: This is the kind of virtual environment to use: * ``kind``: This is the kind of virtual environment to use:
@ -61,6 +63,9 @@ The following parameters are available for HTCondor:
the absolute or relative path to the virtualenv when ``venv`` is used. the absolute or relative path to the virtualenv when ``venv`` is used.
If relative path is used then it should be relative to the YAML. If relative path is used then it should be relative to the YAML.
* ``shell``: This is the shell to use. Only ``bash`` and ``zsh`` are supported
as of now.
* ``mem``: Memory to be used by the job. It must be provided as a string with * ``mem``: Memory to be used by the job. It must be provided as a string with
the units (e.g., ``"2GB"``). the units (e.g., ``"2GB"``).
* ``cpus``: Number of CPUs to be used by the job. It must be provided as an * ``cpus``: Number of CPUs to be used by the job. It must be provided as an
@ -93,10 +98,78 @@ Example in YAML:
env: env:
kind: conda kind: conda
name: junifer name: junifer
shell: zsh
mem: 8G mem: 8G
disk: 2G disk: 2G
collect: "yes" # wrap it in string to avoid boolean collect: "yes" # wrap it in string to avoid boolean
Alternatively, if you use ``venv``, the YAML would look like so:
.. code-block:: yaml
queue:
jobname: TestHTCondorQueue
kind: HTCondor
env:
kind: venv
name: /home/me/junifer-venv # should be at the level above `bin/activate`
shell: zsh
mem: 8G
disk: 2G
collect: "yes" # wrap it in string to avoid boolean
.. _queueing_local:
GNUParallelLocal
----------------
When using GNU Parallel, ``junifer`` will run only in local mode and leverage
maximum computational power of whatever hardware it is run on.
The following parameters are available for GNUParallelLocal:
* ``pre_run``: Extra shell commands to run before ``junifer run``.
* ``pre_collect``: Extra shell commands to run before ``junifer collect``.
* ``env``: Definition of the Python environment. It has the following parameters:
* ``kind``: This is the kind of virtual environment to use:
* ``conda``
* ``venv``
* ``local`` (no virtual environment)
* ``name``: This is the name of the environment to use in case a virtual
environment is used. It should be the name when ``conda`` is used and
the absolute or relative path to the virtualenv when ``venv`` is used.
If relative path is used then it should be relative to the YAML.
* ``shell``: This is the shell to use. Only ``bash`` and ``zsh`` are supported
as of now.
Example in YAML:
.. code-block:: yaml
queue:
jobname: TestGNUParallelLocalQueue
kind: GNUParallelLocal
env:
kind: conda
name: junifer
shell: zsh
Alternatively, if you use ``venv``, the YAML would look like so:
.. code-block:: yaml
queue:
jobname: TestGNUParallelLocalQueue
kind: GNUParallelLocal
env:
kind: venv
name: /home/me/junifer-venv # should be at the level above `bin/activate`
shell: zsh
Once the :ref:`codeless` file is ready, including the ``queue`` section, you can Once the :ref:`codeless` file is ready, including the ``queue`` section, you can
queue the jobs by executing the ``junifer queue`` command. queue the jobs by executing the ``junifer queue`` command.