[ENH]: Allow for local queue using GNU parallel #306
No reviewers
Labels
No labels
CRITICAL
Stale
WIP
bug
concept
coordinate
dataset
dependencies
documentation
duplicate
enhancement
github_actions
good first issue
help wanted
invalid
maintenance
maps
marker
mask
on hold
parcellation
preprocess
question
ready
storage
template-space
triage
wontfix
No milestone
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
juaml/junifer!306
Loading…
Reference in a new issue
No description provided.
Delete branch "feature/local-queue"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Are you requiring a new dataset or marker?
Which feature do you want to include?
So far the
queuesection of the YAML allows to use HTCondor.The feature requested is to include a
localqueue that uses GNU parallel to run, locally, using multicore architectures. This will benefit all users dealing with relatively small datasets on environments without HPC/HTC, like dual-socket and powerful workstations.How do you imagine this integrated in junifer?
By allowing this:
Do you have a sample code that implements this outside of junifer?
No response
Anything else to say?
No response
@ -239,7 +239,7 @@ def queue(if the ``jobdir`` exists and ``overwrite = False``.Why are we changing this? Is this absolutely necessary? If someone updates junifer and then continues with a pre-existing yaml, it will have a weird junifer_jobs directory.
@ -0,0 +1,258 @@"""Define concrete class for generating GNU Parallel (local) assets."""While this syntax is fine for small jobs, it will create a huge command line and does not scale.
A better option is to use the
--arg-fileargument and give a file with the arguments to use.We can directly add this to the command (see my comment above)
Some comments about the
parallelarguments.--bar: ok, shows a bar.--halt: disagree, we should run as much as we can. If one subject fails, continue.--resume: add this so the user sees the full command that can be stopped and continued (with ctrl+C)--resume-failed: add this also, for the same reason. It will not have any effect on the first run. It will have an effect on the subsequent runs.--joblob: perfectSome other to consider:
--output-as-files: to redirect output to a file and keep the stdout clean--delay: prevent having N-jobs doing the same IO operation at the beginning, which will definitely create a bottleneck and maybe a failure.There's no pre-run anymore?
What if the user wants to set environmental variables even in a gnu parallel environment? Keep in mind that GNU parallel can also be used to run over SSH.
I liked the idea of having some sort of self-contained running scripts.
Imagine this scenario:
localand starts running gnu-parallel on the HCP.junifer_jobsdir and runs theparallelcommand again.This should work.
@ -239,7 +239,7 @@ def queue(if the ``jobdir`` exists and ``overwrite = False``.If an user tries out different queueing system but with the same job name (which will most likely be the case), then this is the simplest way imo to handle the assets for each system.
@ -0,0 +1,258 @@"""Define concrete class for generating GNU Parallel (local) assets."""Agreed. Any reason for preferring
--output-as-filesover--results? The latter stores the stdout, stderr and seq value in a structured way.I've decided to revert it in a better way.
Fair argument with the env vars. We'll not let users run GNU parallel over SSH yet as it's not as simple as running locally. You need to tweak file copying, output handling, cleanup and compression, a lot more moving parts. But maybe at some point in the future, we can support it.
@ -239,7 +239,7 @@ def queue(if the ``jobdir`` exists and ``overwrite = False``.From what I know, the junifer queue command does not allow to use an existing
jobdir.@ -0,0 +1,258 @@"""Define concrete class for generating GNU Parallel (local) assets."""Did not know of the existence of
--results. Pick what you think it's the best option.@ -239,7 +239,7 @@ def queue(if the ``jobdir`` exists and ``overwrite = False``.I don't quite understand how that can be a problem as the
jobdirincludes the queue kind dir.For the SSH, not officialy, but someone might "tweak" the parallel command. In any case, even if SSH is not a supported use case, just closing the terminal and opening it again should continue were it was.
That's for the "hacker" spirited-soul :)
Of course.
@ -239,7 +239,7 @@ def queue(if the ``jobdir`` exists and ``overwrite = False``.Updating junifer creates an inconsistent
junifer_jobsdirectory.My own personal cluster (my phd lab during the nights).
@ -239,7 +239,7 @@ def queue(if the ``jobdir`` exists and ``overwrite = False``.We can add it to the release notes and / or issue a note in the documentation. Do you have an alternative for this?
Can relate to the picture, don't have an anecdotal picture though :D
Codecov Report
Attention: Patch coverage is
95.60440%with4 linesin your changes are missing coverage. Please review.Additional details and impacted files
100.00% <ø> (?)88.15% <95.60%> (+0.16%)Flags with carried forward coverage won't be shown. Click here to find out more.
100.00% <100.00%> (ø)99.10% <ø> (ø)93.75% <75.00%> (+0.09%)96.51% <96.51%> (ø)... and 2 files with indirect coverage changes
@ -239,7 +239,7 @@ def queue(if the ``jobdir`` exists and ``overwrite = False``.My take is that we don't need the
kind.lower()part, as you won't be able to runjunifer queueif thejobnamedirectory exists. It will be, technically, a different job with a different name.I don't see any usecase to complicate ourselves like this.
@ -239,7 +239,7 @@ def queue(if the ``jobdir`` exists and ``overwrite = False``.So if a user wants to run a YAML using GNU Parallel (for testing) on juseless for example and then switch to HTCondor (for complete run), you say that either the user changes the job name, overwrites it or creates a new YAML?
@ -239,7 +239,7 @@ def queue(if the ``jobdir`` exists and ``overwrite = False``.The user needs to change the YAML anyways (the
queuesection).Usually the testing is done with
junifer run. You don't need to test the queueing mechanism unless you are doing some advanced things. If you want to test the non-interactive run, even the GNU parallel tool is not the way. You should just runjunifer queuewith--elementto queue only one element.@ -239,7 +239,7 @@ def queue(if the ``jobdir`` exists and ``overwrite = False``.Fair enough.
@fraimondo Do we merge after CI completion?