[ENH]: Introduce junifer.api.generate_yaml #498

Open
synchon wants to merge 21 commits from feat/generate-yaml-api into main
synchon commented 2026-05-29 16:04:42 +00:00 (Migrated from github.com)
  • description of feature/fix
  • tests added/passed
  • add an entry for the latest changes

This PR adds generate_yaml under api to generate feature YAML from metadata. Its primary use-case is in julio's feature addition to registry.

* [x] description of feature/fix * [ ] tests added/passed * [x] add an entry for the latest changes This PR adds `generate_yaml` under `api` to generate feature YAML from metadata. Its primary use-case is in `julio`'s feature addition to registry.
codecov[bot] commented 2026-05-29 16:05:38 +00:00 (Migrated from github.com)

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 100.00%. Comparing base (a8e9f31) to head (ca6cc8a).
⚠️ Report is 8 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff            @@
##              main      #498   +/-   ##
=========================================
  Coverage   100.00%   100.00%           
=========================================
  Files            1         1           
  Lines            1         1           
=========================================
  Hits             1         1           
Flag Coverage Δ
docs 100.00% <ø> (ø)

Flags with carried forward coverage won't be shown. Click here to find out more.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
## [Codecov](https://app.codecov.io/gh/juaml/junifer/pull/498?dropdown=coverage&src=pr&el=h1&utm_medium=referral&utm_source=github&utm_content=comment&utm_campaign=pr+comments&utm_term=juaml) Report :white_check_mark: All modified and coverable lines are covered by tests. :white_check_mark: Project coverage is 100.00%. Comparing base ([`a8e9f31`](https://app.codecov.io/gh/juaml/junifer/commit/a8e9f31c59015f2671e30db6f5bd6b1fce04d17e?dropdown=coverage&el=desc&utm_medium=referral&utm_source=github&utm_content=comment&utm_campaign=pr+comments&utm_term=juaml)) to head ([`ca6cc8a`](https://app.codecov.io/gh/juaml/junifer/commit/ca6cc8a71347e045067768418c240502a8c4a70b?dropdown=coverage&el=desc&utm_medium=referral&utm_source=github&utm_content=comment&utm_campaign=pr+comments&utm_term=juaml)). :warning: Report is 8 commits behind head on main. <details><summary>Additional details and impacted files</summary> [![Impacted file tree graph](https://app.codecov.io/gh/juaml/junifer/pull/498/graphs/tree.svg?width=650&height=150&src=pr&token=5H21JuZXMw&utm_medium=referral&utm_source=github&utm_content=comment&utm_campaign=pr+comments&utm_term=juaml)](https://app.codecov.io/gh/juaml/junifer/pull/498?src=pr&el=tree&utm_medium=referral&utm_source=github&utm_content=comment&utm_campaign=pr+comments&utm_term=juaml) ```diff @@ Coverage Diff @@ ## main #498 +/- ## ========================================= Coverage 100.00% 100.00% ========================================= Files 1 1 Lines 1 1 ========================================= Hits 1 1 ``` | [Flag](https://app.codecov.io/gh/juaml/junifer/pull/498/flags?src=pr&el=flags&utm_medium=referral&utm_source=github&utm_content=comment&utm_campaign=pr+comments&utm_term=juaml) | Coverage Δ | | |---|---|---| | [docs](https://app.codecov.io/gh/juaml/junifer/pull/498/flags?src=pr&el=flag&utm_medium=referral&utm_source=github&utm_content=comment&utm_campaign=pr+comments&utm_term=juaml) | `100.00% <ø> (ø)` | | Flags with carried forward coverage won't be shown. [Click here](https://docs.codecov.io/docs/carryforward-flags?utm_medium=referral&utm_source=github&utm_content=comment&utm_campaign=pr+comments&utm_term=juaml#carryforward-flags-in-the-pull-request-comment) to find out more. </details> <details><summary> :rocket: New features to boost your workflow: </summary> - :snowflake: [Test Analytics](https://docs.codecov.com/docs/test-analytics): Detect flaky tests, report on failures, and find test suite problems. </details>
fraimondo commented 2026-06-03 14:20:23 +00:00 (Migrated from github.com)

I have the impression that this way (dump_exclude) adds a lot of maintenance work: if something changes in the superclass (new internal variable), we need to go and update every subclass, including non-junifer ones.

Can't we just look at what fields are defined in the class (not the superclass) and just pass those ones to the constructor?

I have the impression that this way (`dump_exclude`) adds a lot of maintenance work: if something changes in the superclass (new internal variable), we need to go and update every subclass, including non-junifer ones. Can't we just look at what _fields_ are defined in the class (not the superclass) and just pass those ones to the constructor?
synchon commented 2026-06-08 13:08:26 +00:00 (Migrated from github.com)

I have the impression that this way (dump_exclude) adds a lot of maintenance work: if something changes in the superclass (new internal variable), we need to go and update every subclass, including non-junifer ones.

That's a fair argument and I see your point.

Can't we just look at what fields are defined in the class (not the superclass) and just pass those ones to the constructor?

Not with how I understand the thing works. A model's fields consist of its own fields and superclass' fields (if it has one). So apart from defining what to exclude (or include), I don't see other way. I'll push some updates to make it better.

> I have the impression that this way (`dump_exclude`) adds a lot of maintenance work: if something changes in the superclass (new internal variable), we need to go and update every subclass, including non-junifer ones. That's a fair argument and I see your point. > Can't we just look at what _fields_ are defined in the class (not the superclass) and just pass those ones to the constructor? Not with how I understand the thing works. A model's fields consist of its own fields and superclass' fields (if it has one). So apart from defining what to exclude (or include), I don't see other way. I'll push some updates to make it better.
synchon commented 2026-07-21 12:51:13 +00:00 (Migrated from github.com)

@fraimondo I've updated the datagrabber dumping logic as discussed. Kindly review #499 before this.

@fraimondo I've updated the datagrabber dumping logic as discussed. Kindly review #499 before this.
fraimondo (Migrated from github.com) reviewed 2026-07-21 14:16:05 +00:00
@ -143,6 +143,11 @@ class JuselessUCLA(PatternDataGrabber):
replacements: list[str] = ["subject", "task"] # noqa: RUF012
fraimondo (Migrated from github.com) commented 2026-07-21 14:06:18 +00:00

This should only be types and tasks. The rest is hard-coded in the parameters.

This should only be `types` and `tasks`. The rest is hard-coded in the parameters.
@ -179,0 +179,4 @@
@classmethod
def dump_fields(cls) -> list[str]:
"""Fields to include when dumping model."""
return ["types", "uri", "rootdir"]
fraimondo (Migrated from github.com) commented 2026-07-21 14:14:32 +00:00

this is where it becomes a bit tricky.

If I used a datadir (not temp) and the dataset is dirty, the YAML should account for that? or not?

Maybe we should include some comments in the generated YAMLs indicating stuff like this:

eg. if we have a "dirty" dataset, then add a comment that while the yaml will reproduce the results, the original dataset was "dirty" and so there is no guarantee that the same results will be obtained as there is no strict data provenance.

this is where it becomes a bit tricky. If I used a datadir (not temp) and the dataset is dirty, the YAML should account for that? or not? Maybe we should include some comments in the generated YAMLs indicating stuff like this: eg. if we have a "dirty" dataset, then add a comment that while the yaml will reproduce the results, the original dataset was "dirty" and so there is no guarantee that the same results will be obtained as there is no strict data provenance.
@ -64,0 +64,4 @@
@classmethod
def dump_fields(cls) -> list[str]:
"""Fields to include when dumping model."""
return ["types", "tasks", "phase_encodings", "ica_fix"]
fraimondo (Migrated from github.com) commented 2026-07-21 14:10:13 +00:00

This can also be super(HCP1200) - datadir. Thus any change in the super will also be accounted for here.

This can also be `super(HCP1200) - datadir`. Thus any change in the super will also be accounted for here.
@ -64,0 +65,4 @@
@classmethod
def dump_fields(cls) -> list[str]:
"""Fields to include when dumping model."""
return [
fraimondo (Migrated from github.com) commented 2026-07-21 14:15:21 +00:00

I thinks this should be both "super" fields.

I thinks this should be both "super" fields.
synchon (Migrated from github.com) reviewed 2026-07-21 15:18:54 +00:00
synchon (Migrated from github.com) reviewed 2026-07-21 15:21:03 +00:00
@ -64,0 +64,4 @@
@classmethod
def dump_fields(cls) -> list[str]:
"""Fields to include when dumping model."""
return ["types", "tasks", "phase_encodings", "ica_fix"]
synchon (Migrated from github.com) commented 2026-07-21 15:21:03 +00:00

In that case, the MRO for this class will get dump_fields from DataladDataGrabber which will be incorrect.

In that case, the MRO for this class will get `dump_fields` from `DataladDataGrabber` which will be incorrect.
synchon (Migrated from github.com) reviewed 2026-07-21 15:22:31 +00:00
@ -64,0 +65,4 @@
@classmethod
def dump_fields(cls) -> list[str]:
"""Fields to include when dumping model."""
return [
synchon (Migrated from github.com) commented 2026-07-21 15:22:31 +00:00

It is "both" of them.

It is "both" of them.
synchon (Migrated from github.com) reviewed 2026-07-21 15:25:15 +00:00
@ -179,0 +179,4 @@
@classmethod
def dump_fields(cls) -> list[str]:
"""Fields to include when dumping model."""
return ["types", "uri", "rootdir"]
synchon (Migrated from github.com) commented 2026-07-21 15:25:15 +00:00

We can add a general comment. Making it conditional would be quite tricky.

We can add a general comment. Making it conditional would be quite tricky.
fraimondo (Migrated from github.com) reviewed 2026-07-22 09:05:51 +00:00
@ -179,0 +179,4 @@
@classmethod
def dump_fields(cls) -> list[str]:
"""Fields to include when dumping model."""
return ["types", "uri", "rootdir"]
fraimondo (Migrated from github.com) commented 2026-07-22 09:05:51 +00:00

I would like that the generated YAML is commented.

I would like that the generated YAML is commented.
fraimondo (Migrated from github.com) reviewed 2026-07-22 09:06:40 +00:00
@ -64,0 +65,4 @@
@classmethod
def dump_fields(cls) -> list[str]:
"""Fields to include when dumping model."""
return [
fraimondo (Migrated from github.com) commented 2026-07-22 09:06:40 +00:00

I meant instead of manually placing the fields, to all super. But I understand that might be bothersome.

I meant instead of manually placing the fields, to all super. But I understand that might be bothersome.
fraimondo commented 2026-07-22 09:07:34 +00:00 (Migrated from github.com)

Do we know how this function works when we have external imports (with statements in the YAML)?

Do we know how this function works when we have external imports (`with` statements in the YAML)?
synchon (Migrated from github.com) reviewed 2026-07-22 09:07:56 +00:00
@ -64,0 +65,4 @@
@classmethod
def dump_fields(cls) -> list[str]:
"""Fields to include when dumping model."""
return [
synchon (Migrated from github.com) commented 2026-07-22 09:07:56 +00:00

The MRO would stop at the first dump_fields which would give a partial list.

The MRO would stop at the first `dump_fields` which would give a partial list.
synchon commented 2026-07-22 09:25:25 +00:00 (Migrated from github.com)

Do we know how this function works when we have external imports (with statements in the YAML)?

It copies exactly what is passed. For julio, we copy exactly what is stored in h5.

> Do we know how this function works when we have external imports (`with` statements in the YAML)? It copies exactly what is passed. For julio, we copy exactly what is stored in h5.
fraimondo commented 2026-07-22 09:56:20 +00:00 (Migrated from github.com)

Do we know how this function works when we have external imports (with statements in the YAML)?

It copies exactly what is passed. For julio, we copy exactly what is stored in h5.

In order to create the yaml, we parse the metadata and instantiate the object. I'm not sure this will be possible if the definition of a datagrabber/marker is in an external package which is part of the with statement.

> > Do we know how this function works when we have external imports (`with` statements in the YAML)? > > It copies exactly what is passed. For julio, we copy exactly what is stored in h5. In order to create the yaml, we parse the metadata and instantiate the object. I'm not sure this will be possible if the definition of a datagrabber/marker is in an external package which is part of the `with` statement.
synchon commented 2026-07-22 11:57:46 +00:00 (Migrated from github.com)

Do we know how this function works when we have external imports (with statements in the YAML)?

It copies exactly what is passed. For julio, we copy exactly what is stored in h5.

In order to create the yaml, we parse the metadata and instantiate the object. I'm not sure this will be possible if the definition of a datagrabber/marker is in an external package which is part of the with statement.

Let me lay it down w.r.t. julio:

  • junifer YAML gets parsed by junifer.api.parse_yaml which loads the external modules and registers componenets as needed.
  • If with is present in the junifer YAML, julio adds it directly to the generated YAML.
  • As the external components are already registered, junifer.api.generate_yaml can load them from component registry and dump as needed.

For using junifer.api.generate_yaml without julio, one would need to already register the components beforehand. Now, as it's part of junifer.api, we can presume that the user will do it like so.

Does that solve your concern?

> > > Do we know how this function works when we have external imports (`with` statements in the YAML)? > > > > > > It copies exactly what is passed. For julio, we copy exactly what is stored in h5. > > In order to create the yaml, we parse the metadata and instantiate the object. I'm not sure this will be possible if the definition of a datagrabber/marker is in an external package which is part of the `with` statement. Let me lay it down w.r.t. julio: - junifer YAML gets parsed by `junifer.api.parse_yaml` which loads the external modules and registers componenets as needed. - If `with` is present in the junifer YAML, julio adds it directly to the generated YAML. - As the external components are already registered, `junifer.api.generate_yaml` can load them from component registry and dump as needed. For using `junifer.api.generate_yaml` without julio, one would need to already register the components beforehand. Now, as it's part of `junifer.api`, we can presume that the user will do it like so. Does that solve your concern?
fraimondo commented 2026-07-22 13:22:46 +00:00 (Migrated from github.com)

Do we know how this function works when we have external imports (with statements in the YAML)?

It copies exactly what is passed. For julio, we copy exactly what is stored in h5.

In order to create the yaml, we parse the metadata and instantiate the object. I'm not sure this will be possible if the definition of a datagrabber/marker is in an external package which is part of the with statement.

Let me lay it down w.r.t. julio:

* junifer YAML gets parsed by `junifer.api.parse_yaml` which loads the external modules and registers componenets as needed.

* If `with` is present in the junifer YAML, julio adds it directly to the generated YAML.

* As the external components are already registered, `junifer.api.generate_yaml` can load them from component registry and dump as needed.

For using junifer.api.generate_yaml without julio, one would need to already register the components beforehand. Now, as it's part of junifer.api, we can presume that the user will do it like so.

Does that solve your concern?

junifer.api.generate_yaml has one parameter that is a meta dict. This meta dict can be extracted from and HDF5 file, without parsing any yaml. We need to be able to support generating a YAML from any meta dict, no only after parsing the respective yaml.

Even if this means adding all variables int he meta dict and a huge comment stating that "since this is an external datagrabber/marker/etc not part of junifer core, some of this variables might not be needed or should definitely be removed."

> > > > Do we know how this function works when we have external imports (`with` statements in the YAML)? > > > > > > > > > It copies exactly what is passed. For julio, we copy exactly what is stored in h5. > > > > > > In order to create the yaml, we parse the metadata and instantiate the object. I'm not sure this will be possible if the definition of a datagrabber/marker is in an external package which is part of the `with` statement. > > Let me lay it down w.r.t. julio: > > * junifer YAML gets parsed by `junifer.api.parse_yaml` which loads the external modules and registers componenets as needed. > > * If `with` is present in the junifer YAML, julio adds it directly to the generated YAML. > > * As the external components are already registered, `junifer.api.generate_yaml` can load them from component registry and dump as needed. > > > For using `junifer.api.generate_yaml` without julio, one would need to already register the components beforehand. Now, as it's part of `junifer.api`, we can presume that the user will do it like so. > > Does that solve your concern? `junifer.api.generate_yaml` has one parameter that is a `meta` dict. This meta dict can be extracted from and HDF5 file, without parsing any yaml. We need to be able to support generating a YAML from any meta dict, no only after parsing the respective yaml. Even if this means adding all variables int he meta dict and a huge comment stating that "since this is an external datagrabber/marker/etc not part of junifer core, some of this variables might not be needed or should definitely be removed."
synchon commented 2026-07-22 14:23:54 +00:00 (Migrated from github.com)

Do we know how this function works when we have external imports (with statements in the YAML)?

It copies exactly what is passed. For julio, we copy exactly what is stored in h5.

In order to create the yaml, we parse the metadata and instantiate the object. I'm not sure this will be possible if the definition of a datagrabber/marker is in an external package which is part of the with statement.

Let me lay it down w.r.t. julio:

* junifer YAML gets parsed by `junifer.api.parse_yaml` which loads the external modules and registers componenets as needed.

* If `with` is present in the junifer YAML, julio adds it directly to the generated YAML.

* As the external components are already registered, `junifer.api.generate_yaml` can load them from component registry and dump as needed.

For using junifer.api.generate_yaml without julio, one would need to already register the components beforehand. Now, as it's part of junifer.api, we can presume that the user will do it like so.
Does that solve your concern?

junifer.api.generate_yaml has one parameter that is a meta dict. This meta dict can be extracted from and HDF5 file, without parsing any yaml. We need to be able to support generating a YAML from any meta dict, no only after parsing the respective yaml.

Even if this means adding all variables int he meta dict and a huge comment stating that "since this is an external datagrabber/marker/etc not part of junifer core, some of this variables might not be needed or should definitely be removed."

I don't follow. What happens if one adds "with" key to the "meta" dictionary passed without loading the external components?

> > > > > Do we know how this function works when we have external imports (`with` statements in the YAML)? > > > > > > > > > > > > It copies exactly what is passed. For julio, we copy exactly what is stored in h5. > > > > > > > > > In order to create the yaml, we parse the metadata and instantiate the object. I'm not sure this will be possible if the definition of a datagrabber/marker is in an external package which is part of the `with` statement. > > > > > > Let me lay it down w.r.t. julio: > > ``` > > * junifer YAML gets parsed by `junifer.api.parse_yaml` which loads the external modules and registers componenets as needed. > > > > * If `with` is present in the junifer YAML, julio adds it directly to the generated YAML. > > > > * As the external components are already registered, `junifer.api.generate_yaml` can load them from component registry and dump as needed. > > ``` > > > > > > > > > > > > > > > > > > > > > > > > For using `junifer.api.generate_yaml` without julio, one would need to already register the components beforehand. Now, as it's part of `junifer.api`, we can presume that the user will do it like so. > > Does that solve your concern? > > `junifer.api.generate_yaml` has one parameter that is a `meta` dict. This meta dict can be extracted from and HDF5 file, without parsing any yaml. We need to be able to support generating a YAML from any meta dict, no only after parsing the respective yaml. > > Even if this means adding all variables int he meta dict and a huge comment stating that "since this is an external datagrabber/marker/etc not part of junifer core, some of this variables might not be needed or should definitely be removed." I don't follow. What happens if one adds `"with"` key to the "meta" dictionary passed without loading the external components?
fraimondo commented 2026-07-22 14:41:38 +00:00 (Migrated from github.com)

Do we know how this function works when we have external imports (with statements in the YAML)?

It copies exactly what is passed. For julio, we copy exactly what is stored in h5.

In order to create the yaml, we parse the metadata and instantiate the object. I'm not sure this will be possible if the definition of a datagrabber/marker is in an external package which is part of the with statement.

Let me lay it down w.r.t. julio:

* junifer YAML gets parsed by `junifer.api.parse_yaml` which loads the external modules and registers componenets as needed.

* If `with` is present in the junifer YAML, julio adds it directly to the generated YAML.

* As the external components are already registered, `junifer.api.generate_yaml` can load them from component registry and dump as needed.

For using junifer.api.generate_yaml without julio, one would need to already register the components beforehand. Now, as it's part of junifer.api, we can presume that the user will do it like so.
Does that solve your concern?

junifer.api.generate_yaml has one parameter that is a meta dict. This meta dict can be extracted from and HDF5 file, without parsing any yaml. We need to be able to support generating a YAML from any meta dict, no only after parsing the respective yaml.
Even if this means adding all variables int he meta dict and a huge comment stating that "since this is an external datagrabber/marker/etc not part of junifer core, some of this variables might not be needed or should definitely be removed."

I don't follow. What happens if one adds "with" key to the "meta" dictionary passed without loading the external components?

Let's asume you open and HDF5 file, you load the meta and then you pass it to generate_yaml. At no point there was a parse_yaml call, so all external modules will not be loaded. This code will fail as it needs to instantiate the elements in the meta to generate the yaml.

This use case should be considered. In the case that the object can't be instantiated, the fields should be extracted from the meta dict and a comment in the yaml should be added.

> > > > > > Do we know how this function works when we have external imports (`with` statements in the YAML)? > > > > > > > > > > > > > > > It copies exactly what is passed. For julio, we copy exactly what is stored in h5. > > > > > > > > > > > > In order to create the yaml, we parse the metadata and instantiate the object. I'm not sure this will be possible if the definition of a datagrabber/marker is in an external package which is part of the `with` statement. > > > > > > > > > Let me lay it down w.r.t. julio: > > > ``` > > > * junifer YAML gets parsed by `junifer.api.parse_yaml` which loads the external modules and registers componenets as needed. > > > > > > * If `with` is present in the junifer YAML, julio adds it directly to the generated YAML. > > > > > > * As the external components are already registered, `junifer.api.generate_yaml` can load them from component registry and dump as needed. > > > ``` > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > For using `junifer.api.generate_yaml` without julio, one would need to already register the components beforehand. Now, as it's part of `junifer.api`, we can presume that the user will do it like so. > > > Does that solve your concern? > > > > > > `junifer.api.generate_yaml` has one parameter that is a `meta` dict. This meta dict can be extracted from and HDF5 file, without parsing any yaml. We need to be able to support generating a YAML from any meta dict, no only after parsing the respective yaml. > > Even if this means adding all variables int he meta dict and a huge comment stating that "since this is an external datagrabber/marker/etc not part of junifer core, some of this variables might not be needed or should definitely be removed." > > I don't follow. What happens if one adds `"with"` key to the "meta" dictionary passed without loading the external components? Let's asume you open and HDF5 file, you load the meta and then you pass it to `generate_yaml`. At no point there was a `parse_yaml` call, so all external modules will not be loaded. This code will fail as it needs to instantiate the elements in the meta to generate the yaml. This use case should be considered. In the case that the object can't be instantiated, the fields should be extracted from the meta dict and a comment in the yaml should be added.
synchon commented 2026-07-23 09:33:37 +00:00 (Migrated from github.com)

Do we know how this function works when we have external imports (with statements in the YAML)?

It copies exactly what is passed. For julio, we copy exactly what is stored in h5.

In order to create the yaml, we parse the metadata and instantiate the object. I'm not sure this will be possible if the definition of a datagrabber/marker is in an external package which is part of the with statement.

Let me lay it down w.r.t. julio:

* junifer YAML gets parsed by `junifer.api.parse_yaml` which loads the external modules and registers componenets as needed.

* If `with` is present in the junifer YAML, julio adds it directly to the generated YAML.

* As the external components are already registered, `junifer.api.generate_yaml` can load them from component registry and dump as needed.

For using junifer.api.generate_yaml without julio, one would need to already register the components beforehand. Now, as it's part of junifer.api, we can presume that the user will do it like so.
Does that solve your concern?

junifer.api.generate_yaml has one parameter that is a meta dict. This meta dict can be extracted from and HDF5 file, without parsing any yaml. We need to be able to support generating a YAML from any meta dict, no only after parsing the respective yaml.
Even if this means adding all variables int he meta dict and a huge comment stating that "since this is an external datagrabber/marker/etc not part of junifer core, some of this variables might not be needed or should definitely be removed."

I don't follow. What happens if one adds "with" key to the "meta" dictionary passed without loading the external components?

Let's asume you open and HDF5 file, you load the meta and then you pass it to generate_yaml. At no point there was a parse_yaml call, so all external modules will not be loaded. This code will fail as it needs to instantiate the elements in the meta to generate the yaml.

This use case should be considered. In the case that the object can't be instantiated, the fields should be extracted from the meta dict and a comment in the yaml should be added.

Do the latest commits address your concern?

> > > > > > > Do we know how this function works when we have external imports (`with` statements in the YAML)? > > > > > > > > > > > > > > > > > > It copies exactly what is passed. For julio, we copy exactly what is stored in h5. > > > > > > > > > > > > > > > In order to create the yaml, we parse the metadata and instantiate the object. I'm not sure this will be possible if the definition of a datagrabber/marker is in an external package which is part of the `with` statement. > > > > > > > > > > > > Let me lay it down w.r.t. julio: > > > > ``` > > > > * junifer YAML gets parsed by `junifer.api.parse_yaml` which loads the external modules and registers componenets as needed. > > > > > > > > * If `with` is present in the junifer YAML, julio adds it directly to the generated YAML. > > > > > > > > * As the external components are already registered, `junifer.api.generate_yaml` can load them from component registry and dump as needed. > > > > ``` > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > For using `junifer.api.generate_yaml` without julio, one would need to already register the components beforehand. Now, as it's part of `junifer.api`, we can presume that the user will do it like so. > > > > Does that solve your concern? > > > > > > > > > `junifer.api.generate_yaml` has one parameter that is a `meta` dict. This meta dict can be extracted from and HDF5 file, without parsing any yaml. We need to be able to support generating a YAML from any meta dict, no only after parsing the respective yaml. > > > Even if this means adding all variables int he meta dict and a huge comment stating that "since this is an external datagrabber/marker/etc not part of junifer core, some of this variables might not be needed or should definitely be removed." > > > > > > I don't follow. What happens if one adds `"with"` key to the "meta" dictionary passed without loading the external components? > > Let's asume you open and HDF5 file, you load the meta and then you pass it to `generate_yaml`. At no point there was a `parse_yaml` call, so all external modules will not be loaded. This code will fail as it needs to instantiate the elements in the meta to generate the yaml. > > This use case should be considered. In the case that the object can't be instantiated, the fields should be extracted from the meta dict and a comment in the yaml should be added. Do the latest commits address your concern?
fraimondo (Migrated from github.com) requested changes 2026-07-23 09:38:38 +00:00
fraimondo (Migrated from github.com) commented 2026-07-23 09:36:20 +00:00

I would explicity check for the datagrabber being in the registry that relying on a ValueError.

Could be that because of versions mismatchs, some parameters are renamed and then we do have errors but because of other reasons.

I would explicity check for the datagrabber being in the registry that relying on a ValueError. Could be that because of versions mismatchs, some parameters are renamed and then we do have errors but because of other reasons.
fraimondo (Migrated from github.com) commented 2026-07-23 09:36:33 +00:00

Same here, explicit check

Same here, explicit check
fraimondo (Migrated from github.com) commented 2026-07-23 09:36:42 +00:00

Same here

Same here
fraimondo (Migrated from github.com) commented 2026-07-23 09:37:20 +00:00

This note should only appear if the dataset was dirty (the meta said so)

This note should only appear if the dataset was dirty (the meta said so)
synchon (Migrated from github.com) reviewed 2026-07-23 10:23:02 +00:00
synchon (Migrated from github.com) commented 2026-07-23 10:23:02 +00:00

The check is updated to be precise now. Also, open to go the non-idiomatic route as well.

The check is updated to be precise now. Also, open to go the non-idiomatic route as well.
synchon (Migrated from github.com) reviewed 2026-07-23 10:23:08 +00:00
synchon (Migrated from github.com) commented 2026-07-23 10:23:08 +00:00

Updated now.

Updated now.
fraimondo (Migrated from github.com) requested changes 2026-07-23 10:37:04 +00:00
fraimondo (Migrated from github.com) commented 2026-07-23 10:36:31 +00:00

We still rely on a ValueError. It should be something like

if component is registered:
Instantiate and dump
else:

  • add ALL variables in the meta to the yaml
  • Add the note: " is not a built-in component and thus could not be properly regenerated. Some of these entries in the YAML section might be redundant and not needed. Please check the documentation/implementation of this specific datagrabber and remove the unnecesary entries."
We still rely on a `ValueError`. It should be something like if component is registered: Instantiate and dump else: - add ALL variables in the meta to the yaml - Add the note: "<COMPONENT> is not a built-in component and thus could not be properly regenerated. Some of these entries in the YAML section might be redundant and not needed. Please check the documentation/implementation of this specific datagrabber and remove the unnecesary entries."
synchon commented 2026-07-23 10:59:14 +00:00 (Migrated from github.com)

We still rely on a ValueError.

ValueError is only raised if the component is not registered. try...except is the "idiomatic" way to do it. I understand if that doesn't work and it needs to be superfluous.

> We still rely on a ValueError. `ValueError` is only raised if the component is not registered. `try...except` is the "idiomatic" way to do it. I understand if that doesn't work and it needs to be superfluous.
fraimondo commented 2026-07-23 11:01:04 +00:00 (Migrated from github.com)

We still rely on a ValueError.

ValueError is only raised if the component is not registered. try...except is the "idiomatic" way to do it. I understand if that doesn't work and it needs to be superfluous.

In that case, if any part of the instantiation raises a ValueError (like would happen if a parameter changes options, or using an old junifer version), then we go to:

  • add ALL variables in the meta to the yaml
  • Add the note: " is not a built-in component or it failed to instantiate and thus could not be properly regenerated. Some of these entries in the YAML section might be redundant and not needed. Please check the documentation/implementation of this specific datagrabber and remove the unnecesary
> > We still rely on a ValueError. > > `ValueError` is only raised if the component is not registered. `try...except` is the "idiomatic" way to do it. I understand if that doesn't work and it needs to be superfluous. In that case, if any part of the instantiation raises a ValueError (like would happen if a parameter changes options, or using an old junifer version), then we go to: * add ALL variables in the meta to the yaml * Add the note: " is not a built-in component or it failed to instantiate and thus could not be properly regenerated. Some of these entries in the YAML section might be redundant and not needed. Please check the documentation/implementation of this specific datagrabber and remove the unnecesary
synchon commented 2026-07-23 11:11:04 +00:00 (Migrated from github.com)

We still rely on a ValueError.

ValueError is only raised if the component is not registered. try...except is the "idiomatic" way to do it. I understand if that doesn't work and it needs to be superfluous.

In that case, if any part of the instantiation raises a ValueError (like would happen if a parameter changes options, or using an old junifer version), then we go to:

  • add ALL variables in the meta to the yaml
  • Add the note: " is not a built-in component or it failed to instantiate and thus could not be properly regenerated. Some of these entries in the YAML section might be redundant and not needed. Please check the documentation/implementation of this specific datagrabber and remove the unnecesary

model_construct does not raise an exception so there will be no error during the model construction.

> > > We still rely on a ValueError. > > > > > > `ValueError` is only raised if the component is not registered. `try...except` is the "idiomatic" way to do it. I understand if that doesn't work and it needs to be superfluous. > > In that case, if any part of the instantiation raises a ValueError (like would happen if a parameter changes options, or using an old junifer version), then we go to: > > * add ALL variables in the meta to the yaml > * Add the note: " is not a built-in component or it failed to instantiate and thus could not be properly regenerated. Some of these entries in the YAML section might be redundant and not needed. Please check the documentation/implementation of this specific datagrabber and remove the unnecesary `model_construct` does not raise an exception so there will be no error during the model construction.
fraimondo commented 2026-07-23 11:18:15 +00:00 (Migrated from github.com)

Can we validate? I'm worried about using different junifer versions than the one that generated the meta. Or we either go full strict and not allow any mismatch (which will create a problem with julio later on), or we validate the model. Otherwise, variables that do not match will be "ignored" and not "dumped", which might yield a different yaml than expected.

I prefer to have a YAML with a note saying "check your datagrabber/marker/preprocessor due to possible changes in the API" than one without any message that actually works differently than expected.

Can we validate? I'm worried about using different junifer versions than the one that generated the meta. Or we either go full strict and not allow any mismatch (which will create a problem with julio later on), or we validate the model. Otherwise, variables that do not match will be "ignored" and not "dumped", which might yield a different yaml than expected. I prefer to have a YAML with a note saying "check your datagrabber/marker/preprocessor due to possible changes in the API" than one without any message that actually works differently than expected.
synchon commented 2026-07-23 13:58:50 +00:00 (Migrated from github.com)

Can we validate? I'm worried about using different junifer versions than the one that generated the meta. Or we either go full strict and not allow any mismatch (which will create a problem with julio later on), or we validate the model. Otherwise, variables that do not match will be "ignored" and not "dumped", which might yield a different yaml than expected.

I prefer to have a YAML with a note saying "check your datagrabber/marker/preprocessor due to possible changes in the API" than one without any message that actually works differently than expected.

model_construct is now replaced with model_validate which will fail for most due to the nature of the metadata and the models. Necessary comments will be added on generation.

> Can we validate? I'm worried about using different junifer versions than the one that generated the meta. Or we either go full strict and not allow any mismatch (which will create a problem with julio later on), or we validate the model. Otherwise, variables that do not match will be "ignored" and not "dumped", which might yield a different yaml than expected. > > I prefer to have a YAML with a note saying "check your datagrabber/marker/preprocessor due to possible changes in the API" than one without any message that actually works differently than expected. `model_construct` is now replaced with `model_validate` which will fail for most due to the nature of the metadata and the models. Necessary comments will be added on generation.
fraimondo commented 2026-07-23 14:07:39 +00:00 (Migrated from github.com)

Can we validate? I'm worried about using different junifer versions than the one that generated the meta. Or we either go full strict and not allow any mismatch (which will create a problem with julio later on), or we validate the model. Otherwise, variables that do not match will be "ignored" and not "dumped", which might yield a different yaml than expected.
I prefer to have a YAML with a note saying "check your datagrabber/marker/preprocessor due to possible changes in the API" than one without any message that actually works differently than expected.

model_construct is now replaced with model_validate which will fail for most due to the nature of the metadata and the models. Necessary comments will be added on generation.

I still don't understand why it will fail for "most". As long as you choose the dump_fields and pass it to the constructor, this should recreate the same object.

It should fail in case of:

  • External components
  • Components that changed API and the meta is from a non-compatible version.

All the rest should not fail. Otherwise we are dumping the wrong variables.

> > Can we validate? I'm worried about using different junifer versions than the one that generated the meta. Or we either go full strict and not allow any mismatch (which will create a problem with julio later on), or we validate the model. Otherwise, variables that do not match will be "ignored" and not "dumped", which might yield a different yaml than expected. > > I prefer to have a YAML with a note saying "check your datagrabber/marker/preprocessor due to possible changes in the API" than one without any message that actually works differently than expected. > > `model_construct` is now replaced with `model_validate` which will fail for most due to the nature of the metadata and the models. Necessary comments will be added on generation. I still don't understand why it will fail for "most". As long as you choose the `dump_fields` and pass it to the constructor, this should recreate the same object. It should fail in case of: - External components - Components that changed API and the meta is from a non-compatible version. All the rest should not fail. Otherwise we are dumping the wrong variables.
synchon commented 2026-07-23 14:50:10 +00:00 (Migrated from github.com)

Can we validate? I'm worried about using different junifer versions than the one that generated the meta. Or we either go full strict and not allow any mismatch (which will create a problem with julio later on), or we validate the model. Otherwise, variables that do not match will be "ignored" and not "dumped", which might yield a different yaml than expected.
I prefer to have a YAML with a note saying "check your datagrabber/marker/preprocessor due to possible changes in the API" than one without any message that actually works differently than expected.

model_construct is now replaced with model_validate which will fail for most due to the nature of the metadata and the models. Necessary comments will be added on generation.

I still don't understand why it will fail for "most". As long as you choose the dump_fields and pass it to the constructor, this should recreate the same object.

It should fail in case of:

  • External components
  • Components that changed API and the meta is from a non-compatible version.

All the rest should not fail. Otherwise we are dumping the wrong variables.

Ignore my previous reply's "fail" part, it works as intended.

> > > Can we validate? I'm worried about using different junifer versions than the one that generated the meta. Or we either go full strict and not allow any mismatch (which will create a problem with julio later on), or we validate the model. Otherwise, variables that do not match will be "ignored" and not "dumped", which might yield a different yaml than expected. > > > I prefer to have a YAML with a note saying "check your datagrabber/marker/preprocessor due to possible changes in the API" than one without any message that actually works differently than expected. > > > > > > `model_construct` is now replaced with `model_validate` which will fail for most due to the nature of the metadata and the models. Necessary comments will be added on generation. > > I still don't understand why it will fail for "most". As long as you choose the `dump_fields` and pass it to the constructor, this should recreate the same object. > > It should fail in case of: > > * External components > * Components that changed API and the meta is from a non-compatible version. > > All the rest should not fail. Otherwise we are dumping the wrong variables. Ignore my previous reply's "fail" part, it works as intended.
fraimondo (Migrated from github.com) reviewed 2026-07-24 11:27:56 +00:00
fraimondo (Migrated from github.com) commented 2026-07-24 09:19:29 +00:00

So this tests that the actual function works. Can we test for correctness?

So this tests that the actual function works. Can we test for correctness?
synchon (Migrated from github.com) reviewed 2026-07-24 12:32:55 +00:00
synchon (Migrated from github.com) commented 2026-07-24 12:32:55 +00:00

The latest commit should check for basic correctness.

The latest commit should check for basic correctness.
This pull request can be merged automatically.
This branch is out-of-date with the base branch
You are not authorized to merge this pull request.
View command line instructions

Checkout

From your project repository, check out a new branch and test the changes.
git fetch -u origin feat/generate-yaml-api:feat/generate-yaml-api
git switch feat/generate-yaml-api

Merge

Merge the changes and update on Forgejo.

Warning: The "Autodetect manual merge" setting is not enabled for this repository, you will have to mark this pull request as manually merged afterwards.

git switch main
git merge --no-ff feat/generate-yaml-api
git switch feat/generate-yaml-api
git rebase main
git switch main
git merge --ff-only feat/generate-yaml-api
git switch feat/generate-yaml-api
git rebase main
git switch main
git merge --no-ff feat/generate-yaml-api
git switch main
git merge --squash feat/generate-yaml-api
git switch main
git merge --ff-only feat/generate-yaml-api
git switch main
git merge feat/generate-yaml-api
git push origin main
Sign in to join this conversation.
No reviewers
No milestone
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
juaml/junifer!498
No description provided.