Table of contents
Metadata subsystem
Cerebra extends the repository-based development model with an integrated metadata workflow. In addition to storing project files, a repository can have a structured metadata service with web forms, an API, validation, controlled submission and curation workflows.
This is particularly useful for projects that need to collect consistent information about datasets, samples, instruments, participants, software, publications, workflows, or other research objects. Instead of treating metadata as unstructured text or as a collection of manually maintained spreadsheets, a project can define the information it needs and provide contributors with a guided way to submit and maintain it.
What can it be used for?
The metadata capabilities can support several complementary activities:
- collecting metadata from researchers through generated web forms
- accepting metadata from scripts, pipelines, or other applications through an API
- validating submissions against a project-specific data model
- directing submissions to personal or team inboxes
- separating contribution, review, and curation
- managing read-only, contributor, reviewer, and curator access
- reusing a metadata model across multiple repositories or projects
- extending a shared model with domain-specific concepts
- exposing metadata in a structured, machine-readable form for later reuse
This makes Cerebra suitable not only for documenting files, but also for managing structured information throughout a research or data-engineering process. A project can begin by collecting basic descriptive metadata and later introduce more detailed validation, controlled vocabularies, relationships, and curation stages.
The fundamental architecture
The metadata subsystem has two main provider components that are driven by a common user-provided schema or data model.
The first component is the Metadata User Interface (MUI). Cerebra integrates shacl-vue, which generates browser-based forms from SHACL shapes. A repository can provide a configuration in a compact file that connects it to Cerebra's shared MUI service. The resulting interface can be made available for that repository without developing and deploying a custom frontend.
The second component is the Metadata API (MAPI). Cerebra integrates the Dump Things Service, which provides metadata collections, record storage and retrieval, validation, and an API for programmatic access. The Cerebra integration is documented in the Dumpthings API.
The shared data model that drives both components is authored in LinkML. LinkML provides a readable way to define classes, attributes, identifiers, relationships, and constraints. The model can be converted to SHACL for use by the MUI and MAPI services. Individual projects can create their own schemas, built on the Things schema, which provides reusable concepts for representing identifiable entities and their relationships.
Forgejo authentication and organization teams provide the access-control context. Team membership can be used to determine who may read collections, submit records to personal or team inboxes, or curate metadata.
Advantages
The main advantage is that the user interface, validation behavior, and API do not need to be designed independently. They are derived from the same data model. A change to the schema can therefore be reflected in both the forms used by people and the validation applied to programmatic submissions.
This also distinguishes Cerebra's metadata capabilities from a conventional Git forge. A conventional forge primarily organizes source files and collaboration around those files. In Cerebra, the repository can additionally act as the context for a structured metadata workflow: the project defines its model, contributors submit records, the service validates and stores them, and designated users curate the results.
Because the UI is generated and the metadata service is shared, projects can adopt structured metadata without building a complete application from scratch. At the same time, the schema remains under project control, so different projects can model different domains while still reusing common foundations and concepts.
flowchart LR
subgraph U["User-provided, use-case-specific content"]
S["Things-based<br/>LinkML schema"]
end
subgraph F["Forgejo platform"]
R["Repositories"]
I["Users, organizations,<br/>and teams"]
end
subgraph C["Cerebra integration"]
UI["Metadata user interface<br/>generated from the schema"]
API["Metadata API<br/>collections, records,<br/>and validation"]
end
S -->|"drives"| UI
S -->|"defines metadata<br/>and validation model"| API
R -->|"repository context"| UI
I -->|"authentication and<br/>team-based permissions"| UI
R -->|"repository context"| API
I -->|"authentication and<br/>team-based permissions"| API
UI <-->|"metadata collection,<br/>review, and curation"| API
classDef user fill:#fff3cd,stroke:#d39e00,color:#000
classDef forgejo fill:#e2e3e5,stroke:#6c757d,color:#000
classDef cerebra fill:#d1ecf1,stroke:#0c7489,color:#000
class S user
class R,I forgejo
class UI,API cerebra