993 kB
24 files
Updated 6 days ago
Name
Size
carbon
README.md4.93 kB
xet
faceberg.yml534 Bytes
xet
lineage.yml1.54 kB
xet
provenance.json14.5 kB
xet
README.md

Carbon Catalog

Metadata and provenance catalog for the Carbon genomic pretraining corpus enrichment pipeline.

Overview

The Carbon enrichment pipeline produces multiple datasets across CPU-based biological enrichment, stratified sampling, tokenization, and GPU-based model enrichment. As the pipeline evolves, tracking the origin, structure, and relationships of these artifacts becomes essential for reproducible analysis.

carbon-catalog provides a centralized catalog layer for documenting and organizing metadata associated with the enrichment pipeline.

The catalog is designed to connect data artifacts with their provenance, schemas, and processing context. It establishes a metadata boundary between pipeline execution and downstream dataset consumption.

Rather than storing the enriched genomic records themselves, the catalog serves as a reference layer for understanding how those datasets were produced and how they relate to one another.

Role in the Enrichment Architecture

The catalog is part of the metadata and governance layer of the enrichment pipeline.

Carbon Pretraining Corpus
          │
          ▼
  Enrichment Pipeline
          │
          ├── CPU Enrichment
          ├── Sampling
          ├── Tokenization
          └── GPU Enrichment
                    │
                    ▼
             Dataset Artifacts
                    │
                    ▼
              Carbon Catalog
                    │
          ┌─────────┼─────────┐
          ▼         ▼         ▼
      Metadata  Provenance  Schemas
                    │
                    ▼
       Reproducible Data Analysis

The catalog complements the data processing pipeline by documenting the artifacts it produces and the context in which they were generated.

Objectives

The catalog supports the following objectives:

  • Dataset provenance: Track the origin and lineage of generated data artifacts.
  • Schema documentation: Describe the structure and metadata of enrichment datasets.
  • Artifact organization: Maintain a structured metadata layer for pipeline outputs.
  • Reproducibility: Provide context for understanding how datasets were produced.
  • Pipeline observability: Support inspection of data products and their associated metadata.
  • Downstream discovery: Help analytical workflows identify relevant datasets and their relationships.

Catalog Scope

The catalog is intended to document the datasets and metadata produced throughout the Carbon enrichment workflow.

Potential catalog entries include:

Artifact Metadata scope
CPU-enriched sequences Biological enrichment schema, record identity, and sequence-level metrics.
Sampled cohorts Sampling configuration, cohort definition, and selection context.
Tokenized corpus Tokenization configuration, schema, and processing metadata.
Likelihood statistics Model configuration, inference context, and output schema.
Embeddings Model-derived representation metadata, dimensionality, and output schema.
Common cohort Shared analysis population, interval identity, and taxonomy reference.

The exact artifacts and metadata fields depend on the catalog contents and the pipeline configuration used to generate them.

Provenance

Provenance establishes the relationship between source data, processing stages, and generated artifacts.

For the Carbon enrichment pipeline, provenance can be used to document:

  • Source corpus and dataset identifiers.
  • Processing stage and enrichment method.
  • Input and output dataset relationships.
  • Processing configuration and execution context.
  • Schema and metadata versions.
  • Relevant validation and quality-control information.

A provenance record should make it possible to trace an analytical artifact back to the pipeline stage and source data from which it originated.

Reproducibility

Reproducible analysis requires more than access to the final dataset. It also depends on the ability to understand its source, processing history, and schema.

The catalog is intended to support this requirement by organizing metadata associated with the generated artifacts. Where available, provenance records and schema definitions should be used alongside the corresponding datasets during analysis.

Related Datasets

Total size
993 kB
Files
24
Last updated
Sep 23
Pre-warmed CDN
US EU US EU

Contributors