Buckets:
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| carbon | 20 items | ||
| README.md | 4.93 kB xet | 9dfa204b | |
| faceberg.yml | 534 Bytes xet | ec8324b2 | |
| lineage.yml | 1.54 kB xet | 76adf551 | |
| provenance.json | 14.5 kB xet | b2c16a7a |
Carbon Catalog
Metadata and provenance catalog for the Carbon genomic pretraining corpus enrichment pipeline.
Overview
The Carbon enrichment pipeline produces multiple datasets across CPU-based biological enrichment, stratified sampling, tokenization, and GPU-based model enrichment. As the pipeline evolves, tracking the origin, structure, and relationships of these artifacts becomes essential for reproducible analysis.
carbon-catalog provides a centralized catalog layer for documenting and organizing metadata associated with the enrichment pipeline.
The catalog is designed to connect data artifacts with their provenance, schemas, and processing context. It establishes a metadata boundary between pipeline execution and downstream dataset consumption.
Rather than storing the enriched genomic records themselves, the catalog serves as a reference layer for understanding how those datasets were produced and how they relate to one another.
Role in the Enrichment Architecture
The catalog is part of the metadata and governance layer of the enrichment pipeline.
Carbon Pretraining Corpus
│
▼
Enrichment Pipeline
│
├── CPU Enrichment
├── Sampling
├── Tokenization
└── GPU Enrichment
│
▼
Dataset Artifacts
│
▼
Carbon Catalog
│
┌─────────┼─────────┐
▼ ▼ ▼
Metadata Provenance Schemas
│
▼
Reproducible Data Analysis
The catalog complements the data processing pipeline by documenting the artifacts it produces and the context in which they were generated.
Objectives
The catalog supports the following objectives:
- Dataset provenance: Track the origin and lineage of generated data artifacts.
- Schema documentation: Describe the structure and metadata of enrichment datasets.
- Artifact organization: Maintain a structured metadata layer for pipeline outputs.
- Reproducibility: Provide context for understanding how datasets were produced.
- Pipeline observability: Support inspection of data products and their associated metadata.
- Downstream discovery: Help analytical workflows identify relevant datasets and their relationships.
Catalog Scope
The catalog is intended to document the datasets and metadata produced throughout the Carbon enrichment workflow.
Potential catalog entries include:
| Artifact | Metadata scope |
|---|---|
| CPU-enriched sequences | Biological enrichment schema, record identity, and sequence-level metrics. |
| Sampled cohorts | Sampling configuration, cohort definition, and selection context. |
| Tokenized corpus | Tokenization configuration, schema, and processing metadata. |
| Likelihood statistics | Model configuration, inference context, and output schema. |
| Embeddings | Model-derived representation metadata, dimensionality, and output schema. |
| Common cohort | Shared analysis population, interval identity, and taxonomy reference. |
The exact artifacts and metadata fields depend on the catalog contents and the pipeline configuration used to generate them.
Provenance
Provenance establishes the relationship between source data, processing stages, and generated artifacts.
For the Carbon enrichment pipeline, provenance can be used to document:
- Source corpus and dataset identifiers.
- Processing stage and enrichment method.
- Input and output dataset relationships.
- Processing configuration and execution context.
- Schema and metadata versions.
- Relevant validation and quality-control information.
A provenance record should make it possible to trace an analytical artifact back to the pipeline stage and source data from which it originated.
Reproducibility
Reproducible analysis requires more than access to the final dataset. It also depends on the ability to understand its source, processing history, and schema.
The catalog is intended to support this requirement by organizing metadata associated with the generated artifacts. Where available, provenance records and schema definitions should be used alongside the corresponding datasets during analysis.
Related Datasets
- Total size
- 993 kB
- Files
- 24
- Last updated
- Sep 23
- Pre-warmed CDN
- US EU US EU