chang-m-yun commited on
Commit
f27e66d
·
verified ·
1 Parent(s): 240bca7

Add files using upload-large-folder tool

Browse files
Files changed (1) hide show
  1. README.md +120 -0
README.md ADDED
@@ -0,0 +1,120 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: mit
3
+ library_name: chrombpnet
4
+ tags:
5
+ - encode
6
+ - chrombpnet
7
+ - chromatin-accessibility
8
+ - DNASE
9
+ - renal
10
+ - hg38
11
+ ---
12
+ # ENCODE ChromBPNet Atlas
13
+ As part of the ENCODE 4 Project, we trained ChromBPNet models on 1,512 ENCODE DNAse-seq and ATAC-seq across 408 biosamples. Here, we provide all models for open-source use.
14
+
15
+ For more information about the models, see:
16
+ - Main ENCODE 4 Paper
17
+ - [A unified lexicon of predictive DNA sequence motifs from ENCODE transcription factor binding and chromatin accessibility assays](https://doi.org/10.5281/zenodo.17123347)
18
+ - [ChromBPNet: bias factorized, base-resolution deep learning models of chromatin accessibility reveal cis-regulatory sequence syntax, transcription factor footprints and regulatory variants](https://doi.org/10.1101/2024.12.25.630221)
19
+
20
+ ## ChromBPNet model: DNASE in renal cortex interstitium (ENCSR155NPL)
21
+ - Model: ChromBPNet
22
+ - Assay: DNASE-seq
23
+ - Experiment: [ENCSR155NPL](https://www.encodeproject.org/experiments/ENCSR155NPL/)
24
+ - Model annotation: [ENCSR596EZJ](https://www.encodeproject.org/annotations/ENCSR596EZJ/)
25
+ - Biosample: renal cortex interstitium (Homo sapiens renal cortex interstitium tissue male embryo (113 days))
26
+ - Cell slim(s): None
27
+ - Organ slim(s): kidney
28
+ - Developmental slim(s): mesoderm
29
+ - System slim(s): excretory-system
30
+ - Assembly: hg38
31
+
32
+ ## Directory structure
33
+ - `fold_0`: Model: Cross-validation fold: Fold 0
34
+ - `model.chrombpnet.fold_0.encid.h5`: full chrombpnet model that combines both bias and corrected model in .h5 format
35
+ - `model.chrombpnet_nobias.fold_0.encid.h5`: bias-corrected accessibility model in .h5 format (Use for all biological discovery)
36
+ - `model.bias_scaled.fold_0.encid.h5`: bias model in .h5 format
37
+ - `model.chrombpnet.fold_0.encid.tar`: full chrombpnet model that combines both bias and corrected model in SavedModel format. After being untarred, it results in a directory named "chrombpnet".
38
+ - `model.chrombpnet_nobias.fold_0.encid.tar`: bias-corrected accessibility model in SavedModel format (Use for all biological discovery). After being untarred, it results in a directory named "chrombpnet_wo_bias".
39
+ - `model.bias_scaled.fold_0.encid.tar`: bias model in SavedModel format. After being untarred, it results in a directory named "bias_model_scaled".
40
+ - `logs.models.fold_0.encid`: folder containing log files for training models
41
+ - `fold_1`: Model: Cross-validation fold: Fold 1
42
+ - `fold_2`: Model: Cross-validation fold: Fold 2
43
+ - `fold_3`: Model: Cross-validation fold: Fold 3
44
+ - `fold_4`: Model: Cross-validation fold: Fold 4
45
+
46
+ # Instructions
47
+ ## (1) Pseudocode for loading models in .h5 format
48
+
49
+ (1) Use the code in python after appropriately defining `model_in_h5_format` and `inputs`.
50
+ (2) `inputs` is a one hot encoded sequence of shape (N,2114,4). Here N corresponds to the
51
+ number of tested sequences, 2114 is the input sequence length and 4 corresponds to [A,C,G,T].
52
+
53
+ ```python
54
+ import tensorflow as tf
55
+ from tensorflow.keras.utils import get_custom_objects
56
+ from tensorflow.keras.models import load_model
57
+
58
+ custom_objects={"tf": tf}
59
+ get_custom_objects().update(custom_objects)
60
+
61
+ model=load_model(model_in_h5_format,compile=False)
62
+ outputs = model(inputs)
63
+ ```
64
+
65
+ The list `outputs` consists of two elements. The first element has a shape of (N, 1000) and
66
+ contains logit predictions for a 1000-base-pair output. The second element, with a shape of
67
+ (N, 1), contains logcount predictions. To transform these predictions into per-base signals,
68
+ follow the provided pseudo code lines below.
69
+
70
+ ```python
71
+ import numpy as np
72
+
73
+ def softmax(x, temp=1):
74
+ norm_x = x - np.mean(x,axis=1, keepdims=True)
75
+ return np.exp(temp*norm_x)/np.sum(np.exp(temp*norm_x), axis=1, keepdims=True)
76
+
77
+ predictions = softmax(outputs[0]) * (np.exp(outputs[1])-1)
78
+ ```
79
+
80
+ ## (2) Pseudocode for loading models in .tar format
81
+
82
+ (1) First untar the directory as follows `tar -xvf model.tar`
83
+ (2) Use the code below in python after appropriately defining `model_dir_untared` and `inputs`
84
+ (3) `inputs` is a one hot encoded sequence of shape (N,2114,4). Here N corresponds to the number
85
+ of tested sequences, 2114 is the input sequence length and 4 corresponds to ACGT.
86
+
87
+ Reference: https://www.tensorflow.org/api_docs/python/tf/saved_model/load
88
+
89
+ ```python
90
+ import tensorflow as tf
91
+
92
+ model = tf.saved_model.load('model_dir_untared')
93
+ outputs = model.signatures['serving_default'](**{'sequence':inputs.astype('float32')})
94
+ ```
95
+
96
+ The variable `outputs` represents a dictionary containing two key-value pairs. The first key
97
+ is `logits_profile_predictions`, holding a value with a shape of (N, 1000). This value corresponds
98
+ to logit predictions for a 1000-base-pair output. The second key, named `logcount_predictions``,
99
+ is associated with a value of shape (N, 1), representing logcount predictions. To transform these
100
+ predictions into per-base signals, utilize the provided pseudo code lines mentioned below.
101
+
102
+ ```python
103
+ import numpy as np
104
+ def softmax(x, temp=1):
105
+ norm_x = x - np.mean(x,axis=1, keepdims=True)
106
+ return np.exp(temp*norm_x)/np.sum(np.exp(temp*norm_x), axis=1, keepdims=True)
107
+
108
+ predictions = softmax(outputs["logits_profile_predictions"]) * (np.exp(outputs["logcount_predictions"])-1)
109
+ ```
110
+
111
+ ## Docker image to load and use the models
112
+ https://hub.docker.com/r/kundajelab/chrombpnet-atlas/ (tag:v1)
113
+
114
+ ## Code for ChromBPNet
115
+ - https://github.com/kundajelab/chrombpnet/
116
+
117
+ # License & citation
118
+ External data users may freely download, analyze and publish results based on any ENCODE data without restrictions.
119
+
120
+ Released under the [ENCODE data-use policy](https://www.encodeproject.org/about/data-use-policy/). Please cite the ENCODE Project Consortium and the model software: [ChromBPNet](https://github.com/kundajelab/chrombpnet) (Pampari et al., bioRxiv 2024).