hazemessam commited on
Commit
9f7c1ad
·
verified ·
1 Parent(s): fe0550a

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +25 -25
README.md CHANGED
@@ -8,35 +8,35 @@ Ankh3 is a protein language model that is jointly optimized on two objectives:
8
  * Protein sequence completion.
9
 
10
  1. Masked Language Modeling:
11
- The idea of this task is to intentionally 'corrupt' an input protein sequence by
12
- masking a certain percentage (X%) of its individual tokens (amino acids),
13
- and then train the model to reconstruct the original sequence.
14
-
15
- Example on a protein sequence before and after corruption:
16
- Original protein sequence: MKAYVLINSRGP
17
-
18
- This sequence will be masked/corrupted using sentinel tokens as shown below:
19
- Sequence after corruption: M <extra_id_0> A Y <extra_id_1> L I <extra_id_2> S R G <extra_id_3>
20
-
21
- The decoder learns to correspond each sentinel token to the actual amino acid that was masked.
22
- In this example: <extra_id_0> K means that <extra_id_0> corresponds to the "K" amino acid and so on.
23
- Decoder output: <extra_id_0> K <extra_id_1> V <extra_id_2> N <extra_id_3> P
24
 
25
 
26
 
27
  2. Protein Sequence Completion:
28
- The idea of this task is to cut the input sequence into
29
- two segments, where the first segment is fed to the encoder
30
- and the decoder is tasked to auto-regressively generate the
31
- second segment conditioned on the first segment representation
32
- outputted from the encoder.
33
-
34
- Example on protein sequence completion:
35
-
36
- Original sequence: MKAYVLINSRGP
37
- We will pass "MKAYVL" of it to the encoder, and the decoder is trained
38
- that given the representation of the first part provided by the encoder,
39
- it should output the second part which is: "INSRGP"
40
 
41
 
42
 
 
8
  * Protein sequence completion.
9
 
10
  1. Masked Language Modeling:
11
+ The idea of this task is to intentionally 'corrupt' an input protein sequence by
12
+ masking a certain percentage (X%) of its individual tokens (amino acids),
13
+ and then train the model to reconstruct the original sequence.
14
+
15
+ Example on a protein sequence before and after corruption:
16
+ Original protein sequence: MKAYVLINSRGP
17
+
18
+ This sequence will be masked/corrupted using sentinel tokens as shown below:
19
+ Sequence after corruption: M <extra_id_0> A Y <extra_id_1> L I <extra_id_2> S R G <extra_id_3>
20
+
21
+ The decoder learns to correspond each sentinel token to the actual amino acid that was masked.
22
+ In this example: <extra_id_0> K means that <extra_id_0> corresponds to the "K" amino acid and so on.
23
+ Decoder output: <extra_id_0> K <extra_id_1> V <extra_id_2> N <extra_id_3> P
24
 
25
 
26
 
27
  2. Protein Sequence Completion:
28
+ The idea of this task is to cut the input sequence into
29
+ two segments, where the first segment is fed to the encoder
30
+ and the decoder is tasked to auto-regressively generate the
31
+ second segment conditioned on the first segment representation
32
+ outputted from the encoder.
33
+
34
+ Example on protein sequence completion:
35
+
36
+ Original sequence: MKAYVLINSRGP
37
+ We will pass "MKAYVL" of it to the encoder, and the decoder is trained
38
+ that given the representation of the first part provided by the encoder,
39
+ it should output the second part which is: "INSRGP"
40
 
41
 
42