priyaganesh2050 commited on
Commit
9facd7e
Β·
verified Β·
1 Parent(s): fc7278e

Add the benchmark datasets section to the org card

Browse files
Files changed (1) hide show
  1. README.md +36 -2
README.md CHANGED
@@ -11,7 +11,7 @@ pinned: false
11
 
12
  # Aurigene AI
13
 
14
- We are the AI and computational discovery group at **Aurigene Pharmaceutical Services Limited**. This organization is our public home for open models and browser-based tools that support **AI-driven drug discovery** β€” from picking a target, through generating and triaging chemical matter, to planning the synthesis and reading the literature that justifies all of it.
15
 
16
  Everything here is open, permissively licensed, and mirrored from the original authors with full attribution.
17
 
@@ -87,6 +87,40 @@ Mine the papers that justify the programme.
87
 
88
  ---
89
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
90
  ## ⚑ Quick start
91
 
92
  ```python
@@ -103,7 +137,7 @@ print(embeddings.shape) # torch.Size([2, 768])
103
 
104
  ## πŸ“š Collections
105
 
106
- Browse the catalogue as curated collections: [molecular representation & property prediction](https://huggingface.co/collections/Aurigene-AI/molecular-representation-and-property-prediction-6aa2203a94123133168b49e2), [generative chemistry & synthesis planning](https://huggingface.co/collections/Aurigene-AI/generative-chemistry-and-synthesis-planning-6aa2203c3d5b3527359ee444), [protein & target modeling](https://huggingface.co/collections/Aurigene-AI/protein-and-target-modeling-6aa2203e4e5bdf1f106c7d55), [biomedical language models](https://huggingface.co/collections/Aurigene-AI/biomedical-language-models-6aa2203f4469f54b61382c40), [interactive tools](https://huggingface.co/collections/Aurigene-AI/interactive-drug-discovery-tools-6aa220414e5bdf1f106c7dc0).
107
 
108
  ## πŸ“„ Attribution & licensing
109
 
 
11
 
12
  # Aurigene AI
13
 
14
+ We are the AI and computational discovery group at **Aurigene Pharmaceutical Services Limited**. This organization is our public home for open models, benchmark datasets and browser-based tools that support **AI-driven drug discovery** β€” from picking a target, through generating and triaging chemical matter, to planning the synthesis and reading the literature that justifies all of it.
15
 
16
  Everything here is open, permissively licensed, and mirrored from the original authors with full attribution.
17
 
 
87
 
88
  ---
89
 
90
+ ## πŸ“Š Benchmark datasets
91
+
92
+ Standard evaluation sets, mirrored so every model above has data to train and benchmark against. Row counts on each card are computed from the files themselves.
93
+
94
+ **Lead optimization β€” MoleculeNet ADMET & toxicity**
95
+
96
+ | Dataset | Task | Rows |
97
+ | :-- | :-- | --: |
98
+ | [MoleculeNet_BBBP](https://huggingface.co/datasets/Aurigene-AI/MoleculeNet_BBBP) | Blood-brain barrier penetration | 2,039 |
99
+ | [MoleculeNet_BACE](https://huggingface.co/datasets/Aurigene-AI/MoleculeNet_BACE) | BACE-1 inhibition, an Alzheimer's target | 1,513 |
100
+ | [MoleculeNet_ClinTox](https://huggingface.co/datasets/Aurigene-AI/MoleculeNet_ClinTox) | Clinical toxicity & FDA approval | 1,477 |
101
+ | [MoleculeNet_Tox21](https://huggingface.co/datasets/Aurigene-AI/MoleculeNet_Tox21) | 12 toxicity assays | 7,831 |
102
+ | [MoleculeNet_SIDER](https://huggingface.co/datasets/Aurigene-AI/MoleculeNet_SIDER) | Marketed-drug side effects | 1,427 |
103
+ | [MoleculeNet_ESOL](https://huggingface.co/datasets/Aurigene-AI/MoleculeNet_ESOL) | Aqueous solubility | 1,128 |
104
+ | [MoleculeNet_FreeSolv](https://huggingface.co/datasets/Aurigene-AI/MoleculeNet_FreeSolv) | Hydration free energy | 642 |
105
+ | [MoleculeNet_Lipophilicity](https://huggingface.co/datasets/Aurigene-AI/MoleculeNet_Lipophilicity) | logD at pH 7.4 | 4,200 |
106
+
107
+ **Hit generation β€” virtual screening**
108
+
109
+ | Dataset | Task | Rows |
110
+ | :-- | :-- | --: |
111
+ | [MoleculeNet_HIV](https://huggingface.co/datasets/Aurigene-AI/MoleculeNet_HIV) | HIV replication inhibition | 41,127 |
112
+
113
+ **Evidence & literature β€” medical QA**
114
+
115
+ | Dataset | Task | Rows |
116
+ | :-- | :-- | --: |
117
+ | [MedQA-USMLE-4-options](https://huggingface.co/datasets/Aurigene-AI/MedQA-USMLE-4-options) | US licensing exam, 4-option MCQ | 11,451 |
118
+ | [MedMCQA](https://huggingface.co/datasets/Aurigene-AI/MedMCQA) | Medical entrance exam questions | 193,155 |
119
+
120
+ All eleven are in the [Benchmark Datasets collection](https://huggingface.co/collections/Aurigene-AI/benchmark-datasets-6aa235046a491fa5b8a87fe0).
121
+
122
+ ---
123
+
124
  ## ⚑ Quick start
125
 
126
  ```python
 
137
 
138
  ## πŸ“š Collections
139
 
140
+ Browse the catalogue as curated collections: [molecular representation & property prediction](https://huggingface.co/collections/Aurigene-AI/molecular-representation-and-property-prediction-6aa2203a94123133168b49e2), [generative chemistry & synthesis planning](https://huggingface.co/collections/Aurigene-AI/generative-chemistry-and-synthesis-planning-6aa2203c3d5b3527359ee444), [protein & target modeling](https://huggingface.co/collections/Aurigene-AI/protein-and-target-modeling-6aa2203e4e5bdf1f106c7d55), [biomedical language models](https://huggingface.co/collections/Aurigene-AI/biomedical-language-models-6aa2203f4469f54b61382c40), [interactive tools](https://huggingface.co/collections/Aurigene-AI/interactive-drug-discovery-tools-6aa220414e5bdf1f106c7dc0), [benchmark datasets](https://huggingface.co/collections/Aurigene-AI/benchmark-datasets-6aa235046a491fa5b8a87fe0).
141
 
142
  ## πŸ“„ Attribution & licensing
143