Spaces:
Running
Running
Add the benchmark datasets section to the org card
Browse files
README.md
CHANGED
|
@@ -11,7 +11,7 @@ pinned: false
|
|
| 11 |
|
| 12 |
# Aurigene AI
|
| 13 |
|
| 14 |
-
We are the AI and computational discovery group at **Aurigene Pharmaceutical Services Limited**. This organization is our public home for open models and browser-based tools that support **AI-driven drug discovery** β from picking a target, through generating and triaging chemical matter, to planning the synthesis and reading the literature that justifies all of it.
|
| 15 |
|
| 16 |
Everything here is open, permissively licensed, and mirrored from the original authors with full attribution.
|
| 17 |
|
|
@@ -87,6 +87,40 @@ Mine the papers that justify the programme.
|
|
| 87 |
|
| 88 |
---
|
| 89 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 90 |
## β‘ Quick start
|
| 91 |
|
| 92 |
```python
|
|
@@ -103,7 +137,7 @@ print(embeddings.shape) # torch.Size([2, 768])
|
|
| 103 |
|
| 104 |
## π Collections
|
| 105 |
|
| 106 |
-
Browse the catalogue as curated collections: [molecular representation & property prediction](https://huggingface.co/collections/Aurigene-AI/molecular-representation-and-property-prediction-6aa2203a94123133168b49e2), [generative chemistry & synthesis planning](https://huggingface.co/collections/Aurigene-AI/generative-chemistry-and-synthesis-planning-6aa2203c3d5b3527359ee444), [protein & target modeling](https://huggingface.co/collections/Aurigene-AI/protein-and-target-modeling-6aa2203e4e5bdf1f106c7d55), [biomedical language models](https://huggingface.co/collections/Aurigene-AI/biomedical-language-models-6aa2203f4469f54b61382c40), [interactive tools](https://huggingface.co/collections/Aurigene-AI/interactive-drug-discovery-tools-6aa220414e5bdf1f106c7dc0).
|
| 107 |
|
| 108 |
## π Attribution & licensing
|
| 109 |
|
|
|
|
| 11 |
|
| 12 |
# Aurigene AI
|
| 13 |
|
| 14 |
+
We are the AI and computational discovery group at **Aurigene Pharmaceutical Services Limited**. This organization is our public home for open models, benchmark datasets and browser-based tools that support **AI-driven drug discovery** β from picking a target, through generating and triaging chemical matter, to planning the synthesis and reading the literature that justifies all of it.
|
| 15 |
|
| 16 |
Everything here is open, permissively licensed, and mirrored from the original authors with full attribution.
|
| 17 |
|
|
|
|
| 87 |
|
| 88 |
---
|
| 89 |
|
| 90 |
+
## π Benchmark datasets
|
| 91 |
+
|
| 92 |
+
Standard evaluation sets, mirrored so every model above has data to train and benchmark against. Row counts on each card are computed from the files themselves.
|
| 93 |
+
|
| 94 |
+
**Lead optimization β MoleculeNet ADMET & toxicity**
|
| 95 |
+
|
| 96 |
+
| Dataset | Task | Rows |
|
| 97 |
+
| :-- | :-- | --: |
|
| 98 |
+
| [MoleculeNet_BBBP](https://huggingface.co/datasets/Aurigene-AI/MoleculeNet_BBBP) | Blood-brain barrier penetration | 2,039 |
|
| 99 |
+
| [MoleculeNet_BACE](https://huggingface.co/datasets/Aurigene-AI/MoleculeNet_BACE) | BACE-1 inhibition, an Alzheimer's target | 1,513 |
|
| 100 |
+
| [MoleculeNet_ClinTox](https://huggingface.co/datasets/Aurigene-AI/MoleculeNet_ClinTox) | Clinical toxicity & FDA approval | 1,477 |
|
| 101 |
+
| [MoleculeNet_Tox21](https://huggingface.co/datasets/Aurigene-AI/MoleculeNet_Tox21) | 12 toxicity assays | 7,831 |
|
| 102 |
+
| [MoleculeNet_SIDER](https://huggingface.co/datasets/Aurigene-AI/MoleculeNet_SIDER) | Marketed-drug side effects | 1,427 |
|
| 103 |
+
| [MoleculeNet_ESOL](https://huggingface.co/datasets/Aurigene-AI/MoleculeNet_ESOL) | Aqueous solubility | 1,128 |
|
| 104 |
+
| [MoleculeNet_FreeSolv](https://huggingface.co/datasets/Aurigene-AI/MoleculeNet_FreeSolv) | Hydration free energy | 642 |
|
| 105 |
+
| [MoleculeNet_Lipophilicity](https://huggingface.co/datasets/Aurigene-AI/MoleculeNet_Lipophilicity) | logD at pH 7.4 | 4,200 |
|
| 106 |
+
|
| 107 |
+
**Hit generation β virtual screening**
|
| 108 |
+
|
| 109 |
+
| Dataset | Task | Rows |
|
| 110 |
+
| :-- | :-- | --: |
|
| 111 |
+
| [MoleculeNet_HIV](https://huggingface.co/datasets/Aurigene-AI/MoleculeNet_HIV) | HIV replication inhibition | 41,127 |
|
| 112 |
+
|
| 113 |
+
**Evidence & literature β medical QA**
|
| 114 |
+
|
| 115 |
+
| Dataset | Task | Rows |
|
| 116 |
+
| :-- | :-- | --: |
|
| 117 |
+
| [MedQA-USMLE-4-options](https://huggingface.co/datasets/Aurigene-AI/MedQA-USMLE-4-options) | US licensing exam, 4-option MCQ | 11,451 |
|
| 118 |
+
| [MedMCQA](https://huggingface.co/datasets/Aurigene-AI/MedMCQA) | Medical entrance exam questions | 193,155 |
|
| 119 |
+
|
| 120 |
+
All eleven are in the [Benchmark Datasets collection](https://huggingface.co/collections/Aurigene-AI/benchmark-datasets-6aa235046a491fa5b8a87fe0).
|
| 121 |
+
|
| 122 |
+
---
|
| 123 |
+
|
| 124 |
## β‘ Quick start
|
| 125 |
|
| 126 |
```python
|
|
|
|
| 137 |
|
| 138 |
## π Collections
|
| 139 |
|
| 140 |
+
Browse the catalogue as curated collections: [molecular representation & property prediction](https://huggingface.co/collections/Aurigene-AI/molecular-representation-and-property-prediction-6aa2203a94123133168b49e2), [generative chemistry & synthesis planning](https://huggingface.co/collections/Aurigene-AI/generative-chemistry-and-synthesis-planning-6aa2203c3d5b3527359ee444), [protein & target modeling](https://huggingface.co/collections/Aurigene-AI/protein-and-target-modeling-6aa2203e4e5bdf1f106c7d55), [biomedical language models](https://huggingface.co/collections/Aurigene-AI/biomedical-language-models-6aa2203f4469f54b61382c40), [interactive tools](https://huggingface.co/collections/Aurigene-AI/interactive-drug-discovery-tools-6aa220414e5bdf1f106c7dc0), [benchmark datasets](https://huggingface.co/collections/Aurigene-AI/benchmark-datasets-6aa235046a491fa5b8a87fe0).
|
| 141 |
|
| 142 |
## π Attribution & licensing
|
| 143 |
|