README / README.md
priyaganesh2050's picture
Lay the Space buttons out in a row and fix the clipped title
2bd469c verified
|
Raw History Blame
7.95 kB
metadata
title: README
emoji: 🧬
colorFrom: green
colorTo: indigo
sdk: static
pinned: false

Aurigene AI β€” open models and interactive tools for AI-driven drug discovery

Aurigene AI

We are the AI and computational discovery group at Aurigene Pharmaceutical Services Limited. This organization is our public home for open models and browser-based tools that support AI-driven drug discovery β€” from picking a target, through generating and triaging chemical matter, to planning the synthesis and reading the literature that justifies all of it.

Everything here is open, permissively licensed, and mirrored from the original authors with full attribution.


🧭 Start here β€” three tools that run in your browser

No GPU, no login, no install. These are static Spaces, so they never sleep and load instantly.

Molecule Explorer Protein Target Explorer Drug Discovery Model Hub

πŸ§ͺ  Molecule Explorer  β†’

Paste a SMILES string or a compound name. Get the 2D structure, physicochemical descriptors, Lipinski / Veber / Ghose / Egan / Muegge rules, an approximate QED, highlighted structural alerts, nearest approved drugs by Tanimoto similarity, and batch profiling with CSV export. RDKit runs as WebAssembly, so nothing is uploaded.

🧬  Protein Target Explorer  β†’

Enter a UniProt accession, a human gene symbol or a PDB ID. Get the target's annotation, its AlphaFold model coloured by per-residue confidence, every experimental structure with bound ligands, DrugBank drugs that hit it, and full sequence physicochemistry.

πŸ—ΊοΈ  Drug Discovery Model Hub  β†’

A live dashboard of this whole catalogue, mapped onto the five stages of discovery, with copy-paste transformers snippets and adoption statistics pulled from the Hub API.


πŸ”¬ Models by discovery stage

Stage 01 β€” Target identification

Understand the protein before you try to drug it.

  • ESM-2 650M β€” Meta's protein language model. Per-residue embeddings that transfer to binding-site prediction, variant-effect scoring and structure-aware featurisation.
  • ESMFold v1 β€” end-to-end structure prediction from a single sequence, no MSA step. Folds orphan sequences and designed constructs in seconds.
  • MAMMAL biomed multi-alignment 458M β€” IBM's multimodal model trained on over 2 billion biological samples; handles drug–target interaction and binding-affinity prediction.

Stage 02 β€” Hit generation

Generate and screen chemical matter.

  • MoLFormer-XL β€” IBM's linear-attention SMILES encoder pretrained on 1.1 billion molecules from ZINC and PubChem. The workhorse for embedding a library.
  • ChemFM-1B β€” a 1 B-parameter chemistry foundation model trained on UniChem, for generation, property prediction and reaction tasks.
  • GPT-2 ZINC 87M β€” small autoregressive SMILES generator trained on ~480 M ZINC molecules; samples de novo structures on a laptop CPU.

Stage 03 β€” Lead optimization

Predict properties, refine the series.

  • ChemBERTa-2 77M MTR β€” DeepChem's chemical language model pretrained with multitask regression over 200 RDKit descriptors. A strong, cheap ADMET baseline.
  • MMELON multi-view 84M β€” fuses SMILES, 2D graph and rendered image views of a molecule into one embedding for property prediction and virtual screening.

Stage 04 β€” Synthesis planning

Can we actually make it?

Stage 05 β€” Evidence & literature

Mine the papers that justify the programme.

  • BiomedBERT (PubMedBERT) β€” Microsoft's encoder pretrained from scratch on PubMed abstracts plus PMC full text.
  • BioGPT β€” generative biomedical LM for relation extraction and question answering over 15 M PubMed abstracts.
  • BioMistral-7B β€” Mistral-7B further pretrained on PubMed Central Open Access; chat-capable across ten languages.
  • Biomedical NER (107 entities) β€” DistilBERT token classifier covering diseases, drugs, dosages, signs and lab values.
  • MolT5-large SMILES β†’ text β€” writes a natural-language description of a molecule from its SMILES.

⚑ Quick start

from transformers import AutoModel, AutoTokenizer

model_id = "Aurigene-AI/MoLFormer-XL-both-10pct"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModel.from_pretrained(model_id, trust_remote_code=True, deterministic_eval=True)

smiles = ["CC(=O)Oc1ccccc1C(=O)O", "CC(C)Cc1ccc(cc1)C(C)C(=O)O"]
embeddings = model(**tokenizer(smiles, padding=True, return_tensors="pt")).pooler_output
print(embeddings.shape)   # torch.Size([2, 768])

πŸ“š Collections

Browse the catalogue as curated collections: molecular representation & property prediction, generative chemistry & synthesis planning, protein & target modeling, biomedical language models, interactive tools.

πŸ“„ Attribution & licensing

Every model here is a mirror of an open upstream release. The original authors β€” IBM Research, Meta AI, Microsoft Research, DeepChem, ChemFM, BioMistral and others β€” retain all credit, and each repository keeps the upstream model card and licence intact. Please cite the original work.

Models and tools are provided for research use. Rule-based filters and predictions are triage heuristics, not statements about safety or efficacy, and nothing here is a medical device or clinical advice.

🌐 aurigeneservices.com