Spaces:
Running
Running
|
Download README.md from Aurigene-AI/README: direct link, hf CLI and curl.
- Browser
- Download file 7.95 kB
-
https://huggingface.co/spaces/Aurigene-AI/README/resolve/4b19f853a060c189ed29bb7dd5855e8e496ab172/README.md
- Command line
-
hf download hf://spaces/Aurigene-AI/README@4b19f853a060c189ed29bb7dd5855e8e496ab172/README.md
-
curl -L -o README.md https://huggingface.co/spaces/Aurigene-AI/README/resolve/4b19f853a060c189ed29bb7dd5855e8e496ab172/README.md
7.95 kB
| title: README | |
| emoji: 𧬠| |
| colorFrom: green | |
| colorTo: indigo | |
| sdk: static | |
| pinned: false | |
|  | |
| # Aurigene AI | |
| We are the AI and computational discovery group at **Aurigene Pharmaceutical Services Limited**. This organization is our public home for open models and browser-based tools that support **AI-driven drug discovery** β from picking a target, through generating and triaging chemical matter, to planning the synthesis and reading the literature that justifies all of it. | |
| Everything here is open, permissively licensed, and mirrored from the original authors with full attribution. | |
| --- | |
| ## π§ Start here β three tools that run in your browser | |
| No GPU, no login, no install. These are static Spaces, so they never sleep and load instantly. | |
| <table> | |
| <tr> | |
| <td align="center"><a href="https://huggingface.co/spaces/Aurigene-AI/molecule-explorer"><img src="./btn-molecule.png" width="316" alt="Molecule Explorer"></a></td> | |
| <td align="center"><a href="https://huggingface.co/spaces/Aurigene-AI/protein-target-explorer"><img src="./btn-protein.png" width="316" alt="Protein Target Explorer"></a></td> | |
| <td align="center"><a href="https://huggingface.co/spaces/Aurigene-AI/drug-discovery-model-hub"><img src="./btn-hub.png" width="316" alt="Drug Discovery Model Hub"></a></td> | |
| </tr> | |
| </table> | |
| ### π§ͺ [Molecule Explorer β](https://huggingface.co/spaces/Aurigene-AI/molecule-explorer) | |
| Paste a SMILES string or a compound name. Get the 2D structure, physicochemical descriptors, Lipinski / Veber / Ghose / Egan / Muegge rules, an approximate QED, highlighted structural alerts, nearest approved drugs by Tanimoto similarity, and batch profiling with CSV export. RDKit runs as WebAssembly, so nothing is uploaded. | |
| ### 𧬠[Protein Target Explorer β](https://huggingface.co/spaces/Aurigene-AI/protein-target-explorer) | |
| Enter a UniProt accession, a human gene symbol or a PDB ID. Get the target's annotation, its AlphaFold model coloured by per-residue confidence, every experimental structure with bound ligands, DrugBank drugs that hit it, and full sequence physicochemistry. | |
| ### πΊοΈ [Drug Discovery Model Hub β](https://huggingface.co/spaces/Aurigene-AI/drug-discovery-model-hub) | |
| A live dashboard of this whole catalogue, mapped onto the five stages of discovery, with copy-paste `transformers` snippets and adoption statistics pulled from the Hub API. | |
| --- | |
| ## π¬ Models by discovery stage | |
| ### Stage 01 β Target identification | |
| Understand the protein before you try to drug it. | |
| - **[ESM-2 650M](https://huggingface.co/Aurigene-AI/esm2_t33_650M_UR50D)** β Meta's protein language model. Per-residue embeddings that transfer to binding-site prediction, variant-effect scoring and structure-aware featurisation. | |
| - **[ESMFold v1](https://huggingface.co/Aurigene-AI/esmfold_v1)** β end-to-end structure prediction from a single sequence, no MSA step. Folds orphan sequences and designed constructs in seconds. | |
| - **[MAMMAL biomed multi-alignment 458M](https://huggingface.co/Aurigene-AI/biomed.omics.bl.sm.ma-ted-458m)** β IBM's multimodal model trained on over 2 billion biological samples; handles drugβtarget interaction and binding-affinity prediction. | |
| ### Stage 02 β Hit generation | |
| Generate and screen chemical matter. | |
| - **[MoLFormer-XL](https://huggingface.co/Aurigene-AI/MoLFormer-XL-both-10pct)** β IBM's linear-attention SMILES encoder pretrained on 1.1 billion molecules from ZINC and PubChem. The workhorse for embedding a library. | |
| - **[ChemFM-1B](https://huggingface.co/Aurigene-AI/ChemFM-1B)** β a 1 B-parameter chemistry foundation model trained on UniChem, for generation, property prediction and reaction tasks. | |
| - **[GPT-2 ZINC 87M](https://huggingface.co/Aurigene-AI/gpt2_zinc_87m)** β small autoregressive SMILES generator trained on ~480 M ZINC molecules; samples de novo structures on a laptop CPU. | |
| ### Stage 03 β Lead optimization | |
| Predict properties, refine the series. | |
| - **[ChemBERTa-2 77M MTR](https://huggingface.co/Aurigene-AI/ChemBERTa-77M-MTR)** β DeepChem's chemical language model pretrained with multitask regression over 200 RDKit descriptors. A strong, cheap ADMET baseline. | |
| - **[MMELON multi-view 84M](https://huggingface.co/Aurigene-AI/biomed.sm.mv-te-84m)** β fuses SMILES, 2D graph and rendered image views of a molecule into one embedding for property prediction and virtual screening. | |
| ### Stage 04 β Synthesis planning | |
| Can we actually make it? | |
| - **[ReactionT5 v2 β forward](https://huggingface.co/Aurigene-AI/ReactionT5v2-forward)** β predicts products from reactants and reagents, trained on the Open Reaction Database. | |
| - **[ReactionT5 v2 β retrosynthesis](https://huggingface.co/Aurigene-AI/ReactionT5v2-retrosynthesis)** β single-step retrosynthesis. Chain it with the forward model to score a proposed route. | |
| ### Stage 05 β Evidence & literature | |
| Mine the papers that justify the programme. | |
| - **[BiomedBERT (PubMedBERT)](https://huggingface.co/Aurigene-AI/BiomedNLP-BiomedBERT-base-uncased-abstract-fulltext)** β Microsoft's encoder pretrained from scratch on PubMed abstracts plus PMC full text. | |
| - **[BioGPT](https://huggingface.co/Aurigene-AI/biogpt)** β generative biomedical LM for relation extraction and question answering over 15 M PubMed abstracts. | |
| - **[BioMistral-7B](https://huggingface.co/Aurigene-AI/BioMistral-7B)** β Mistral-7B further pretrained on PubMed Central Open Access; chat-capable across ten languages. | |
| - **[Biomedical NER (107 entities)](https://huggingface.co/Aurigene-AI/biomedical-ner-all)** β DistilBERT token classifier covering diseases, drugs, dosages, signs and lab values. | |
| - **[MolT5-large SMILES β text](https://huggingface.co/Aurigene-AI/molt5-large-smiles2caption)** β writes a natural-language description of a molecule from its SMILES. | |
| --- | |
| ## β‘ Quick start | |
| ```python | |
| from transformers import AutoModel, AutoTokenizer | |
| model_id = "Aurigene-AI/MoLFormer-XL-both-10pct" | |
| tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True) | |
| model = AutoModel.from_pretrained(model_id, trust_remote_code=True, deterministic_eval=True) | |
| smiles = ["CC(=O)Oc1ccccc1C(=O)O", "CC(C)Cc1ccc(cc1)C(C)C(=O)O"] | |
| embeddings = model(**tokenizer(smiles, padding=True, return_tensors="pt")).pooler_output | |
| print(embeddings.shape) # torch.Size([2, 768]) | |
| ``` | |
| ## π Collections | |
| Browse the catalogue as curated collections: [molecular representation & property prediction](https://huggingface.co/collections/Aurigene-AI/molecular-representation-and-property-prediction-6aa2203a94123133168b49e2), [generative chemistry & synthesis planning](https://huggingface.co/collections/Aurigene-AI/generative-chemistry-and-synthesis-planning-6aa2203c3d5b3527359ee444), [protein & target modeling](https://huggingface.co/collections/Aurigene-AI/protein-and-target-modeling-6aa2203e4e5bdf1f106c7d55), [biomedical language models](https://huggingface.co/collections/Aurigene-AI/biomedical-language-models-6aa2203f4469f54b61382c40), [interactive tools](https://huggingface.co/collections/Aurigene-AI/interactive-drug-discovery-tools-6aa220414e5bdf1f106c7dc0). | |
| ## π Attribution & licensing | |
| Every model here is a mirror of an open upstream release. The original authors β IBM Research, Meta AI, Microsoft Research, DeepChem, ChemFM, BioMistral and others β retain all credit, and each repository keeps the upstream model card and licence intact. Please cite the original work. | |
| Models and tools are provided for **research use**. Rule-based filters and predictions are triage heuristics, not statements about safety or efficacy, and nothing here is a medical device or clinical advice. | |
| π [aurigeneservices.com](https://www.aurigeneservices.com) | |