Instructions to use ctheodoris/Geneformer with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ctheodoris/Geneformer with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("fill-mask", model="ctheodoris/Geneformer")# Load model directly from transformers import AutoTokenizer, AutoModelForMaskedLM tokenizer = AutoTokenizer.from_pretrained("ctheodoris/Geneformer") model = AutoModelForMaskedLM.from_pretrained("ctheodoris/Geneformer", device_map="auto") - Inference
- Notebooks
- Google Colab
- Kaggle
Commit History
tokenizer zarr integration (#561) 91215c4 verified
add input_identifier to tokenize specific matched files ac59e36
Christina Theodoris commited on
update with V2 models d319fef
Christina Theodoris commited on
Add checks for custom attributes and n_counts prior to sum ensembl id (#461) 09de197 verified
Update geneformer/tokenizer.py (#450) 664f71e verified
Update geneformer/tokenizer.py (#415) 63275a8 verified
edit docs formatting ef094b2
update tokenizer to defaults for 95M models for special token and input size da8cf3d verified
precommit formatting f07bfd7
Add function for summing of Ensembl IDs (#377) 1e18102 verified
move dicts to init ea428cb
update tokenizer to include eos token ead0550
Christina Theodoris commited on
fix cell state gene embeddings bug (#345) c0e7b19 verified
patch datasets save_to_disk 75c67a1
Christina Theodoris commited on
correct typo 5a43832 verified
Update readthedocs for classifier f75f5ac
Christina Theodoris commited on
Get the gene keys and gene list keys from the token dictionary instead of medians (#304) b294421 verified
Add classifier module and examples 9e9cca9
Christina Theodoris commited on
Fix typo (#301) 075bd53 verified
Add option for variable input_size and to add CLS/SEP Tokens (#299) aa25cd2 verified
edit docstring format to highlight options e3330a6
Christina Theodoris commited on
change doc formatting 17f036a
Christina Theodoris commited on
add sphinx docs 2a0dcbe
Christina Theodoris commited on
Add option for modified batch size for loom tokenizer 0960cf6
Christina Theodoris commited on
Add option for modifying chunk size for anndata tokenizer fd93ebf
Christina Theodoris commited on
tokenizer-uncropped-input_ids (#275) 8df5dc1
anndata_tokenizer (#170) 4302f48
Add error for no files found and suppress loompy import warning abdf980
Christina Theodoris commited on
Update tokenizer to allow tokenization without custom cell attributes 57b9778
Christina Theodoris commited on
Modify tokenizer to allow renaming attr names btwn loom and .dataset e78c44d
Christina Theodoris commited on
Add further explanation regarding input file format for transcriptome tokenizer c34ead6
Christina Theodoris commited on
Add further explanation to tokenizer example script and updated tokenizer to match loompy raised error 78dd83b
Christina Theodoris commited on
Fix bug with metadata when processing multiple .loom files (#3) 044d737
Add data collator for cell classification and example for cell classification 088ea6e
Christina Theodoris commited on
Add Geneformer tokenizer and updated model card 5426788
Christina Theodoris commited on