Text Generation
Transformers
Safetensors
PyTorch
English
gpt2
trained-from-scratch
text-completion
english
sangraha
text-generation-inference
Instructions to use sraivante/Custom-GPT-40M-Base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use sraivante/Custom-GPT-40M-Base with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="sraivante/Custom-GPT-40M-Base")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("sraivante/Custom-GPT-40M-Base") model = AutoModelForCausalLM.from_pretrained("sraivante/Custom-GPT-40M-Base", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use sraivante/Custom-GPT-40M-Base with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "sraivante/Custom-GPT-40M-Base" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sraivante/Custom-GPT-40M-Base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/sraivante/Custom-GPT-40M-Base
- SGLang
How to use sraivante/Custom-GPT-40M-Base with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "sraivante/Custom-GPT-40M-Base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sraivante/Custom-GPT-40M-Base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "sraivante/Custom-GPT-40M-Base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sraivante/Custom-GPT-40M-Base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use sraivante/Custom-GPT-40M-Base with Docker Model Runner:
docker model run hf.co/sraivante/Custom-GPT-40M-Base
Download example_generate.py from sraivante/Custom-GPT-40M-Base: direct link, hf CLI and curl.
- Browser
- Download file 1.77 kB
-
https://huggingface.co/sraivante/Custom-GPT-40M-Base/resolve/main/example_generate.py
- Command line
-
hf download hf://sraivante/Custom-GPT-40M-Base/example_generate.py
-
curl -L -o example_generate.py https://huggingface.co/sraivante/Custom-GPT-40M-Base/resolve/main/example_generate.py
1.77 kB
| """Generate a short English continuation with Custom GPT 40M. | |
| Copyright (c) 2026 sraivante. SPDX-License-Identifier: Apache-2.0 | |
| """ | |
| import argparse | |
| import torch | |
| from transformers import AutoTokenizer, AutoModelForCausalLM | |
| def main(): | |
| parser=argparse.ArgumentParser(description=__doc__) | |
| parser.add_argument("--model",default="sraivante/Custom-GPT-40M-Base") | |
| parser.add_argument("--revision",default="main",help="Pin a Hub commit for reproducibility.") | |
| parser.add_argument("--prompt",default="The future of artificial intelligence") | |
| parser.add_argument("--max-new-tokens",type=int,default=80) | |
| parser.add_argument("--seed",type=int,default=42) | |
| parser.add_argument("--greedy",action="store_true") | |
| parser.add_argument("--device",default="cpu",choices=["cpu","cuda"]) | |
| args=parser.parse_args() | |
| if not args.prompt.strip():parser.error("Provide a non-empty English prompt.") | |
| if args.max_new_tokens<1:parser.error("--max-new-tokens must be positive.") | |
| tokenizer=AutoTokenizer.from_pretrained(args.model,revision=args.revision) | |
| model=AutoModelForCausalLM.from_pretrained(args.model,revision=args.revision).to(args.device).eval() | |
| inputs=tokenizer(args.prompt,return_tensors="pt",add_special_tokens=False).to(args.device) | |
| remaining=model.config.n_positions-inputs["input_ids"].shape[1] | |
| if remaining<1:parser.error("The prompt must be shorter than 256 tokens.") | |
| torch.manual_seed(args.seed) | |
| options={} if args.greedy else {"temperature":0.8,"top_k":50} | |
| with torch.inference_mode(): | |
| output=model.generate(**inputs,max_new_tokens=min(args.max_new_tokens,remaining),do_sample=not args.greedy,**options) | |
| print(tokenizer.decode(output[0],skip_special_tokens=True)) | |
| if __name__=="__main__":main() | |