Instructions to use aisingapore/SEA-LION-v1-7B-IT with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use aisingapore/SEA-LION-v1-7B-IT with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="aisingapore/SEA-LION-v1-7B-IT", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("aisingapore/SEA-LION-v1-7B-IT", trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained("aisingapore/SEA-LION-v1-7B-IT", trust_remote_code=True, device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use aisingapore/SEA-LION-v1-7B-IT with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "aisingapore/SEA-LION-v1-7B-IT" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "aisingapore/SEA-LION-v1-7B-IT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/aisingapore/SEA-LION-v1-7B-IT
- SGLang
How to use aisingapore/SEA-LION-v1-7B-IT with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "aisingapore/SEA-LION-v1-7B-IT" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "aisingapore/SEA-LION-v1-7B-IT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "aisingapore/SEA-LION-v1-7B-IT" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "aisingapore/SEA-LION-v1-7B-IT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use aisingapore/SEA-LION-v1-7B-IT with Docker Model Runner:
docker model run hf.co/aisingapore/SEA-LION-v1-7B-IT
Unclear instructions
#3
by agoudarzi - opened
I deployed this model on the Vertex AI but I get Incomplete generation error:
prompt = """
### USER:\n Apa sentimen dari kalimat berikut ini?
Kalimat: Buku ini sangat membosankan.
Jawaban: \n\n### RESPONSE:\n
"""
messages = { "instances" : [ { "inputs" : prompt}],
"parameters" : { "max_new_tokens" : 100} }
headers = {
"Content-Type": "application/json",
"Authorization" : f"Bearer {token}"
}
print( headers)
url = f"https://us-central1-aiplatform.googleapis.com/v1/projects/{PROJECT_ID}/locations/us-central1/endpoints/{ENDPOINT_ID}:predict"
print(url)
t0 = time.time()
results = requests.post(url, json=messages, headers=headers)
t1 = time.time()
print(results.content)
print(t1-t0)
Dear @agoudarzi ,
Thank you for your interests in SEA-LION.
Could you elaborate more on the error you have encountered by sharing any error messages you have received when running this code?
Or do you meant the generated text was incomplete? If yes, could you kindly increase the value for max_new_tokens to a higher value (e.g. 1000) and see if the model is able to completely answer your query?
Thank you
Raymond