Instructions to use TheBloke/TinyLlama-1.1B-Chat-v1.0-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use TheBloke/TinyLlama-1.1B-Chat-v1.0-GGUF with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("TheBloke/TinyLlama-1.1B-Chat-v1.0-GGUF", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use TheBloke/TinyLlama-1.1B-Chat-v1.0-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf TheBloke/TinyLlama-1.1B-Chat-v1.0-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf TheBloke/TinyLlama-1.1B-Chat-v1.0-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf TheBloke/TinyLlama-1.1B-Chat-v1.0-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf TheBloke/TinyLlama-1.1B-Chat-v1.0-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf TheBloke/TinyLlama-1.1B-Chat-v1.0-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf TheBloke/TinyLlama-1.1B-Chat-v1.0-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf TheBloke/TinyLlama-1.1B-Chat-v1.0-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf TheBloke/TinyLlama-1.1B-Chat-v1.0-GGUF:Q4_K_M
Use Docker
docker model run hf.co/TheBloke/TinyLlama-1.1B-Chat-v1.0-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use TheBloke/TinyLlama-1.1B-Chat-v1.0-GGUF with Ollama:
ollama run hf.co/TheBloke/TinyLlama-1.1B-Chat-v1.0-GGUF:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use TheBloke/TinyLlama-1.1B-Chat-v1.0-GGUF with Docker Model Runner:
docker model run hf.co/TheBloke/TinyLlama-1.1B-Chat-v1.0-GGUF:Q4_K_M
- Lemonade
How to use TheBloke/TinyLlama-1.1B-Chat-v1.0-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull TheBloke/TinyLlama-1.1B-Chat-v1.0-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.TinyLlama-1.1B-Chat-v1.0-GGUF-Q4_K_M
List all available models
lemonade list
- Atomic Chat
Python bindings not working
I have never been able to get a decent response out of any library other than kobold or llama cpp (not llama cpp python) and since I work with python a lot I tried ctransformers as well which is the worst to be used (in my experience). However, I am finding it very difficult to community with model using kobold or llama cpp (server) when used as drop in replacement for openai's api.
here is what i tried:
1> run server with command: ./server -m tinyllama.gguf
2> (on different cmd/tab) run openai's replacement: python api_like_OAI.py # must have flask installed
or just run kobold you'll have an endpoint
3> use following code:
from langchain_openai import OpenAI
from langchain.prompts import PromptTemplate
from langchain.chains import LLMChain
llm = OpenAI(openai_api_base="http://10.192.4.242:8081/v1", openai_api_key="somethig")
question = "How many planets are there in our solar system?"
template = """Question: {question}
Answer: Let's think step by step."""
prompt = PromptTemplate(template=template, input_variables=["question"])
llm_chain = LLMChain(prompt=prompt, llm=llm)
response = llm_chain.invoke(question, max_tokens=10)
print(response)
If however i use any model with llama_cpp_python then I get very weird output, which i tried with different models (all quantized) with different prompts. nothing worked :'(