Instructions to use AGofficial/AgGPT-9 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use AGofficial/AgGPT-9 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf AGofficial/AgGPT-9 # Run inference directly in the terminal: llama cli -hf AGofficial/AgGPT-9
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf AGofficial/AgGPT-9 # Run inference directly in the terminal: llama cli -hf AGofficial/AgGPT-9
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf AGofficial/AgGPT-9 # Run inference directly in the terminal: ./llama-cli -hf AGofficial/AgGPT-9
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf AGofficial/AgGPT-9 # Run inference directly in the terminal: ./build/bin/llama-cli -hf AGofficial/AgGPT-9
Use Docker
docker model run hf.co/AGofficial/AgGPT-9
- LM Studio
- Jan
- Ollama
How to use AGofficial/AgGPT-9 with Ollama:
ollama run hf.co/AGofficial/AgGPT-9
- Unsloth Desktop
- Docker Model Runner
How to use AGofficial/AgGPT-9 with Docker Model Runner:
docker model run hf.co/AGofficial/AgGPT-9
- Lemonade
How to use AGofficial/AgGPT-9 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull AGofficial/AgGPT-9
Run and chat with the model
lemonade run user.AgGPT-9-{{QUANT_TAG}}List all available models
lemonade list
- Atomic Chat
File size: 833 Bytes
3f21cf1 0f7f181 3f21cf1 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 | from llama_cpp import Llama
def run_local_llm():
print("Loading AgGPT-9... (This may take a moment)")
model_path = "./AgGPT-9.gguf"
model = Llama(model_path=model_path, n_ctx=2048, n_gpu_layers=35)
print("Model loaded. Type 'exit' to quit.")
while True:
prompt = input("\nEnter your prompt: ")
if prompt.lower() == 'exit':
break
messages = [
{"role": "system", "content": "You are AgGPT-9, an advanced AI assistant created by AG, the 9th series of the AgGPT models."},
{"role": "user", "content": prompt}
]
output = model.create_chat_completion(messages, max_tokens=550, temperature=0.7)
print("\nGenerated text:")
print(output["choices"][0]["message"]["content"])
if __name__ == "__main__":
run_local_llm()
|