How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf neopolita/trilm_3.9b_unpacked-gguf:
# Run inference directly in the terminal:
llama cli -hf neopolita/trilm_3.9b_unpacked-gguf:
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf neopolita/trilm_3.9b_unpacked-gguf:
# Run inference directly in the terminal:
llama cli -hf neopolita/trilm_3.9b_unpacked-gguf:
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf neopolita/trilm_3.9b_unpacked-gguf:
# Run inference directly in the terminal:
./llama-cli -hf neopolita/trilm_3.9b_unpacked-gguf:
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf neopolita/trilm_3.9b_unpacked-gguf:
# Run inference directly in the terminal:
./build/bin/llama-cli -hf neopolita/trilm_3.9b_unpacked-gguf:
Use Docker
docker model run hf.co/neopolita/trilm_3.9b_unpacked-gguf:
Quick Links

GGUF quants for SpectraSuite/TriLM_3.9B_Unpacked using llama.cpp

Terms of Use: Please check the original model

cthulhu

Quants

  • q2_k: Uses Q4_K for the attention.vw and feed_forward.w2 tensors, Q2_K for the other tensors.
  • q3_k_s: Uses Q3_K for all tensors
  • q3_k_m: Uses Q4_K for the attention.wv, attention.wo, and feed_forward.w2 tensors, else Q3_K
  • q3_k_l: Uses Q5_K for the attention.wv, attention.wo, and feed_forward.w2 tensors, else Q3_K
  • q4_0: Original quant method, 4-bit.
  • q4_1: Higher accuracy than q4_0 but not as high as q5_0. However has quicker inference than q5 models.
  • q4_k_s: Uses Q4_K for all tensors
  • q4_k_m: Uses Q6_K for half of the attention.wv and feed_forward.w2 tensors, else Q4_K
  • q5_0: Higher accuracy, higher resource usage and slower inference.
  • q5_1: Even higher accuracy, resource usage and slower inference.
  • q5_k_s: Uses Q5_K for all tensors
  • q5_k_m: Uses Q6_K for half of the attention.wv and feed_forward.w2 tensors, else Q5_K
  • q6_k: Uses Q8_K for all tensors
  • q8_0: Almost indistinguishable from float16. High resource use and slow. Not recommended for most users.
Downloads last month
53
GGUF
Model size
4B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including neopolita/trilm_3.9b_unpacked-gguf