Instructions to use nisten/qwenv2-7b-inst-imatrix-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use nisten/qwenv2-7b-inst-imatrix-gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf nisten/qwenv2-7b-inst-imatrix-gguf:F16_OUTPUTF # Run inference directly in the terminal: llama cli -hf nisten/qwenv2-7b-inst-imatrix-gguf:F16_OUTPUTF
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf nisten/qwenv2-7b-inst-imatrix-gguf:F16_OUTPUTF # Run inference directly in the terminal: llama cli -hf nisten/qwenv2-7b-inst-imatrix-gguf:F16_OUTPUTF
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf nisten/qwenv2-7b-inst-imatrix-gguf:F16_OUTPUTF # Run inference directly in the terminal: ./llama-cli -hf nisten/qwenv2-7b-inst-imatrix-gguf:F16_OUTPUTF
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf nisten/qwenv2-7b-inst-imatrix-gguf:F16_OUTPUTF # Run inference directly in the terminal: ./build/bin/llama-cli -hf nisten/qwenv2-7b-inst-imatrix-gguf:F16_OUTPUTF
Use Docker
docker model run hf.co/nisten/qwenv2-7b-inst-imatrix-gguf:F16_OUTPUTF
- LM Studio
- Jan
- Ollama
How to use nisten/qwenv2-7b-inst-imatrix-gguf with Ollama:
ollama run hf.co/nisten/qwenv2-7b-inst-imatrix-gguf:F16_OUTPUTF
- Unsloth Desktop
- Docker Model Runner
How to use nisten/qwenv2-7b-inst-imatrix-gguf with Docker Model Runner:
docker model run hf.co/nisten/qwenv2-7b-inst-imatrix-gguf:F16_OUTPUTF
- Lemonade
How to use nisten/qwenv2-7b-inst-imatrix-gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull nisten/qwenv2-7b-inst-imatrix-gguf:F16_OUTPUTF
Run and chat with the model
lemonade run user.qwenv2-7b-inst-imatrix-gguf-F16_OUTPUTF
List all available models
lemonade list
- Atomic Chat
Upload 9 files
Browse filesAll my experiments that produced good results
- .gitattributes +9 -0
- qwen7bf16.gguf +3 -0
- qwen7bq4kembeddingbf16outputbf16.gguf +3 -0
- qwen7bq4koutput8bit.gguf +3 -0
- qwen7bq4xs.gguf +3 -0
- qwen7bq4xsembedding5bitkoutput8bit.gguf +3 -0
- qwen7bq4xsoutput8bit.gguf +3 -0
- qwen7bq5km.gguf +3 -0
- qwenq8bitimatrix.dat +3 -0
- qwenq8v2.gguf +3 -0
.gitattributes
CHANGED
|
@@ -33,3 +33,12 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
+
qwen7bf16.gguf filter=lfs diff=lfs merge=lfs -text
|
| 37 |
+
qwen7bq4kembeddingbf16outputbf16.gguf filter=lfs diff=lfs merge=lfs -text
|
| 38 |
+
qwen7bq4koutput8bit.gguf filter=lfs diff=lfs merge=lfs -text
|
| 39 |
+
qwen7bq4xs.gguf filter=lfs diff=lfs merge=lfs -text
|
| 40 |
+
qwen7bq4xsembedding5bitkoutput8bit.gguf filter=lfs diff=lfs merge=lfs -text
|
| 41 |
+
qwen7bq4xsoutput8bit.gguf filter=lfs diff=lfs merge=lfs -text
|
| 42 |
+
qwen7bq5km.gguf filter=lfs diff=lfs merge=lfs -text
|
| 43 |
+
qwenq8bitimatrix.dat filter=lfs diff=lfs merge=lfs -text
|
| 44 |
+
qwenq8v2.gguf filter=lfs diff=lfs merge=lfs -text
|
qwen7bf16.gguf
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:bcfd06d4f9911b033ef0574937de75878e4d7719bd89966203162a31b1a39f6f
|
| 3 |
+
size 15237850656
|
qwen7bq4kembeddingbf16outputbf16.gguf
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:74f0117f8970e9eb4c040296b773268db4c72827ea5546a6a146c48e56a3a1ea
|
| 3 |
+
size 6109431552
|
qwen7bq4koutput8bit.gguf
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:de2963991d78d1937c31a580539e3bf541e06f1ab6d198c5172997a176381155
|
| 3 |
+
size 4815062784
|
qwen7bq4xs.gguf
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:de3b2365e61a7834e8403516954a6665363c66ce0f4db8af7b4e77bb9a37c00a
|
| 3 |
+
size 4218470144
|
qwen7bq4xsembedding5bitkoutput8bit.gguf
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:f23da1e0d3619ea20ebc77ec0627fe278e9164b255bf98a6d9fae713475c29ad
|
| 3 |
+
size 4639991552
|
qwen7bq4xsoutput8bit.gguf
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:8b6722c2295e8d4bdea93990f2ef3f746366fad81631dcb5ab52759d8525a504
|
| 3 |
+
size 4350461696
|
qwen7bq5km.gguf
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:c50c9226270edee91afbb988f84bbe0456a7e78bdfdfb6127af1ae22cca3bf15
|
| 3 |
+
size 5576820480
|
qwenq8bitimatrix.dat
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:690c315c1cac97e20ce87473b9c98df2b6c4c26f60c0a1d9564b63aeb4f33003
|
| 3 |
+
size 4536673
|
qwenq8v2.gguf
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:16b4385ad242aa8276e9ec865c6500552fb5d5fa80b62dd77eee4ce8eb7e9a20
|
| 3 |
+
size 8098522656
|