Feature Extraction
Safetensors
sentence-transformers
Transformers
furiosa-llm
qwen3
furiosa-ai
harrier-oss-v1
mteb
text-embeddings-inference
Instructions to use furiosa-ai/harrier-oss-v1-0.6b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use furiosa-ai/harrier-oss-v1-0.6b with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("furiosa-ai/harrier-oss-v1-0.6b") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Transformers
How to use furiosa-ai/harrier-oss-v1-0.6b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="furiosa-ai/harrier-oss-v1-0.6b")# Load model directly from transformers import AutoTokenizer, AutoModel tokenizer = AutoTokenizer.from_pretrained("furiosa-ai/harrier-oss-v1-0.6b") model = AutoModel.from_pretrained("furiosa-ai/harrier-oss-v1-0.6b", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Publish harrier-oss-v1-0.6b-b4e5762f15-7f200bf07a-2607200600.fxb
Browse files- README.md +206 -203
- harrier-oss-v1-0.6b-b4e5762f15-7f200bf07a-2607200600.fxb +3 -0
README.md
CHANGED
|
@@ -1,238 +1,241 @@
|
|
| 1 |
---
|
| 2 |
-
|
| 3 |
-
- mteb
|
| 4 |
-
- sentence-transformers
|
| 5 |
-
- transformers
|
| 6 |
language:
|
| 7 |
-
- multilingual
|
| 8 |
-
- af
|
| 9 |
-
- am
|
| 10 |
-
- ar
|
| 11 |
-
- as
|
| 12 |
-
- az
|
| 13 |
-
- be
|
| 14 |
-
- bg
|
| 15 |
-
- bn
|
| 16 |
-
- br
|
| 17 |
-
- bs
|
| 18 |
-
- ca
|
| 19 |
-
- cs
|
| 20 |
-
- cy
|
| 21 |
-
- da
|
| 22 |
-
- de
|
| 23 |
-
- el
|
| 24 |
-
- en
|
| 25 |
-
- eo
|
| 26 |
-
- es
|
| 27 |
-
- et
|
| 28 |
-
- eu
|
| 29 |
-
- fa
|
| 30 |
-
- fi
|
| 31 |
-
- fr
|
| 32 |
-
- fy
|
| 33 |
-
- ga
|
| 34 |
-
- gd
|
| 35 |
-
- gl
|
| 36 |
-
- gu
|
| 37 |
-
- ha
|
| 38 |
-
- he
|
| 39 |
-
- hi
|
| 40 |
-
- hr
|
| 41 |
-
- hu
|
| 42 |
-
- hy
|
| 43 |
-
- id
|
| 44 |
-
- is
|
| 45 |
-
- it
|
| 46 |
-
- ja
|
| 47 |
-
- jv
|
| 48 |
-
- ka
|
| 49 |
-
- kk
|
| 50 |
-
- km
|
| 51 |
-
- kn
|
| 52 |
-
- ko
|
| 53 |
-
- ku
|
| 54 |
-
- ky
|
| 55 |
-
- la
|
| 56 |
-
- lo
|
| 57 |
-
- lt
|
| 58 |
-
- lv
|
| 59 |
-
- mg
|
| 60 |
-
- mk
|
| 61 |
-
- ml
|
| 62 |
-
- mn
|
| 63 |
-
- mr
|
| 64 |
-
- ms
|
| 65 |
-
- my
|
| 66 |
-
- ne
|
| 67 |
-
- nl
|
| 68 |
-
- 'no'
|
| 69 |
-
- om
|
| 70 |
-
- or
|
| 71 |
-
- pa
|
| 72 |
-
- pl
|
| 73 |
-
- ps
|
| 74 |
-
- pt
|
| 75 |
-
- ro
|
| 76 |
-
- ru
|
| 77 |
-
- sa
|
| 78 |
-
- sd
|
| 79 |
-
- si
|
| 80 |
-
- sk
|
| 81 |
-
- sl
|
| 82 |
-
- so
|
| 83 |
-
- sq
|
| 84 |
-
- sr
|
| 85 |
-
- su
|
| 86 |
-
- sv
|
| 87 |
-
- sw
|
| 88 |
-
- ta
|
| 89 |
-
- te
|
| 90 |
-
- th
|
| 91 |
-
- tl
|
| 92 |
-
- tr
|
| 93 |
-
- ug
|
| 94 |
-
- uk
|
| 95 |
-
- ur
|
| 96 |
-
- uz
|
| 97 |
-
- vi
|
| 98 |
-
- xh
|
| 99 |
-
- yi
|
| 100 |
-
- zh
|
| 101 |
license: mit
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 102 |
---
|
|
|
|
| 103 |
|
| 104 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 105 |
|
| 106 |
-
|
| 107 |
-
The models use decoder-only architectures with last-token pooling and L2 normalization to produce dense text embeddings.
|
| 108 |
-
They can be applied to a wide range of tasks, including but not limited to **retrieval**, **clustering**, **semantic similarity**, **classification**, **bitext mining**, and **reranking**.
|
| 109 |
-
The models achieve state-of-the-art results on the [Multilingual MTEB v2](https://huggingface.co/spaces/mteb/leaderboard) benchmark as of the release date.
|
| 110 |
|
| 111 |
-
|
| 112 |
-
|
| 113 |
-
|
| 114 |
-
|
| 115 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 116 |
|
| 117 |
-
|
|
|
|
|
|
|
|
|
|
| 118 |
|
| 119 |
-
|
| 120 |
-
The 270m and 0.6b variants are additionally trained with knowledge distillation from larger embedding models.
|
| 121 |
|
| 122 |
-
|
| 123 |
|
| 124 |
-
|
| 125 |
|
| 126 |
-
|
|
|
|
| 127 |
|
| 128 |
-
|
| 129 |
-
from sentence_transformers import SentenceTransformer
|
| 130 |
|
| 131 |
-
model
|
|
|
|
|
|
|
|
|
|
| 132 |
|
| 133 |
-
|
| 134 |
-
"how much protein should a female eat",
|
| 135 |
-
"summit define",
|
| 136 |
-
]
|
| 137 |
-
documents = [
|
| 138 |
-
"As a general guideline, the CDC's average requirement of protein for women ages 19 to 70 is 46 grams per day. But, as you can see from this chart, you'll need to increase that if you're expecting or training for a marathon. Check out the chart below to see how much protein you should be eating each day.",
|
| 139 |
-
"Definition of summit for English Language Learners. : 1 the highest point of a mountain : the top of a mountain. : 2 the highest level. : 3 a meeting or series of meetings between the leaders of two or more governments."
|
| 140 |
-
]
|
| 141 |
|
| 142 |
-
|
| 143 |
-
document_embeddings = model.encode(documents)
|
| 144 |
|
| 145 |
-
|
| 146 |
-
|
|
|
|
| 147 |
```
|
| 148 |
|
| 149 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 150 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 151 |
|
| 152 |
-
|
|
|
|
| 153 |
|
| 154 |
```python
|
| 155 |
-
import
|
| 156 |
-
import torch.nn.functional as F
|
| 157 |
-
|
| 158 |
-
from torch import Tensor
|
| 159 |
-
from transformers import AutoTokenizer, AutoModel
|
| 160 |
-
|
| 161 |
-
|
| 162 |
-
def last_token_pool(last_hidden_states: Tensor, attention_mask: Tensor) -> Tensor:
|
| 163 |
-
left_padding = (attention_mask[:, -1].sum() == attention_mask.shape[0])
|
| 164 |
-
if left_padding:
|
| 165 |
-
return last_hidden_states[:, -1]
|
| 166 |
-
else:
|
| 167 |
-
sequence_lengths = attention_mask.sum(dim=1) - 1
|
| 168 |
-
batch_size = last_hidden_states.shape[0]
|
| 169 |
-
return last_hidden_states[torch.arange(batch_size, device=last_hidden_states.device), sequence_lengths]
|
| 170 |
-
|
| 171 |
-
|
| 172 |
-
def get_detailed_instruct(task_description: str, query: str) -> str:
|
| 173 |
-
return f'Instruct: {task_description}\nQuery: {query}'
|
| 174 |
-
|
| 175 |
-
|
| 176 |
-
# Each query must come with a one-sentence instruction that describes the task
|
| 177 |
-
task = 'Given a web search query, retrieve relevant passages that answer the query'
|
| 178 |
-
queries = [
|
| 179 |
-
get_detailed_instruct(task, 'how much protein should a female eat'),
|
| 180 |
-
get_detailed_instruct(task, 'summit define')
|
| 181 |
-
]
|
| 182 |
-
# No need to add instruction for retrieval documents
|
| 183 |
-
documents = [
|
| 184 |
-
"As a general guideline, the CDC's average requirement of protein for women ages 19 to 70 is 46 grams per day. But, as you can see from this chart, you'll need to increase that if you're expecting or training for a marathon. Check out the chart below to see how much protein you should be eating each day.",
|
| 185 |
-
"Definition of summit for English Language Learners. : 1 the highest point of a mountain : the top of a mountain. : 2 the highest level. : 3 a meeting or series of meetings between the leaders of two or more governments."
|
| 186 |
-
]
|
| 187 |
-
input_texts = queries + documents
|
| 188 |
-
|
| 189 |
-
tokenizer = AutoTokenizer.from_pretrained('microsoft/harrier-oss-v1-0.6b')
|
| 190 |
-
model = AutoModel.from_pretrained('microsoft/harrier-oss-v1-0.6b', dtype='auto')
|
| 191 |
-
model.eval()
|
| 192 |
-
model.cuda()
|
| 193 |
-
|
| 194 |
-
max_length = 32768
|
| 195 |
-
# Tokenize the input texts
|
| 196 |
-
batch_dict = tokenizer(input_texts, max_length=max_length, padding=True, truncation=True, return_tensors='pt')
|
| 197 |
-
batch_dict = {k: v.cuda() for k, v in batch_dict.items()}
|
| 198 |
-
|
| 199 |
-
outputs = model(**batch_dict)
|
| 200 |
-
embeddings = last_token_pool(outputs.last_hidden_state, batch_dict['attention_mask'])
|
| 201 |
-
|
| 202 |
-
# normalize embeddings
|
| 203 |
-
embeddings = F.normalize(embeddings, p=2, dim=1)
|
| 204 |
-
scores = (embeddings[:2] @ embeddings[2:].T) * 100
|
| 205 |
-
print(scores.tolist())
|
| 206 |
-
```
|
| 207 |
|
| 208 |
-
|
| 209 |
|
| 210 |
-
|
| 211 |
-
|
| 212 |
-
|
| 213 |
-
|
| 214 |
-
|
| 215 |
|
| 216 |
-
|
|
|
|
|
|
|
|
|
|
| 217 |
|
| 218 |
-
|
| 219 |
-
|
|
|
|
| 220 |
|
| 221 |
-
##
|
| 222 |
|
| 223 |
-
|
|
|
|
|
|
|
| 224 |
|
| 225 |
-
|
| 226 |
-
|
| 227 |
-
This is a way to customize text embeddings for different scenarios through natural language instructions.
|
| 228 |
|
| 229 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 230 |
|
| 231 |
-
|
|
|
|
|
|
|
| 232 |
|
| 233 |
-
|
|
|
|
|
|
|
| 234 |
|
| 235 |
-
|
| 236 |
|
| 237 |
-
|
| 238 |
-
|
|
|
|
|
|
| 1 |
---
|
| 2 |
+
base_model: microsoft/harrier-oss-v1-0.6b
|
|
|
|
|
|
|
|
|
|
| 3 |
language:
|
| 4 |
+
- multilingual
|
| 5 |
+
- af
|
| 6 |
+
- am
|
| 7 |
+
- ar
|
| 8 |
+
- as
|
| 9 |
+
- az
|
| 10 |
+
- be
|
| 11 |
+
- bg
|
| 12 |
+
- bn
|
| 13 |
+
- br
|
| 14 |
+
- bs
|
| 15 |
+
- ca
|
| 16 |
+
- cs
|
| 17 |
+
- cy
|
| 18 |
+
- da
|
| 19 |
+
- de
|
| 20 |
+
- el
|
| 21 |
+
- en
|
| 22 |
+
- eo
|
| 23 |
+
- es
|
| 24 |
+
- et
|
| 25 |
+
- eu
|
| 26 |
+
- fa
|
| 27 |
+
- fi
|
| 28 |
+
- fr
|
| 29 |
+
- fy
|
| 30 |
+
- ga
|
| 31 |
+
- gd
|
| 32 |
+
- gl
|
| 33 |
+
- gu
|
| 34 |
+
- ha
|
| 35 |
+
- he
|
| 36 |
+
- hi
|
| 37 |
+
- hr
|
| 38 |
+
- hu
|
| 39 |
+
- hy
|
| 40 |
+
- id
|
| 41 |
+
- is
|
| 42 |
+
- it
|
| 43 |
+
- ja
|
| 44 |
+
- jv
|
| 45 |
+
- ka
|
| 46 |
+
- kk
|
| 47 |
+
- km
|
| 48 |
+
- kn
|
| 49 |
+
- ko
|
| 50 |
+
- ku
|
| 51 |
+
- ky
|
| 52 |
+
- la
|
| 53 |
+
- lo
|
| 54 |
+
- lt
|
| 55 |
+
- lv
|
| 56 |
+
- mg
|
| 57 |
+
- mk
|
| 58 |
+
- ml
|
| 59 |
+
- mn
|
| 60 |
+
- mr
|
| 61 |
+
- ms
|
| 62 |
+
- my
|
| 63 |
+
- ne
|
| 64 |
+
- nl
|
| 65 |
+
- 'no'
|
| 66 |
+
- om
|
| 67 |
+
- or
|
| 68 |
+
- pa
|
| 69 |
+
- pl
|
| 70 |
+
- ps
|
| 71 |
+
- pt
|
| 72 |
+
- ro
|
| 73 |
+
- ru
|
| 74 |
+
- sa
|
| 75 |
+
- sd
|
| 76 |
+
- si
|
| 77 |
+
- sk
|
| 78 |
+
- sl
|
| 79 |
+
- so
|
| 80 |
+
- sq
|
| 81 |
+
- sr
|
| 82 |
+
- su
|
| 83 |
+
- sv
|
| 84 |
+
- sw
|
| 85 |
+
- ta
|
| 86 |
+
- te
|
| 87 |
+
- th
|
| 88 |
+
- tl
|
| 89 |
+
- tr
|
| 90 |
+
- ug
|
| 91 |
+
- uk
|
| 92 |
+
- ur
|
| 93 |
+
- uz
|
| 94 |
+
- vi
|
| 95 |
+
- xh
|
| 96 |
+
- yi
|
| 97 |
+
- zh
|
| 98 |
license: mit
|
| 99 |
+
pipeline_tag: feature-extraction
|
| 100 |
+
library_name: furiosa-llm
|
| 101 |
+
tags:
|
| 102 |
+
- furiosa-ai
|
| 103 |
+
- harrier-oss-v1
|
| 104 |
+
- qwen3
|
| 105 |
+
- mteb
|
| 106 |
+
- sentence-transformers
|
| 107 |
+
- transformers
|
| 108 |
---
|
| 109 |
+
# harrier-oss-v1-0.6b
|
| 110 |
|
| 111 |
+
This repository contains [`microsoft/harrier-oss-v1-0.6b`](https://huggingface.co/microsoft/harrier-oss-v1-0.6b)
|
| 112 |
+
together with a Furiosa Executable Bundle (FXB) for running it on
|
| 113 |
+
[FuriosaAI RNGD](https://furiosa.ai) with [Furiosa-LLM](https://developer.furiosa.ai/latest/en/furiosa_llm/intro.html).
|
| 114 |
+
The same model also runs on other frameworks (such as Sentence Transformers and
|
| 115 |
+
Transformers); for usage with those, see the upstream
|
| 116 |
+
[`microsoft/harrier-oss-v1-0.6b`](https://huggingface.co/microsoft/harrier-oss-v1-0.6b) model card.
|
| 117 |
|
| 118 |
+
## Overview
|
|
|
|
|
|
|
|
|
|
| 119 |
|
| 120 |
+
Harrier OSS v1 is a family of multilingual text-embedding models developed by
|
| 121 |
+
Microsoft. The 0.6B model uses a dense, decoder-only Qwen3 architecture, but it
|
| 122 |
+
is trained with Harrier's own multilingual, instruction-aware embedding recipe
|
| 123 |
+
rather than the Qwen3-Embedding training recipe. It produces 1,024-dimensional
|
| 124 |
+
embeddings through last-token pooling and L2 normalization. It is designed for
|
| 125 |
+
retrieval, clustering, semantic similarity, classification, bitext mining, and
|
| 126 |
+
reranking. Its intended use is the same as the upstream
|
| 127 |
+
[`microsoft/harrier-oss-v1-0.6b`](https://huggingface.co/microsoft/harrier-oss-v1-0.6b),
|
| 128 |
+
and it is released under the [MIT License](https://opensource.org/license/mit).
|
| 129 |
|
| 130 |
+
- **Architecture:** Qwen3 (dense), `Qwen3Model`
|
| 131 |
+
- **Input / Output:** Text / Embeddings (vector)
|
| 132 |
+
- **Supported Inference Engine:** Furiosa LLM
|
| 133 |
+
- **Supported Hardware:** FuriosaAI RNGD
|
| 134 |
|
| 135 |
+
### Quantization
|
|
|
|
| 136 |
|
| 137 |
+
No quantization — the model runs in its native BF16 precision.
|
| 138 |
|
| 139 |
+
### Parallelism Strategy
|
| 140 |
|
| 141 |
+
On RNGD, harrier-oss-v1-0.6b runs with a **tensor-parallel size of 8 PEs**, which
|
| 142 |
+
maps to a **single RNGD card** (8 PEs per card).
|
| 143 |
|
| 144 |
+
## Usage
|
|
|
|
| 145 |
|
| 146 |
+
To run this model with Furiosa-LLM, follow the examples below after
|
| 147 |
+
[installing Furiosa-LLM and its prerequisites](https://developer.furiosa.ai/latest/en/get_started/furiosa_llm.html#installing-furiosa-llm).
|
| 148 |
+
You can use the model either online through the OpenAI-compatible server or
|
| 149 |
+
offline through the Furiosa-LLM Python API.
|
| 150 |
|
| 151 |
+
### Launch the server
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 152 |
|
| 153 |
+
Serve the model by passing its `furiosa-ai/<repo>` identifier:
|
|
|
|
| 154 |
|
| 155 |
+
```sh
|
| 156 |
+
# Launch the server, listening on port 8000 by default
|
| 157 |
+
furiosa-llm serve furiosa-ai/harrier-oss-v1-0.6b
|
| 158 |
```
|
| 159 |
|
| 160 |
+
When the server is ready, you will see:
|
| 161 |
+
|
| 162 |
+
```sh
|
| 163 |
+
INFO: Started server process [27507]
|
| 164 |
+
INFO: Waiting for application startup.
|
| 165 |
+
INFO: Application startup complete.
|
| 166 |
+
INFO: Uvicorn running on http://0.0.0.0:8000 (Press CTRL+C to quit)
|
| 167 |
+
```
|
| 168 |
|
| 169 |
+
### Basic Usage
|
| 170 |
+
|
| 171 |
+
The server exposes an OpenAI-compatible `/v1/embeddings` endpoint. Harrier is
|
| 172 |
+
instruction-aware: prepend a one-sentence task description to each query in the
|
| 173 |
+
`Instruct: ...\nQuery: ...` format, and do not add the instruction to documents.
|
| 174 |
+
For more details, see the
|
| 175 |
+
[base model card](https://huggingface.co/microsoft/harrier-oss-v1-0.6b).
|
| 176 |
+
Request embeddings with `curl`:
|
| 177 |
+
|
| 178 |
+
```sh
|
| 179 |
+
curl http://localhost:8000/v1/embeddings \
|
| 180 |
+
-H "Content-Type: application/json" \
|
| 181 |
+
-d '{
|
| 182 |
+
"model": "furiosa-ai/harrier-oss-v1-0.6b",
|
| 183 |
+
"input": [
|
| 184 |
+
"Instruct: Given a web search query, retrieve relevant passages that answer the query\nQuery: summit define",
|
| 185 |
+
"Definition of summit: the highest point of a mountain."
|
| 186 |
+
]
|
| 187 |
+
}' \
|
| 188 |
+
| python -m json.tool
|
| 189 |
+
```
|
| 190 |
|
| 191 |
+
Because the endpoint is OpenAI-compatible, you can also use the OpenAI Python
|
| 192 |
+
client:
|
| 193 |
|
| 194 |
```python
|
| 195 |
+
from openai import OpenAI
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 196 |
|
| 197 |
+
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
|
| 198 |
|
| 199 |
+
query = (
|
| 200 |
+
"Instruct: Given a web search query, retrieve relevant passages that answer the query\n"
|
| 201 |
+
"Query: summit define"
|
| 202 |
+
)
|
| 203 |
+
document = "Definition of summit: the highest point of a mountain."
|
| 204 |
|
| 205 |
+
response = client.embeddings.create(
|
| 206 |
+
model="furiosa-ai/harrier-oss-v1-0.6b",
|
| 207 |
+
input=[query, document],
|
| 208 |
+
)
|
| 209 |
|
| 210 |
+
for data in response.data:
|
| 211 |
+
print(f"Index {data.index}: {len(data.embedding)} dimensions")
|
| 212 |
+
```
|
| 213 |
|
| 214 |
+
### Advanced Usage
|
| 215 |
|
| 216 |
+
For offline use, load the model with the `LLM` constructor (the FXB shipped in
|
| 217 |
+
the repo is discovered automatically) and call `embed` to obtain L2-normalized
|
| 218 |
+
dense vectors. Their dot product is therefore the cosine similarity:
|
| 219 |
|
| 220 |
+
```python
|
| 221 |
+
from furiosa_llm import LLM
|
|
|
|
| 222 |
|
| 223 |
+
query = (
|
| 224 |
+
"Instruct: Given a web search query, retrieve relevant passages that answer the query\n"
|
| 225 |
+
"Query: summit define"
|
| 226 |
+
)
|
| 227 |
+
document = "Definition of summit: the highest point of a mountain."
|
| 228 |
|
| 229 |
+
with LLM("furiosa-ai/harrier-oss-v1-0.6b") as llm:
|
| 230 |
+
outputs = llm.embed([query, document])
|
| 231 |
+
embeddings = [output.outputs.embedding for output in outputs]
|
| 232 |
|
| 233 |
+
similarity = sum(a * b for a, b in zip(*embeddings, strict=True))
|
| 234 |
+
print(f"Cosine similarity: {similarity:.4f}")
|
| 235 |
+
```
|
| 236 |
|
| 237 |
+
## Learn more
|
| 238 |
|
| 239 |
+
* [Furiosa-LLM Server (`furiosa-llm serve`)](https://developer.furiosa.ai/latest/en/furiosa_llm/furiosa-llm-serve.html) — full OpenAI-compatible API reference, including the Embeddings API
|
| 240 |
+
* [Furiosa-LLM](https://developer.furiosa.ai/latest/en/furiosa_llm/intro.html) — Furiosa-LLM documentation and API reference
|
| 241 |
+
* [`microsoft/harrier-oss-v1-0.6b`](https://huggingface.co/microsoft/harrier-oss-v1-0.6b) — upstream model card
|
harrier-oss-v1-0.6b-b4e5762f15-7f200bf07a-2607200600.fxb
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:4c90ec2ad6ac655c376e07c0e82b7bd6a41d33dba1fad49d434522350465ea16
|
| 3 |
+
size 19796989
|