kite-67-12m / README.md
qikp's picture
Upload README.md with huggingface_hub
425d66e verified
|
Raw
History Blame Contribute Delete
559 Bytes
metadata
license: cc0-1.0
datasets:
  - microsoft/ms_marco
language:
  - en
pipeline_tag: text-generation
library_name: transformers

Kite

🎉 You are looking at Kite 6(.)7, which was trained on over 3x the data!

Kite is a small, trained, 12 million parameter language model.

Training

It was trained on 250K selected passages of MS MARCO, using 1 epoch, 32 batch size, 1e-3 learning rate, and the MicroSupra tokenizer.

Limitations

Due to its size, the model is not suitable for production workloads.