Instructions to use xformAI/opt-6.7b-ub-16-qcqa-best-for-KV-cache with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use xformAI/opt-6.7b-ub-16-qcqa-best-for-KV-cache with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("xformAI/opt-6.7b-ub-16-qcqa-best-for-KV-cache", device_map="auto") - Notebooks
- Google Colab
- Kaggle
This is a QCQA version of the original model facebook/opt-125m. In this version, the original MHA architecture is preserved but instead of having a single K/V head, different K/V heads corresponding to the same group have the same mean-pooled K or V values. It has upto 16 groups of KV heads per layer instead of the original 32 KV heads in the MHA implementation. This implementation is supposed to more efficient than corresponding GQA one. This has been optimized for KV-cache.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support