Instructions to use TigreGotico/opus-mt-ja-en-onnx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use TigreGotico/opus-mt-ja-en-onnx with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "translation" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # 'pip install "transformers<5.0.0' from transformers import pipeline pipe = pipeline("translation", model="TigreGotico/opus-mt-ja-en-onnx")# Load model directly from transformers import AutoTokenizer, AutoModelForSeq2SeqLM tokenizer = AutoTokenizer.from_pretrained("TigreGotico/opus-mt-ja-en-onnx") model = AutoModelForSeq2SeqLM.from_pretrained("TigreGotico/opus-mt-ja-en-onnx", device_map="auto") - Notebooks
- Google Colab
- Kaggle
opus-mt-ja-en-onnx
ONNX export (fp32 + dynamic int8 quantized) of Helsinki-NLP/opus-mt-ja-en, a Marian translation model from the Helsinki-NLP OPUS-MT project.
License: apache-2.0 (inherited from the base model; verify at the source link above).
Export
optimum-cli export onnx --model Helsinki-NLP/opus-mt-ja-en --task text2text-generation-with-past <out>
Quantized to int8 with onnxruntime.quantization.quantize_dynamic (QUInt8 weights).
File layout
./ fp32 ONNX graphs (encoder_model.onnx, decoder_model.onnx, decoder_with_past_model.onnx) + tokenizer files
./int8/ int8 dynamic-quantized ONNX graphs
fp32 size: ~1187.5 MB | int8 size: ~552.0 MB
Parity check
Compared PyTorch (MarianMTModel) vs ONNX fp32 (ORTModelForSeq2SeqLM) on 2 sentences, greedy and beam=4 (max_new_tokens=64). Overall: greedy PASS, beam4 PASS.
- src: これは本です。
- pytorch greedy: This is a book.
- onnx fp32 greedy: This is a book. (match)
- pytorch beam4: This is a book.
- onnx fp32 beam4: This is a book. (match)
- onnx int8 greedy: This is a book.
- src: 私は日本語を勉強しています。
- pytorch greedy: I'm studying Japanese.
- onnx fp32 greedy: I'm studying Japanese. (match)
- pytorch beam4: I'm studying Japanese.
- onnx fp32 beam4: I'm studying Japanese. (match)
- onnx int8 greedy: I'm studying Japanese.
Usage
from optimum.onnxruntime import ORTModelForSeq2SeqLM
from transformers import AutoTokenizer
repo = "TigreGotico/opus-mt-ja-en-onnx"
tok = AutoTokenizer.from_pretrained(repo)
model = ORTModelForSeq2SeqLM.from_pretrained(repo) # fp32; pass subfolder="int8" for the quantized graphs
inputs = tok("これは本です。", return_tensors="pt")
out = model.generate(**inputs, num_beams=4, max_new_tokens=64)
print(tok.decode(out[0], skip_special_tokens=True))
Exported for the OVOS / TigreGotico offline translation stack.
- Downloads last month
- 9
Model tree for TigreGotico/opus-mt-ja-en-onnx
Base model
Helsinki-NLP/opus-mt-ja-en