Instructions to use tencent/HY-MT1.5-1.8B-FP8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use tencent/HY-MT1.5-1.8B-FP8 with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "translation" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # pip install "transformers<5.0.0" from transformers import pipeline pipe = pipeline("translation", model="tencent/HY-MT1.5-1.8B-FP8")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("tencent/HY-MT1.5-1.8B-FP8") model = AutoModelForCausalLM.from_pretrained("tencent/HY-MT1.5-1.8B-FP8", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Help Request Draft: HY-MT1.5-7B-FP8 OpenAI-Compatible Response Shape / Japanese-Dominant HTTP 500
Help Request Draft: HY-MT1.5-7B-FP8 OpenAI-Compatible Response Shape / Japanese-Dominant HTTP 500
Purpose: external help request draft for Hugging Face Discussion or GitHub Issue. This document is intentionally privacy-preserving: it does not include private audio/text payloads, server IPs, device identifiers, raw tracebacks, credentials, or raw ASR/MT/TTS content.
Recommended Title
Japanese-dominant HTTP 500 in OpenAI-compatible HY-MT1.5-7B-FP8 integration: possible non-string or malformed message.content / MT candidate shape?
English Post
Hi, we are integrating tencent/HY-MT1.5-7B-FP8 as the MT component in a real-time speech translation pipeline, and we are seeing a recurring server-side HTTP 500 pattern that is strongly associated with Japanese output cases.
We are not claiming this is a model bug. We are trying to determine whether this is:
a known HY-MT1.5 output / response-shape behavior,
a known OpenAI-compatible serving / vLLM response-shape issue,
a prompt / token / decoding parameter issue,
or simply an integration bug in our server-side parser / postprocessor.
Runtime Context
Model: tencent/HY-MT1.5-7B-FP8
Local served model name observed by our stack: hymt15_7b_fp8
Interface style: OpenAI-compatible /v1/chat/completions
Task: Chinese source text to target-language translation, then TTS
Product path: streaming ASR -> MT -> TTS -> client playback
We use HY-MT as the MT step only; ASR and TTS are separate components.
Prompt / Request Shape
Our HY-MT request uses a lightweight user-only translation prompt. We do not use a system role for this model path.
Conceptually, the user prompt asks the model to translate Chinese source text into the target language and output only the translation. We also enforce a local prompt budget guard before sending the request.
We would like to know whether HY-MT1.5 expects any stricter prompt template, chat message structure, or decoding parameters for Japanese (ja) output.
Observed Symptom
In targeted live validations, the external client path reaches our baseline gateway/workers correctly, but some MT rows fail before the streaming response is fully established.
The error is observed as server-side HTTP 500 from our real-time translation worker path. Our safe/redacted origin metadata consistently points to:
AttributeError / backend.main / _translate_text / worker_mt / backend_main_mt
The distribution is Japanese-dominant:
One targeted validation observed HTTP 500 in 9/16 rows, with 9/10 selected Japanese (ja) cases affected.
Earlier aggregation also showed Japanese dominance, with occasional non-Japanese rows such as vi or hi.
We do not currently have evidence that the issue is Japanese-only.
What We Found in Our Integration
We found and patched two server-side parser/postprocessor assumptions:
choices[0].message.content could exist but not be a Python string. Our old code could call .strip() directly and raise AttributeError.
Later MT candidate values could enter postprocess/repair/finalization helpers as non-string shapes. Our old code could again hit .strip() assumptions and raise AttributeError.
We changed those cases to enter our existing HTTP 502 malformed-output path instead of uncaught AttributeError. Normal string output behavior was not changed.
However, after these guards, targeted validation still shows Japanese-dominant HTTP 500 residuals. We have added additional safe origin metadata, but we have not yet proven the exact inner helper or exact model/service response shape for the residual cases.
Questions
Has anyone observed HY-MT1.5-7B-FP8 or the HY-MT1.5 family returning an OpenAI-compatible chat completion where choices[0].message.content is not a plain string?
Are there known cases where the served model returns structured content, empty content, null, a list/dict payload, tool-like content, or another non-string shape through /v1/chat/completions?
Are there known Japanese (ja) translation quirks for HY-MT1.5 that can produce unusual response shapes or trigger serving-side errors, especially for short or low-content streaming fragments?
Is the recommended prompt template for HY-MT1.5-7B-FP8 strictly user-only, or should we include any particular system message / instruction format / target-language marker for Japanese?
Are there recommended decoding parameters for stable translation-only output, for example temperature, top_p, max_tokens, stop sequences, or repetition controls?
If this model is served through vLLM or another OpenAI-compatible server, are there known version-specific issues where response content shape differs from the expected string format?
For production integrations, should callers defensively treat non-string or missing message.content as a normal malformed-model-output case rather than an unexpected exception?
What Would Help
Any of the following would be useful:
confirmation of the expected response schema for HY-MT1.5 under /v1/chat/completions,
recommended prompt template for Chinese -> Japanese translation,
recommended decoding parameters for translation-only output,
known issues around Japanese output, short fragments, or multilingual drift,
examples of safe server-side parsing / validation patterns for this model,
known vLLM / OpenAI-compatible serving versions that work reliably with this model.
Thank you. We can provide additional redacted metadata if useful, but we cannot post raw audio, private source text, translated text, server IPs, or raw tracebacks publicly.
中文说明
这份帖子不是把问题定性为“模型 bug”。它的目标是向模型作者/社区确认:HY-MT1.5-7B-FP8 在 OpenAI-compatible /v1/chat/completions 服务形态下,是否可能返回非字符串 choices[0].message.content,或者日语翻译是否有已知的输出形状/服务兼容问题。
当前项目内证据只能说明:
HTTP 500 发生在 MT worker path(机器翻译 worker 路径)。
安全来源聚合落在 backend.main::_translate_text / worker_mt / backend_main_mt。
ja 日语高发,但不是已证明 ja-only。
R410/R414 已经修过两类服务器解析/后处理的字符串假设:
choices[0].message.content 非字符串;
MT candidate(机器翻译候选)在后处理链路中非字符串。
残余 HTTP500 仍未关闭;当前不能说是模型根因,也不能说模型无关。
Optional Redacted Details To Add Later
Only add these if the model maintainers ask for more detail:
Served model name: hymt15_7b_fp8
Interface: OpenAI-compatible /v1/chat/completions
Failure class: AttributeError
Safe origin class: backend.main / _translate_text / worker_mt / backend_main_mt
Affected target distribution in a bounded validation: Japanese-dominant, e.g. ja 9/10 selected cases in one targeted run
Non-Japanese residual observations seen in earlier diagnostics: occasional vi / hi
Do not post raw ASR text, raw MT output, audio, private packets, device identifiers, LAN IPs, API keys, or full raw traceback.
Links For Posting Context
Hugging Face model page: https://huggingface.co/tencent/HY-MT1.5-7B-FP8
GitHub repository: https://github.com/Tencent-Hunyuan/HY-MT