Qwen3-4B Genie β€” Snapdragon 8 Elite for Galaxy β€” 12K

A Qualcomm Genie / QAIRT deployment build of Qwen3-4B, compiled for the Samsung Galaxy S25 family and Qualcomm Snapdragon 8 Elite for Galaxy.

Build configuration

Property Value
Base model Qwen/Qwen3-4B
Model behaviour Thinking-capable
Quantisation W4A16
Runtime GenieX QAIRT
Target device Samsung Galaxy S25 family
Target chipset Snapdragon 8 Elite for Galaxy
Context length 12,288 tokens
Prompt-processor sequence length 128
Decode sequence length 1
QAIRT compilation version 2.45.0.260326154327

Status

The Qualcomm AI Hub export and compilation completed successfully with exit status 0.

On-device validation is pending. This repository should currently be treated as a compiled deployment candidate rather than a confirmed performance release.

Important compatibility information

This is not a standard Transformers checkpoint.

The .bin files are compiled Qualcomm QNN/Genie deployment contexts and require the Qualcomm Genie SDK and an appropriate QAIRT Android runtime.

The Qualcomm SDK and runtime libraries are not distributed in this repository.

Included files

  • Four compiled model context binaries
  • Genie runtime configuration
  • HTP backend configuration
  • Qwen tokenizer files
  • Export metadata
  • Build and checksum records

The four partN_of_4.bin files must remain together with the configuration and tokenizer files.

Intended use

This build is intended for on-device text generation and reasoning experiments on compatible Snapdragon 8 Elite for Galaxy Android devices.

It was produced for the My Mettle Android project but may also be useful to other developers testing Qwen3 through Qualcomm Genie.

Limitations

  • Device execution has not yet been validated.
  • Performance, memory consumption and thermal behaviour are not yet measured.
  • Compatibility with devices outside the Samsung Galaxy S25 family is not guaranteed.
  • A matching or compatible QAIRT/Genie runtime is required.
  • The 12,288-token context may exceed practical memory limits on some devices.

Attribution

  • Original model: Qwen3-4B by the Qwen team
  • Original model licence: Apache License 2.0
  • Deployment compilation: Qualcomm AI Hub Workbench
  • Packaging and publication: MagneRex

Qualcomm Genie, QAIRT and related SDK components are not included.

Downloads last month
9
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for MagneRex/Qwen3-4B-Genie-Snapdragon-8-Elite-12K

Finetuned
Qwen/Qwen3-4B
Quantized
(310)
this model