Manga character parts β mobile inference pack
Optional components for conservative frontal-head hair/skin guidance. Download size: 396,724,626 bytes (both foreground precisions included) before device-specific Core ML compilation. This is a multi-model research integration, not a universally accurate manga segmenter or a standalone colorizer.
Components
head.onnx: DeepGHS head_detect_v2.0_s; author-compatible 640-square RGB bicubic input.foreground.onnx: original FP32 SkyTNT ISNetIS (176,069,933 bytes); foreground segmentation at 1024.foreground-fp16.onnx: optional FP16 derivative (88,070,593 bytes), same architecture and float32 public input/output.Parts.mlpackage: original StyleAnime anime parser converted to FP16 Core ML; externally ImageNet-normalized float32 RGB[1,3,512,512], int32 labels[1,512,512](19 classes).generator512.onnx: shape-only derivative of Faridzar's manga-colorization-v2 FP16 export, fixed[1,5,512,512]. All weights and nodes preserved; output RGB is already 0β1. It does not replace the variable-size whole-page model.
See manifest.json for immutable checksums and SOURCES.md for pinned upstream sources, input contracts and evaluation limits. The application adds anatomy/identity gates, foreground/stroke/dialogue exclusions, separate reference palettes and native-resolution part composition. Masks alone do not establish reliable ownership. Profiles, closed eyes, overlapping people and hair outside head crops can be left unassigned.
Attribution and terms
Each component retains its upstream terms; no blanket license is assigned to this collection. Head detection's model card declares MIT; foreground declares Apache-2.0. Full available notices and pinned cards accompany the pack under notices/.
The parser architecture is MIT (copyright 2019 zll), but the original StyleAnime checkpoint has no explicit redistribution license in the inspected source. The generator export mirror declares MIT; its upstream pretrained weights likewise have no explicit license in the inspected repository. These labels do not establish unrestricted rights to the underlying weights. Permissions for those weights remain unresolved, and this experimental pack is not a claim of commercial clearance. Upstream authors retain their rights.
Validation
Exact input/resampling checks, original-vs-export comparisons, actual Swift inference, and a public pink-reference/line-art end-to-end fixture were evaluated locally. Conversion agreement is not semantic accuracy. Physical-device quality and independent human scoring remain unmeasured.
Optional foreground precision
Both weights are retained. The application exposes Fast (FP16) and Full precision (FP32), keeping FP32 as the existing default. Upgrading an existing pack needs only the additional 88 MB file; all unchanged files are checksum-verified and reused. Model precision has its own mask, palette and rendered-page cache identities.
The FP16 conversion uses onnx 1.17.0 and onnxconverter-common 1.16.0, keep_io_types=True, min_positive_val=0 and max_finite_val=65504, with the converter's default blocked operations. Original foreground source: skytnt/anime-seg revision 493cb60893f47441b26ec4fb9a306bce9e342982, isnetis.onnx. Its Apache-2.0 notices remain applicable. Reproduction script: convert_foreground_fp16.py.
Local Apple Silicon native ORT/Core ML measurements over eight public head crops: warmed foreground calls approximately 0.34β0.35 seconds FP32 versus 0.215β0.225 seconds FP16; process peak RSS approximately 629β635 MB versus 466 MB. One pixel across those eight 512-square masks changed sides of the 0.5 threshold; maximum probability error 0.00854. Two end-to-end public pages kept identical character review outcomes; one had exactly identical rendered pixels, the other a mean absolute channel error of 0.001267/255 and maximum 3/255. These are limited local comparisons, not an accuracy guarantee for arbitrary comics.
CPU-only inference also completed, but FP16 was not faster (4.02 versus 3.94 seconds) and used more process memory (1.01 GB versus 745 MB) on the measured fixture. Prefer FP32 for CPU-only use. iPhone/iPad runtime and quality measurements remain pending.
- Downloads last month
- 11