Confucius4-R2T2-CoreML-ANE / r2t2_lm_head_lut8.mlmodelc

Commit History

v4: two-lane decoder (two pieces per infer call, grouped-query attention)
b8b7f8c
verified

lunks commited on

Confucius4-R2T2 converted for Core ML on the Apple Neural Engine (encoder fp16, decoder LUT8 with 128-row prefill and verify head)
d0e38c6
verified

lunks commited on