Confucius4-R2T2-CoreML-ANE / embed_tokens.f16.bin

Commit History

Confucius4-R2T2 converted for Core ML on the Apple Neural Engine (encoder fp16, decoder LUT8 with 128-row prefill and verify head)
d0e38c6
verified

lunks commited on