Confucius4-R2T2 converted for Core ML on the Apple Neural Engine (encoder fp16, decoder LUT8 with 128-row prefill and verify head) d0e38c6 verified lunks commited on 10 days ago