Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
🔄
In a Training Loop
1300.9
TFLOPS
Andy Chen
andynoodles
14
34
Follow
dipankarsarkar's profile picture
1 follower
·
24 following
yi-hsiang-chen-tw
AI & ML interests
Feel free to contact me through Linkedin
Recent Activity
replied
to
onekq
's
post
2 days ago
My take on device-side inference: it's all about high bandwidth memory (thinking about it, this holds for the cloud too). MacBooks enjoy incidental capacity of apple silicon, but per-device RAM is too low (16 to 24GB), only sufficient for a decent SLM. 512GB is the highest you can go (Kimi K2*). Counting MLX downloads of Kimi K2* on Huggingface, I estimate the user base to be <25K. On the other hand, the newly debuted DGX station (Nvidia) has 748GB, which can fit in the latest Kimi, DS, and Qwen. Also the quantization options of CUDA is way better than MLX. For high-end inferencing, I place my bet on workstations over Macs.
liked
a model
3 days ago
zai-org/GLM-5.3
liked
a model
3 days ago
zai-org/GLM-5.3-Flash
View all activity
Organizations
andynoodles
's models
2
Sort:Â Recently updated
andynoodles/Qwen3.6-35B-A3B-NVFP4-sharded
22B
•
Updated
May 15
•
45
andynoodles/LLCG-OCI
1B
•
Updated
Aug 24, 2025
•
8