Model Quantization — Roadmap Step & Resources
Reducing numerical precision of weights to shrink models and speed up inference
Open on Cached Info