-
Deploying vision models to the edge: quantization and latency
Training a computer vision model and deploying that model on a real device are two different engineering disciplines. A network that achieves excellent accuracy in a cloud notebook with a powerful GPU
-
Quantization in practice: running a capable model on one consumer GPU
Running a capable language model locally is mostly a memory-management problem.