I run the AI for Fring, a free wardrobe app, on a second-hand Tesla P40 in my homelab. The virtual try-on uses Leffa, a diffusion model. For a month every render took about 21 minutes, and I assumed that was just what a 2016 card could do. It wasn't.
The model was loaded in fp16, because "use half precision to save VRAM" is the default advice everywhere. It is right on almost every recent card. It is wrong on consumer Pascal: the P40 (GP102, compute capability 6.1) runs fp16 at roughly 1/64 of fp32 throughput. The code had said fp16 since day one; the warning had been sitting in a ticket for a month. Nobody had read the two together.
Measured on 23 September 2026, on the same 20-step render.
fp16 — 63.8 s per diffusion step, 6.6 GB of VRAM, 1241 s for the full render.
fp32 — 4.7 s per step, 9.8 GB of VRAM, 85 s for the full render.
The images were practically identical (mean difference 0.77/255 per pixel). And the card never needed the VRAM savings: the model used 6.4 GB out of 23. The saving fp16 bought was worth nothing, and it cost a factor of 14.
A CPU-only torch wheel. The Dockerfile installed torch from the CPU index, inherited from an older server whose 4 GB card could not hold the model. On the P40, `nvidia-smi` saw the GPU while `torch.cuda.is_available()` returned False, and renders took hours on the CPU. Check inside the container, not on the host.
An unpinned dependency. A mediapipe upgrade removed `mediapipe.solutions`. The health check only tested the import, so it kept reporting pose estimation as healthy while the service silently fell back to a centred paste. A probe must test the attribute actually used, not the package.
The service now picks fp16 only when the GPU actually accelerates it (Volta and newer, or P100), or when fp32 would not fit in VRAM — on a 4 GB card, fp16 is still what makes the model fit. An environment variable can force either. A virtual try-on now takes one to two minutes, queue included.
For anyone reviving an old datacenter card: measure before applying the usual advice. A five-minute measurement would have saved a month of 21-minute renders.
Fring is a free digital wardrobe: photograph a garment, the AI fills in the details, the app builds outfits from what you own and tracks each piece's cost per wear.