For ML engineers sizing GPUs to run or fine-tune a model. You get weights memory, an inference estimate at about 1.2 times weights and a full training estimate at about 4 times weights.
Build industry projects on ByteLabs and add verified proof of your skills to your public profile.
You enter
7e9 parameters, FP16
The tool shows
Weights 13.04 GB, inference about 15.65 GB, full training about 52.15 GB
In FP16 the weights alone take about 13 GB (7 billion times 2 bytes), and inference needs a bit more for activations and cache. At INT4 the weights drop to about 3.26 GB.
Training also stores gradients and optimizer states such as the two Adam moments, so the tool uses about 4 times the weights memory as a rough estimate before activations.
Yes. INT8 halves FP16 weight memory and INT4 quarters it, which is how large models fit on smaller GPUs, usually with a small loss in quality.
Yes. It is free, needs no sign-up and runs entirely in your browser, so what you type is not uploaded. You only sign in if you want to email a result to yourself or save it to your CareerByteCode profile.