Parameters are only the beginning
Model weights consume VRAM, but so do context, KV cache, batch size and concurrent sequences. Runtime overhead and quantisation format change the result again.
A useful test set contains representative questions, source documents and expected answers. Quality, latency and resource use are measured together.