We can Load weights In cpu ram than in Vram of Gpu but slow.. CPU Offloading

 VRAM Size: Binary Code vs. Weights Percentage Split

If we only stored the binary instructions of the model in VRAM and kept the weights out, the reduction in VRAM size would be roughly 99.999%.
The program instructions (the loop to do the math) take up only a few kilobytes or megabytes of binary code. The weights of a modern model (like Llama or GPT) take up gigabytes or terabytes of binary numbers.
So, can a person choose to let the weights sit in the laptop's RAM and only load the binary code into VRAM?
Yes, you can absolutely do this today! This technique is called CPU Offloading or running on Shared System Memory. Programs like llama.cpp do this so people can run large AI models on cheap laptops without massive GPUs.

Comments

Popular posts from this blog

how to add all current and future projects of android studio to allow in windows firewall security..

adjusting width of explorrer in xamp project folder file names - php LocalHost