we can run any AI model in less Vram GPU by CPU offloading method -- as that is handled by the AI Compiler
2. Can You Buy a Small GPU to Only Run Binary Code?
Yes, you can run any AI model this way, but reducing the number of cores to avoid them "sitting idle" defeats the purpose of buying a GPU.
If you buy a small GPU with very little VRAM (e.g., 4GB) and try to run a massive 70-billion parameter AI model, the software tools (like llama.cpp) will automatically chop the model up. It will put 100% of the binary instructions into the VRAM, fit a tiny fraction of the weights into the remaining VRAM, and leave the remaining 95% of the weights in your laptop's regular RAM.
Any standard AI model supports this out of the box because it is handled by the AI Compiler / Runtime software, not the model itself. However, because you cut down the cores and VRAM, the system will perform at the slow speed of a CPU anyway.
Comments
Post a Comment