~/topics/local-models
Models on my own hardware
What actually fits in a 16 GB card, and what gets lost along the way.
Almost everything published about open models is measured on datacenter cards. The question here is a different one: if the file weighs more than the memory available, does it still run, and at what cost? These pieces share a bench — the same 16 GB GPU, the same frozen inputs — so the numbers can be placed side by side instead of read in isolation.
11 pieces










