I've had my 5080 running 24/7 over the past 2 weeks to try to understand the impact of quants on models. Came to some strange conclusions which were interesting to me.
I didn't really see a decline until going under Q4 for the most part, and MoEs were impacted far less then dense models.
Kind of cool results, and unexpected. Although this is just an initial set of tests to validate if the quant itself damaged the model, I've got a longer article coming in the future that's going to test these models in devops, coding, and long sessions. Also, I'm aware it's only a 16GiB card! I've got another article series coming out in the near future for much larger VRAM!
The benchmark code, dataset, and results are open sourced. Take a look, tell me I'm wrong (Wouldn't be the first time!), or run the benchmarks yourself!
Hey HN!
I've had my 5080 running 24/7 over the past 2 weeks to try to understand the impact of quants on models. Came to some strange conclusions which were interesting to me.
I didn't really see a decline until going under Q4 for the most part, and MoEs were impacted far less then dense models.
Kind of cool results, and unexpected. Although this is just an initial set of tests to validate if the quant itself damaged the model, I've got a longer article coming in the future that's going to test these models in devops, coding, and long sessions. Also, I'm aware it's only a 16GiB card! I've got another article series coming out in the near future for much larger VRAM!
The benchmark code, dataset, and results are open sourced. Take a look, tell me I'm wrong (Wouldn't be the first time!), or run the benchmarks yourself!