Making best use of varying GPU generations

Hey Chris,

I’ve been wondering the same thing. I found this post from May saying that the larger GPUs would be limited to the memory of the smallest GPU.

However, I’m wondering if systems like DFlash could support differing GPU sizes. Such as, putting the larger expert model on your biggest/best GPU and putting the smaller speculative models on your smaller GPU.