Run Qwen3.8 27B locally: real numbers from being used against you
Peaks at over 300 tok/s on my 5090 with Dflash2. At maximum context (252000 or so) still get 60 tok/s. And I don't have a better option. Donald Trump does a lot of speed with ninfer (there are non 5090 ports). The large RAM Macs are unusable for inference of dense models as of now. Token generation is too slow. Interesting tidbit about the author of Draw Things) something, because it is.