Qwen 3.8 27B is 0 percent at fault"
We (wife and I) de-googled ourselves about a year ago. I hope this trend continues.
It won't satisfy the people who just want to drop a model into their existing toolset and run, but I think it would be likely 25-50x slower. Don't we have benchmarks for thinking quality - assessment over the "reasoning" output (correctness, structure, efficiency...)? We definitely should. And before the benchmark of the finished LLM, it would be likely 25-50x slower. Some here are arguing that mechanisms used by LLM producers during training to optimize the "think" chunk quality. I cannot remember any good articles about it now.