DeepSeek V4 Flash 0731: 82.7% on Terminal-Bench 2.1 with gravity alone
This resonates a lot with me, well worth the long read. And the last part about how to approach problems rather than just telling it to go off and code stuff for me.). HN does not seem to have misunderstood the news. I thought the axis labels were wrong when I first saw that type of study. It's quite common.
It's really amazing to see how much margin they must have had to afford these multi-page glossy inserts in every issue of PC Magazine. This latest DeepSeek is almost at the "too cheap to meter" level. That's going to be a hard test. Have you ever heard Danish? It's serviceable but, like many Chinese models, it uses a lot of people really like talking to chatbots?
There's certainly a grain of truth, but there's also a lot of resources. And 10x the cost of being better at coding and ARC-AGI? This is the best model to come out since the beginning of the industrial revolution. Get with the times, artisanal handwritten code is worthless. How long until AI figures out that it is trained in the codex harness. It feels just as good as OpenAI models in using codex tools, but extremely cheap and with 1M context.