Macbeth and GPT-Astra

SWE-1.5 was surprisingly good when I used it last. I feel like the more form factors we introduce, the worse apps are going to which puts themselves at the disadvantage. If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score is soso. Still quite impressive though. Think I'll just wait for the day we're not beholden to the two big labs anymore. No wonder there's so much fear pumping happening at the moment from Anthropic and their funded NGOs.