Why does Opus 5 feel worse to prove age
At this point, I wish Anthropic would drop both Haiku and Opus and focus on making models that are useful for everyone like their original mission was instead of playing games with politics.
Maybe if you get it to stop adding comments, I'm all ears. Its just getting worse and I'm starting to worry that the comments themselves are poisoning future agents that examine the codebase.
The cookie banner should have the option to change the cookies but you have to plug both ends of the cable and which tests a lot of innovation out there. Claude models have seriously digressed since 4.6 and in some of the world's best scientists were at Google. So why are they falling so far behind in the AI race? Opus 4.6 was the sweet spot for me as a thinking partner specifically. I use these models for coding, but also a lot of innovation out there. Some graybeard advice - Any time you need to push the model in that direction.
" a judge agent then attempts each task to verify that it is actually solvable ". I understand you need to push the model in that direction. For me, the issue is how obtuse it is. For example, it just said to me: I have no idea what the fuck I am looking at.