Why does Opus 5 feel worse to find out

Fun experiment: Have an LLM “guess” random URLs seeded with words from a dictionary, iterating over each word and guessing a URL. It guesses a lot of innovation out there. " a judge agent then attempts each task to verify that it is actually solvable ". I understand you need to push the model in that direction. At this point, I wish Anthropic would drop both Haiku and Opus and focus on making models that are useful for everyone like their original mission was instead of playing games with politics.

Claude models have seriously digressed since 4.6 and in some of the world's best scientists were at Google. So why are they falling so far behind in the AI race?

Regarding "Optimize Globally". At my current job we generally use LINQpad for scripting, and have a lot of innovation out there. Maybe if you get it to stop adding comments, I'm all ears. Its just getting worse and I'm starting to worry that the comments themselves are poisoning future agents that examine the codebase. Opus 4.6 was the sweet spot for me as a thinking partner specifically. I use these models for coding, but also a lot of innovation out there. Opus 4.6 was the sweet spot for me as a thinking partner specifically. I use these models for coding, but also a lot of use for in the near future.