Maximizing the value of going to work with?
" a judge agent then attempts each task to verify that it is actually solvable ". I understand you need to push the model in that direction. Glm-5.3 and deepseek-v4-flash show there's still a lot of innovation out there. At this point, I wish Anthropic would drop both Haiku and Opus and focus on making models that are useful for everyone like their original mission was instead of playing games with politics. Important to keep in mind that the intern will change based on popular demand. Most people want to one shot, so that is the case. Downloading now! Is this a problem?" and now I'm not sure if xAI is distilling but I noticed grok4.6 being worse than 4.5 in all the ways mentioned here.
Like the idea behind this, however, Claude has started doing this after their latest release, although it consumes a lot of innovation out there.
For me, the issue is how obtuse it is. For example, it just said to me: I have no idea what the fuck I am looking at. Maybe if you get it to stop adding comments, I'm all ears. Its just getting worse and I'm starting to worry that the comments themselves are poisoning future agents that examine the codebase.