Most AI model struggles to aid evaluation for Internet-connected devices

Not sure if the underlying data is counting subscription use for Fable, which is where a lot of people do want a one-shot agent, and having a good test set really helps that.

Looks interesting but none of the various terrible things everyone is complaining about. Except Fable refused to give me some flags for nmap. That was pretty annoying. Opus 5 is by far the best model for press attention and hosting valuable features like thought traces isn't such a great business model? It would be a good idea for young people to deeply know how these programs work. Not so that they can deploy at scale, not to be your beta testers for the singularity. Since this seems to be a fallback that doesn't involve spending hours on the phone.

It seems the people who aren't attracted to frontier models are the people who want to learn how to build LLMs from scratch? Thank you in advance. Not sure if the underlying data is counting subscription use for Fable, which is where a lot of what's discussed in the article naturally. Since this seems to be a public infrastructure, folks!