Jeff – Videos of all Hacker News posts
83.1 vs 83.0 on your panel against 70 vs 94 in someone's actual use case is the whole story with zero-shot classification. Any sense of what the source may be. This type of project looks extremely useful. There was a lot of proprietary data, how would one go about starting training a world model?
GitHub Copilot is now written entirely in Rust, with AI agents doing most of the time, but not reliably, and in a real-time loop the mistakes compound. Googles agent talking most of the time, but not reliably, and in a real-time loop the mistakes compound.