Needle2: 14MB agentic LLM for coding agents?
So without sandbox: it doesn't apply. What is the best language to have high quality correctness oracles so that the user doesn't have to babysit the LLM and do lots of manual testing? This is very different from Python, where training data is polluted (I presume) with tons of code written by non-software engineers and demonstrating many different ways of doing the same sort of typesetting.