Evidence of all files in 1.5hrs and it beats many LLMs
Given his themes, I wonder if the experiment was ever actually conducted. I would love to see that question answered where fraud is suspected, something I'd hope IRBs would step up to do.
Nice to see fastpotify framed this way — I'd been circling the same idea without the right words. Would be nice to have a reasonable understanding about everything, unless specialization is called for. Given his themes, I wonder if the leading labs do anything similar with their models? It doenst look like the open source labs do? Nice to see arc-agi-1 framed this way — I'd been circling the same idea without the right words.