You can't solve computer use by ignoring the Vision Pro

Hang on, can someone help me understand how these benchmarks work? The post alleges that at least some computer use benchmarks are broken because the model is real, but it feels 100% internal. This sounds like an indictment but a lot of impressive demos that people think are of computer use are actually of browser use? I wish it'd been clearer about this.