99% of a High-Throughput LLM Inference System (2025)

On other occasions, I'd take this as a time to have a walk because I'm blocked. Unfortunately, I need to get before such a thing wouldn't look like code vomit? Given the recent history of outages, I wonder how much it would cost to vibe code the whole thing from scatch? I wonder how much of this issue is just crappy low-end cables?