I have the following questions for those who really understand LLMs and have programmed them:
1. How interchangeable are the services really? Let's say you're set up on Azure cloud, and Google Cloud or Coreweave's bare metal offering become way cheaper. How much of an effort is it to transition to the cheaper offering? What specifically would you need to change?
2. For inference, how much of a difference is there between Azure, Google Cloud, or random bare metal rental services like vast.ai?
3. My understanding is that even if Vera Rubin chips start shipping tomorrow, it takes time for training algorithms to move to the new hardware, so they won't be immediately valuable. There will be some kind of delay. How much of a delay are we talking?
4. How much of an economic advantage is there between Vera Rubin and H100's? In other words, since NVidia is moving to once/year releases, and it takes considerable time (?) to port code over to the new generation of GPUs, does it make sense to just skip a generation and just wait until the next year to port once to Vera Rubin instead of porting twice? I guess I don't have a good handle on the ROI of porting efforts.
Part of the reason I'm asking this is because I see rental prices for an H100 on Azure for let's say $8/hr and then I see random loose H100's on [vast.ai](http://vast.ai) for $1/hr, and that just seems like an enormous difference.
If Microsoft really can get that $8/hr all day long and has unlimited demand, their AI spending might actually pay off.
But if H100's in today's frontier training data centers get moved to inference, I'm thinking the economics in the long run will trend towards $1/hr levels regardless of if they're in a fully networked massive data center or just random loose GPUs.
Is this correct?
So much of the AI bubble hinges on these questions...