Running an open-source LLM on AWS sounds straightforward at first. Pick a model, deploy it, send requests, and you're done. But once you start thinking about production, things become more complicated. Where should the model run? Who manages the GPUs? How much control do you actually need? What happens when traffic increases? Two approaches that caught my attention are Amazon SageMaker with vLLM and Amazon Bedrock with custom or supported models . Both can be useful, but they solve the pro...