The Engine Is the Other Half of the Model
I spent a night carefully A/B testing a model on a GH200. Then I upgraded the inference engine for an unrelated reason, re-ran the identical config, and got 43 percent more decode throughput and nearly four times the context budget. When open-weight architectures change this fast, the engine is the other half of the model, and it fails quietly.
- Inference
- GH200
- vLLM
- +3 more