IBM's nine-minute explainer on distributed inference: data, pipeline, tensor and expert parallelism, KV cache and prefill/decode bottlenecks.