Technical blog

Booting Fast and Slow

At H Company, our agents are served by a fleet of vLLM replicas running on GPU nodes in Kubernetes.

Illustration of one ice-covered GPU beside one glowing, flame-covered GPU.