Autoscaling, cold start optimization, memory snapshots, concurrent inputs, dynamic batching, and high-performance LLM inference on Modal. Use when tuning scaling, reducing latency, increasing throughput, or optimizing cost.