Lyft Moves Streaming Fleet to Apache Flink Kubernetes Operator
Lyft migrated hundreds of production Flink jobs from a homegrown Kubernetes operator to the Apache Flink Kubernetes Operator, gaining last-state upgrades, in-place autoscaling, and resource autotuning. The legacy operator's savepoint-trigger lacked retry logic and idempotency, causing deploy failures on large-state jobs, and its single memory knob wasted resources. Lyft adopted the open-source operator's BlueGreen deployment CRD (shipped in v1.14.0), contributed an upstream fix for a configuration-rename bug, and upgraded to Flink 1.19 to enable in-place scaling and the KinesisStreamsSource for autoscaler backlog metrics.