Rook and Ceph: What the Operator Actually Does

▶ Watch (1:29)

Rook is a Kubernetes operator that installs and manages Ceph storage using CRDs. Once Ceph is running, a CSI layer provisions and mounts volumes for application pods. Applications then communicate directly with the Ceph data layer. Ceph itself provides block storage via RBD (read-write-once), shared file storage via CephFS (read-write-many), and S3-compatible object storage via RGW. The whole stack is Apache 2.0 licensed. Rook turned 10 years old at this conference.

Cluster Topology and Zero-Downtime Maintenance

▶ Watch (9:10)

With replication size three across availability zones A, B, and C, Ceph places exactly one copy per zone. One zone going offline leaves the cluster fully readable and writable. Two zones down stops IO to protect the last surviving copy. Rook handles upgrades through pod disruption budgets that keep Kubernetes aware of Ceph topology, ensuring only one zone is ever down at a time. A kubectl plugin shortcuts common operations: checking cluster status, removing OSDs, and restoring monitor quorum.

120 Petabytes: SAP’s Migration to Rook Ceph

▶ Watch (11:06)

SAP’s cloud infrastructure stores 120 petabytes across 15 regions and 26 availability zones. The old stack used proprietary appliances and OpenStack Swift, which lacks read-after-write guarantees that modern applications expect. Migration to Rook and Ceph started two years ago. The first region went live in May 2024. Today 10 regions run with 37 petabytes of raw capacity. Clusters range from 15 to 60 nodes, each with 16 NVMe drives and 100 Gb networking. The team provisioned each region using a Helm chart base blueprint with per-region overlays, enabling single-click region provisioning.

Ceph-CSI: Drivers, Backends, and New Features

▶ Watch (16:16)

Ceph-CSI ships as a single container image supporting RBD, CephFS, NFS, and NVMe-oF backends. RBD suits databases and small-file workloads formatted with ext4 or xfs. CephFS handles read-write-many file access but slows on metadata operations over millions of small files. Volume group snapshots pause IO across all labeled PVCs simultaneously, then resume, giving consistency-group semantics for multi-volume applications. Change block tracking, requested by KubeVirt, enables differential backups by returning only the blocks modified between two snapshots.

Erasure Coding: Cutting Storage Overhead from 3x to 1.5x

▶ Watch (27:06)

Three-way replication stores every byte three times. A 4+2 erasure coding profile stores four data blocks plus two parity chunks. Any two of those six pieces can be lost and Ceph still reconstructs the data. That drops effective overhead from 3x to roughly 1.5x while keeping the same failure tolerance. Benchmark results from Cephalocon in October 2024 show random-write latency for erasure-coded pools nearly matching replication. The trade-off is CPU cost for parity calculation. Rook simplifies setup: define the pool and Rook configures failure domains automatically.

Notable Quotes

so far uh there is no um uh customerf facing incidents or downtime. Artem Torubarov · ▶ 15:43

today we can actually provision uh the whole region with hardware in one click. Artem Torubarov · ▶ 15:05

This is very inefficient and cost like three times the hard disks that you would expected to use. Niels de Vos · ▶ 27:27

If you do it manually it’s rather tricky. Niels de Vos · ▶ 29:01

Key Takeaways

  • Rook 1.18 enables the Ceph-CSI operator by default; 1.20 targets full CRD-only CSI configuration.
  • SAP’s 10-region Rook deployment reached 37 petabytes with no customer-facing downtime across OS, Kubernetes, and Ceph upgrades.
  • Erasure coding 4+2 profiles cut storage overhead to roughly 1.5x replication cost with comparable random-write performance.

About the Speaker(s)

Niels de Vos is a core developer and maintainer for Ceph-CSI, employed by IBM and working with the Red Hat team that supports OpenShift Data Foundation. His main focus is improving storage support for containers across Ceph, Rook, and OpenShift.

Deepika Upadhyay is a Ceph Engineer at Clyso and active Rook contributor, specializing in large-scale Rook Ceph deployments for enterprises. Her background spans RADOS and RBD storage engineering.

Artem Torubarov is a Senior Software Engineer at Clyso GmbH with over 10 years of experience, focused on distributed backend systems and storage technologies including Ceph.