5 Cassandra Enhancement Proposals Operators Should Be Watching

Wait 5 sec.

Anyone who’s run Cassandra at real scale has a repair story, and inevitably it starts somewhere around 2am. The good news is that a chunk of the work moving through the Cassandra Enhancement Proposal process right now is aimed at stories like that. What jumps out to me is where the effort is going, into the operational corners of Cassandra that its power users know intimately.A CEP is how meaningful change to the open source database gets designed and argued into shape out in the open, with someone signing up to do the work and win community agreement along the way. Several have been accepted lately, and five in particular stuck out to me because, between them, they cover performance consistency, data placement, day-to-day admin, and how well Cassandra utilizes the hardware you’re paying for.The CEPs are at varying completion stages: CEP-49 is already merged into trunk and slated for 7.0, CEP-45 has substantial code written on a dedicated feature branch, and the other three are accepted with implementation underway. Even where the code isn't written yet, the shape of these proposals tells you plenty about where the project is putting its energy.CEP-45 tracks writes instead of comparing everythingI’ll start with the one that’ll be appreciated by anyone who’s had to babysit a repair. Cassandra has two ways to catch a write that didn’t make it to every replica, repair and read repair, and both do it the expensive way (by pulling data off multiple nodes and comparing it). Repair ships whole partitions around when it finds a mismatch, driving up streaming and compaction and turning very large partitions into a headache. Read repair only patches the slice of a partition that your query touched, leaving a write half-applied and quietly breaking the partition-level write atomicity you assumed you had.CEP-45 comes at this from an entirely different angle. The coordinator stamps every write with a unique ID. That ID rides along to the replicas, with each replica keeping a record of which IDs it has applied. On a read, one replica sends back the data plus a summary of its applied IDs, while others send just the summary. Do the summaries line up? Great, the data is good, and any gap gets filled by shipping over only the specific writes that are missing. A background process continues reconciling those IDs and marks a lower bound (a sort of watermark) below which old log entries are safe to clear out. It turns on per keyspace or table through a new replication-type setting and reuses the addressable commit log from Accord so that individual entries can be pulled by ID. The end game here is that reconciling one write at a time (as opposed to hauling whole partitions across the network) should take a big bite out of streaming and compaction costs and loosen those partition-size limits.CEP-60 (Flexible Placements) loosens data placement from the token ringWith Cassandra today, a node sitting on the token ring decides what data it owns. That comes with a set of annoyances (e.g. growing a cluster on the cheap usually means doubling it). While a new node joins, its range gets served by an extra replica for a while, which piles on load…right when you least want it. And with vnodes you cannot move tokens, so an unbalanced ring becomes yours to sort out by hand, with no clean way to plan one big change as a single operation or chop a long one into smaller steps you can retry.Enter CEP-60, which cuts the tie between placement and token position. The unit of ownership becomes the tablet, a range for a specific keyspace and table pair, and replicas get assigned directly. It builds on the transactional cluster metadata from CEP-21, so that bootstrap and streaming run as smaller, resumable steps (and it decides where data lives from real per-range load and capacity numbers, not a percentage of the ring).It’s worth stressing that this isn't the end of tokens. Ranges still have token bounds, and tokens still drive routing, streaming, and repair. What changes is that owning a token no longer determines what data you own.Cassandra operators will feel the payoff as density. A cluster that stays close to balanced at any size lets you run nodes hotter and steadier, different from the sawtooth you get from token-based growth where you provision for the peak and pay for idle boxes. Flexible placements are enabled per keyspace, so you can try it on one keyspace, and the CEP includes tooling to roll a cluster back to token or vnode ownership.CEP-38 runs administration through CQLMost of what you do to administer Cassandra, be that snapshots or compactions, runs through JMX MBeans under the proverbial hood. Tools you already lean on, like nodetool and Sidecar, all speak JMX. Over the years the ecosystem has papered over that by wrapping JMX in REST or slipping past it with Java agents, but all of that sits on an internal API not designed to be a stable contract. You inherit JMX’s security exposure and operational overhead, not to mention the running cost of keeping nodetool alive and a server that doesn’t tell you much about the commands you run.CEP-38 makes CQL itself a management interface, so you can issue admin commands through the language your devs already use instead of standing up JMX access and its bolted-on tooling for every client. Each command gets defined once in a single registry with real metadata, which gives the agents and REST layers something solid to sit on and keeps JMX, the CLI, and any REST interface from drifting apart. It also moves command execution toward an async, observable model where you submit a command, get back an identifier, and check the result when it’s ready (as opposed to firing it off and blocking).Nodetool and cqlsh turn into thin front doors over the same command registry, and a dedicated admin port gives the control plane a clean boundary for orchestration to build on later. Importantly, though, none of this rips out the existing MBeans or CLI. JMX continues working, though the change nudges the project toward a day when retiring it is realistic.CEP-62 changes configuration without hand-editing filesMuch of what shapes a Cassandra node lives in files you can’t touch while running. Memtable settings, SSTable options, storage_compatibility_mode and the like all sit in cassandra.yaml, and the startup tuning, like heap size and garbage collection, lives in JVM options files. Runtime settings can nudge through JMX, but for these on-disk files there has never been a programmatic way in, so Cassandra operators edit them by hand or wire up their own scripts. An earlier change taught the Sidecar to start and stop instances, and it still couldn’t touch the configuration those instances read at boot.CEP-62 solves for this with a Sidecar REST API for reading and changing cassandra.yaml and the JVM options. It keeps a sparse overlay of just your explicit changes on top of a base template and merges the two into what the node runs, and a pluggable provider can hold those overlays locally or in something central like etcd or Consul. A version-aware check throws out any setting a given Cassandra version wouldn’t recognize, so typos or unsupported options aren’t likely to leave you stranded with a node that won’t start. Changes take effect on the next restart, and all of it lives in the Sidecar, with Cassandra left alone. If you’re managing config across a big fleet or feeding it from centralized tooling, this is the one you’ll want. It is off by default and purely additive, so no one who isn’t running Sidecar has to care, and it opens the door to driving upgrades through the Sidecar down the line.CEP-49 lets the hardware do the compressionCompression is one of the bigger CPU sinks in Cassandra. The database ships with four compressors (LZ4, Zstd, Deflate, and Snappy), and the work of squeezing and unsqueezing data eats cycles during flush and compaction, and then again on the commitlog and on data moving across the network. Meanwhile, a lot of newer processors carry dedicated accelerators exactly like this, such as Intel’s QuickAssist on Xeon chips, which knows how to accelerate three of Cassandra’s four compressors.CEP-49 hands off the compression work to that hardware wherever it exists, freeing the CPU for things only the CPU can do and speeding up compression itself. The framework reaches for the accelerator when it’s there and falls back to the software compressor when it isn’t, with room to add other accelerators later. Backends ship as plugins that Cassandra locates at startup, and if one fails it drops back to the standard compressor. This one is furthest along of the five; the implementation landed in trunk is on track to be available in a future release, so if you’re running compression-heavy workloads on capable hardware, this is close to found money! The one caveat is that the hardware has to be set up correctly first. Anything not working right falls back to software compression, which is safe, but just means you paid for silicon you’re not using.Why these five, togetherEvery one of these targets something operational with Cassandra, the stuff that decides whether running Cassandra at scale is pleasant or…less so. A few of them also open doors for future work, so some of their biggest payoff may show up a release or two after they land. But if you run Cassandra, whether self-hosted or on a managed platform, Cassandra’s current CEP process is the clearest window you have into where it’s going, and these are the five I would be watching especially closely.