
Dragonfly 2.0 vs Redis 8 and Valkey 9: with 2 CPUs it wins big, with 4 it depends on how you configure the others
I measured Dragonfly 2.0, Redis 8.10 and Valkey 9.1 in containers with 2 and 4 CPUs. With io-threads on, Redis and Valkey catch up with Dragonfly without pipelining and beat it at pipeline 16; with 2 CPUs and no pipelining, Dragonfly wins by 59% to 99% over stock Redis. It uses 8 to 15% less memory per key.
On 18 September, heise announced Dragonfly 2.0 under the headline “Redis und Valkey im Visier” (Redis and Valkey in the crosshairs). I found no mention of that version on Hacker News or in the forums I follow. The notes of the v2.0.0 release, from 16 September, say it brings no major new features and marks a maturity milestone. So there was no novelty to test: there was an old promise to check.
The promise is that an in-memory store with one thread per core is much faster than Redis. Redis 8 and Valkey 9 already have I/O threads, so the right question is, with the same CPU budget: how much does Dragonfly beat Redis and Valkey by when they also use the cores you give them?
With 4 CPUs and no pipelining all three pass 480,000 operations per second and the bench cannot rank them. With pipeline 16, Redis and Valkey with io-threads win (4.21 and 4.04 million against 2.73). With 2 CPUs and no pipelining, Dragonfly wins by 59% to 99% over stock Redis. And it saves 8 to 15% of memory per key.
What Dragonfly is
It is an in-memory store compatible with the Redis and Memcached protocols, written in C++ and licensed under BSL 1.1, which turns into Apache 2.0 on 1 November 2030. What sets it apart is the architecture: one thread per core, each owning a slice of the keys, sharing no state. Redis serves commands on a single thread and since version 6 can spread socket reads and writes over I/O threads. Valkey, the Redis fork, did the same.
I measured three servers, each alone in its container: Dragonfly v2.0.0, Redis 8.10.1 and Valkey 9.1.2. Redis 8.10.2 came out on 17 September, but Docker’s redis:8 image still shipped 8.10.1 when I measured, and the 8.10.2 tag did not exist.
Same client, same key space (1M keys), same limits. What changes is the engine, its thread setting and, for Dragonfly, the I/O backend.
The bench and its stumbles
Everything ran on fedora, on a Ryzen 7 7800X3D. Each server sits in a container with equal --cpuset-cpus and --cpus (4 or 2 cores), 2 GB and no persistence. The client is memtier_benchmark 2.5.1, in another container with its own cores, over a bridge network. One million keys, 256-byte values, random keys, 20 seconds per run and 5 runs per cell. For each Redis and Valkey setup I measured two variants: stock (one thread) and with --io-threads equal to the number of CPUs.
Four things that went wrong or are worth knowing:
- Dragonfly fell back to epoll without me asking. Docker’s default seccomp profile blocks
io_uring. The log says so: “Check if io_uring is disabled via /proc/sys/kernel/io_uring_disabled. Switching to epoll.” The host sysctl was at 0, and withseccomp=unconfinedthe same binary opened the io_uring ring. I measured both variants:dragonfly-epollon a normal Docker anddragonfly-uringwith--security-opt seccomp=unconfinedon that container only. - The Redis 8 image ships with modules loaded (ReJSON and search). It is what anyone gets from
docker run redis:8, and I did not turn them off. - My first 4 KiB run was invalid. One million 4 KiB values are 4 GB and the container has 2. The client crashed with Redis and Valkey, and with Dragonfly a number came out over an incomplete dataset. I set those runs aside and repeated with 200,000 keys.
- The first Dragonfly snapshot measurement did not exist. I started it with an empty file name and it answered
ERR filename is not specified, so my script timed a save that never happened. I redid it with the file configured and 2 million keys.
The host had only about 5 GB of free RAM and swap in use by another process unrelated to the bench, and the CPU governor was on powersave. None of the three servers came near those limits, but I say so because they are conditions of the measurement.
With 4 CPUs: no pipelining hits a ceiling, pipelining does not
Reads and writes 10 to 1, 200 connections (4 threads of 50), no pipelining. Stock Redis and Valkey, with one thread, did about 170,000 operations per second. With --io-threads 4, Redis 490,555 and Valkey 481,269. Dragonfly, 491,375 with epoll and 489,604 with io_uring. The p99 latency of the three with threads was 0.6 ms, against 2.3 ms for the single-thread servers.
With pipeline 16 the order changes: Redis with io-threads reached 4,214,020 operations per second, Valkey 4,037,297 and Dragonfly 2,727,721 with epoll and 2,738,973 with io_uring. That is, on that load Redis was 1.54 times faster than Dragonfly.
Real time. The three with threads arrive together: it is the ceiling of the bench, explained below.
At four times real time. Redis and Valkey with threads win.
Real time. Here Dragonfly pulls away, especially with io_uring.
At four times real time. The three with threads tie.
Real time, 200,000 keys, no pipelining.
Each lane fills in the time that server took to serve one million requests, using the median of its 5 runs. Switch scenarios with the buttons.
The bench limit: four servers, the same figure
That Redis, Valkey and Dragonfly give almost the same without pipelining on 4 CPUs made me suspect the bench before the servers. memtier keeps 200 requests in flight: the next one leaves when a reply arrives. In a bench like that, operations per second are the requests in flight divided by the mean latency. They are not two measurements: it is the same one seen from two sides.
- Redis stock200 in flight ÷ 1.154 ms = 173,310/smemtier reports 173,194/s
- Valkey stock200 in flight ÷ 1.202 ms = 166,389/smemtier reports 166,312/s
- Redis with io-threads 4200 in flight ÷ 0.407 ms = 491,400/smemtier reports 490,555/s
- Valkey with io-threads 4200 in flight ÷ 0.415 ms = 481,928/smemtier reports 481,269/s
- Dragonfly, epoll200 in flight ÷ 0.407 ms = 491,400/smemtier reports 491,375/s
- Dragonfly, io_uring200 in flight ÷ 0.408 ms = 490,196/smemtier reports 489,604/s
The four servers with threads sit at 0.41 ms mean latency. The client used about 3 cores, with 70% of its time in the kernel, and the servers between 2.4 and 3.0 of 4 (measured in a rerun of 3 runs, with the cgroup). No component I measure was saturated, so what limits is the latency of the whole path (client, Docker bridge network and server), which I cannot separate.
That is why I treat the 490,000 as “at least” and not as equal capacity. All the measurement says is that, with 4 CPUs and no pipelining, all three pass 480,000 and the bench does not reach any higher.
With that in mind, the pipeline 16 result does separate them: there Redis and Valkey with threads did about 4 million and Dragonfly 2.7, so the bench had room to see the difference.
With 2 CPUs: here Dragonfly does win
To take the spotlight off the client I cut the server down to 2 cores and gave the client 6 (6 threads of 34 connections: 204). Without pipelining:
- Stock Redis, 172,081 operations per second. With
--io-threads 2, 174,196 (+1.2%). - Stock Valkey, 171,208. With
--io-threads 2, 169,653 (−0.9%). - Dragonfly with epoll, 273,104 (+58.7% over stock Redis). With io_uring, 342,601 (+99.1%).
With 2 CPUs, io-threads did not move the throughput without pipelining, although they did move the consumption: Valkey went from 0.66 to 1.64 cores used to deliver the same. With pipeline 16 the three multithreaded servers tied: 1,475,365 with Redis, 1,443,686 with Valkey and 1,459,164 with Dragonfly, about 20% above the single-thread servers (1.21 million). I cannot explain that tie with CPU: the servers used between 0.93 and 1.6 of 2 cores.
4 KiB values
With 4 KiB values, 200,000 keys and 4 CPUs, no pipelining: Redis with io-threads 343,776, Valkey 335,449, Dragonfly with epoll 340,113 and with io_uring 403,036. It is the only load where Dragonfly with io_uring pulls away from the other two with threads: about 17% above Redis. With epoll it stays level with them.
Memory
With 1 million 256-byte keys, the container cgroup grew 326.5 bytes per key with Dragonfly, against 378.2 with stock Redis and 384.1 with Valkey. With 2 million keys loaded, Dragonfly took 635 to 667 MiB and Redis and Valkey 727 to 738 MiB (three loads per server; the 667 MiB one was Dragonfly with epoll). That is a saving of 8 to 15% depending on the dataset size. With 1 million keys it was a single load per server: I take it as an order of magnitude, not an exact figure.
- Dragonfly 2.0.0326 BThe same with epoll and io_uring.
- Redis 8.10.1378 B383 B with io-threads 4.
- Valkey 9.1.2384 B384 B with io-threads 4.
Growth of the container's memory.current after loading the keys. One load per server.
Snapshot under load
With 2 million keys and a 1 to 1 load, I ran BGSAVE 12 s into a 40 s run and compared against the same run with no snapshot. Three repetitions per server, 4 CPUs.
Stock Redis and Valkey did not change at all: their p99.9 was already at 2.4 to 2.5 ms without a snapshot and stayed there with one. With --io-threads 4 it does show. Redis’s p99.9 went from 0.97 to 2.11 ms (2.2 times) and Valkey’s from 1.20 to 3.49 ms (2.9 times). Dragonfly went from 0.81 to 1.00 ms with epoll and from 0.84 to 1.06 with io_uring.
Memory tells the same story. Over what the loaded dataset took, the container peak was +33 to +83 MiB in Redis with threads and +91 to +125 MiB in Valkey with threads. In Dragonfly, at most +13 to +26 MiB: the cgroup peak is a running maximum and does not separate the save from the writes of the load. Redis saves with fork and its own INFO reports the copy-on-write (current_cow_size); in Dragonfly I did not see that cost and did not check why.
- Dragonfly, epoll1.00 msWithout a snapshot, 0.81 ms.
- Dragonfly, io_uring1.06 msWithout a snapshot, 0.84 ms.
- Redis with io-threads 42.11 msWithout a snapshot, 0.97 ms.
- Valkey with io-threads 43.49 msWithout a snapshot, 1.20 ms.
2 million keys, reads and writes 1 to 1, BGSAVE 12 s into a 40 s run. Without a snapshot: 0.97 Redis, 1.20 Valkey, 0.81 and 0.84 Dragonfly.
The time the save takes under load is similar in all of them: 2.5 to 10 s. At rest, with the same 2 million keys, Dragonfly with epoll took 16 and 17 s, and with io_uring 2 and 3 s (two runs of each, according to its own last_success_save_duration_sec). I do not know why at rest with epoll it is about four times slower than under load and about six times slower than with io_uring.
What I cannot claim
- That stock Redis is limited by its single thread. It used 0.65 cores of what it had. Dragonfly did 2.8 times more than stock Redis with 4 CPUs, but my data do not show why the single-thread server stays at 173,000, only that it does.
- Which is faster with 4 CPUs and no pipelining. There the ceiling of the bench rules.
- The latency numbers as production latency. The bench is closed-loop: each latency is that of a server under 200 concurrent requests.
- That io_uring is always better. It contributed a lot with 2 CPUs and with 4 KiB values, and nothing with 4 CPUs, 256 bytes and no pipelining.
- Anything about persistence, clusters, replicas or commands other than GET and SET. I measured one case.
It is the same care I took measuring Wild vs mold and CUDA in Rust. In the second, the 3.2% in Rust’s favour stopped holding once I checked that the two programs did not compute the same thing.
My opinion
I was surprised that Dragonfly’s “several times faster” edge is in large part the difference between a server with one thread and one with several. With io-threads on, Redis 8 and Valkey 9 get very close, and at pipeline 16 with 4 CPUs they win. If you already have Redis or Valkey, turning on --io-threads and measuring with your load is worth more than migrating.
Where Dragonfly clearly wins in my measurements is with few cores, no pipelining and io_uring, and in memory per key. If your server has 2 vCPUs and your traffic is single requests, that is a concrete reason. And the snapshot cost it little: 0.2 ms of p99.9, against 1.1 to 2.3 ms for Redis and Valkey with threads.
When I would use it
- Yes: a small server, 2 to 4 CPUs, with many connections and single requests; data that takes a lot of memory and where 8 to 15% per key counts.
- Yes, carefully: if you are going to use io_uring in Docker. You have to open seccomp for that container and decide whether you accept it.
- Not yet: if you depend on commands or modules I did not test, or if your load is large pipelines on 4 or more cores: there Redis and Valkey with io-threads gave more.
Sources
- Dragonfly 2.0 ist da, heise online, 18 September 2026 (German). Source of the news.
- Dragonfly v2.0.0, release notes from 16 September 2026.
- Redis 8.10.1, from 17 August 2026. 8.10.2 came out on 17 September.
- Valkey 9.1.2, from 1 September 2026.
- memtier_benchmark 2.5.1, from 16 July 2026.
- Little’s law, for reading a closed-loop bench.
Comments
No comments yet. The first one is yours.