Spate Benchmark
Every system here consumes the same topic of Confluent-framed Avro, flattens each
message's events array into one row per event, and inserts those rows into
ClickHouse. Same bytes, same envelope, same hardware, same delivery guarantee.
One sortable row per arm, one table per comparability context, and nothing ranked that the fairness contract does not allow to be ranked. Here are the results; make your own determination.
Kafka → Avro → ClickHouse. 32 CPU and 96 GiB of data plane per system, at-least-once. Every system consumes the same topic, decodes and flattens each message, applies the same two filters and two derived columns, and lands the surviving rows in ClickHouse.
6 arms · 5 systems · 1 measurement context across 1 environment. Contexts are never compared with one another — what invalidates a comparison.
Spate is run by the author of this benchmark. Every row of that system carries a dagger † saying so, and no published number is reported by the system that produced it.
How to read this
The mark is the uncertainty. The capsule spans the smallest to the largest repetition and the notch is the median — at three repetitions those are the three measurements, and nothing is modelled. Measuring one arm twice under two labels in these sweeps moved it by 7.2%, so a difference narrower than that is the rig rather than the system.
Grey is shown, not ranked. A tuned or stripped arm is drawn on the same axis, because quantifying its difference is the only reason it exists — but it gets no rank and it never sets the scale.
An empty lane is a disowned number. An infra-bound figure keeps its digits and its reason but not its position — a position on a shared axis is itself a claim, and that number describes ClickHouse rather than the system.
A dagger † is a conflict of interest. It marks a system run by the author of this benchmark, on every row that system has, and it renders from the descriptor rather than from anything the site knows about that name. No published number is reported by the system that produced it.
Every axis starts at zero and prints its real end value, so no difference is magnified by a cropped scale. Scales belong to one column of one group and are never shared across groups. No system has a colour, and nothing is coloured by whether its number is good.
c8gd-metal-24xl-ec2-dockerdrain6 arms
decode, flatten, filter, derive — one row per surviving event · drain — how fast the system can go through a fixed corpus · harness v2 · corpus d2-60d7e5bb2a82
| # | System · arm | Throughput per corehigher is better →02.00Mrows/s per core | Throughputhigher is better →025.00Mrows/s | Cores used← lower is better040.00cores | Peak memory← lower is better080.00 GBbytes | Duplicate rows← lower is better01rows | CPU per row← lower is better020.000 µs | Cores, data plane← lower is better040.00cores | Peak memory, data plane← lower is better080.00 GBbytes | Peak charged memory← lower is better080.00 GBbytes | Throttled← lower is better080.00 s | ClickHouse CPU per row← lower is better03.000 µs | GC pause p99← lower is better080.0 ms | detail |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Spate — run by the vendor of this benchmarkNative | 1.59Mrange 3.4% | 20.37Mrange 2.5% | 12.79range 0.9% | 21.84 GBtiesrange 0.7% | 0tiesno spread (3 reps) | 0.628 µsrange 3.3% | 12.79range 0.9% | 21.84 GBtiesrange 0.7% | 21.91 GBtiesrange 0.8% | 0.000 µsno spread (3 reps) | 0.485 µsrange 1.8% | not measured | Detail for Spate Native |
Spate · Nativeclose
Full profile for Spate — every arm, its configuration, and how to tell us we got it wrong. | ||||||||||||||
| 2 | ClickHouse Kafka table engineDistributed forward | 1.13Mrange 0.2% | 18.61Mrange 4.1% | 16.52range 3.9% | 20.93 GBrange 4.7% | 0tiesno spread (3 reps) | 0.887 µsrange 0.2% | 16.52range 3.9% | 20.93 GBrange 4.7% | 21.09 GBrange 5.2% | 408 msrange 366.6% | 0.320 µsrange 0.3% | not measured | Detail for ClickHouse Kafka table engine Distributed forward |
ClickHouse Kafka table engine · Distributed forwardclose
Full profile for ClickHouse Kafka table engine — every arm, its configuration, and how to tell us we got it wrong. | ||||||||||||||
| 3 | Apache FlinkRowBinary | 547krange 4.3% | 5.74Mrange 2.8% | 10.41range 3.8% | 22.44 GBrange 0.2% | 0tiesno spread (3 reps) | 1.828 µsrange 4.3% | 10.37range 3.8% | 21.83 GBtiesrange 0.2% | 22.61 GBrange 0.2% | 16.42 srange 53.4% | 1.721 µsrange 0.7% | 60.0 msrange 8.6% | Detail for Apache Flink RowBinary |
Apache Flink · RowBinaryclose
Full profile for Apache Flink — every arm, its configuration, and how to tell us we got it wrong. | ||||||||||||||
| 4 | Kafka Connect + clickhouse-kafka-connectclickhouse-kafka-connect · RowBinary + MV | 179krange 8.7% | 4.83Mrange 6.5% | 25.57range 7.2% | 68.12 GBrange 0.0% | 0no spread (3 reps) | 5.582 µsrange 8.0% | 25.57range 7.2% | 68.12 GBrange 0.0% | 68.28 GBrange 0.0% | 1.97 srange 63.6% | 0.790 µsrange 1.8% | 36.3 msrange 7.7% | Detail for Kafka Connect + clickhouse-kafka-connect clickhouse-kafka-connect · RowBinary + MV |
Kafka Connect + clickhouse-kafka-connect · clickhouse-kafka-connect · RowBinary + MVclose
Full profile for Kafka Connect + clickhouse-kafka-connect — every arm, its configuration, and how to tell us we got it wrong. | ||||||||||||||
| 5 | VectorJSONEachRow | 66krange 1.0% | 2.10Mrange 0.9% | 31.87range 0.2% | 51.41 GBrange 12.0% | 0tiesno spread (3 reps) | 15.231 µsrange 1.0% | 31.87range 0.2% | 51.41 GBrange 12.0% | 51.61 GBrange 11.9% | 45.39 srange 15.2% | 2.620 µsrange 0.1% | not measured | Detail for Vector JSONEachRow |
Vector · JSONEachRowclose
Full profile for Vector — every arm, its configuration, and how to tell us we got it wrong. | ||||||||||||||
| Spate — run by the vendor of this benchmarkRowBinaryinfra-bound | 1.66Mrange 0.1% | 10.24Mrange 4.2% | 6.15range 4.2% | 15.75 GBrange 1.4% | 0no spread (3 reps) | 0.602 µsrange 0.1% | 6.15range 4.2% | 15.75 GBrange 1.4% | 15.89 GBrange 1.8% | 0.000 µsno spread (3 reps) | 1.465 µsrange 0.5% | not measured | Detail for Spate RowBinary | |
Spate · RowBinaryclose
Full profile for Spate — every arm, its configuration, and how to tell us we got it wrong. | ||||||||||||||
Ranked positions are given only to realistic arms that passed the infrastructure-headroom limit. 1 arm is shown without one. Every column has its own scale and no scale is shared with any other group.
The systems
- Spate native · Apache-2.0 · activerun by the vendor
- Apache Flink jvm · Apache-2.0 · active
- Vector native · MPL-2.0 · active
- Kafka Connect + clickhouse-kafka-connect jvm · Apache-2.0 · active
- ClickHouse Kafka table engine native · Apache-2.0 · active
- Redpanda Connect go · BUSL-1.1 · planned
- Apache Iggy connectors native · Apache-2.0 · planned
Who runs this
Spate is measured here, and Spate is mine. That is a conflict of interest, and the only useful response to one is to make it impossible to hide.
- No published number is reported by the system that produced it. Throughput
is
SELECT count()against ClickHouse. CPU and memory are cgroup v2 counters read by a sidecar container. Correctness is a query against the rows that actually landed. - Every competitor configuration is in the repository, in full — and so is the search that chose it.
- Where Spate loses, that is published with the same prominence as where it wins. A comparison containing only wins is read as marketing and convinces nobody.
If you think an arm is configured badly, that is a bug and I want the pull request.
How to read a number here
Every result carries the version of the system that produced it, the exact image digest, the environment it ran on, and the date. None of that is optional and none of it is typed in by hand — a run whose image digest cannot be read is recorded as failed rather than published.
Three things invalidate comparison outright, and this site refuses to draw records across them rather than quietly averaging:
| If this differs | Then |
|---|---|
harness_version | The measurement protocol changed. Not comparable. |
dataset_version | The corpus or schema changed. Not comparable. |
env_id | Different hardware. Not comparable. |
Softer differences — a ClickHouse patch release, a compiler version — are recorded and shown as a footnote rather than treated as disqualifying.
Runs only ever append. There is no code path in the driver that truncates a results file, and re-running one system does not re-run or overwrite any other. A number later found to be wrong is corrected in a commit of its own, so what changed and why is in the repository's history.
The table shows each arm's most recent reading, so a re-measurement supersedes
the one before it rather than sitting beside it and being ranked against it. That
is a choice about display and not about retention: every sitting ever taken is
still in results/, and nothing selected against here has been deleted. If an
arm's latest sweep produced no number, it is listed as a gap rather than falling
back to the last figure it managed.
What this does not tell you
Read the limitations before citing anything here. The short version: one workload, one machine, one shape of data, and a benchmark whose author has a stake in the outcome.