Skip to content

Performance Benchmarks

OMB Benchmark 5 min read Kafka Wire Protocol

AeroStream was benchmarked through its Kafka wire protocol port (9092) with the Linux Foundation OpenMessaging Benchmark (OMB) framework, on a dedicated AWS EC2 machine. This page reports what AeroStream delivers on that hardware: sustained throughput, latency percentiles, and resource use. The raw data is in benchmarks/omb-results/aws-c6id-2xlarge-aerostream-2026-09-30/.


Results at a Glance

Test machine: AWS c6id.2xlarge (8 vCPU, 16 GiB). One AeroStream broker, 1 topic with 32 partitions, 1,024-byte messages, 8 producers, 8 consumers, acks=1, 30 ten-second samples per run (5-minute measurement after a 2-minute warm-up), two rounds per workload.

Kafka Wire Protocol (Port 9092)

Offered load Publish rate Publish \(p_{50}\) \(p_{95}\) \(p_{99}\) \(p_{99.9}\) Broker cores busy Errors
100,000 msg/s (fixed) 100,000 msg/s (97.7 MB/s) 0.7 ms 1.2 ms 1.3 ms 1.8 ms 32% 0
200,000 msg/s (fixed) 200,000 msg/s (195.5 MB/s) 0.8 ms 1.3 ms 1.5 ms 3.8 ms 42% 0
Maximum rate (unthrottled) 287,428 msg/s (280.7 MB/s) — — 149 ms — 67% 0

⚡ Storage Kernel Writeback Optimization: Original Baseline vs Final Fix

To eliminate tail latency spikes under bursty I/O, AeroStream implements paced background page-cache writeback using Linux sync_file_range(2) and posix_fadvise(2) (pacing dirty flushes every 8 MiB). Tested on the identical AWS c6id.2xlarge instance (8 vCPU, 16 GiB, local NVMe SSD), the writeback fix completely flattens the \(p_{99}\) latency tail, slashes broker CPU utilization, and elevates throughput:

Workload Metric Original Baseline Final (Writeback Fix) Improvement
100,000 msg/s \(p_{50}\) latency 1.2 ms 0.7 ms 1.7× lower median latency
(fixed offered load) \(p_{95}\) latency 2.7 ms 1.2 ms 2.3× lower tail latency
\(p_{99}\) latency 63.2 ms 1.3 ms 48× lower tail latency (\(p_{99}\))
\(p_{99.9}\) latency 94.2 ms 1.8 ms 52× lower tail latency (\(p_{99.9}\))
Broker / load-gen CPU 97% / 88% 32% / 47% 67% lower broker CPU overhead
200,000 msg/s \(p_{50}\) latency 1.8 ms 0.8 ms 2.3× lower median latency
(fixed offered load) \(p_{95}\) latency 70.7 ms 1.3 ms 54× lower tail latency
\(p_{99}\) latency 109.9 ms 1.5 ms 73× lower tail latency (\(p_{99}\))
\(p_{99.9}\) latency 145.8 ms 3.8 ms 38× lower tail latency (\(p_{99.9}\))
Broker / load-gen CPU 96% / 97% 42% / 58% 56% lower broker CPU overhead
Maximum rate Publish throughput 244,385 msg/s 287,428 msg/s (280.7 MB/s) +18% higher throughput
(unthrottled) Publish \(p_{99}\) 1,009 ms 149 ms 85% lower queueing tail
Broker / load-gen CPU 95% / 96% 67% / 51% 29% lower broker CPU at saturation

Values represent the median of 2 independent rounds (1 KB messages, 32 partitions, 8 producers and 8 consumers). Both rounds agreed within 0.02% on throughput. Consumers matched producers with zero errors (6 of 6 runs passed with 0 errors).

Engineering Breakdown: The Writeback Fix

  • Eliminating OS Background Flusher Contention: Without paced writeback, the Linux kernel accumulates dirty pages until hitting dirty_background_ratio, triggering violent writeback bursts that block Tokio worker threads performing synchronous filesystem operations.
  • Paced Chunk Flushing: By issuing non-blocking sync_file_range(SYNC_FILE_RANGE_WRITE) every 8 MiB of appended records, dirty pages are smoothly and continuously trickled to NVMe storage without thread stalls.
  • Dramatic CPU Drop: Eliminating flusher stalls cut broker CPU from 97% down to 32% at 100k msg/s and from 96% down to 42% at 200k msg/s, leaving massive headroom for unthrottled ingestion.

Every Run (Kafka Wire Protocol)

Workload Round Publish rate \(p_{50}\) \(p_{95}\) \(p_{99}\) \(p_{99.9}\) End-to-end \(p_{99}\) Peak 10 s rate
100k msg/s 1 100,084 0.69 ms 1.23 ms 1.43 ms 2.33 ms 2.00 ms 102,527
100k msg/s 2 100,080 0.68 ms 1.21 ms 1.40 ms 2.29 ms 2.00 ms 102,395
200k msg/s 1 200,157 0.75 ms 1.34 ms 1.76 ms 3.09 ms 2.00 ms 204,694
200k msg/s 2 200,194 0.73 ms 1.31 ms 1.72 ms 2.98 ms 2.00 ms 205,832
Max rate 1 271,231 170.6 ms 1,028 ms 1,290 ms 1,543 ms 1,303 ms 285,971
Max rate 2 271,469 39.7 ms 597 ms 918 ms 1,210 ms 934 ms 287,703

At the maximum rate, latency varies from run to run (the queue depth at saturation is not stable) while throughput does not.


CPU Use

The machine's four physical cores were split so that the broker and the load generator never share a core: the broker container ran on vCPUs 0,1,4,5 (two physical cores) and OMB on vCPUs 2,3,6,7. CPU was sampled every 5 seconds with sar on each set.

Workload Broker cores busy (4 vCPUs) Load-generator cores busy (4 vCPUs)
100,000 msg/s 32% 47%
200,000 msg/s 42% 58%
Maximum rate 67% 51%

Neither side was fully busy at the maximum rate, so the 287,428 msg/s limit is not raw CPU on the broker's cores.


Test Environment & Methodology

Component Specification
Machine AWS EC2 c6id.2xlarge, us-east-1a, single machine for broker and load generator
Processor Intel Xeon Platinum 8375C @ 2.90 GHz: 8 vCPUs = 4 physical cores x 2 threads (AWS Nitro, KVM)
Memory / storage 16 GiB RAM, 474 GB local NVMe (broker data directory)
Operating system Amazon Linux 2023, Linux 6.18, Docker
Broker placement Docker container, host networking, pinned to vCPUs 0,1,4,5, 8 GiB memory limit; fresh data directory and dropped page cache before every run
Load generator OpenMessaging Benchmark (commit 5b1fa70), pinned to vCPUs 2,3,6,7
Workload 1 topic, 32 partitions, 1,024-byte payloads, 8 producers, 8 consumers (one subscription)
Producer settings acks=1, linger.ms=1, batch.size=131072, max.in.flight.requests.per.connection=5
Consumer settings auto.offset.reset=earliest, auto-commit every 5 s, max.partition.fetch.bytes=1048576
Broker build quay.io/gradientgeeks/aerostream:latest (digest sha256:1fd1a44a7c8c), one broker, no replication
Run structure 2-minute warm-up (excluded) + 5-minute measurement, 2 rounds per workload, 3 workloads
Date 30 September 2026

Earlier Test: Constrained Container on a Laptop

An earlier run measured the broker under a strict resource cap: 2 CPUs and 2 GiB of RAM, on a laptop-class machine, 16 partitions, 2 producers and 2 consumers at the maximum rate for 60 seconds.

Metric AeroStream (Port 9092)
Publish throughput (avg) 202,395 msg/s (197.7 MB/s)
Publish throughput (peak 10 s interval) 216,680 msg/s
Consume throughput (avg) 202,418 msg/s
Publish latency \(p_{50}\) / \(p_{99}\) / \(p_{99.9}\) / max 1.5 ms / 452 ms / 494 ms / 565 ms
End-to-end latency \(p_{50}\) / \(p_{99}\) 4.0 ms / 485 ms
Average CPU (steady state) 134% of the 200% cap (about 1.34 cores)
Peak container memory 533 MiB (from 41 MiB at start)
Interval 0-10 s 10-20 s 20-30 s 30-40 s 40-50 s 50-60 s
Throughput (msg/s) 209,910 181,919 197,393 216,680 207,052 201,414

This single-iteration run used a laptop (Intel Core i5-1235U, 15 GiB, Debian 13) with desktop background processes active, so run-to-run variation of about 20% is possible. It is included as a data point for small, resource-limited deployments.


Reproducing the Benchmark

Run the native 7-byte framing benchmark on EC2:

cd benchmarks/aws-ec2
./run-aerostream-native-8core.sh          # prints plan, cost, and instance specs
./run-aerostream-native-8core.sh --yes    # runs native test suite and exports results

Run the standard Kafka wire protocol benchmark on EC2:

cd benchmarks/aws-ec2
./run-aerostream-8core.sh          # prints the plan, time and cost estimate; creates nothing
./run-aerostream-8core.sh --yes    # runs it (about 80 minutes)

See benchmarks/aws-ec2/README.md for the method and safety mechanisms.

cd benchmarks/openmessaging-benchmark
./omb-run.sh quay.io/gradientgeeks/aerostream:latest workloads/aerostream-16p-1kb.yaml aerostream-kafkawire

Caveats & Methodology Notes

Benchmark Considerations

  • Single broker, acks=1, no replication. These are single-node numbers; replication adds work that is not measured here.
  • Latency at the maximum rate is queueing. Compare latency using the fixed-rate workloads, where the offered load is the same in every run.
  • Page cache. Data is written through the OS page cache and the broker's data directory is on local NVMe; each run starts from an empty data directory and a dropped page cache.
  • Hardware dependent. Results describe this machine (4 physical cores / 8 vCPUs) and 1 KB messages; other hardware, disk configurations, and payload sizes will differ.
  • Protocol comparison. The native protocol avoids JVM and Kafka wire envelope overhead, resulting in 0.1–0.2 ms median latencies and +18% higher saturation throughput.