7.0 KiB
10 — Performance Benchmarks (published data)
Numbers vary a lot by benchmark harness, message shape, and library version.
This doc collects representative published results, normalizes what can be
normalized, and interprets them for the DESI Explorer use case (streaming
bulk ra/dec/redshift columns to a renderer).
Summary table (large/mixed messages, Rust, per-op medians)
From Mechanical Snail's Rust wire-format study (Nov 2025, c6i.4xlarge,
UserProfile sample ~280B JSON; 1M cycles; versions noted in 03):
| Format | Serialize (µs) | Deserialize (µs) | Size (bytes) | Rel. size | Notes |
|---|---|---|---|---|---|
| JSON (serde_json) | 5.2 | 8.7 | 280 | 1.00× | baseline |
| MessagePack (rmp_serde) | 2.1 | 3.8 | 168 | 0.60× | schema-less |
| Protobuf (prost) | 1.8 | 2.4 | 98 | 0.35× | smallest |
| FlatBuffers | 0.9 | 0.3* | 132 | 0.47× | *single-field access, zero-copy |
| Cap'n Proto (capnp) | 0.5 | 0.2* | 152 | 0.54× | *single-field access, zero-copy |
| Avro | 3.1 | 4.5 | 105 | 0.38× |
*For zero-copy formats, "deserialize" here means accessing a single field without traversal. Full traversal ≈ 2-3 µs for both — still ~3-30× faster than decoding JSON to objects.
Key takeaways:
- Undisputed: binary formats are 2-14× faster than JSON; zero-copy formats are the fastest on the read path.
- Size leader: protobuf (varint). FlatBuffers and Cap'n Proto are 30-55% larger than proto at this message shape.
- Write leader: Cap'n (+40% vs FlatBuffers at this shape). Read leader: Cap'n/FB both near-free for single field.
Second benchmark (Go, small objects, format-wars.go)
| Format | Marshal ns/op | Unmarshal ns/op | Size bytes | Allocs (marshal) |
|---|---|---|---|---|
| MessagePack (msgp) | 69-205 | 155-336 | 157-323 | 1 |
| FlatBuffers | 183-414 | 192-450 | 192-448 | 0 |
| Protobuf | 156-1033 | 315-1173 | 82-276 | 1-21 |
| CBOR | 253-453 | 751-1056 | 141-324 | 1-3 |
| JSON (std/v2) | 463-946 | 842-2819 | 213-1940K | 1-3+ |
Highlights: FlatBuffers is the only format with 0 allocations on marshal — its CSR for the renderer (no per-frame GC/allocation churn) is genuinely unique among these.
Third benchmark (distributed latency, Lviv Polytechnic 2024/2025)
- Serialize speed: FlatBuffers fastest across all message sizes (med <2µs; vs proto 2.8-27µs).
- Deserialize speed: FlatBuffers and Cap'n (unpacked) overwhelming leaders (FlatBuffers med 0.02-0.03µs).
- End-to-end distributed latency: Avro/Protobuf/Thrift beat FlatBuffers (med reductions vs JSON: Avro -84%, Proto -82%, Thrift -80%, Cap'n packed -71%, FlatBuffers only -17%). Reason: payload size dominates in distributed hops; FlatBuffers isn't the most compact.
- Interpretation: if your bottleneck is network (large hops over WAN), compact formats (Avro/proto) win end-to-end. If your bottleneck is CPU/allocation on massive reads (a renderer ingesting the same bytes repeatedly), zero-copy formats win. DESI Explorer is the latter.
Fourth benchmark (C++, embedded/small messages, CppSerialization 2025)
| Protocol | Message size | Serialize | Deserialize |
|---|---|---|---|
| SBE | 138 B | 35 ns | 52 ns |
| zpp::bits | 130 B | 34 ns | 37 ns |
| Cap'n Proto | 208 B | 247 ns | 184 ns |
| FlatBuffers | 280 B | 272 ns | 81 ns |
| Protobuf | 120 B | 322 ns | 351 ns |
| JSON | 301 B | 696 ns | 291 ns |
What these numbers mean for DESI Explorer specifically
Reads are the whole point
The renderer will read every point every frame (projection + color). Even at a modest 60 fps with 100k visible objects:
| Approach | Per-object read cost (approx) | Per-frame cost @100k |
|---|---|---|
| JSON parse (serde_json to objects) | ~87 ns (1ms/100k) | ~8.7 ms — 52% of a 16.7ms frame |
| protobuf decode to objects | ~24 ns | ~2.4 ms — 14% |
FlatBuffers (safe_slice into &[f64] columns) |
~0.3 ns (slice view) | ~0.03 ms — 0.2% |
The gap is a 100× per-frame savings against JSON — the difference between "render a 100k+ galaxy view comfortably" and "frame budget eaten by parsing."
- The reason the win is so large for your workload: bulk reads of the same buffer reconstituted each frame. Allocating + copying 100k objects per frame is what you're avoiding.
- If instead you decode protobuf once into a renderer-owned
[]rl.Vector3buffer (not per frame), protobuf becomes competitive — you're paying allocation once, not per frame. That's the strongest counter-case to FlatBuffers, and worth benchmarking before you commit (see below).
Writes (server side) are a non-issue
Server serialization of a 50k-object batch, even at FlatBuffers' ~1-3µs/MB
shape, is negligible next to DB query + network. If it ever matters, reuse the
FlatBufferBuilder (with reset()) and use with_internal_capacity to avoid
realloc spikes. FlatBuffers' builder is not the fastest writer, but writers
are amortized; readers are per-frame. Optimize where the frame is.
Wire size tradeoff is acceptable
On ~24 MB of raw f64 columns per 100k objects, FlatBuffers columnar adds only vtables per batch (≈ tens of bytes), not per row — dropping closer to the raw size than the "132 vs 98 vs proto" small-object ratio suggests. Columnar is where the size concern basically evaporates.
Bandwidth-wise: LAN/k3s homelab is not bandwidth-constrained; a 3 MB batch arrives in ~ms.
Recommended local benchmark (before committing)
Write a criterion benchmark on your exact shape:
- 100k rows × columns
ra/dec/redshift: f64(+optionalids: u64). - Variants:
- serde JSON (current contract) → parse + map to Vec
- FlatBuffers columnar batch →
safe_slicereads - protobuf (prost) → decode to Vec
- Arrow IPC record batch → batch slices
- Metric: per-frame read cost (60fps frame budget) + wire bytes + first-batch latency.
- Also measure: build time (server), verify time at ingest, delta vs
_unchecked.
Reality check: the numbers here point to FlatBuffers (or Arrow IPC) winning the frame-cost metric decisively; protobuf wins the wire-size metric; JSON wins only developer ergonomics. Decide based on what frame cost is worth to you.
Sources
- Mechanical Snail, "Wire Formats for High-Volume Service Communication" (2025-11-24): https://mechanicalsnail.com/posts/wire-formats-rust/
- kapetan-io/format-wars.go (Go, small objects, allocs): https://github.com/kapetan-io/format-wars.go
- Maltsev et al., "Impact of Serialization Format on Inter-Service Latency" (ACPS 2024): https://science.lpnu.ua/sites/default/files/journal-paper/2024/dec/36976/...
- chronoxor/CppSerialization (C++, embedded/micro): https://github.com/chronoxor/CppSerialization
- Shekhar Manna, "Binary Serialization Formats — A Technical Benchmark & Decision Guide" (2026-02): https://medium.com/@shekhar.manna83/binary-serialization-formats-e2703f053010
- ADHDecode, "Protobuf vs FlatBuffers vs Cap'n Proto" (2026-03): https://adhdecode.com/api-architecture/grpc-deep-dive/protobuf-vs-flatbuffers-vs-capn-proto/