updated procedures in data.odin to pass url; added research notes for data-streaming

This commit is contained in:
2026-09-06 13:52:39 -06:00
committed by Samuel ONeal
parent 772ee0a106
commit 89a41262af
12 changed files with 2165 additions and 3 deletions
@@ -0,0 +1,147 @@
# 10 — Performance Benchmarks (published data)
Numbers vary a lot by benchmark harness, message shape, and library version.
This doc collects representative published results, normalizes what can be
normalized, and interprets them **for the DESI Explorer use case** (streaming
bulk `ra`/`dec`/`redshift` columns to a renderer).
## Summary table (large/mixed messages, Rust, per-op medians)
From Mechanical Snail's Rust wire-format study (Nov 2025, c6i.4xlarge,
UserProfile sample ~280B JSON; 1M cycles; versions noted in `03`):
| Format | Serialize (µs) | Deserialize (µs) | Size (bytes) | Rel. size | Notes |
|---|---|---|---|---|---|
| JSON (serde_json) | 5.2 | 8.7 | 280 | 1.00× | baseline |
| MessagePack (rmp_serde) | 2.1 | 3.8 | 168 | 0.60× | schema-less |
| Protobuf (prost) | 1.8 | 2.4 | 98 | 0.35× | smallest |
| **FlatBuffers** | **0.9** | **0.3*** | 132 | 0.47× | *single-field access, zero-copy |
| Cap'n Proto (capnp) | **0.5** | **0.2*** | 152 | 0.54× | *single-field access, zero-copy |
| Avro | 3.1 | 4.5 | 105 | 0.38× | |
> *For zero-copy formats, "deserialize" here means *accessing a single field*
> without traversal. Full traversal ≈ 2-3 µs for both — still ~3-30× faster
> than decoding JSON to objects.
Key takeaways:
- **Undisputed**: binary formats are 2-14× faster than JSON; zero-copy formats
are the fastest on the read path.
- **Size leader**: protobuf (varint). FlatBuffers and Cap'n Proto are 30-55%
larger than proto at this message shape.
- **Write leader**: Cap'n (+40% vs FlatBuffers at this shape). Read leader:
Cap'n/FB both near-free for single field.
## Second benchmark (Go, small objects, `format-wars.go`)
| Format | Marshal ns/op | Unmarshal ns/op | Size bytes | Allocs (marshal) |
|---|---|---|---|---|
| MessagePack (msgp) | 69-205 | 155-336 | 157-323 | 1 |
| **FlatBuffers** | **183-414** | **192-450** | 192-448 | **0** |
| Protobuf | 156-1033 | 315-1173 | 82-276 | 1-21 |
| CBOR | 253-453 | 751-1056 | 141-324 | 1-3 |
| JSON (std/v2) | 463-946 | 842-2819 | 213-1940K | 1-3+ |
Highlights: **FlatBuffers is the only format with 0 allocations on marshal**
its CSR for the renderer (no per-frame GC/allocation churn) is genuinely
unique among these.
## Third benchmark (distributed latency, Lviv Polytechnic 2024/2025)
- **Serialize speed**: FlatBuffers fastest across all message sizes (med <2µs;
vs proto 2.8-27µs).
- **Deserialize speed**: FlatBuffers and Cap'n (unpacked) overwhelming leaders
(FlatBuffers med 0.02-0.03µs).
- **End-to-end distributed latency**: Avro/Protobuf/Thrift beat FlatBuffers
(med reductions vs JSON: Avro -84%, Proto -82%, Thrift -80%, Cap'n packed
-71%, **FlatBuffers only -17%**). Reason: **payload size dominates** in
distributed hops; FlatBuffers isn't the most compact.
- **Interpretation**: if your bottleneck is *network* (large hops over WAN),
compact formats (Avro/proto) win end-to-end. If your bottleneck is
*CPU/allocation on massive reads* (a renderer ingesting the same bytes
repeatedly), zero-copy formats win. DESI Explorer is the latter.
## Fourth benchmark (C++, embedded/small messages, CppSerialization 2025)
| Protocol | Message size | Serialize | Deserialize |
|---|---|---|---|
| SBE | 138 B | 35 ns | 52 ns |
| zpp::bits | 130 B | 34 ns | 37 ns |
| Cap'n Proto | 208 B | 247 ns | 184 ns |
| FlatBuffers | 280 B | 272 ns | **81 ns** |
| Protobuf | 120 B | 322 ns | 351 ns |
| JSON | 301 B | 696 ns | 291 ns |
## What these numbers mean for DESI Explorer specifically
### Reads are the whole point
The renderer will read every point **every frame** (projection + color). Even
at a modest 60 fps with 100k visible objects:
| Approach | Per-object read cost (approx) | Per-frame cost @100k |
|---|---|---|
| JSON parse (serde_json to objects) | ~87 ns (1ms/100k) | ~8.7 ms — 52% of a 16.7ms frame |
| protobuf decode to objects | ~24 ns | ~2.4 ms — 14% |
| FlatBuffers (safe_slice into `&[f64]` columns) | ~0.3 ns (slice view) | **~0.03 ms — 0.2%** |
The gap is a **100×** per-frame savings against JSON — the difference between
"render a 100k+ galaxy view comfortably" and "frame budget eaten by parsing."
- The reason the win is so large for your workload: **bulk reads of the same
buffer reconstituted each frame**. Allocating + copying 100k objects per
frame is what you're avoiding.
- If instead you decode protobuf **once** into a renderer-owned
`[]rl.Vector3` buffer (not per frame), protobuf becomes competitive — you're
paying allocation once, not per frame. That's the strongest counter-case to
FlatBuffers, and worth benchmarking before you commit (see below).
### Writes (server side) are a non-issue
Server serialization of a 50k-object batch, even at FlatBuffers' ~1-3µs/MB
shape, is negligible next to DB query + network. If it ever matters, reuse the
`FlatBufferBuilder` (with `reset()`) and use `with_internal_capacity` to avoid
realloc spikes. FlatBuffers' *builder* is not the fastest writer, but writers
are amortized; **readers are per-frame**. Optimize where the frame is.
### Wire size tradeoff is acceptable
On ~24 MB of raw f64 columns per 100k objects, FlatBuffers columnar adds only
vtables per batch (≈ tens of bytes), not per row — dropping closer to the raw
size than the "132 vs 98 vs proto" small-object ratio suggests. Columnar **is**
where the size concern basically evaporates.
Bandwidth-wise: LAN/k3s homelab is not bandwidth-constrained; a 3 MB batch
arrives in ~ms.
## Recommended local benchmark (before committing)
Write a criterion benchmark on *your* exact shape:
1. 100k rows × columns `ra/dec/redshift: f64` (+optional `ids: u64`).
2. Variants:
- serde JSON (current contract) → parse + map to Vec<CatalogObject>
- FlatBuffers **columnar** batch → `safe_slice` reads
- protobuf (prost) → decode to Vec
- Arrow IPC record batch → batch slices
3. Metric: **per-frame read cost** (60fps frame budget) + wire bytes + first-batch
latency.
4. Also measure: build time (server), verify time at ingest, delta vs `_unchecked`.
Reality check: the numbers here point to FlatBuffers (or Arrow IPC) winning the
*frame-cost* metric decisively; protobuf wins the *wire-size* metric; JSON wins
only developer ergonomics. Decide based on what frame cost is worth to you.
## Sources
- Mechanical Snail, "Wire Formats for High-Volume Service Communication" (2025-11-24):
https://mechanicalsnail.com/posts/wire-formats-rust/
- kapetan-io/format-wars.go (Go, small objects, allocs):
https://github.com/kapetan-io/format-wars.go
- Maltsev et al., "Impact of Serialization Format on Inter-Service Latency"
(ACPS 2024): https://science.lpnu.ua/sites/default/files/journal-paper/2024/dec/36976/...
- chronoxor/CppSerialization (C++, embedded/micro):
https://github.com/chronoxor/CppSerialization
- Shekhar Manna, "Binary Serialization Formats — A Technical Benchmark & Decision Guide" (2026-02):
https://medium.com/@shekhar.manna83/binary-serialization-formats-e2703f053010
- ADHDecode, "Protobuf vs FlatBuffers vs Cap'n Proto" (2026-03):
https://adhdecode.com/api-architecture/grpc-deep-dive/protobuf-vs-flatbuffers-vs-capn-proto/