Benchmarks
Similar bulk speed at both widths; extra cost on short inputs. These runs measure the current digest on the hosts below. Use them to compare within a run, then measure your own workload.
64 vs 128 bits · WebAssembly · 128-bit comparison · ChibiHash · Reproduce
For a quick check on your machine, run the browser speed test.
| Apple M1 Pro | P-core, bare metal. Apple clang 21
-O3 -mcpu=native. |
|---|---|
| AMD Ryzen AI 9 HX PRO 370 | Zen 5 P-core 0, bare metal. GCC 16,
-O3 -march=native. |
| AMD EPYC 9655 | Zen 5 core 0, KVM guest without frequency control. GCC 13,
-O3 -march=native. Its absolute rate is
indicative; within-host ratios are useful. |
Cost of the 128-bit result
Both widths share one pass over the input. Each cell shows 64-bit / 128-bit results: the median of nine calibrated rounds of about 40 ms on one pinned core.
| host | 8 B chained (ns) | 8 B independent (ns) | 1 MiB (GB/s) |
|---|---|---|---|
| Apple M1 Pro | 7.88 / 9.25 | 2.70 / 3.81 | 30.73 / 30.76 |
| Ryzen AI 9 HX PRO 370 | 4.29 / 4.90 | 1.97 / 2.67 | 61.47 / 61.31 |
| EPYC 9655 (KVM) | 4.92 / 5.66 | 2.14 / 3.27 | 54.15 / 53.99 |
Chained calls feed each result into the next seed and measure latency. Independent calls can overlap. Lower ns is better; higher GB/s is better.
Baseline wasm32 on the M1 Pro
Zig 0.16 compiled all rows into one wasm32-wasi module
with -O3, no SIMD, and no wide multiply. The timing
loops ran inside wasm under Node 26 / V8, so the JS boundary is
outside every measurement.
| hash | 8 B chained (ns/hash) | 1 MiB (GB/s) |
|---|---|---|
| hayahash128 | 10.6 | 23.34 |
| hayahash64 | 7.7 | 23.55 |
| ChibiHash v2 | 10.4 | 18.57 |
| XXH3-64 | 8.2 | 17.41 |
| XXH64 | 5.9 | 14.71 |
| rapidhash v3 | 22.6 | 6.43 |
hayahash128 retains 99% of hayahash64's bulk rate here. All other rows return 64 bits.
SMHasher3 128-bit shootout
The current digest is compared using SMHasher3
commit 51d3cd1a. Speed cells are medians of three
independent, round-robin processes on the two bare-metal hosts.
Small-key latency is the corrected 1 to 31 byte average; bulk is
the fixed 256 KiB average. The full suite ran separately on the
EPYC 9655 with 128-bit-wide expectations.
| 128-bit hash | M1 small (cy) | M1 bulk (B/cy) | Zen 5 small (cy) | Zen 5 bulk (B/cy) | full suite | needs for peak speed |
|---|---|---|---|---|---|---|
| hayahash128 | 38.53 | 9.77 | 13.71 | 31.02 | pass | ordinary scalar source; auto-vectorized Zen 5 bulk |
| MuseAir-128 | 24.26 | 8.65 | 7.44 | 22.80 | pass | 64×64→128 multiply |
| a5hash-128 | 22.29 | 10.92 | 6.00 | 22.99 | pass | 64×64→128 multiply |
| MeowHash | - | - | 28.64 | 32.22 | pass | x86 AES instructions |
| XXH3-128 | 30.73 | 12.64 | 11.92 | 48.91 | fail (26) | wide multiply; SIMD for peak bulk |
| t1ha2-128 | 64.58 | 5.86 | 21.03 | 16.24 | pass | 64×64→128 multiply |
| SpookyHash2-128 | 53.37 | 4.20 | 24.39 | 15.35 | fail (10) | ordinary 64-bit operations |
| prvhash-128 | 67.45 | 1.00 | 25.49 | 2.58 | pass (187/187) | ordinary 64-bit operations |
| FarmHash-128.CC.seed1 | 60.09 | 5.64 | 21.09 | 16.45 | pass | ordinary 64-bit operations |
Small-key values are dependent latency, not independent throughput. Their raw process averages are corrected to one call-overhead baseline per host; the raw records contain every process value and the calculation. MeowHash has no M1 row because this implementation needs x86 AES instructions.
XXH3-128 has the highest bulk throughput on both hosts and fails 26 test groups. Among hashes that pass, a5hash has the lowest small-key latency. hayahash128's Zen 5 bulk result uses compiler auto-vectorization.
hayahash128 leads the passing ordinary-scalar implementations in this comparison. prvhash has 187 applicable groups; the others have 188. See the raw records and measurement method.
64-bit coverage
The native 64-bit comparison against wide-multiply and SIMD hashes has not been rerun since the digest changed. Current comparisons are limited to baseline wasm and ChibiHash.
Comparison: ChibiHash
ChibiHash uses the
same portable arithmetic, making it a useful baseline. Reproduce
this comparison with make -C tests run-bench; both
ChibiHash versions are vendored in tests/.
Apple M1, clang -O3 -mcpu=native
| size (B) | chibihash v1 | chibihash v2 | hayahash64 | hayahash128 |
|---|---|---|---|---|
| 64 | 7.70 | 10.04 | 13.00 | 10.39 |
| 256 | 15.84 | 18.00 | 20.67 | 19.03 |
| 1024 | 15.62 | 19.72 | 27.55 | 26.28 |
| 16384 | 14.76 | 19.00 | 30.41 | 30.38 |
| 1048576 | 14.73 | 18.99 | 30.73 | 30.76 |
| len (B) | chibihash v1 | chibihash v2 | hayahash64 | hayahash128 |
|---|---|---|---|---|
| 4 | 9.62 | 10.00 | 7.88 | 9.23 |
| 8 | 6.32 | 9.54 | 7.88 | 9.25 |
| 16 | 6.70 | 9.77 | 7.87 | 9.25 |
| 32 | 12.13 | 11.34 | 8.41 | 10.52 |
| 64 | 13.98 | 12.91 | 9.68 | 11.74 |
| 128 | 18.27 | 16.52 | 12.18 | 14.46 |
| len (B) | chibihash v1 | chibihash v2 | hayahash64 | hayahash128 |
|---|---|---|---|---|
| 4 | 7.13 | 4.21 | 2.76 | 3.83 |
| 8 | 4.41 | 4.42 | 2.70 | 3.81 |
| 16 | 5.05 | 4.88 | 2.70 | 3.81 |
| 32 | 6.81 | 5.17 | 3.72 | 4.98 |
| 64 | 8.38 | 6.38 | 4.92 | 6.15 |
| 128 | 11.69 | 8.90 | 7.38 | 8.53 |
hayahash64 is faster than ChibiHash v2 at every size shown. ChibiHash v1 wins some short-input timings but fails SMHasher3. Fixed-size measurements do not capture mixed-size workloads; see dispatch tradeoffs.
AMD Zen 5, GCC 16 -march=native
| measure | chibihash v2 | hayahash64 | hayahash128 |
|---|---|---|---|
| 128 B, independent hashes (ns, lower is better) | 6.48 | 5.31 | 6.06 |
| 128 B, bulk speed (GB/s, higher is better) | 19.84 | 24.28 | 21.26 |
| 512 B, bulk speed (GB/s) | 29.32 | 42.06 | 38.01 |
| 1 MiB, bulk speed (GB/s) | 31.13 | 61.47 | 61.31 |
| 1 MiB, bulk speed, build without AVX-512DQ (GB/s) | 31.2 | ~35 | not measured |
The 1 MiB result uses AVX-512DQ auto-vectorization. Without it, hayahash64 reaches about 35 GB/s from the same portable source. Both builds produce identical hashes.
Reproduce it
# The ChibiHash comparison. The references are vendored in tests/.
make -C tests run-bench
# SMHasher3, pinned to an exact upstream commit.
make -C tests/smhasher3 run
# The wasm32 shootout and the wasm-against-native bit-exactness check.
make -C tests/wasm run-kat run-bench
See the SMHasher3 guide for the full build matrix and timing corrections. Raw records include hosts, compilers, dispatch shapes, and source checksums.