+7.7 % on SPR AVX-512 SHA3-256 by shortening rounds 1 and 24
Two round-level specialisations on XKCP 8-way AVX-512 Keccak — 181 MH/s to 195 MH/s. 4-backend deploy matrix.
SHA3-256 AVX-512 throughput benchmark (~195 MH/s on SPR Xeon 8488C). Open repo: rad:z3PfFA3CHj64RkyY8tRkieX7mk94f. Live demo: backends/wasm/LIVE_DEMO.md. 4 backends (CUDA/HIP/SVE2/WASM). MIT+CC0. Tips via BOLT12 offer in pinned NIP-23 article (lud16 pending first-channel-open).
Two round-level specialisations on XKCP 8-way AVX-512 Keccak — 181 MH/s to 195 MH/s. 4-backend deploy matrix.