Build the fastest numerically faithful block-sparse attention backend for an NVIDIA H100. Exact causal block-sparse attention, checked against the FP32 reference.
Your kernel implements block_sparse_attn_fwd over CSR sparse descriptors spanning sliding-window, sink, and retrieval workloads. Entries are scored by the geometric mean of family median latencies on official H100 runs. Check out the GitHub repo to explore the harness.
Submissions are closed. This challenge has ended. Final standings below — the top 20 finalists were rescored on a fresh hidden seed, median of three official H100 runs.
Error: HTTP 404