GPU Benchmark Results

GPU Information

CPU Information

Median time is reported in milliseconds for ok records. Inputs are prepared on the GPU before timed runs; initial host-to-device transfer is outside the timed region. Timed runs include the host API call and backend-native device synchronization. tenferro-rs CUDA uses the explicit tenferro-rs synchronize API without downloading result tensors in the timed region. Dense and einsum inputs use the same deterministic benchmark generator in the Rust and Python runners; host-to-device layout conversion remains outside the timed region. tenferro-rs uses native column-major GPU tensors; PyTorch and vendor-wrapper columns use their native row-major framework tensors unless noted. The cuSOLVER column is torch.linalg with preferred_linalg_library=cusolver; for SVD it pins driver=gesvd as a QR-based cuSOLVER comparison. tenferro-rs CUDA SVD uses its backend default driver policy, currently gesvdj for matrices with both dimensions at most 1024 and gesvd otherwise. Non-ok cells show the structured backend status.

gpu/einsum / einsum / allocating output

Problem tenferro-rs CUDA trace tenferro-rs CUDA eager PyTorch CUDA cuBLASLt CUTLASS cuSOLVER cuSPARSE Ginkgo
einsum_bin_matmul_3072_f64 3.200 3.165 3.173 3.181 3.479 unsupported unsupported unsupported