GPU Benchmark Results

GPU Information

CPU Information

Median time is reported in milliseconds for ok records. Inputs are prepared on the GPU before timed runs; initial host-to-device transfer is outside the timed region. Timed runs include the host API call and backend-native device synchronization. tenferro-rs CUDA uses the explicit tenferro-rs synchronize API without downloading result tensors in the timed region. Dense and einsum inputs use the same deterministic benchmark generator in the Rust and Python runners; host-to-device layout conversion remains outside the timed region. tenferro-rs uses native column-major GPU tensors; PyTorch and vendor-wrapper columns use their native row-major framework tensors unless noted. The cuSOLVER column is torch.linalg with preferred_linalg_library=cusolver; for SVD it pins driver=gesvd as a QR-based cuSOLVER comparison. tenferro-rs CUDA SVD uses its backend default driver policy, currently gesvdj for matrices with both dimensions at most 1024 and gesvd otherwise. Non-ok cells show the structured backend status.

gpu/sparse / spmm / allocating output

Problem tenferro-rs CUDA trace tenferro-rs CUDA eager PyTorch CUDA cuBLASLt CUTLASS cuSOLVER cuSPARSE Ginkgo
sparse_synthetic_64k_4m_spmm_f64_rhs1024 unsupported unsupported 14.737 unsupported unsupported unsupported 14.984 verification failed

gpu/sparse / spmv / allocating output

Problem tenferro-rs CUDA trace tenferro-rs CUDA eager PyTorch CUDA cuBLASLt CUTLASS cuSOLVER cuSPARSE Ginkgo
sparse_synthetic_4m_64m_spmv_f64 unsupported unsupported 2.334 unsupported unsupported unsupported 2.328 verification failed