GPU Benchmark Results

GPU Information

CPU Information

Median time is reported in milliseconds for ok records. Inputs are prepared on the GPU before timed runs; initial host-to-device transfer is outside the timed region. Timed runs include the host API call and backend-native device synchronization. tenferro-rs CUDA uses the explicit tenferro-rs synchronize API without downloading result tensors in the timed region. Dense and einsum inputs use the same deterministic benchmark generator in the Rust and Python runners; host-to-device layout conversion remains outside the timed region. tenferro-rs uses native column-major GPU tensors; PyTorch and vendor-wrapper columns use their native row-major framework tensors unless noted. The cuSOLVER column is torch.linalg with preferred_linalg_library=cusolver; for SVD it pins driver=gesvd as a QR-based cuSOLVER comparison. tenferro-rs CUDA SVD uses its backend default driver policy, currently gesvdj for matrices with both dimensions at most 1024 and gesvd otherwise. Non-ok cells show the structured backend status.

gpu/dense / batched_matmul / allocating output

Problem tenferro-rs CUDA trace tenferro-rs CUDA eager PyTorch CUDA
dense_batched_matmul_f64_b1024_256 83.469 83.386 83.352

gpu/dense / eigh / allocating output

Problem tenferro-rs CUDA trace tenferro-rs CUDA eager PyTorch CUDA
dense_eigh_f64_1024 44.518 44.504 44.335

gpu/dense / matmul / allocating output

Problem tenferro-rs CUDA trace tenferro-rs CUDA eager PyTorch CUDA
dense_matmul_f64_3072 141.040 140.960 140.872

gpu/dense / qr / allocating output

Problem tenferro-rs CUDA trace tenferro-rs CUDA eager PyTorch CUDA
dense_qr_f64_1536 55.834 55.802 55.065

gpu/dense / solve / allocating output

Problem tenferro-rs CUDA trace tenferro-rs CUDA eager PyTorch CUDA
dense_solve_f64_1024_rhs16 8.552 8.386 7.877
dense_solve_f64_2048_rhs128 32.789 32.754 31.822
dense_solve_f64_2048_rhs16 28.300 28.088 27.568
dense_solve_f64_4096_rhs16 148.177 148.064 146.961
dense_solve_f64_512_rhs16 3.009 3.265 2.757

gpu/dense / svd / allocating output

SVD note: SVD rows use synchronized timed regions and matched Rust/Python input generators. tenferro-rs CUDA uses its backend default driver policy, currently gesvdj for matrices with both dimensions at most 1024 and gesvd otherwise. The cuSOLVER column pins torch.linalg.svd driver=gesvd as a QR-based cuSOLVER comparison; PyTorch's default row may use a different SVD driver and row-major framework layout.

Problem tenferro-rs CUDA trace tenferro-rs CUDA eager PyTorch CUDA
dense_svd_f64_256 77.527 77.490 79.282