nvidia-gpugpu/tensornetworkbenchmarks/gpu/tensornetwork.yaml2026-10-06T19:36:19.553911+00:005cf78c7ec0ad9516dd78bab546d5e7bd42fa3102cuda:0NVIDIA A100 80GB PCIeGPU-530977e1-4968-9283-4129-9fbec3e6654280 GiB580.126.0913.012.992700AMD EPYC 7713P 64-Core ProcessorAuthenticAMD6416411Linux-6.8.0-101-generic-x86_64-with-glibc2.39Median time is reported in milliseconds for ok records.
Inputs are prepared on the GPU before timed runs; initial host-to-device transfer is outside the timed region.
Timed runs include the host API call and backend-native device synchronization. tenferro-rs CUDA uses the explicit tenferro-rs synchronize API without downloading result tensors in the timed region.
Dense and einsum inputs use the same deterministic benchmark generator in the Rust and Python runners; host-to-device layout conversion remains outside the timed region.
tenferro-rs uses native column-major GPU tensors; PyTorch and vendor-wrapper columns use their native row-major framework tensors unless noted.
The cuSOLVER column is torch.linalg with preferred_linalg_library=cusolver; for SVD it pins driver=gesvd as a QR-based cuSOLVER comparison. tenferro-rs CUDA SVD uses its backend default driver policy, currently gesvdj for matrices with both dimensions at most 1024 and gesvd otherwise.
Non-ok cells show the structured backend status.
| Problem | tenferro-rs CUDA trace | tenferro-rs CUDA eager | PyTorch CUDA |
|---|---|---|---|
| tensornetwork_permutation_optimized_f32 | 56.340 | 63.180 | 111.084 |