nvidia-gpugpu/elementwisebenchmarks/gpu/elementwise.yaml2026-10-06T19:43:50.100370+00:005cf78c7ec0ad9516dd78bab546d5e7bd42fa3102cuda:0NVIDIA A100 80GB PCIeGPU-530977e1-4968-9283-4129-9fbec3e6654280 GiB580.126.0913.012.992700AMD EPYC 7713P 64-Core ProcessorAuthenticAMD6416411Linux-6.8.0-101-generic-x86_64-with-glibc2.39Median time is reported in milliseconds for ok records.
Inputs are prepared on the GPU before timed runs; initial host-to-device transfer is outside the timed region.
Timed runs include the host API call and backend-native device synchronization. tenferro-rs CUDA uses the explicit tenferro-rs synchronize API without downloading result tensors in the timed region.
Dense and einsum inputs use the same deterministic benchmark generator in the Rust and Python runners; host-to-device layout conversion remains outside the timed region.
tenferro-rs uses native column-major GPU tensors; PyTorch and vendor-wrapper columns use their native row-major framework tensors unless noted.
The cuSOLVER column is torch.linalg with preferred_linalg_library=cusolver; for SVD it pins driver=gesvd as a QR-based cuSOLVER comparison. tenferro-rs CUDA SVD uses its backend default driver policy, currently gesvdj for matrices with both dimensions at most 1024 and gesvd otherwise.
Non-ok cells show the structured backend status.
| Problem | tenferro-rs CUDA trace | tenferro-rs CUDA eager | PyTorch CUDA | cuBLASLt | CUTLASS | cuSOLVER | cuSPARSE | Ginkgo |
|---|---|---|---|---|---|---|---|---|
| elementwise_chain_f64_1k | 0.129 | 0.429 | 0.133 | unsupported | unsupported | unsupported | unsupported | unsupported |
| elementwise_chain_f64_1m | 0.219 | 0.396 | 0.319 | unsupported | unsupported | unsupported | unsupported | unsupported |