GPU Benchmark Results
- Target profile:
nvidia-gpu
- Suite:
gpu/dense
- Suite file:
benchmarks/gpu/dense.yaml
- Timestamp:
2026-10-08T12:59:42.398069+00:00
- tenferro-rs commit:
b3f47296244ff7b7c55ac0a75f782cb0835418c1
GPU Information
- Device:
cuda:0
- Name:
NVIDIA L4
- UUID:
GPU-a2742c63-0edb-1b0a-9cc7-5b8184ed07e0
- Memory:
23034 MiB
- Driver version:
580.126.20
- CUDA version:
13.0
- CUDA runtime:
12.8
- cuDNN version:
92000
CPU Information
- Model:
AMD EPYC 9254 24-Core Processor
- Vendor:
AuthenticAMD
- Logical CPUs:
48
- Sockets:
1
- Cores per socket:
24
- Threads per core:
2
- NUMA nodes:
1
- Python platform:
Linux-6.8.0-106-generic-x86_64-with-glibc2.39
Median time is reported in milliseconds for ok records.
Inputs are prepared on the GPU before timed runs; initial host-to-device transfer is outside the timed region.
Timed runs include the host API call and backend-native device synchronization. tenferro-rs CUDA uses the explicit tenferro-rs synchronize API without downloading result tensors in the timed region.
Dense and einsum inputs use the same deterministic benchmark generator in the Rust and Python runners; host-to-device layout conversion remains outside the timed region.
tenferro-rs uses native column-major GPU tensors; PyTorch and vendor-wrapper columns use their native row-major framework tensors unless noted.
The cuSOLVER column is torch.linalg with preferred_linalg_library=cusolver; for SVD it pins driver=gesvd as a QR-based cuSOLVER comparison. tenferro-rs CUDA SVD uses its backend default driver policy, currently gesvdj for matrices with both dimensions at most 1024 and gesvd otherwise.
Non-ok cells show the structured backend status.
gpu/dense / batched_matmul / allocating output
| Problem |
tenferro-rs CUDA trace |
tenferro-rs CUDA eager |
PyTorch CUDA |
| dense_batched_matmul_f64_b1024_256 |
83.469 |
83.386 |
83.352 |
gpu/dense / eigh / allocating output
| Problem |
tenferro-rs CUDA trace |
tenferro-rs CUDA eager |
PyTorch CUDA |
| dense_eigh_f64_1024 |
44.518 |
44.504 |
44.335 |
gpu/dense / matmul / allocating output
| Problem |
tenferro-rs CUDA trace |
tenferro-rs CUDA eager |
PyTorch CUDA |
| dense_matmul_f64_3072 |
141.040 |
140.960 |
140.872 |
gpu/dense / qr / allocating output
| Problem |
tenferro-rs CUDA trace |
tenferro-rs CUDA eager |
PyTorch CUDA |
| dense_qr_f64_1536 |
55.834 |
55.802 |
55.065 |
gpu/dense / solve / allocating output
| Problem |
tenferro-rs CUDA trace |
tenferro-rs CUDA eager |
PyTorch CUDA |
| dense_solve_f64_1024_rhs16 |
8.552 |
8.386 |
7.877 |
| dense_solve_f64_2048_rhs128 |
32.789 |
32.754 |
31.822 |
| dense_solve_f64_2048_rhs16 |
28.300 |
28.088 |
27.568 |
| dense_solve_f64_4096_rhs16 |
148.177 |
148.064 |
146.961 |
| dense_solve_f64_512_rhs16 |
3.009 |
3.265 |
2.757 |
gpu/dense / svd / allocating output
SVD note: SVD rows use synchronized timed regions and matched Rust/Python input generators. tenferro-rs CUDA uses its backend default driver policy, currently gesvdj for matrices with both dimensions at most 1024 and gesvd otherwise. The cuSOLVER column pins torch.linalg.svd driver=gesvd as a QR-based cuSOLVER comparison; PyTorch's default row may use a different SVD driver and row-major framework layout.
| Problem |
tenferro-rs CUDA trace |
tenferro-rs CUDA eager |
PyTorch CUDA |
| dense_svd_f64_256 |
77.527 |
77.490 |
79.282 |