CPU FFT Benchmark Results
- Suite:
cpu/fft
- Target profile:
amd-cpu
- Suite file:
benchmarks/cpu/fft.yaml
- Run metadata:
data/results/amd-cpu/cpu/fft/20261008_060740/run.yaml
- Timestamp:
20261008_060740
Latest run: ./scripts/run_cpu_fft.sh 1 4.
This file is generated from sequential CPU FFT runs under data/results/amd-cpu/cpu/fft/20261008_060740.
- tenferro-rs commit:
5cf78c7ec0ad9516dd78bab546d5e7bd42fa3102
CPU Information
- Model:
AMD Ryzen 9 9955HX 16-Core Processor
- Vendor:
AuthenticAMD
- Logical CPUs:
32
- Sockets:
1
- Cores per socket:
16
- Threads per core:
2
- NUMA nodes:
1
- Python platform:
Linux-7.0.0-38-generic-x86_64-with-glibc2.39
Thread Environments
Threads: 1
- Run metadata:
data/results/amd-cpu/cpu/fft/20261008_060740/run_t1.yaml
- OMP_NUM_THREADS:
1
- OMP_THREAD_LIMIT:
1
- OMP_DYNAMIC:
FALSE
- RAYON_NUM_THREADS:
1
- OPENBLAS_NUM_THREADS:
1
- GOTO_NUM_THREADS:
1
- MKL_NUM_THREADS:
1
- VECLIB_MAXIMUM_THREADS:
1
- VECLIB_NUM_THREADS:
1
- NUMEXPR_NUM_THREADS:
1
- BLIS_NUM_THREADS:
1
- XLA_FLAGS:
--xla_cpu_multi_thread_eigen=false --xla_cpu_experimental_ynn_fusion_type= intra_op_parallelism_threads=1
Threads: 4
- Run metadata:
data/results/amd-cpu/cpu/fft/20261008_060740/run_t4.yaml
- OMP_NUM_THREADS:
4
- OMP_THREAD_LIMIT:
4
- OMP_DYNAMIC:
FALSE
- RAYON_NUM_THREADS:
4
- OPENBLAS_NUM_THREADS:
4
- GOTO_NUM_THREADS:
4
- MKL_NUM_THREADS:
4
- VECLIB_MAXIMUM_THREADS:
4
- VECLIB_NUM_THREADS:
4
- NUMEXPR_NUM_THREADS:
4
- BLIS_NUM_THREADS:
4
- XLA_FLAGS:
--xla_cpu_multi_thread_eigen=true --xla_cpu_experimental_ynn_fusion_type= intra_op_parallelism_threads=4
Timing Discipline
- Input tensors are created outside the timed region.
- Each timed call creates the output tensor.
- tenferro-rs immediate rows require BENCH_INCLUDE_SETUP_DIAGNOSTICS=1 and use the separate cpu/fft_setup suite.
- tenferro-rs read rows require BENCH_INCLUDE_SETUP_DIAGNOSTICS=1 because the API performs one-shot planning.
- tenferro-rs cached rows reuse a caller-owned
FftExecutor across warmups and timed runs.
- tenferro-rs eager rows reuse one
EagerRuntime and input EagerTensor; setup is outside timing.
- tenferro-rs trace rows construct and compile
TracedTensorFftExt graphs outside timing and reuse the compiled program.
- PyTorch rows use
torch.fft after warmup, allowing PyTorch internal planning/cache behavior.
- The primary fair comparison is cached
FftExecutor versus warmed torch.fft; one-shot diagnostics are opt-in and never operation comparisons.
- This initial suite only measures 1D transforms to avoid row-major/column-major batched-axis layout artifacts.
Threads: 1 4
- CSV:
data/results/amd-cpu/cpu/fft/20261008_060740/cpu_fft_t1_20261008_060740.csv
- CSV:
data/results/amd-cpu/cpu/fft/20261008_060740/cpu_fft_t4_20261008_060740.csv
- Source table:
data/results/amd-cpu/cpu/fft/20261008_060740/cpu_fft_20261008_060740.md
CPU FFT Benchmark Items
Median ± IQR (ms). Missing backends are shown as -.
Timing audit: PyTorch 2.12 CPU torch.fft creates and commits a DFTI descriptor per call even after warmup (upstream source). Its rows include setup and are noncompliant as operation-only performance evidence; warmed execution does not remove that setup. Input tensors are untimed and allocation-returning calls retain output allocation. Traced calls also include internal session work. Use the cached native oneMKL follow-up for audited current-main operation evidence. No historical timing values have been edited.
Operation execution
| suite |
benchmark |
dtype |
threads |
shape |
tenferro-rs one-shot diagnostic (ms) |
tenferro-rs TensorRead API (ms) |
tenferro-rs cached primary (ms) |
tenferro-rs eager mode (ms) |
tenferro-rs trace mode (ms) |
PyTorch torch.fft (ms) |
| cpu/fft |
fft |
c32 |
1 |
1d_n1024 |
skipped |
skipped |
0.003 ± 0.000 |
skipped |
skipped |
0.007 ± 0.000 |
| cpu/fft |
fft |
c32 |
1 |
1d_n1048576 |
skipped |
skipped |
5.122 ± 0.065 |
4.909 ± 0.082 |
4.884 ± 0.244 |
9.227 ± 0.095 |
| cpu/fft |
fft |
c32 |
4 |
1d_n1024 |
skipped |
skipped |
0.003 ± 0.000 |
skipped |
skipped |
0.007 ± 0.000 |
| cpu/fft |
fft |
c32 |
4 |
1d_n1048576 |
skipped |
skipped |
5.008 ± 0.068 |
4.913 ± 0.107 |
5.798 ± 0.955 |
4.597 ± 2.181 |
| cpu/fft |
fft |
c64 |
1 |
1d_n1024 |
skipped |
skipped |
0.004 ± 0.000 |
skipped |
skipped |
0.010 ± 0.000 |
| cpu/fft |
fft |
c64 |
1 |
1d_n1048576 |
skipped |
skipped |
9.375 ± 0.112 |
15.364 ± 0.142 |
15.424 ± 0.182 |
15.284 ± 0.202 |
| cpu/fft |
fft |
c64 |
4 |
1d_n1024 |
skipped |
skipped |
0.004 ± 0.000 |
skipped |
skipped |
0.010 ± 0.000 |
| cpu/fft |
fft |
c64 |
4 |
1d_n1048576 |
skipped |
skipped |
9.280 ± 0.068 |
19.152 ± 3.643 |
11.171 ± 5.322 |
15.169 ± 0.171 |
| cpu/fft |
ifft |
c32 |
1 |
1d_n1024 |
skipped |
skipped |
0.003 ± 0.000 |
skipped |
skipped |
0.007 ± 0.000 |
| cpu/fft |
ifft |
c32 |
1 |
1d_n1048576 |
skipped |
skipped |
5.204 ± 0.051 |
5.114 ± 0.111 |
5.124 ± 0.233 |
9.479 ± 0.078 |
| cpu/fft |
ifft |
c32 |
4 |
1d_n1024 |
skipped |
skipped |
0.003 ± 0.000 |
skipped |
skipped |
0.007 ± 0.000 |
| cpu/fft |
ifft |
c32 |
4 |
1d_n1048576 |
skipped |
skipped |
5.176 ± 0.079 |
5.335 ± 0.063 |
5.650 ± 0.444 |
2.105 ± 0.034 |
| cpu/fft |
ifft |
c64 |
1 |
1d_n1024 |
skipped |
skipped |
0.004 ± 0.000 |
skipped |
skipped |
0.010 ± 0.000 |
| cpu/fft |
ifft |
c64 |
1 |
1d_n1048576 |
skipped |
skipped |
15.696 ± 0.204 |
15.755 ± 0.220 |
15.708 ± 0.154 |
15.519 ± 0.244 |
| cpu/fft |
ifft |
c64 |
4 |
1d_n1024 |
skipped |
skipped |
0.004 ± 0.000 |
skipped |
skipped |
0.010 ± 0.000 |
| cpu/fft |
ifft |
c64 |
4 |
1d_n1048576 |
skipped |
skipped |
15.657 ± 0.099 |
15.710 ± 0.300 |
10.432 ± 0.356 |
15.680 ± 0.217 |
| cpu/fft |
irfft |
c32 |
1 |
1d_n1024 |
skipped |
skipped |
0.003 ± 0.000 |
skipped |
skipped |
0.007 ± 0.000 |
| cpu/fft |
irfft |
c32 |
1 |
1d_n1048576 |
skipped |
skipped |
4.897 ± 0.155 |
5.063 ± 0.169 |
5.011 ± 0.044 |
17.899 ± 0.027 |
| cpu/fft |
irfft |
c32 |
4 |
1d_n1024 |
skipped |
skipped |
0.003 ± 0.000 |
skipped |
skipped |
0.007 ± 0.000 |
| cpu/fft |
irfft |
c32 |
4 |
1d_n1048576 |
skipped |
skipped |
4.850 ± 0.056 |
5.043 ± 0.085 |
6.123 ± 0.925 |
3.714 ± 0.018 |
| cpu/fft |
irfft |
c64 |
1 |
1d_n1024 |
skipped |
skipped |
0.005 ± 0.000 |
skipped |
skipped |
0.008 ± 0.000 |
| cpu/fft |
irfft |
c64 |
1 |
1d_n1048576 |
skipped |
skipped |
9.131 ± 0.082 |
9.802 ± 0.109 |
9.780 ± 0.084 |
18.218 ± 0.159 |
| cpu/fft |
irfft |
c64 |
4 |
1d_n1024 |
skipped |
skipped |
0.005 ± 0.000 |
skipped |
skipped |
0.008 ± 0.000 |
| cpu/fft |
irfft |
c64 |
4 |
1d_n1048576 |
skipped |
skipped |
9.821 ± 0.135 |
9.796 ± 0.108 |
10.208 ± 0.695 |
13.875 ± 0.366 |
| cpu/fft |
rfft |
f32 |
1 |
1d_n1024 |
skipped |
skipped |
0.003 ± 0.000 |
skipped |
skipped |
0.006 ± 0.000 |
| cpu/fft |
rfft |
f32 |
1 |
1d_n1048576 |
skipped |
skipped |
4.307 ± 0.028 |
4.463 ± 0.282 |
4.433 ± 0.093 |
17.272 ± 0.042 |
| cpu/fft |
rfft |
f32 |
4 |
1d_n1024 |
skipped |
skipped |
0.003 ± 0.000 |
skipped |
skipped |
0.006 ± 0.000 |
| cpu/fft |
rfft |
f32 |
4 |
1d_n1048576 |
skipped |
skipped |
4.340 ± 0.040 |
4.471 ± 0.014 |
5.079 ± 0.702 |
3.549 ± 0.010 |
| cpu/fft |
rfft |
f64 |
1 |
1d_n1024 |
skipped |
skipped |
0.004 ± 0.000 |
skipped |
skipped |
0.007 ± 0.000 |
| cpu/fft |
rfft |
f64 |
1 |
1d_n1048576 |
skipped |
skipped |
8.524 ± 0.074 |
9.461 ± 0.183 |
8.605 ± 0.080 |
18.660 ± 0.198 |
| cpu/fft |
rfft |
f64 |
4 |
1d_n1024 |
skipped |
skipped |
0.004 ± 0.000 |
skipped |
skipped |
0.007 ± 0.000 |
| cpu/fft |
rfft |
f64 |
4 |
1d_n1048576 |
skipped |
skipped |
8.442 ± 0.071 |
9.482 ± 0.230 |
9.366 ± 0.232 |
16.396 ± 0.109 |
Notes:
- pytorch-cpu: torch.fft warmup before measured runs; input allocation outside timed region; operations_per_sample=128; median_batch_ms=0.776872
- pytorch-cpu: torch.fft warmup before measured runs; input allocation outside timed region; operations_per_sample=128; median_batch_ms=0.780960
- pytorch-cpu: torch.fft warmup before measured runs; input allocation outside timed region; operations_per_sample=128; median_batch_ms=0.836555
- pytorch-cpu: torch.fft warmup before measured runs; input allocation outside timed region; operations_per_sample=128; median_batch_ms=0.837166
- pytorch-cpu: torch.fft warmup before measured runs; input allocation outside timed region; operations_per_sample=128; median_batch_ms=0.888724
- pytorch-cpu: torch.fft warmup before measured runs; input allocation outside timed region; operations_per_sample=128; median_batch_ms=0.893562
- pytorch-cpu: torch.fft warmup before measured runs; input allocation outside timed region; operations_per_sample=128; median_batch_ms=0.907148
- pytorch-cpu: torch.fft warmup before measured runs; input allocation outside timed region; operations_per_sample=128; median_batch_ms=0.914602
- pytorch-cpu: torch.fft warmup before measured runs; input allocation outside timed region; operations_per_sample=128; median_batch_ms=0.942234
- pytorch-cpu: torch.fft warmup before measured runs; input allocation outside timed region; operations_per_sample=128; median_batch_ms=0.956371
- pytorch-cpu: torch.fft warmup before measured runs; input allocation outside timed region; operations_per_sample=128; median_batch_ms=1.013709
- pytorch-cpu: torch.fft warmup before measured runs; input allocation outside timed region; operations_per_sample=128; median_batch_ms=1.037714
- pytorch-cpu: torch.fft warmup before measured runs; input allocation outside timed region; operations_per_sample=128; median_batch_ms=1.224987
- pytorch-cpu: torch.fft warmup before measured runs; input allocation outside timed region; operations_per_sample=128; median_batch_ms=1.250925
- pytorch-cpu: torch.fft warmup before measured runs; input allocation outside timed region; operations_per_sample=128; median_batch_ms=1.266054
- pytorch-cpu: torch.fft warmup before measured runs; input allocation outside timed region; operations_per_sample=128; median_batch_ms=1.285400
- pytorch-cpu: torch.fft warmup before measured runs; input allocation outside timed region; operations_per_sample=1; median_batch_ms=13.874824
- pytorch-cpu: torch.fft warmup before measured runs; input allocation outside timed region; operations_per_sample=1; median_batch_ms=15.169161
- pytorch-cpu: torch.fft warmup before measured runs; input allocation outside timed region; operations_per_sample=1; median_batch_ms=15.284128
- pytorch-cpu: torch.fft warmup before measured runs; input allocation outside timed region; operations_per_sample=1; median_batch_ms=15.518599
- pytorch-cpu: torch.fft warmup before measured runs; input allocation outside timed region; operations_per_sample=1; median_batch_ms=15.680474
- pytorch-cpu: torch.fft warmup before measured runs; input allocation outside timed region; operations_per_sample=1; median_batch_ms=16.395870
- pytorch-cpu: torch.fft warmup before measured runs; input allocation outside timed region; operations_per_sample=1; median_batch_ms=17.272000
- pytorch-cpu: torch.fft warmup before measured runs; input allocation outside timed region; operations_per_sample=1; median_batch_ms=17.899392
- pytorch-cpu: torch.fft warmup before measured runs; input allocation outside timed region; operations_per_sample=1; median_batch_ms=18.218141
- pytorch-cpu: torch.fft warmup before measured runs; input allocation outside timed region; operations_per_sample=1; median_batch_ms=18.659634
- pytorch-cpu: torch.fft warmup before measured runs; input allocation outside timed region; operations_per_sample=1; median_batch_ms=2.105434
- pytorch-cpu: torch.fft warmup before measured runs; input allocation outside timed region; operations_per_sample=1; median_batch_ms=3.548942
- pytorch-cpu: torch.fft warmup before measured runs; input allocation outside timed region; operations_per_sample=1; median_batch_ms=3.713562
- pytorch-cpu: torch.fft warmup before measured runs; input allocation outside timed region; operations_per_sample=1; median_batch_ms=4.596916
- pytorch-cpu: torch.fft warmup before measured runs; input allocation outside timed region; operations_per_sample=1; median_batch_ms=9.226883
- pytorch-cpu: torch.fft warmup before measured runs; input allocation outside timed region; operations_per_sample=1; median_batch_ms=9.478798
- tenferro-fft-eager: API includes per-call planning or session entry; use shared cached executor for short operation timing
- tenferro-fft-eager: EagerTensorFftExt with one reused eager runtime and input; operations_per_sample=1; median_batch_ms=15.363547
- tenferro-fft-eager: EagerTensorFftExt with one reused eager runtime and input; operations_per_sample=1; median_batch_ms=15.709549
- tenferro-fft-eager: EagerTensorFftExt with one reused eager runtime and input; operations_per_sample=1; median_batch_ms=15.755144
- tenferro-fft-eager: EagerTensorFftExt with one reused eager runtime and input; operations_per_sample=1; median_batch_ms=19.152100
- tenferro-fft-eager: EagerTensorFftExt with one reused eager runtime and input; operations_per_sample=1; median_batch_ms=4.463344
- tenferro-fft-eager: EagerTensorFftExt with one reused eager runtime and input; operations_per_sample=1; median_batch_ms=4.470728
- tenferro-fft-eager: EagerTensorFftExt with one reused eager runtime and input; operations_per_sample=1; median_batch_ms=4.908642
- tenferro-fft-eager: EagerTensorFftExt with one reused eager runtime and input; operations_per_sample=1; median_batch_ms=4.912680
- tenferro-fft-eager: EagerTensorFftExt with one reused eager runtime and input; operations_per_sample=1; median_batch_ms=5.042805
- tenferro-fft-eager: EagerTensorFftExt with one reused eager runtime and input; operations_per_sample=1; median_batch_ms=5.063033
- tenferro-fft-eager: EagerTensorFftExt with one reused eager runtime and input; operations_per_sample=1; median_batch_ms=5.113618
- tenferro-fft-eager: EagerTensorFftExt with one reused eager runtime and input; operations_per_sample=1; median_batch_ms=5.334985
- tenferro-fft-eager: EagerTensorFftExt with one reused eager runtime and input; operations_per_sample=1; median_batch_ms=9.461114
- tenferro-fft-eager: EagerTensorFftExt with one reused eager runtime and input; operations_per_sample=1; median_batch_ms=9.482013
- tenferro-fft-eager: EagerTensorFftExt with one reused eager runtime and input; operations_per_sample=1; median_batch_ms=9.796315
- tenferro-fft-eager: EagerTensorFftExt with one reused eager runtime and input; operations_per_sample=1; median_batch_ms=9.802036
- tenferro-fft-executor-cached: FftExecutor reused across warmups and measured runs; operations_per_sample=128; median_batch_ms=0.340030
- tenferro-fft-executor-cached: FftExecutor reused across warmups and measured runs; operations_per_sample=128; median_batch_ms=0.343857
- tenferro-fft-executor-cached: FftExecutor reused across warmups and measured runs; operations_per_sample=128; median_batch_ms=0.400975
- tenferro-fft-executor-cached: FftExecutor reused across warmups and measured runs; operations_per_sample=128; median_batch_ms=0.401717
- tenferro-fft-executor-cached: FftExecutor reused across warmups and measured runs; operations_per_sample=128; median_batch_ms=0.419721
- tenferro-fft-executor-cached: FftExecutor reused across warmups and measured runs; operations_per_sample=128; median_batch_ms=0.422917
- tenferro-fft-executor-cached: FftExecutor reused across warmups and measured runs; operations_per_sample=128; median_batch_ms=0.436071
- tenferro-fft-executor-cached: FftExecutor reused across warmups and measured runs; operations_per_sample=128; median_batch_ms=0.436792
- tenferro-fft-executor-cached: FftExecutor reused across warmups and measured runs; operations_per_sample=128; median_batch_ms=0.478080
- tenferro-fft-executor-cached: FftExecutor reused across warmups and measured runs; operations_per_sample=128; median_batch_ms=0.483681
- tenferro-fft-executor-cached: FftExecutor reused across warmups and measured runs; operations_per_sample=128; median_batch_ms=0.516432
- tenferro-fft-executor-cached: FftExecutor reused across warmups and measured runs; operations_per_sample=128; median_batch_ms=0.517865
- tenferro-fft-executor-cached: FftExecutor reused across warmups and measured runs; operations_per_sample=128; median_batch_ms=0.564032
- tenferro-fft-executor-cached: FftExecutor reused across warmups and measured runs; operations_per_sample=128; median_batch_ms=0.568781
- tenferro-fft-executor-cached: FftExecutor reused across warmups and measured runs; operations_per_sample=128; median_batch_ms=0.577677
- tenferro-fft-executor-cached: FftExecutor reused across warmups and measured runs; operations_per_sample=128; median_batch_ms=0.588138
- tenferro-fft-executor-cached: FftExecutor reused across warmups and measured runs; operations_per_sample=1; median_batch_ms=15.656869
- tenferro-fft-executor-cached: FftExecutor reused across warmups and measured runs; operations_per_sample=1; median_batch_ms=15.695903
- tenferro-fft-executor-cached: FftExecutor reused across warmups and measured runs; operations_per_sample=1; median_batch_ms=4.307450
- tenferro-fft-executor-cached: FftExecutor reused across warmups and measured runs; operations_per_sample=1; median_batch_ms=4.340383
- tenferro-fft-executor-cached: FftExecutor reused across warmups and measured runs; operations_per_sample=1; median_batch_ms=4.849732
- tenferro-fft-executor-cached: FftExecutor reused across warmups and measured runs; operations_per_sample=1; median_batch_ms=4.897130
- tenferro-fft-executor-cached: FftExecutor reused across warmups and measured runs; operations_per_sample=1; median_batch_ms=5.007849
- tenferro-fft-executor-cached: FftExecutor reused across warmups and measured runs; operations_per_sample=1; median_batch_ms=5.121753
- tenferro-fft-executor-cached: FftExecutor reused across warmups and measured runs; operations_per_sample=1; median_batch_ms=5.176226
- tenferro-fft-executor-cached: FftExecutor reused across warmups and measured runs; operations_per_sample=1; median_batch_ms=5.204479
- tenferro-fft-executor-cached: FftExecutor reused across warmups and measured runs; operations_per_sample=1; median_batch_ms=8.442085
- tenferro-fft-executor-cached: FftExecutor reused across warmups and measured runs; operations_per_sample=1; median_batch_ms=8.523559
- tenferro-fft-executor-cached: FftExecutor reused across warmups and measured runs; operations_per_sample=1; median_batch_ms=9.131433
- tenferro-fft-executor-cached: FftExecutor reused across warmups and measured runs; operations_per_sample=1; median_batch_ms=9.280454
- tenferro-fft-executor-cached: FftExecutor reused across warmups and measured runs; operations_per_sample=1; median_batch_ms=9.374811
- tenferro-fft-executor-cached: FftExecutor reused across warmups and measured runs; operations_per_sample=1; median_batch_ms=9.821412
- tenferro-fft-immediate: API includes per-call planning or session entry; use shared cached executor for short operation timing
- tenferro-fft-read: API includes per-call planning or session entry; use shared cached executor for short operation timing
- tenferro-fft-trace: API includes per-call planning or session entry; use shared cached executor for short operation timing
- tenferro-fft-trace: TracedTensorFftExt graph compiled once; compiled program reused across runs; operations_per_sample=1; median_batch_ms=10.208190
- tenferro-fft-trace: TracedTensorFftExt graph compiled once; compiled program reused across runs; operations_per_sample=1; median_batch_ms=10.432463
- tenferro-fft-trace: TracedTensorFftExt graph compiled once; compiled program reused across runs; operations_per_sample=1; median_batch_ms=11.171204
- tenferro-fft-trace: TracedTensorFftExt graph compiled once; compiled program reused across runs; operations_per_sample=1; median_batch_ms=15.424001
- tenferro-fft-trace: TracedTensorFftExt graph compiled once; compiled program reused across runs; operations_per_sample=1; median_batch_ms=15.707916
- tenferro-fft-trace: TracedTensorFftExt graph compiled once; compiled program reused across runs; operations_per_sample=1; median_batch_ms=4.432926
- tenferro-fft-trace: TracedTensorFftExt graph compiled once; compiled program reused across runs; operations_per_sample=1; median_batch_ms=4.884086
- tenferro-fft-trace: TracedTensorFftExt graph compiled once; compiled program reused across runs; operations_per_sample=1; median_batch_ms=5.011195
- tenferro-fft-trace: TracedTensorFftExt graph compiled once; compiled program reused across runs; operations_per_sample=1; median_batch_ms=5.078632
- tenferro-fft-trace: TracedTensorFftExt graph compiled once; compiled program reused across runs; operations_per_sample=1; median_batch_ms=5.123587
- tenferro-fft-trace: TracedTensorFftExt graph compiled once; compiled program reused across runs; operations_per_sample=1; median_batch_ms=5.649728
- tenferro-fft-trace: TracedTensorFftExt graph compiled once; compiled program reused across runs; operations_per_sample=1; median_batch_ms=5.798117
- tenferro-fft-trace: TracedTensorFftExt graph compiled once; compiled program reused across runs; operations_per_sample=1; median_batch_ms=6.123380
- tenferro-fft-trace: TracedTensorFftExt graph compiled once; compiled program reused across runs; operations_per_sample=1; median_batch_ms=8.604711
- tenferro-fft-trace: TracedTensorFftExt graph compiled once; compiled program reused across runs; operations_per_sample=1; median_batch_ms=9.365644
- tenferro-fft-trace: TracedTensorFftExt graph compiled once; compiled program reused across runs; operations_per_sample=1; median_batch_ms=9.780284