CPU ≥1.2× follow-up: current public operation MWEs

2026-10-08, AMD Ryzen 9 9955HX, Linux CPU MKL devcontainer. 43 operation workloads were confirmed and reported in 9 new tenferro-rs Issues. One workload is already covered by open #2021. Two workloads remain inconclusive; seven did not remain ≥1.2× suspects in the operation-only 1T/4T scan. A scan below threshold is not a confirmed parity claim.

Selection: 53 operation groups / 139 original condition rows from the earlier inventory without open owners, selected because at least one original ratio was ≥1.2. All 53 have runnable public-API reproducers. The earlier table rows are screening evidence only. Metadata fixtures are deliberately smaller; their original large-fixture timings are not confirmed. Broadcast #2021 was identified by a renewed full-body open-Issue search before filing and excluded from new reports.

Library: b3f47296244ff7b7c55ac0a75f782cb0835418c1, pulled main then deliberately pinned for the campaign. tenferro BLAS linkage and native FFT reference: oneMKL 2026.0.1; tenferro FFT itself uses cached RustFFT 6.4.1 plans. PyTorch 2.12.0+cpu, wheel MKL 2024.2. JAX/jaxlib 0.10.1, CPU. Julia 1.13.1. strided-rs 12ff2de906f4d894a4329f7dbf195b0f874f846e. No CUDA, no runner taskset/numactl pinning; tenferro public backend-managed placement is retained. Provider/config metadata and hardware.

Setup is untimed on every arm, including calibration: fixture creation/conversion, wrapping, contexts/sessions, input views/descriptors, FFT planning, JIT and first completed initialization. Timers contain only the declared operation, intrinsic allocation, output retention assignment and native completion. Results stay alive past the clock; downloads/signatures/validation/destruction are untimed. Direct/eager paths hold one entered session. Metadata is session-free. MWE setup/commands, timing audit.

Confirmation: 3 warmups and 15 batched samples per process; four A/A pairs and four independent balanced comparison pairs AB/BA/BA/AB. Statistic: median of four round-median ratios. Declared gate: ratio ≥1.2, A/A median within 10%, every sample-set CoV ≤20%, numerical validation passed. The corrected/longer-batch phases also balance A/A and require each A/A round within 10%. No threshold was estimated from candidate results or loosened. Scan selects the strongest eligible 1T/4T condition; the table does not claim both thread counts are slow.

Batch targets are 10 ms or 50 ms for longer-batch retries, bounded by 512 MiB retained tensor payloads/owned metadata inputs. One output is still retained for the 2 GiB rotation. The final 2×2 metadata-transpose reference uses ~1.86 million operations in ~1.5 ms at its memory cap; normalized sub-ns values are specialized bulk metadata throughput, not isolated call latency. Absolute tenferro metadata costs are hundreds of ns. Output shape/count, aggregates and ~130 deterministic probes agree; activation/permutation outputs additionally pass full scalar/odometer oracles. FFT checks input preservation as well.

Compiled JAX uses dynamic arguments and completed execution. Its 1/4 active Eigen workers were verified with untimed probes. Native layouts are preserved with equivalent logical values; no one-sided layout conversion is timed. FFT compares cached native public oneMKL DFTI, not PyTorch FFT: PyTorch CPU FFT replans per call, so its historical rows are not operation-only evidence.

New reports

Issue family confirmed workloads ratio range
#2032 activations 5 5.359–12.479×
#2033 analytic 9 1.223–5.394×
#2034 reductions 8 1.332–2.974×
#2035 norm 1 11.330–11.330×
#2036 indexing 3 1.369–1.758×
#2037 permutation 9 1.233–2.573×
#2038 diagonal 1 2.295–2.295×
#2039 fft 4 1.695–3.469×
#2040 metadata 3 38.290–305.287×

All selected workloads

Times are per operation in µs. Confirmed rows show independent paired results; inconclusive ratios are diagnostic and were not filed. Screen-only rows show the strongest eligible scan condition. The exact physical permutation and source strides are in the fixture definitions. The default MWE fixture catalog preserves original observation identities.

case ID input shape dtype T path reference tenferro µs reference µs ratio disposition / owner campaign
activation.erf 1024×64 f32 4 eager-shared pytorch 581.891 51.6569 11.415× confirmed / #2032 20261008_confirmed_cpu
activation.gelu_tanh 1024×64 f32 4 eager-shared pytorch 688.397 55.1557 12.479× confirmed / #2032 20261008_confirmed_cpu
activation.sigmoid 1024×64 f32 4 eager-shared pytorch 266.034 45.4312 5.850× confirmed / #2032 20261008_confirmed_cpu
activation.silu 1024×64 f32 4 eager-shared pytorch 328.55 44.8623 7.262× confirmed / #2032 20261008_confirmed_cpu
activation.softplus 1024×64 f32 4 eager-shared pytorch 471.544 87.8062 5.359× confirmed / #2032 20261008_confirmed_cpu
complex.exp 4194304 c64 1 direct-shared julia-base 97699.6 60599.3 1.614× confirmed / #2033 20261008_corrected_confirm
complex.norm_fro 2048×1536 c64 4 direct-shared jax 21181 1867.51 11.330× confirmed / #2035 20261008_confirmed_cpu
fft.fft 1048576 c64 1 direct-shared mkl-dfti 18873.5 10034.6 1.876× confirmed / #2039 20261008_corrected_confirm
fft.ifft 1048576 c32 1 direct-shared mkl-dfti 8659.5 5108.87 1.695× confirmed / #2039 20261008_corrected_confirm
fft.irfft 524289 c32 1 direct-shared mkl-dfti 8431.51 2428.67 3.469× confirmed / #2039 20261008_corrected_confirm
fft.rfft 1048576 f32 1 direct-shared mkl-dfti 7913.9 3052.44 2.596× confirmed / #2039 20261008_corrected_confirm
index.concatenate 1048576 f64 1 direct-shared pytorch 4035.38 3636.29 1.110× screened-below-threshold / — 20261008_cpu_wheel
index.dynamic_slice 4194304 f64 1 direct-shared pytorch 3742.24 3772.91 0.992× screened-below-threshold / — 20261008_materialize_views
index.dynamic_update_slice 2097152 f64 1 direct-shared jax 4873.18 3575.65 1.369× confirmed / #2036 20261008_long_batches
index.gather 262144 f64 4 direct-shared jax 411.038 238.391 1.758× confirmed / #2036 20261008_confirmed_cpu
index.pad 2097152 f64 1 direct-shared pytorch 4794.64 4402.36 1.089× screened-below-threshold / — 20261008_cpu_wheel
index.reverse 2097152 f64 1 direct-shared pytorch 3997.48 3685.51 1.085× screened-below-threshold / — 20261008_cpu_wheel
index.scatter 262144 f64 4 direct-shared pytorch 1035.85 609.441 1.689× confirmed / #2036 20261008_long_batches
index.slice 4194304 f64 1 direct-shared pytorch 4891.07 4727.51 1.037× inconclusive / — 20261008_materialize_views
linalg.slogdet 1024×1024 f64 4 direct-shared pytorch 9255.66 6908.24 1.327× inconclusive / — 20261008_long_batches
metadata.reshape_view 1024 f64 1 metadata julia-base 0.306697 0.00803898 38.290× confirmed / #2040 20261008_corrected_confirm
metadata.slice_view 4096 f64 1 metadata julia-base 0.281946 0.00237346 118.995× confirmed / #2040 20261008_corrected_confirm
metadata.transpose_view 2×2 f64 1 metadata julia-base 0.21361 0.000686809 305.287× confirmed / #2040 20261008_long_batches
perm.cyclic_15d_3 3×3×3×3×3×3×3×3×3×3×3×3×3×3×3 f64 1 materialize strided-rs 42416.5 31016.6 1.372× confirmed / #2037 20261008_confirmed_cpu
perm.reverse_15d_3 3×3×3×3×3×3×3×3×3×3×3×3×3×3×3 f64 1 materialize strided-rs 172714 67133.4 2.573× confirmed / #2037 20261008_confirmed_cpu
perm.reverse_23d_2 2×2×2×2×2×2×2×2×2×2×2×2×2×2×2×2×2×2×2×2×2×2×2 f64 1 materialize strided-rs 66670.6 38759 1.722× confirmed / #2037 20261008_confirmed_cpu
perm.rotation_6d_32_32_32_32_16_16 32×32×32×32×16×16 f64 1 materialize strided-rs 787149 638174 1.233× confirmed / #2037 20261008_confirmed_cpu
perm.tn_light_415_24d_contiguous_same_perm 2×2×2×2×2×2×2×2×2×2×2×2×2×2×2×2×2×2×2×2×2×2×2×2 f64 1 materialize strided-rs 44882.8 36671.4 1.233× confirmed / #2037 20261008_confirmed_cpu
perm.tn_light_415_24d_scattered_to_colmajor 2×2×2×2×2×2×2×2×2×2×2×2×2×2×2×2×2×2×2×2×2×2×2×2 f64 1 materialize strided-rs 48412.6 37944.2 1.285× confirmed / #2037 20261008_confirmed_cpu
perm.transpose_3d_256_102 256×256×256 f64 1 materialize strided-rs 57109.5 42287.9 1.342× confirmed / #2037 20261008_confirmed_cpu
perm.transpose_3d_256_201 256×256×256 f64 1 materialize strided-rs 80989.8 53208.6 1.530× confirmed / #2037 20261008_confirmed_cpu
real.cos 8388608 f64 4 direct-shared pytorch 30132.1 12291.1 2.452× confirmed / #2033 20261008_confirmed_cpu
real.exp 8388608 f64 1 direct-shared jax 49809.3 21654 2.306× confirmed / #2033 20261008_confirmed_cpu
real.expm1 4194304 f64 1 direct-shared jax 29716.2 13863.7 2.141× confirmed / #2033 20261008_confirmed_cpu
real.log 8388608 f64 4 direct-shared pytorch 12779.2 10362.8 1.223× confirmed / #2033 20261008_confirmed_cpu
real.log1p 4194304 f64 1 direct-shared pytorch 40440.3 16051.2 2.526× confirmed / #2033 20261008_confirmed_cpu
real.pow 4194304 f64 1 direct-shared pytorch 64619.4 31880.7 2.029× confirmed / #2033 20261008_confirmed_cpu
real.reduce_max_all 8192×4096 f64 1 direct-shared julia-base 18311.5 11836.1 1.547× confirmed / #2034 20261008_corrected_confirm
real.reduce_max_axis0 2048×2048 f64 1 direct-shared julia-base 2308.95 1253.99 1.846× confirmed / #2034 20261008_corrected_confirm
real.reduce_max_axis1 2048×2048 f64 1 direct-shared julia-base 2059.21 692.641 2.974× confirmed / #2034 20261008_corrected_confirm
real.reduce_min_all 8192×4096 f64 1 direct-shared julia-base 16380 12294 1.332× confirmed / #2034 20261008_corrected_confirm
real.reduce_min_axis0 2048×2048 f64 1 direct-shared julia-base 2109.01 1314.79 1.604× confirmed / #2034 20261008_corrected_confirm
real.reduce_min_axis1 4096×4096 f64 1 direct-shared julia-base 8259.17 4690.38 1.765× confirmed / #2034 20261008_corrected_confirm
real.reduce_prod_axis1 2048×2048 f64 4 direct-shared pytorch 281.396 129.751 2.159× confirmed / #2034 20261008_confirmed_cpu
real.reduce_sum_axis1 2048×2048 f64 4 direct-shared pytorch 269.979 126.847 2.155× confirmed / #2034 20261008_confirmed_cpu
real.sin 8388608 f64 4 direct-shared pytorch 30035 12600.7 2.382× confirmed / #2033 20261008_confirmed_cpu
real.tanh 8388608 f64 1 direct-shared jax 110805 20462.6 5.394× confirmed / #2033 20261008_confirmed_cpu
structural.broadcast_in_dim 8192×1 f64 4 direct-shared pytorch 23163.9 18196.8 1.273× covered-by-open-issue / #2021 20261008_materialize_views
structural.cast_f64_f32 33554432 f64->f32 4 direct-shared pytorch 13453.2 12643.5 1.064× screened-below-threshold / — 20261008_cpu_wheel
structural.extract_diagonal 8388608×2×2 f64 4 direct-shared jax 31310.3 13586.6 2.295× confirmed / #2038 20261008_materialize_views
structural.transpose 4096×4096 f64 1 direct-shared pytorch 85731.4 45676.2 1.875× confirmed / #2037 20261008_materialize_views
structural.tril 4096×4096 f64 4 direct-shared pytorch 12601.9 12942 0.974× screened-below-threshold / — 20261008_cpu_wheel
structural.triu 4096×4096 f64 4 direct-shared pytorch 10888.5 12221.5 0.891× screened-below-threshold / — 20261008_cpu_wheel

The maintainer-requested pinned Host DynRank/static-rank extension for #2040 is reported separately in CPU metadata Host comparisons; it remeasures all representations in balanced rounds and preserves the original TensorValue observations above.

Evidence and supersession

The selected prepared-trace paths are explicitly scope-ineligible, since their public API cannot borrow an execution session and includes internal admission/session work. They are not counted as confirmed, parity or covered. This confirmation campaign uses the independent crate and dedicated registered-case runner. The root CPU-provider migration landed separately on main during the campaign; the original measurement revisions are preserved. The follow-up does not refresh the complete run_all suite.

Canonical precedence is encoded in the report formatter. Original confirmations involving Julia first-calibration JIT or PyTorch FFT planning are superseded, never used as operation evidence. The old transpose-metadata 32×32 case was inconclusive and is replaced by the 2×2 long batch. Earlier PyTorch materialization closures with timed input-view construction are superseded. Earlier scatter replacement semantics were corrected to additive scatter in both arms before eligible confirmation. Preliminary CUDA-wheel, constant-captured JAX, unspecialized Julia and incorrect Fortran-symbol DFTI probes are not publication evidence. No measured timing/status values were edited.

raw campaign purpose and eligibility
20261008_cpu_wheel 53-operation 1T/4T screen with CPU wheel; Julia/FFT/materialization subsets superseded below
20261008_scatter_add fair additive-scatter screen
20261008_confirmed_cpu first confirmation: use non-Julia/non-FFT cases, excluding materialization rows superseded below
20261008_corrected_refs untimed exact Julia batch JIT; metadata/reduction/complex-exp screen only, FFT probe failures excluded
20261008_cached_mkl_fft cached public native oneMKL FFT screen, input-preservation validation
20261008_corrected_confirm balanced A/A corrected Julia and FFT confirmation; transpose metadata superseded below
20261008_long_batches 50 ms target and balanced A/A retry: metadata transpose/scatter/dynamic update confirmed; slogdet noisy
20261008_materialize_views untimed input-view preparation; diagonal/transpose confirmed; slice remains inconclusive; broadcast owned by #2021

Per-process JSON is preserved losslessly in each campaign’s processes.jsonl.gz (restore instructions); it includes exact per-process commands, source revisions, counts, durations and validation signatures. Submission identities retain MWE source commits and case membership. Workloads and reference arms are added to cpu/perf_issues at opening (manifest v4; Host extensions in v5); the exact existing activation cases are annotated instead of duplicated. Use scripts/run_cpu_followup.sh inside the Linux MKL devcontainer to run registered latest-main cases. No automated audit, registry, scheduled run or mandatory baseline campaign is introduced.