Run 20261008_host_views, AMD Ryzen 9 9955HX, Linux MKL devcontainer, CPU 1 thread, no external affinity pinning. Pure metadata needs no session or BLAS work. tenferro-rs is explicitly pinned to b3f47296244ff7b7c55ac0a75f782cb0835418c1, as requested in the maintainer reply; this is not a latest-main run.
MWE source, predeclared experiment, paired summary, environment/providers, lossless process records. The original TensorValue evidence remains valid and unchanged.
Every batch constructs separate, equivalent column-major f64 owners outside the timer. All typed as_view() inputs and output-retention capacity are prepared before timing. The timer includes only the declared metadata transformation and storing each returned descriptor. Input generation, wrapping, input-view construction, validation and destruction are untimed. All outputs are retained through clock stop. For typed arms, every owner is retained until after all its input/output views are destroyed; for TensorValue the consuming output retains the underlying owner.
The typed arms borrow TypedTensorView<f64, R, Host>; the original arm consumes owned, dtype-erased, dynamic-rank TensorValue. Their ownership/lifetime contracts differ and their costs must not be treated as interchangeable. Host DynRank keeps dynamic input/output rank. Static slice uses Rank<1>, and static transpose uses Rank<2>, preserving rank. Static-input reshape uses Rank<1> but returns DynRank, not Rank<2>. Metadata descriptors match Julia; every typed logical output element and its address within the original owner are checked untimed, confirming aliasing rather than a copy.
Each process uses 3 warmups and 15 batched samples. Four comparison rounds run Julia → owned → Host DynRank → Host static, then the reverse, reverse, forward. Each Rust arm has four independent A/A pairs in AB/BA/BA/AB order. Gates declared before collection: median paired ratio ≥1.2, every A/A pair within 10%, maximum per-process sample CoV ≤20%, correctness passed. Target interval is 50 ms; the 512 MiB retained-memory budget and 2,000,000-operation cap can make batches shorter. Values are batched throughput normalized per operation, not isolated call latency. Julia transpose specializes the fixed permutation and can reduce the loop to compact metadata stores; its sub-nanosecond normalized time is not a physical single-call latency.
Times are medians of four process medians, in ns/op. Ratios are medians of the four paired round ratios; they need not equal the quotient of the displayed aggregate times.
| operation / fixture | Rust representation | Rust ns/op | Julia ns/op | Rust / Julia | paired ratio range | max A/A deviation | max CoV | conclusion |
|---|---|---|---|---|---|---|---|---|
| reshape [1024] → [32,32] | TensorValue (consuming owned) | 336.726 | 7.766625 | 43.267× | 39.874–43.457× | 5.8% | 16.9% | confirmed-slower |
| reshape [1024] → [32,32] | Host borrowed / DynRank | 365.380 | 7.766625 | 47.085× | 42.362–48.227× | 4.5% | 16.9% | confirmed-slower |
| reshape [1024] → [32,32] | Host borrowed / static input rank | 316.467 | 7.766625 | 40.761× | 36.857–41.034× | 2.6% | 16.9% | confirmed-slower |
| slice [4096], 128:3968:2 → [1920] | TensorValue (consuming owned) | 278.965 | 2.335538 | 120.291× | 112.209–125.362× | 6.6% | 15.1% | confirmed-slower |
| slice [4096], 128:3968:2 → [1920] | Host borrowed / DynRank | 292.649 | 2.335538 | 124.890× | 123.265–132.593× | 1.8% | 15.1% | confirmed-slower |
| slice [4096], 128:3968:2 → [1920] | Host borrowed / static input rank | 226.906 | 2.335538 | 96.700× | 95.555–100.886× | 3.0% | 15.1% | confirmed-slower |
| transpose [2,2], axes [1,0] | TensorValue (consuming owned) | 214.643 | 0.735334 | 293.049× | 255.186–327.567× | 5.3% | 14.8% | confirmed-slower |
| transpose [2,2], axes [1,0] | Host borrowed / DynRank | 324.871 | 0.735334 | 444.300× | 393.051–492.424× | 3.2% | 14.8% | confirmed-slower |
| transpose [2,2], axes [1,0] | Host borrowed / static input rank | 242.615 | 0.735334 | 331.828× | 291.647–372.086× | 9.7% | 14.8% | confirmed-slower |
Owned / Host >1 means the borrowed Host arm took less time. These compare the explicitly different contracts above. The static slice improvement passes the same 1.2 threshold and both Rust arms’ noise gates; small differences below that threshold are not confirmed improvement claims.
| operation | owned / Host DynRank | owned / Host static input rank |
|---|---|---|
| reshape [1024] → [32,32] | 0.920× | 1.064× |
| slice [4096], 128:3968:2 → [1920] | 0.943× | 1.239× |
| transpose [2,2], axes [1,0] | 0.657× | 0.878× |
Host specialization does not close the Julia gap in these fixtures. Static-rank slicing improves relative to TensorValue, while reshape and transpose do not show a ≥1.2 improvement over TensorValue. These measurements narrow the discussion to the measured view paths; they do not identify shared internal fields as the cause or establish that a broader representation redesign is needed.
Ranges below cover the four comparison processes per arm; durations are their median measured batch durations.
| case | arm | operations per sample | batch duration (ms) |
|---|---|---|---|
metadata.reshape_view |
julia-base | 63,550–63,550 | 0.492–0.543 |
metadata.reshape_view |
metadata | 57,456–57,456 | 19.243–19.581 |
metadata.reshape_view |
metadata-host-dyn | 54,032–54,032 | 19.508–20.257 |
metadata.reshape_view |
metadata-host-static | 55,097–55,097 | 17.356–17.508 |
metadata.slice_view |
julia-base | 16,256–16,256 | 0.037–0.039 |
metadata.slice_view |
metadata | 15,827–15,827 | 4.225–4.554 |
metadata.slice_view |
metadata-host-dyn | 15,556–15,556 | 4.531–4.668 |
metadata.slice_view |
metadata-host-static | 15,701–15,701 | 3.486–3.636 |
metadata.transpose_view |
julia-base | 1,864,135–1,864,135 | 1.255–1.529 |
metadata.transpose_view |
metadata | 262,144–262,144 | 54.875–57.798 |
metadata.transpose_view |
metadata-host-dyn | 262,144–262,144 | 84.459–86.886 |
metadata.transpose_view |
metadata-host-static | 262,144–262,144 | 62.715–65.653 |
These are the recorded collection commands. To collect fresh results, check out the linked MWE source commit and the pinned library commit, rebuild, and use a new output directory. Use the default Linux MKL devcontainer. The MWE’s private system-mkl feature maps to the pinned upstream blas API and links the installed MKL; it is separate from the root harness’s migrated CPU feature names.
docker exec -u vscode -w /workspaces/tenferro-benchmark 14099c68ff8e bash -lc '
CARGO_TARGET_DIR="$PWD/target/followup-b3f47296" CARGO_BUILD_JOBS=8 \
cargo build --release --locked --manifest-path mwe/cpu_followup/Cargo.toml'
python3 mwe/cpu_followup/collect_metadata_host.py --output data/results/amd-cpu/cpu/metadata_host/20261008_host_views
python3 scripts/format_metadata_host_results.py --raw data/results/amd-cpu/cpu/metadata_host/20261008_host_views
The collector sets OMP/MKL/Rayon/OpenBLAS/Julia thread counts to 1 via configure_cpu_thread_env 1, sets CPU_FOLLOWUP_TARGET_NS=50000000, and exports the installed MKL/compiler library paths. Host and container idle guards remain enabled. Exact process commands, timings and signatures are in the compressed records; restore instructions apply unchanged. Both typed arms and the existing Julia reference are registered under #2040 in cpu/perf_issues (manifest v5).