Explicit Backend Session Boundary
Status: implemented for #1929’s reduced 2026-09-27 delivery contract. This record includes a historical source inventory and step-by-step design at the stated baseline; its line numbers, provisional statuses and measurement runbook do not describe the current source. See the current-state summary below and the handoff. The work belongs to #1926 under #1929.
Inventory baseline: a45833d4f (origin/main at the time of writing, the revision named by #1926).
Related authorities and records:
exec-session.md— what a session is and how each backend maps it.session-oriented-concrete-apis.md— the session-explicit concrete surface (#1673). Its “the one-shot API stays available” clause is narrowed by this document; see Supersession.cpu-session-open-cost.md— measured per-call cost of opening a CPU session.cpu-backend-execution.md—CpuOperationEntry, execution-owner/reentrancy contract.api-and-convention-freeze.md— why removing implicit entry points is in scope before the release freeze closes.
Current delivery boundary
Phase B removed the 31 paired owner one-shots and moved CPU, CUDA and WebGPU operation bodies to session types. Runtime and extension helpers use borrowed sessions, and compiled instructions share a session across their terminal-value probe, execution and last-use reclaim. EagerRuntime::with_eager_session now lends a runtime-bound EagerSession for forward operations. The former EagerTensor operation methods and arithmetic overloads were removed, not replaced with implicit-entry shims; value/trace handles remain, as do named backward and runtime boundaries. Gradient-slot accumulation uses one borrowed session at the backward boundary. Ordinary eager linalg and FFT operations use borrowed extension traits. Tensor-owned linalg solve is intentionally kept for its calling-thread no_grad behavior, and consuming in-place FFT preserves its exclusive-ownership contract. Native-context or owned-only extension fallback runs in a separate top-level region, never inside a borrowed session.
#1946 closed the remaining self-entering surfaces: eager einsum/tensordot run only on a borrowed EagerSession (F5); backend owner types implement no execution, canonicalization, fusion or buffer trait, so CpuBackend, CudaBackend and WebGpuBackend are session hosts plus the explicit TensorDeviceTransfer boundary (F6); and the owner-context fallback accepts extension ops only (F9). Nested entry from inside any session is rejected with a typed SessionEntryError instead of blocking (F1). The audit allowlist carries no pending entries and --check rejects one. The owner-only BLAS linalg mode helper was deleted with the owner entry points. Production linalg already ran through the session’s engine mode; the two select different ParallelMode values only on an executor without Rayon inner parallelism, where every tenferro consumer of the mode (faer parallelism, native thread count, strided context, lane fan-out) already requires Rayon and the LAPACK path reads no mode, so no built-in behavior differs.
A2 (with_evaluation_scope plus AD wiring) is deferred, not blocked: the 2026-09-27 contract leaves its backend ownership decision and performance assessment to #1938 / #1927 / #1904. The session-entry audit keeps its A2 allowlist at zero. No new benchmark campaign or published improvement is a #1929 closure requirement. See the handoff for exact artifact revisions, focused validation and remaining CPU/GPU capabilities.
Problem
BackendSession is documented as the execution-time primitive surface and the _in/_read session methods exist, but the source still contains several distinct ways for ordinary operations to reach backend execution without the caller ever naming an execution boundary:
- Operation-level public APIs that open a session internally. Every
EagerTensormethod, and the one-shotTensor*trait methods on the concrete CPU backend, create a session (or an equivalent permit + pool context) per call. - A second CPU entry mechanism that is not a session at all.
CpuBackend’sinstall_with_pool_context*family installs a permit,CpuOperationEntry, and a buffer-pool loan, but hands the closure aCpuExecutionContext+BufferPoolinstead of aBackendSession. - Extension owner routes registered as
execute = ..._ownerthat callwith_backend_sessionon a&mut B: TensorBackend, even though a borrowed-session sibling already exists. - Generic helpers that reach a session-opening method on a
&mut Btype parameter rather than through a passed&mut dyn BackendSession.
The consequences are the ones the umbrella describes: an unbounded number of fixed per-call entries that cannot be amortized by a caller, and no mechanical way to tell a legitimate top-level entry from an accidental nested one. The design is also not uniformly stated — #1673 kept the one-shot surface as a deliberate convenience, while #1926 requires it to be removed.
Definitions
- Execution boundary — a public entry that is allowed to create backend execution state (session, permit, pool context, device ordering). Boundaries are named, few, documented, and enumerable.
- Hidden entry — any other function that creates backend execution state.
- Session surface — operations that receive a borrowed
&mut dyn BackendSession(or a concrete session reference) and never create one.
Target contract
R1. A caller opens an execution boundary and passes the borrowed session through user algorithms and helpers. A helper that receives a session never creates another one.
R2. Operation-level public APIs do not open sessions. The one-shot spelling of an operation is replaced by boundary(|session| op_in(session)), not kept as a second implicit entry.
R3. The set of boundaries is exactly the allowlist in the audit gate. A crate is not exempt as a unit; each allowed entry is named individually.
R4. Preserved during the whole migration: - CPU Send bounds on with_backend_session (a managed CPU session may run the closure on a Rayon worker thread; see session-oriented-concrete-apis.md). - Provider exclusion, resource-permit lifetime, and the reentrancy contract from cpu-backend-execution.md. - GPU device ordering and enqueue-vs-synchronize semantics. - Typed errors, validation, dtype promotion, broadcasting, and AD semantics — value/error parity for every migrated route. - Cache ownership and prepared-plan lifetime are independent of the session. Removing an implicit wrapper must not discard prepared plans, force materialization, or change which cache slot a call uses.
R5. Compiled-program execution (Runtime::run_compiled) and its segmented region boundaries remain valid top-level boundaries. They create compatible execution regions, which is a scheduler decision at an explicit entry, not a nested per-op entry. The same holds for the eager boundaries (with_eager_session, with_execution_session, with_extension_execution_context).
Inventory at a45833d4f
Column meanings: entry is the mechanism that creates execution state; status is one of boundary (allowed), remove (operation-level implicit entry to be deleted), thread (helper to receive a session instead of creating one), or decide (needs an explicit recorded decision).
A. Allowed boundaries
| Public entry | Definition | Entry mechanism | Backends |
|---|---|---|---|
BackendSessionHost::with_backend_session / with_backend_session_cached |
default bodies tenferro-tensor/src/backend.rs:3837, :3848 → default_backend_session :4218 |
session factory | all |
CpuBackend host impl |
tenferro-cpu/src/backend.rs:3791 → run_backend_session_cached (:3742) |
permit + CpuOperationEntry + CpuExecSession |
CPU |
CudaBackend host impl |
tenferro-gpu/src/cubecl/exec_session.rs:728 |
CudaExecSession + debug entry guard |
CUDA/CubeCL |
WebGpuBackend host impl |
tenferro-gpu/src/webgpu/exec_session.rs:265 |
WebGpuExecSession |
WebGPU |
EagerRuntime::with_eager_session |
tenferro-ad/src/eager.rs:988 |
forwards to with_backend_session |
all |
EagerRuntime::with_execution_session |
tenferro-ad/src/eager.rs:1863 |
forwards to with_backend_session |
all |
EagerRuntime::with_extension_execution_context |
tenferro-ad/src/eager.rs:1934 |
session + extension cache, one lifetime | all |
Runtime::run_compiled / run_compiled_values |
tenferro-runtime/src/runtime/snapshot.rs:1004, :1145 |
scheduler-owned regions (exec.rs, segment.rs) |
all |
with_backend_session_cached is the runtime-cache-aware entry used by the scheduler and eager paths; with_backend_session is the canonical user entry. Both are boundaries. The distinction between them is cache access, not session existence.
B. CPU backend operation-level entries
Two mechanisms coexist. run_backend_session_cached produces a CpuExecSession; the install_with_* family produces a permit plus a CpuExecutionContext/BufferPool loan with no session. Both are full execution entries — install_with_pool_context_unmarked (tenferro-cpu/src/backend.rs:2720, alongside install_with_pool_unmarked :2700, install_with_pool :2770, install_with_pool_context :2780, and install_with_indexed_pool_context :2790) acquires admission, constructs CpuOperationEntry, and calls entry.enter(...) exactly like the session path.
| Family | Trait impls and sites | Entry | Status | Replacement route |
|---|---|---|---|---|
| BLAS-1 | impl BackendSession for CpuBackend :2919 — vdot_read :2921, norm_squared_read :2925, axpby_read_into_accum :2935 |
run_backend_session_cached |
boundary | these are session methods; the CpuBackend-as-session impl stays |
| Dense contraction | impl TensorDot for CpuBackend :3451 — dot_general :3458, dot_general_read :3467, dot_general_read_into :3479, dot_general_read_into_accum :3492, dot_general_with_conj :3505 |
run_backend_session_cached |
decide | BackendSession::dot_general*_read* on a caller session |
| Dense contraction (cached) | impl BackendCachedDot for CpuBackend :3511 — dot_general_cached :3520, dot_general_with_conj_cached :3535, dot_general_read_into_accum_cached :3550, grouped_gemm_cached :3571 |
run_backend_session_cached(Some(cache)) |
decide | with_backend_session_cached + session _cached methods |
| Elementwise | impl TensorElementwise for CpuBackend :2953 (~30 methods) |
install_with_pool_context{,_unmarked} |
decide | session-provided pool/context |
| Reductions | impl TensorReduction for CpuBackend :3382 (5 methods) |
install_with_pool_context |
decide | session-provided pool/context |
| Indexing | impl TensorIndexing for CpuBackend :3577 (~10 methods) |
install_with_indexed_pool_context / install_with_pool_context |
decide | session-provided pool/context + indexed plan cache |
| Fusion | impl TensorFusion for CpuBackend :3882 (3 methods) |
install_with_pool_context |
decide | session-provided pool/context |
| Linalg pool helper | CpuBackend::with_linalg_pool :2830 |
install_with_pool_context |
boundary | documented external-implementation boundary for tenferro-linalg |
impl TensorAnalytic :3199 and impl TensorStructural :3281 do not call install_with_*; they are host-side and are not entries.
Compatibility decision (recorded). The impl Tensor* for CpuBackend blocks listed as decide are public backend SPI. Removing them from the public surface requires either (a) making the session method the only spelling and migrating every caller, or (b) keeping one named one-shot boundary per family. This document records the target as (a): the session surface (BackendSession) becomes the only operation primitive, and with_backend_session[_cached] the only entry. Option (b) is retained only for with_linalg_pool, which exists specifically so an operation-family crate can own its backend implementation while sharing the CPU pool. That is a named, documented boundary, not a crate-wide exemption.
Consequence for the audit gate: install_with_* becomes private to the CpuBackend boundary implementations themselves, and run_backend_session_cached is called only from BackendSessionHost/BackendSession.
C. Eager per-op entries (tenferro-ad)
Every public EagerTensor operation reaches a session through this chain. The public entry points are eager_ops.rs methods such as add :161, mul :202, reduce_sum :278, dot_general :335, matmul :544, transpose :595, reshape :634.
| Site | Entry | Status | Replacement route |
|---|---|---|---|
eager_exec.rs:461 in exec_standard_op_on_tensor_reads (:450) |
backend.with_backend_session |
thread | existing exec_standard_op_on_tensor_reads_in_session (:467) |
eager_exec.rs:261 in exec_dot_general_with_conj_on_tensor_reads (:254) |
backend.with_backend_session |
thread | session dot_general_with_conj_read |
eager_exec.rs:319 in exec_op_on_tensor_reads_with_runtime (:302) |
backend.with_backend_session |
thread | session to_contiguous_read / concrete_tensor_reads |
eager_exec.rs:733 in exec_standard_op_on_tensors (:726) |
backend.with_backend_session |
thread | session variant already exists inline |
eager.rs:2009 exec_outputs_read → eager.rs:4456 exec_single_output_read |
transitively opens a session per eager op | remove | with_eager_session / with_execution_session + session op |
eager.rs:2548, :2556 in EagerRuntime::store_grads |
backend.with_backend_session |
decide | AD-boundary session (see below) |
eager.rs:2032 exec_standard_graph_outputs |
backend.with_backend_session |
thread | test-only (#[cfg(test)]) |
eager_backend.rs:430, :581 |
with_backend_session dispatch wrappers |
boundary | these are the backend-erasure host plumbing |
Eager AD decision (explicit, not an exemption). The eager AD surface (EagerTensor in eager_ops.rs, store_grads) is migrated the same way as non-AD eager code: the AD boundary is the existing with_eager_session/with_extension_execution_context entry, and the differentiation entry point documents that it is the boundary for a whole forward/backward group. The migration must therefore establish that gradient accumulation in store_grads runs inside one boundary rather than one session per slot. If a single boundary for a forward/backward group cannot be expressed without changing AD value lifetime, that is a scope change and goes back to the issue — it is not silently exempted.
Eager migration contract (delivered; historical implementation notes follow). An eager operation must receive a runtime-bound, borrowed session created by the owning EagerRuntime at an explicit caller boundary. A plain &mut dyn BackendSession is insufficient: two eager runtimes may select the same backend type but own different provider/device state, while the eager value and trace are tied to one runtime. The boundary therefore carries both the originating runtime identity and the concrete borrowed session, and operations reject foreign eager values before dispatch. The public eager operation spelling moves to methods on this runtime-bound EagerSession (for example, session.neg(&tensor)), while EagerTensor remains the value/trace handle; remove the old implicit operation methods rather than forwarding them through a new entry. The session holds the backend lock in the existing lock/permit order; operation helpers receive it and do not lock the backend or enter again. EagerBackend is now only a session host/owner, not a composite Tensor* operation delegate: its former callers use the concrete borrowed backend session. CUDA/WebGPU owners no longer provide the deleted one-shot methods, and the tenferro-ad feature check now compiles. The former implicit EagerTensor operation methods were migrated and deleted; do not replace them with per-operation entry. Session-aware materialization must preserve placement, errors and retained-value lifetime; do not call the current duplicate_value from inside the session, because its fallback re-enters with_execution_session. Semantic recording also materializes untracked operands when a tracked operation consumes them: that copy must use the active session for core, view, and extension ops, rather than re-entering the backend during trace construction. CPU callbacks may run on a worker thread; the eager session entry therefore carries the calling thread’s no_grad / capture_trace depths into the callback for its duration (#1938 D12), while a guard started inside the callback stays local to it. This does not implement A2’s deferred AD wiring. Host-only leaf import can remain session-free, but a leaf constructor, copy, or retained-value materialization that actually executes on a backend also receives the borrowed session instead of hiding an entry.
The eager forward op surface (including non-AD use) migrates together; do not leave an old one-shot wrapper that secretly opens a session. Operation-family crates use traits or functions taking the borrowed eager session rather than adding another owner-entry method to EagerTensor. Eager backward’s compiled-program execution retains its named top-level runtime region, then accumulates gradient slots under one borrowed session opened at the backward boundary (the store_grads slice now does this). A prepared extension with a session executor uses that borrowed session and the runtime’s extension cache; its implementation must not lock the eager backend again. FFT’s EagerSessionFftExt implements the FFT operation-family borrowed route (fft/ifft/rfft/irfft); tensor-owned EagerTensorFftExt retains only the consuming in-place operations that require exclusive ownership and cannot use an ordinary borrowed input. A legacy native-context executor cannot be called while a session already holds that backend: it stays behind an explicitly named top-level runtime execution region invoked after releasing the eager session. No EagerTensor convenience operation may silently open that fallback region. Preserve the existing fallback’s errors and output semantics until #1938 makes a separate SPI decision. Acceptance requires numeric/error parity, same-runtime and foreign-runtime cases, nested-entry and provider-exclusion coverage, and CUDA / WebGPU feature builds with device tests where available. A2’s evaluation-scope hook is not a substitute for this API migration and remains deferred.
D. Extension owner routes
Each *_owner route has a borrowed-session sibling already. The _owner variant is registered as the extension’s execute/execute_reads fallback, which is why it still opens sessions.
| Crate | Owner route (entry) | Session sibling | Status |
|---|---|---|---|
tenferro-fft |
execute_fft_extension_reads_owner lib.rs:1462 (registration :1547–:1548) |
execute_fft_extension_reads_session :1473 → _on_session :1514 |
remove/thread |
tenferro-linalg |
execute_linalg_extension_reads_owner extension.rs:904 (registration :1118–:1119) |
execute_linalg_extension_reads :896 → _on_session :916 |
remove/thread |
tenferro-einsum |
execute_einsum_extension extension.rs:1021, execute_einsum_extension_reads :1050 (registration :1015–:1016) |
execute_einsum_extension_session_reads :1089 |
remove/thread |
tenferro-linalg helper |
matmul_preserve_trailing_batch tensor_ext.rs:2812 (dot_general_read :2824), linalg_matmul_read :2831 (:2846) call backend.dot_general_read(...) on &mut B: LinalgBackend |
session dot_general_read |
thread |
The registration macro already supports a session-capable route (execute_..._session); the migration is to make the session route the primary one and keep the owner route only as the scheduler’s explicit region fallback, with the region construction owned by the runtime rather than by the extension.
E. GPU backend entries
CudaBackend and WebGpuBackend implement the Tensor* operation traits by calling the same functions as their BackendSession impls. There is no distinct one-shot session struct: the backend is effectively its own session.
| Backend | Tensor* impls |
BackendSession impl |
Status |
|---|---|---|---|
| CUDA/CubeCL | TensorElementwise cubecl/mod.rs:4719, TensorAnalytic :5601, TensorStructural :5879, TensorReduction :6669, TensorDot :6902, TensorIndexing :6959, TensorFusion :7895 |
:8109 |
remove the owner impls |
| WebGPU | TensorElementwise webgpu/mod.rs:758, TensorAnalytic :836, TensorStructural :878, TensorReduction :952, TensorDot :970, TensorIndexing :992, TensorFusion :1047 |
:1077 |
remove the owner impls |
The delegate! shim is transitional and goes with them. Both GPU exec sessions implement the operation traits with a delegate! macro whose bodies are self.backend.<method>(..) (cubecl/exec_session.rs:444, webgpu/exec_session.rs:76), so the session is a thin forwarder and the real implementation lives on the owner. That is precisely the owner-as-session shape B3 removes, and it cannot survive the removal of the owner impls: once impl TensorElementwise for CudaBackend is gone, the shim has nothing to forward to.
Decision: B3(b) deletes the shim together with the owner impls, and the GPU operation bodies are expressed once as module-level functions over the backend — the shape gemm::*, structural::* and dispatch::* already use — with the session methods calling those functions. The session stays the only entry, the owner keeps only BackendRuntimeCache, TensorDeviceTransfer and BackendSessionHost, and no implementation is duplicated between an owner trait impl and a session trait impl.
The GPU-specific risk is different from CPU: a one-shot Tensor* call on the GPU backend does not open a session, it runs unbatched. The migration question is therefore whether the unbatched spelling is a boundary or a hidden entry. This document records it as a hidden entry: it must be spelled as with_backend_session(|s| op_in(s)) so that the ordering/lifetime contract is explicit at the call site. Nested-entry detection for the GPU overrides was debug-only when this was written; since #1938 D6 the portable guard rejects a nested entry with SessionEntryError::Reentered in every build profile.
F. Explicitly not entries
The single documented exception to the entry inventory is with_cpu_exec_session (crates/tenferro-cpu): it is a capability bridge for crates that must run a CPU-specific step inside a session they already own, not an alternative execution entry. It is retained deliberately on the out-of-scope list of #1929 (together with with_backend_session returning Result<R>, and release-mode nested-entry detection on the GPU, both since implemented by #1938 D6; D7 moved the bridge onto the opaque NativeSessionRef), and any new use of it must justify why the ordinary session route does not apply. Every other entry below is excluded because it is not an operation entry at all.
Runtime::run_compiledregion construction and the scheduler’s session regions (exec.rs:633–:796,segment.rs:313–:750).tenferro-fft’s plan/cache types and*_on_sessionbodies.tenferro-ad’sto_contiguous_host_read(eager.rs:1881), which is deliberately session-free and returnsNonewhen no such path exists.- Device transfer, allocation-domain, and buffer-retirement helpers, which do not execute tensor-sized ops.
Migration slices
This is the original implementation ordering, not a claim that every slice has landed: eager threading (slice 2) remains unresolved, and measurement (slice 7) is no longer a #1929 closure gate. The slices were ordered so each could be reviewed and reverted independently.
- Design + inventory (this document). No behavior change.
- Eager threading. Convert
eager_exechelpers to session-taking (*_in_session) and route every publicEagerTensoroperation through the caller’s boundary. Add nested-entry tests; verify value/error parity on the eager AD paths. Highest volume, and the precondition for removing the per-op entry. - Extension owner routes. Make the
_sessionroute primary; keep the owner route as an explicit runtime region fallback with the region formed by the runtime. Cover einsum, linalg (includingtensor_ext.rshelpers), and FFT. - CPU backend SPI. Collapse
install_with_*onto the session-provided pool/context and reduce theimpl Tensor* for CpuBackendblocks to session methods (or the recorded named boundary). Migratetenferro-linalg’sLinalgBackend-generic helpers. - GPU backend. Apply the same split to
CudaBackend/WebGpuBackend; add release-mode nested-entry enforcement for the overrides. - Audit gate. Add the deterministic check and its allowlist (below), then delete the removed public spellings.
- Measurement. The matched one-shot /
with_execution_scope/ one-shared-session comparison (below), reported separately from the API change.
Slices 2–5 may be split further per crate, but each must land with its own nested-entry test and parity evidence. 6 depends on 2–5.
Audit gate design
The gate must reject session factories in operation or helper implementations, not merely the text with_backend_session. Requirements taken from #1926 and made concrete here:
- Allowlist by exact source location, not by crate or by regex. Each allowed entry is one function body; adding one requires a reviewed edit to the allowlist file.
- Indirect detection. The check must follow the entry mechanism, not only the public name:
run_backend_session_cached,install_with_pool_context*,install_with_indexed_pool_context*,default_backend_session,with_backend_session[_cached], and the GPU exec-session constructors. A helper that only calls an alias must still fail. - Exclusions that are not exemptions. Test/bench/example code, and the bodies of the allowed boundary implementations themselves, are excluded because they are the boundary or are non-library code — the exclusion is source-location-scoped, and moving the code into the library must fail.
- Reachability sanity. The gate must include at least one indirect/aliased case in its own tests so it cannot silently degrade to a text grep, plus a negative test that a renamed alias is still caught.
- Integration follows the existing repository-rules review path (
scripts/repository-rules-review.pyrouting,REPOSITORY_RULES.mdsection). No new CI workflow. - The gate text must be mirrored by a normative bullet in
REPOSITORY_RULES.mdso the rule has one owner.
The gate cannot land before slice 5, because at a45833d4f it would fail on the inventory in sections B–E.
Audit gate: implemented
scripts/audit-session-entry.py implements the gate described above, earlier than slice 5 because it is useful as a freeze rather than as a cleanup: it records today’s legitimate entry functions in scripts/session-entry-allowlist.json and fails on any new one.
- Mechanisms tracked:
default_backend_session,with_session_entry_guard,install_with_pool_context[_fresh],install_with_indexed_pool_context[_unmarked],run_backend_session_cached, andCpuExecSession/CudaExecSession/WebGpuExecSessionconstruction. - The allowlist is keyed by mechanism and holds
path::functionentries, so a reviewed edit is required to add one, and a removal is recorded with--bless. - Aliases are resolved (
use ...::{Type as Alias},use ... as alias), so a renamed import is still reported;--checkruns the negative tests (aliased import, unrelated struct literal, allowlisted boundary, method call) on every invocation, so the checker cannot silently degrade into a text grep. - Library scope only:
tests/,benches/andexamples/are excluded as the boundary’s users, and the exclusion is path-based rather than symbol-based. scripts/check-pr-fast.shruns the audit for code changes, andREPOSITORY_RULES.mdnames it as the single owner of the rule (routed byscripts/repository-rules-review.pyfor backend/session paths).
The frozen inventory started at 76 path::function entries, dominated by the CPU owner entries that B3 removes (install_with_pool_context in 27 functions, run_backend_session_cached in 14). B3’s commits shrank it to 16, and any entry that reappears outside the allowlist fails the local gate:
| Mechanism | Entries |
|---|---|
CpuExecSession construction |
2 |
CudaExecSession construction |
1 |
WebGpuExecSession construction |
1 |
install_with_pool_context |
3 |
run_backend_session_cached |
3 |
with_session_entry_guard |
6 |
default_backend_session, install_with_indexed_pool_context*, install_with_pool_context_fresh |
0 |
The zero rows are kept deliberately: the mechanism stays tracked, so one reappearing entry is reported rather than silently untracked.
B3(iii) blocks on an acceptance-surface decision
A third attempt deleted the eight CPU owner implementations and migrated the resulting 317 errors (about 140 wrapped mechanically, the cached family moved to with_backend_session_cached, and the contract tests retargeted to exec_session.rs). The CPU crate’s own suite passed, and the session-entry allowlist shrank from 76 entries to 17, which is the shape B3(iii) is supposed to have.
It then failed in tenferro-ad, and the reason is a design question rather than a migration gap:
eager::tests::untracked_nary_ops_consume_lazy_views_without_materializing_inputsreduces a lazy view and now seesUnsupported { op: "reduce_sum", message: "backend does not accept borrowed tensor views at this execution boundary" }. The ad layer reached the CPU owner, whose read half materializes a view; the CPU session deliberately rejects one instead. Both behaviours are documented (the CPU session’s read halves reject borrowed views, the owner’s accepted them), and Step-1’s contract test pins the session behaviour.eager::tests::eager_backend_session_identity_projects_to_ownercompares session identities and now sees the same id on both sides, because the ad layer’s arrangement assumes the owner is its own session.
So “one implementation set on the session” is not behaviour-preserving for consumers that relied on the owner’s wider acceptance: either the session’s accepted input surface is widened to the owner’s (changing the session contract that Step 1 pinned and that the session-route benchmark measures), or those consumers materialize explicitly before the call (a real change in the ad layer). The issue’s direction resolves it: #1926 keeps the session interface, and the session’s read halves are the contract Step 1 pinned and that docs/testing/session-route-baseline.json measures, so the session’s accepted input surface is authoritative and the consumers that relied on the owner’s materialization have to materialize explicitly (the ad layer already has to_contiguous_read for exactly that). That is a behaviour-preserving migration of the consumer, not a widening of the session contract, and it is the path B3(iii) should take. The eager_backend_session_identity_projects_to_owner test then needs to state the new arrangement (the owner is the session provider, not a session).
The attempt was reverted because that consumer migration is a real change of its own and had not been reviewed yet; the tree stays at B3(i)+(ii)+B4.
The rest of the CPU-side work is then mechanical: the eight impl ... for CpuBackend blocks, the now-dead CpuBackendSessionMarker and install_with_indexed_pool_context*, the source-text contract retargets to exec_session.rs, and the allowlist re-bless.
B3(iii): the exact consumer that needs the decision
A fourth attempt reached the same point with the CPU-side work complete (the eight owner impls, the dead marker and the indexed pool helpers deleted; the session-entry allowlist down to 17 entries; the three source-text contracts retargeted to exec_session.rs; the CPU suite green) and stopped in tenferro-ad again. The failing path is now pinned:
EagerTensor::reduce_sumbuildsStdTensorOp::ReduceSumand runs it through the untracked eager path, which reached the CPU owner; the owner’sreduce_sum_readmaterialized a borrowed view before delegating, while the CPU session’sreduce_sum_readrejects one (crates/tenferro-ad/src/eager_ops.rsunary_op→ the untracked n-ary execution → the backend entry). The traced path is unaffected because the runtime materializes slots beforeeager_exec.rscallsexec.reduce_sum_read(:524).eager_backend_session_identity_projects_to_ownercompares session identities and assumes the owner is its own session.
So the consumer-side fix is not a single mechanical edit; it is a policy choice about where the ad layer materializes: always, on the untracked path only, or per operation (only for the read halves whose session contract rejects views, which is the reductions and the same family the trait’s provided defaults reject). Each option has a different cost profile for the untracked path, which is exactly what the Phase-C baseline measures, so it should be decided with the issue rather than picked inside a migration slice.
B3(iii) status: CPU and WebGPU done, CUDA outlined
The earlier “acceptance surface” blocker was a misdiagnosis and is resolved: the failure came from routing the composite EagerBackend’s BackendSessionHost::with_backend_session to with_session_entry_guard(|| f(self)). The enum is not a session; forwarding to the concrete backend’s session (as before) keeps the untracked eager path on the same accepted input surface. With that, B3(iii) needed no consumer policy change at all.
Landed so far:
- CPU (
007da98a7): the eight owner impls, the dead marker and the indexed pool helpers are gone, ~380 call sites migrated, three source-text contracts retargeted, the allowlist down to 16 entries, workspace tests green. - WebGPU (
ad6b41720): the operation bodies moved fromimpl Tensor* for WebGpuBackendintoimpl Tensor* for WebGpuExecSession<'_>(thedelegate!invocations for the operation families became real impls), the owner’sBackendSession/BackendCachedDotimpls and the session marker deleted, andSessionCachedDotimplemented directly on the session (WebGPU has no runtime cache, so the trait defaults are what the owner’s blanket impl provided). Verified with and without thewebgpufeature.
CUDA is the remaining half, and the procedure is now known:
- move each
impl Tensor* for CudaBackendbody intoimpl Tensor* for CudaExecSession<'_>, rewriting receivers toself.backendandselfpassed as an argument (structural::transpose(self, ..),gemm::dot_general(self, ..),promotion::*(self, ..)), which is the part a naive rewrite misses; - add the imports the moved bodies need in
cubecl/exec_session.rs(dispatch,elementwise,gemm,permutation,fusion, the promotion helpers,DType, …); - fix the E0599 calls that were owner methods (
to_contiguous_read,dot_general_with_conj, …) to the session form, and addimpl SessionCachedDot for CudaExecSession<'_> {}in place ofdelegate_cached!; - delete
impl BackendSession for CudaBackend/BackendCachedDot for CudaBackendand the marker, then retarget the CUDA source-text contracts (cuda_launch_contract,public_surface_contract,session_contract,backend_read_contract) that namecubecl/mod.rssections.
Attempted twice and reverted. What the attempts measured, in addition to the 1113 errors of a move-only pass (572 E0425, 302 E0599, 201 E0433, 36 E0277):
- the receiver rewrite must also handle
selffollowed by a newline before the dot (about 130 sites per pass; a literalself.rewrite misses them), andselfpassed as a bare argument (structural::transpose(self, ..)); - the moved bodies invoke the
dispatch::launch_*macros, which expand at the call site, socubecl/exec_session.rsneeds the macro bodies’ names in scope (CubeclCudaRuntime,ArrayArg,ComputeClient, the promotion helpers, …), not just the module path; - the CUDA-specific
*_typedhelpers (gather_typed,dynamic_slice_typed,scatter_float_typed, …) are inherentimpl CudaBackendmethods incubecl/mod.rsand stay; only the trait entries move; delegate_cached!becomesimpl SessionCachedDot for CudaExecSession<'_> {}like WebGPU’s, since the CUDA owner’s cached behaviour was the trait default plus the provider’s own caching.
**The module-function shape landed, and it is what the attempt history argued. The CUDA operation bodies now live in crates/tenferro-gpu/src/cubecl/ops.rs as 57 free functions over &mut CudaBackend, the seven owner impls and the owner BackendSession/BackendCachedDot impls are deleted, and exec_session.rs keeps one delegation macro per direction:
- each owner method became
pub(super) fn <op>(backend: &mut CudaBackend, <args>) -> crate::Result<...> { <body> }with a plainself->backendrename, which is also what keepsself-taking macro invocations valid; delegate_ops!generates the session impls and forwardsops::<op>(self.backend, <args>), so the session’s method table still comes from one list per family;delegate!continues to serve the families that stay on the owner (TensorBuffer,TensorDeviceTransfer);impl SessionCachedDot for CudaExecSession<'_>keeps the two device-path_read_cachedoverrides and takes the trait defaults for the rest, matching WebGPU.
The receiver rename is mechanical and reviewable per function, and step 2 keeps one implementation per operation, which is what B3 requires. The measured fallout supports the choice: the in-place attempts produced 1113 errors, while this shape left six — four to_contiguous_read calls in blas1.rs and an inherent helper, the owner BackendCachedDot impl, and one helper that needs a session-typed receiver.
Two traps worth recording, both caught by the compiler rather than by review:
TensorElementwise::elementwise_read_intocannot move to a free function: its allocating fallback is generic overTensorElementwise, which only the session implements once the owner impl is gone. It stays a hand-written session method, anddelegate_ops!grew an optionaloverride { ... }group for it;- that first version silently dropped the native read-into kernels, which showed up only as three
never usedwarnings (elementwise_read_into_nativeand two launch helpers). The override now tries the native path first and falls back to the allocating helper, so no work moved onto the allocating path.
The CUDA tests and example migrated through the same wrapper as CPU and WebGPU. Two shapes needed hand work: upload(&gpu, ..) inside a session closure borrows gpu immutably while the session needs it mutably (E0502), so the upload is hoisted into a let before the enclosing statement, and two object-safety tests that wrote let exec: &mut dyn BackendSession = &mut gpu; now obtain the erased session from with_backend_session, which still proves the same interface.
The source-text contracts were retargeted to the module that now holds the bodies: cubecl_launch_contract (eight anchors), public_surface_contract (two) and backend_read_contract, whose session half became a compile-time assertion that CudaExecSession<'static> implements all six operation traits — the same shape webgpu_backend_contract uses, and stronger than the source scan it replaced. backend_read_contract is now #![cfg(feature = "cuda")], because that assertion names the CUDA session type.
Verification on this host: cargo check --workspace --all-targets and cargo check -p tenferro-gpu --features cuda --all-targets are error- and warning-free; the CUDA suite reports 108 passed / 189 ignored in the lib and 73 passed in the integration target; the WebGPU suite reports 89 + 31 + 3 + 2 passed; and scripts/audit-session-entry.py --check stays green without an allowlist change. What this host cannot do is run the hardware-gated kernels, so their behaviour remains with the hosted GPU matrix; the bodies themselves are verbatim moves, and no receiver rewrite changes an argument.
Local gate state (mid-Phase-B)
bash scripts/check-pr-fast.sh --no-fetch --coverage-reviewed --test 'cargo test -p tenferro-gpu --features webgpu --lib --tests' passes on the current branch after the fixes it surfaced, which are worth remembering:
- the standalone
ext/tenferro-cpu-tblismanifest is outside the root workspace, so its provider test kept the deleted one-shot spelling; docs/guides/devices-and-gpu.mdis generated from a snippet source and was stale for the same reason (check-doc-snippets.pysyncs it);- clippy (
-D warnings) caught seven orphaned# Errorsdoc blocks left behind by deleted trait items and the needless&borrows the migration introduced.
Run it with RUSTC_WRAPPER="" locally: the kache wrapper’s path remapping breaks trybuild .stderr comparisons, which is also why the pre-existing tenferro-ad::eager_backend_capability_contract fixture mismatches here while the tenferro-gpu session contracts pass. That fixture mismatch is span-only: the same E0432 is reported, with a wider underline than the recorded .stderr, so it is a rustc-rendering difference on this toolchain rather than a behaviour change.
Two other fixtures of that contract were invalidated by this refactor and are blessed with it. Both reported E0308/E0576 for the removed owner-projection APIs, and the only change in their .stderr is the removal of the compiler’s suggestion that the expected type could become a session, which is false now that the owner is not a session. The third, span-only fixture is deliberately left untouched, so after this change the contract reports exactly the same single mismatch it reports on the baseline worktree — the pre-existing claim above is a comparison, not an assumption.
Phase-B completion audit
Each Phase-B requirement mapped to the artifact that satisfies it. “Evidence” means a file, a command output, or a test that runs in CI, not an intention.
| Requirement | Evidence |
|---|---|
B1 fail: owner add/add_read on three backends |
tenferro-cpu/src/lib.rs crate docs pin the deleted owner spellings for add, mul, exp, reduce_sum, transpose, dot_general; cubecl/exec_session.rs and webgpu/mod.rs pin that the CUDA/WebGPU owners do not implement an operation trait (a call cannot resolve for the same reason) |
B1 fail: dyn BackendSession old add |
tenferro-tensor/src/backend.rs::BackendSession compile_fail example |
B1 fail: BackendCachedDot |
tenferro-cpu/src/lib.rs owner-bound compile_fail example |
B1 fail: default_backend_session |
tenferro-tensor/src/backend.rs::BackendSessionHost compile_fail example |
B1 pass: Tensor::add(.., session), cached operations, typed-view operations, scheduler/extension dyn BackendSession, scope nesting |
five trybuild fixtures in tenferro-runtime/tests/ui/session_surface/pass/, driven by session_surface_contract.rs, which runs under the CI nextest profile (verified with cargo nextest run -p tenferro-runtime --test session_surface_contract) |
| B2 delegation inverted, one-shot spelling deleted | 31 per-operation deletion commits; _read/_into are required items |
| B2 acceptance ranges unchanged | backend_default_read_tests.rs: default_read_methods_delegate_owned_tensors_and_reject_views, structural_runtime_materialization_rejects_views_by_default |
| B3 supertraits trimmed | pub trait TensorBackend: BackendRuntimeCache + TensorDeviceTransfer + BackendSessionHost |
| B3 one operation implementation set per backend | CPU CpuExecSession; CUDA ops.rs bodies with delegate_ops!; WebGPU real impls on WebGpuExecSession |
| B3 deletions | owner impls, the non-public install_with_pool_context*, every BackendCachedDot impl, default_backend_session, the backend-as-session markers and factory |
| B3 CPU linalg policy preserved | preferred_linalg_mode consumed at tenferro-cpu/src/backend.rs:2752, defined in provider.rs |
B3 GPU delegate! shim decided and documented |
CUDA operations moved to ops.rs and deleted from the shim; delegate! remains only for TensorBuffer/TensorDeviceTransfer |
| B3 generic-bound and call-site migration (~215 sites, doctests, examples) | cargo check --workspace --all-targets is clean, which is the machine-checkable form of the migration |
| B4 extension owner path is session-primary | define_extension_runtime!’s owner entry forms the session and hands extensions &mut dyn BackendSession; owner routes are limited to the runtime-formed fallback (execute_owner_extension_fallback); the linalg helper is session-typed |
| B5 deterministic audit gate | scripts/audit-session-entry.py + scripts/session-entry-allowlist.json, run by scripts/check-pr-fast.sh; tracked mechanisms and the 16-entry inventory are in the table above |
| B5 alias negative test, deny on a hidden entry, pass on the boundary | the --check self-test plus the recorded demonstration below (exit 1 with the hidden entry, exit 0 without) |
| B5 single rules owner and doc consistency | REPOSITORY_RULES.md “Backend Session Entry” section, routed by scripts/repository-rules-review.py; docs/design/explicit-session-boundary.md and docs/design/index.md updated together |
B5 with_cpu_exec_session as the single documented exception |
explicit-session-boundary.md “capability bridge” entry plus the allowlist row |
| Final state: one-shot spellings do not compile | the compile_fail fixtures above |
Final state: with_backend_session/with_execution_scope are the only boundaries |
the 16-entry allowlist, whose remaining rows are session construction and the runtime cache entry |
Final state: operations are not reachable from TensorBackend |
the trimmed supertrait list together with the owner compile_fail fixtures |
| Evidence: build/check and tests pass | the table under “Phase-B closing state” |
| Evidence: doctests compile and run, deleted-API doctests moved | cargo test --doc for the affected crates, including the fail fixtures that now run as doctests |
| Evidence: pushed to the work branch | refactor/1929-session-route-unification |
| Evidence: one batched local gate at the end | scripts/check-pr-fast.sh and scripts/repository-rules-review.py --dry-run |
Outside this goal’s scope, by the task’s own split: the harness change forced by the deletions (before-only arms cannot outlive the spellings) means the recorded baseline is the before-reference and the certification run that re-measures the baseline side, including the CUDA target, is the umbrella’s Phase C.
Phase-B closing state
With the CUDA body move in, Phase B is closed on this host:
| Check | Result |
|---|---|
scripts/check-pr-fast.sh --no-fetch --coverage-reviewed --test 'cargo test -p tenferro-gpu --features cuda --test integration' |
pass (fast PR checks passed) |
cargo check --workspace --all-targets |
0 error, 0 warning |
cargo check -p tenferro-gpu --features cuda --all-targets |
0 error, 0 warning |
cargo test --workspace --no-fail-fast |
all targets pass except one fixture of the pre-existing tenferro-ad trybuild contract, verified pre-existing by running the same test on the baseline worktree |
cargo test --manifest-path ext/tenferro-cpu-tblis/Cargo.toml |
5 + 3 passed |
scripts/audit-session-entry.py --check |
pass, 16 entries, no allowlist change |
scripts/repository-rules-review.py --dry-run |
pass |
What Phase B does not include is evidence that the refactor is performance-neutral. That is Phase C: its recapture and first matched comparison are recorded below, and they show a real cost on small operations that Phase C has to own.
Phase-C pre-flight (tooling verified, measurement pending)
scripts/compare-session-route-baseline.py was run against the recorded baseline and the Phase-A candidate logs to validate the comparison path before the real measurement. Result: 98 PAIRED_OK, 1 NOISY, 0 REGRESSION, 0 DELETED, and two expected failures that prove the fail-closed behaviour:
tenferro-gpu|route_matrix_gpuhad no candidate log in those older logs. The CUDA benchmark does run on this host (13 cases in the recapture below); the missing log was a property of that log set, not of the host.session_chain/broadcast/execution_scopeis absent from those older candidate logs, and the comparator reports a missing non-deleted baseline case instead of skipping it.
So the harness, the capture/comparison pair and the thresholds work; what remains for Phase C is the measurement itself, which needs the documented quiet window (load below the recorded threshold with no live cargo/rustc, and a free core — cpu=0 is occupied on this host).
Harness identity bug found during the first candidate run
The first candidate run compared route-matrix cases against routes the candidate no longer has. route_matrix still registered oneshot arms for dot_general and reduce_sum, whose one-shot spellings Step 2 deleted, and later runs found the same for slice and cast, whose arms the Step-2 migration had already rewritten to with_backend_session — every one of them timed the session route under a before-route name, so the comparator compared them against owner-route baseline numbers. route_matrix_gpu had the same shape for dot_general: its oneshot helper had become a byte-for-byte duplicate of its session helper.
All before-only arms and their helpers are now removed from the campaign harnesses, so no live registration uses an oneshot/one_shot label. That is what the baseline’s deleted-route predicate expects: the recorded numbers stay the before reference and the comparator reports those cases as DELETED rather than as regressions. The comparable pairs are session/* and scope/* against their own recorded values, and the price of the unification is read off the recorded oneshot/* numbers against the candidate’s session/*.
The pruning was then verified without re-benchmarking, by deleting the oneshot blocks from an existing candidate log of route_matrix and re-running the comparator on a scratch logs directory: deleted_route=16 and no regression came from a pruned arm. The three +5.6%/+6.4% entries it still reported are on session/scope cases of that loaded diagnostic capture, which is exactly the noise this host cannot resolve; they are not a result, only evidence that the classification path works. One caveat found while doing this: the comparator reads the gate’s merged stdout/stderr form (Benchmarking <case>: Analyzing), so a stdout-only capture (cargo bench ... > log) cannot feed it and must be re-run through scripts/run-session-route-performance-gate.sh.
Recapture performed, and the first matched comparison
The baseline side has since been re-collected with the current harness, which is what the harness-identity rule asks for once the before-only arms are pruned: seven targets, 96 cases, docs/testing/session-route-baseline-recaptured.json. The capture guard fired first (expected 47 cases, parsed 31), so the recapture is the deliberate EXPECTED_CASES update in af916dd87; the frozen session-route-baseline.json keeps its 125 rows and the 29 before-only references.
The measurement used cpu=1, because cpu=0 — the recorded runs’ core — is 100% busy for 30-second averages, held by a foreign long-running julia process on this shared host. Three comparator runs followed:
| Comparison | Result |
|---|---|
| matched: candidate vs recaptured baseline (one core, one window) | paired_ok=49 noisy=27 regressions=20 |
| historical: candidate vs recorded baseline | paired_ok=67 noisy=21 regressions=8 deleted_route=29 |
| A/A: candidate pass 2 vs candidate pass 1 | paired_ok=53 noisy=7 regressions=0 |
The A/A pass covers the targets carrying the matched regressions and flags none of them, so those deltas are not host noise: one multi-millisecond backward case (+14.9%) and ten small-operation dispatch cases (+5.0%…+7.2%), consistent with more session entries per unit of user work after B3/B4. That is the umbrella’s Phase C optimization input, recorded in docs/worklogs/2026-09-26-session-route-recapture-and-matched-comparison.md; this document keeps it as a Phase-C finding rather than a Phase-B correctness issue.
Under the original measurement contract, three-pair certification would have needed a consistently free core (cpu=0 is unavailable on this host). The 2026-09-27 revision removed that certificate as a #1929 closure gate.
Still open: one session per operation in the eager/tape path
What remains regressed is the eager path, and it is a different mechanism: B3 removed the owner route, so every tape node now constructs its own session. That predicts, and the measurements match, a per-operation cost of roughly one session floor:
| target | case | baseline | after the compiled-path fix |
|---|---|---|---|
linalg_vjp_gate |
triangular_solve_vjp/{8,16} |
320.95 / 319.77 µs | 387.47 / 388.06 µs (+21%) |
linalg_vjp_gate |
svd_values_vjp/{8,16} |
336.29 / 333.32 µs | 363.88 / 364.52 µs (+8-9%) |
eager_dispatch_baseline |
materialized/reduce_sum_f64/1 |
13.61 µs | 14.07 µs (+3.5%) |
A triangular-solve VJP is on the order of ten linalg operations, so ten session constructions at ~7 µs predict +70 µs, which is the observed +68 µs. The greedy probe/reclaim entries tested earlier are not the cause: gating them changed the linalg numbers by less than the run-to-run spread (a recorded negative result).
The matching fix is to evaluate a whole eager op sequence, or a whole backward pass, inside one execution scope so the per-operation entries reuse one permit — which is exactly the pattern session_chain measures and shows no regression for. Two constraints make this a design decision rather than a local edit:
- nesting a session entry inside an open session is prohibited and detected (the scope API is the sanctioned way to reuse a permit, and the
session_scope_nestingfixture pinswith_execution_scopewrappingwith_backend_session); with_execution_scopeis a CPU-specific API, while the eager evaluator is backend-generic, so the AD layer needs a generic mechanism — a scope hook on the backend contract, or a reusable session — before it can use this.
That is an API decision for the umbrella, not an unattended correctness fix, so it is recorded here rather than implemented.
The scope hypothesis is confirmed (bench-only intervention)
The council and the Astra review fixed the bar before measuring: recover at least 80% of the linalg_vjp_gate delta, with one scope per timed iteration. The intervention lives only in crates/tenferro-linalg/benches/linalg_vjp_gate.rs (the fixture keeps a CpuBackend clone so each iteration can open its own scope); no library file changed. Both variants are cases of the same benchmark run, so the comparison is matched by construction, and the pinned core was idle before and after every run.
| case | baseline | no scope | with scope | vs baseline |
|---|---|---|---|---|
triangular_solve_vjp/8 |
320.95 µs | 379.84 µs (+18.3%) | 178.79 µs | 1.80× |
svd_values_vjp/8 |
336.29 µs | 370.81 µs (+10.3%) | 182.78 µs | 1.84× |
triangular_solve_vjp/16 |
319.77 µs | 383.43 µs (+19.9%) | 185.01 µs | 1.73× |
svd_values_vjp/16 |
333.32 µs | 372.11 µs (+11.6%) | 185.15 µs | 1.80× |
The scope does not merely recover the regression: it makes these VJPs about 1.8× faster than the pre-unification baseline, which says the eager path was paying much of the same per-operation session cost before the unification too. So this is a cost reduction available on top of the contracts, not only a restoration.
The session-entry counts are unchanged by the scope (2000 outer and 12000 inner entries per profiler window in both variants); what changes is the cost of an outer entry, 18.9 µs → 7.9 µs, once the permit and pool loan are already held. Session construction inside the scope measures ~32 ns per call. Details, the profiler caveat that its sections do not cover the whole ~200 µs saving, and the remaining questions are in docs/worklogs/2026-09-27-scope-intervention-experiment.md.
Execution scopes as an evaluation boundary: ownership protocol
Status (2026-09-27 revision). This section is a deferred design, not part of the #1929 completion contract. The maintainer’s 2026-09-27 revision removed the benchmark-first order, the exhaustive completion checklist and the per-step performance gates; the evaluation-scope hook and its AD wiring are performance optimization and move to the post-integration work in #1938, tracked by #1927 and #1904. Everything below is retained as the ownership analysis and fallback contract that work will need. Nothing here is implemented, and
with_evaluation_scopestays at zero allowlisted entries in the session-entry audit.
Both intervention halves confirm the direction, so the next artefact is the protocol that a scope-as-evaluation-boundary must satisfy. This is the design to agree on before any public hook is frozen or shipped; nothing here is implemented.
What a scope owns. One execution permit, acquired once from the backend’s domain, plus the CpuOperationEntry and the resource set the permit keys (BufferPool, the GEMM analysis cache, the indexed plan cache). Sessions opened inside the scope reuse that permit and those resources rather than acquiring their own — measured as ~32 ns of session construction per inner entry, against a fresh outer entry of 18.9 µs before the scope and 7.9 µs with it.
Holder and release. The scope holds the permit/entry/resources for exactly the duration of its callback and releases them on every exit path: normal return, early return through ?, and unwind (the state is dropped). A session must never outlive its scope, which the callback’s lifetime already constrains; the implementation must also not return a borrow of the scope state to the caller.
Joining, and never a new failure mode. The hook is “ensure a scope for this domain, then run”, not “always open one”:
| state when the hook is entered | required behaviour |
|---|---|
| no scope, managed domain, no active execution | open one scope for the callback, then release |
| a scope is already active for this domain (user code, or a nested evaluation) | run the callback inside it; do not acquire a second permit |
a scope or CPU execution is active for another domain, or the domain is externally managed (Error::Unsupported), or the backend has no scope concept |
run the callback plainly; the optimization is skipped |
The last row is the load-bearing rule: the hook may only make the common case faster and must never turn a working call into an error. Today with_execution_scope returns Error::RuntimeState for a second scope or active execution and Error::Unsupported for an externally managed domain; the hook must absorb both and fall back.
No waiting. Contention is handled by falling back, never by blocking: a hook that waited for a permit would introduce waiting cycles and delay other users, which the review flagged as a risk. Consequence: a scope’s duration must stay bounded (one evaluation or one backward pass, not a user session), and a contention measurement is part of acceptance — with a second thread active, the other user’s latency must stay within a recorded bound or the fallback must trigger.
Backend-generic shape. A defaulted method on the backend contract, so backends without a scope keep the current behaviour:
/// Run `f` with one execution permit/resource set shared by every session entry
/// opened inside it. Default: no scope. Never fails because a scope is
/// unavailable; it falls back to running `f` directly.
fn with_evaluation_scope<R: Send>(&mut self, f: impl FnOnce(&mut Self) -> R + Send) -> R {
f(self)
}CPU implements it by splitting the existing scope machinery: acquire the scope state from &self (permit, entry, resources), then run f(self) with &mut self and release on exit. The existing user-facing CpuBackend::with_execution_scope(&self, …) stays as it is; the hook must not require Clone on the backend or on the caller’s tensors. CUDA and WebGPU keep the default: their session entry is a struct plus an entry guard with no permit, pool loan or plan-cache lookup, and the GPU campaign target shows no regression, so a GPU scope would buy nothing today. If one is ever added it must preserve enqueue order and the synchronize points exactly.
Invariants the hook must not touch. Values, dtypes, validation order, typed errors and their sources, failure-output immutability, AD semantics, provider exclusion, permit lifetime semantics (acquired once per evaluation instead of once per operation, released at scope exit), GPU device ordering, cache ownership and plan lifetime. A scope is evidence that a permit is held, not a new kind of session: one session at a time inside it stays the rule, and the existing nested-entry detection is unchanged.
Audit gate (implemented now, before the hook exists). The session-entry audit tracks mechanisms that create execution state, and a scope is one of them even though it is not a session entry: it holds an execution permit and the resource set that permit keys. Two mechanisms were therefore added to scripts/audit-session-entry.py, and REPOSITORY_RULES.md names both in its “Backend Session Entry” section:
with_execution_scope— allowlisted at exactly one site, the API definition intenferro-cpu(the reviewed boundary). Adding the mechanism immediately surfaced that site, which is the point of the gate.with_evaluation_scope— the planned hook, tracked with zero allowlisted entries. Library code cannot call it without failing the gate, so the hook cannot appear, or spread to a second call site, without a reviewed allowlist change.
Demonstrated end to end: inserting an aliased with_evaluation_scope call in crates/tenferro-ad/src/eager_exec.rs fails the check with with_evaluation_scope: unallowlisted session entry at crates/tenferro-ad/src/eager_exec.rs::<fn> (exit 1), and the reverted tree passes again. The audit’s own negative tests now cover both scope mechanisms, so the check cannot silently degrade into name matching.
Public benchmark harness compatibility is follow-up work, not a #1929 gate. The two repositories retain their own issues (#107 and #41). The maintainer’s 2026-09-27 revision removed their completion and publication from #1929. In an isolated worktree of tenferro-benchmark at origin/main (2a8469f), pointed at this branch:
cargo build --release --features cpu-faer --binsfails in four binaries with deleted owner spellings onCpuBackend(reduce_sum,mul,exp, etc.) and separate config-type errors (DotGeneralConfigexpectsSmallVec<[usize; 4]>rather thanVec<usize>). The complete benchmark harness needs an API update.- after migrating two
reduce_sumcalls toreduce_sum_readwithTensorRead::from_tensor,small_work_casebuilds. The separatebenchmark_cpu_sessionbinary also needsdot_general_readand config updates; buildingsmall_work_casedoes not establish a working session lane. - the runner records the exact
tenferro_rscommit anddirtyflag, enabled features, BLAS implementation, and thread environment. It has raw-sample storage, but an empty sample file is not a measured result. - a host-only diagnostic invocation with
BENCH_INSTANCE=add_f64_concrete_freshproduced zero rows:selected_casesfilters out-freshcases unlessBENCH_INCLUDE_SETUP_DIAGNOSTICS=1, even when explicitly selected. This was selection filtering, not evidence that the faer provider cannot run the case. The invocation is not a publishable Linux result: the benchmark repo requires Linux collection inside its devcontainer. - what remains under
tenferro-benchmark#107 (not #1929 closure): explicit expected/selected/executed/unsupported/failed/noisy case identities in run metadata, nonempty matching samples, versionedquick/fullmanifests with measured wall times, and the session/public-route latency lane. TENFERRO_CPU_FEATURES=cpu-faeravoids the OpenBLAS prefix requirement that the Linux runner defaults to; this is useful for compilation diagnostics, not a substitute for the repository’s Linux publication provider policy.
Deferred: the eager backend’s ownership shape (verified in code, hook reverted). A first implementation of the hook landed and compiled — a defaulted with_evaluation_scope(&self, f: impl FnOnce() -> R + Send) -> R on BackendSessionHost, the CPU override built on a split of the existing scope machinery (open_evaluation_scope acquiring the permit/entry/resources, and a run that keeps the callback available when entry admission fails), and the EagerBackend forward. It was reverted unused, because no safe call site exists in the current eager runtime:
- a CPU permit was acquired with a
compare_exchange, so a second concurrent acquisition panicked (BACKEND_REENTRY_PANIC) instead of waiting (since #1938 D6, a permit held by another thread is waited for in FIFO order and same-thread reentry is a typedSessionEntryError::Reentered). Opening a scope outside the eager backend’s mutex and then locking inside it inverts the lock order: the scope holds the permit and waits for the mutex while a concurrent same-runtime operation holds the mutex and takes the permit. Today that pair serializes. CpuOperationEntry::entermay install its callback into the domain executor, so the callback must beSend. The eager path’s closure captures the backendMutexGuard, which is!Send, so it cannot be passed into a scope.EagerBackendis a composite enum behind that mutex. It cannot forward a&mut self-callback hook (the match arm reborrowsselfwhile the callback needs it again), and the&self+FnOnce() -> Rshape needs a second handle of the same domain, which the enum cannot produce: it does not implementClone, and the test-onlyRecordingBackendis not semantically cloneable.- a thread-bound (
!Send) scope entry does not help either: the outermost scope’s callback is the eager body, andentermay still need to install it.
So the hook needs an ownership decision before it can be implemented, with one viable path: evaluate an eager operation (or a backward pass) on a concrete handle instead of the composite enum, so no mutex guard crosses the scope and today’s lock and permit ordering is preserved. That changes how the eager runtime owns its backend, which is why it is not an unattended edit. The measured win awaiting it is the linalg 2.0–2.1× and the eager small-op −14…−18% from the intervention experiment. Under the 2026-09-27 revision the hook is deferred rather than decided: the audit freeze keeps with_evaluation_scope at zero allowlisted entries, so the implementation cannot appear unreviewed, and the ownership question moves to #1938 together with the rest of the post-integration performance work.
Deliverable boundary after the 2026-09-27 revision. The active #1929 assignment is reduced to: deliver the #1926 explicit-session and route/API migration for its agreed scope, run the required CI and focused tests for the changed behaviour, and leave a short handoff. Phase B landed the backend owner/session route migration; the maintainer subsequently approved migrating eager entry too. That migration now exposes borrowed EagerSession operations and removes the ordinary per-operation eager entry surface. store_grads borrows one session opened at the backward boundary instead of entering separately for each slot. Owned-only extension fallback remains a separate named top-level boundary. This is independent of the deferred evaluation-scope hook. The hook, its AD wiring, the re-verification campaign and the three alternating pairs are not performance closure gates. Neither is completing tenferro-benchmark (#107) or strided-rs-benchmark-suite (#41): their manifests, full coverage and publication move to the post-integration work in #1938, and the per-issue CPU/GPU optimization workstreams (#1927, #1928 and their focused issues) keep their own ownership. The kept harness work is recorded in the handoff, and unbounded measurement is explicitly out of scope.
The following was the pre-revision hook proposal, not an implemented or approved ownership decision; #1938 must re-evaluate it against the current eager runtime before proceeding:
- Hook home and shape. A defaulted method on the existing
BackendSessionHosttrait intenferro-tensor, namedwith_evaluation_scope, taking&selfand aFnOnce() -> R + Sendcallback with a default that simply calls it. The caller opens the scope on a handle and runs the evaluation on another handle of the same domain (the pattern the intervention measured), so the hook needs noClonebound and no interior mutability. - AD granularity. One scope per eager evaluation and per
backward()call: the single-operation entry path wraps its operation, andbackward()wraps its tape replay. The join rule makes the nested case free, and this is the granularity the measurements used. - Residue. Operations executed outside an evaluation (runtime-dispatched compiled runs, direct session use) keep today’s behaviour. Any residue that survives the hook is recorded per case with a bound rather than silently accepted.
Deferred plan (recorded, not a closure gate). The steps below are what the post-integration work in #1938 would follow; the 2026-09-27 revision removed them as #1929 requirements.
- The API is tracked under the existing issue #1926 / umbrella #1929; no new intake. The audit freeze above is already in place.
- Implement the hook plus the AD call sites: one scope per eager evaluation and per
backward()(the granularity the experiments measured), with the fallback table above. - Tests: parity of values/dtypes/errors with and without the scope; release on normal,
Errand panic paths, with a following scope opening normally; joining a caller’s open scope; externally managed domain falls back silently; other backends keep the default; contention/latency bound; the compiled-path and GPU non-regression rows. - Open questions that must be settled before any public hook is frozen: the hook’s trait and name, the AD granularity (per evaluation, per backward, or both), whether the residue that a scope does not cover (ops executed outside an evaluation) is accepted with a recorded bound, and the ownership decision above. The allowlist question is settled: the hook’s call sites enter the audit as reviewed allowlist entries, and the freeze stays at zero until then.
Consequences for the design decision: the direction is worth adopting, and Astra’s ordering stands — the execution-ownership protocol (who holds the permit and the entered execution, where it is released on normal, error and unwind paths, how a caller’s already-open scope is joined, how other backends and external executors behave, and the waiting relationships under contention) must be settled before any public hook is frozen or shipped. The eager_dispatch_baseline small-op cases are the other half of the residue and were not re-measured under a scope; they belong in the same follow-up.
The criterion settings stay at the pinned defaults if the campaign is ever run; cheaper settings are acceptable for a diagnostic pass only, because they change the confidence intervals the comparator uses to separate NOISY from REGRESSION.
Certification runbook (not required under the reduced contract)
The campaign below is the full seven-target, three-alternating-pair certificate the old contract asked for. The 2026-09-27 revision removed it as a #1929 closure gate and moved broad performance assessment to #1938; it is kept here because the scripts and the frozen baseline still exist and a later bounded pass can reuse them.
The pieces below exist and were verified to the point this host allows. No quiet window, free cpu=0 or full campaign is required to close #1929. If a later bounded assessment reuses the campaign script, it pins the criterion settings, the 1T thread environment and the CPU affinity:
- Baseline side. In the baseline worktree (
bench/1929-route-baseline, which holds25da8d431), apply the current harness source — the seven campaign bench files — and confirm it builds:cargo bench --no-run -p tenferro-cpu --bench route_matrix,-p tenferro-runtime --bench session_chain,-p tenferro-gpu --features cuda --bench route_matrix_gpu. This was verified here: all three targets compile against the pinned code. - Pairs. Alternate three times, each pair writing to its own output directory: baseline
bash scripts/run-session-route-performance-gate.sh --mode run --label baseline --output-dir target/pair-N/baselinein the baseline worktree, then candidate with--label candidateand--output-dir target/pair-N/candidatein the refactor worktree. - Compare. For each pass,
python3 scripts/compare-session-route-baseline.py --logs-dir target/pair-N/candidate --label candidateagainst the recordeddocs/testing/session-route-baseline.json. Theoneshot/*andone_shotrows are before-only and must come back asDELETED; a baseline row that is neither deleted-route nor present in the candidate log is a fail-closed error, not a skip. - GPU. The campaign’s
route_matrix_gputarget needs a device; it ran on this host (13 cases), and the protocol requires reporting GPU enqueue and synchronized-completion cost separately. - Record. Write the comparator report and the alternating-pair logs into
docs/testing/and a worklog entry, then remove the harness files from the baseline worktree (git -C .worktrees/issue-1929-bench checkout -- .) so the frozen harness is restored.
The recorded baseline keeps the before-only rows: re-capturing produces session/scope rows only, and the before-reference stays the JSON that was captured while the deleted spellings still existed.
Measurement protocol
This protocol is guidance for the deferred performance work (above), not a #1929 closure gate. Under the 2026-09-27 revision the bounded post-integration assessment in #1938 chooses the decisive metrics instead of running this per-case matrix.
Removing syntax does not by itself save time; #1926 requires measurement separately. The comparison must hold kernel, cache state, thread count, and output contract fixed, and report the session-entry count alongside timing:
- cases: (i) current one-shot per-op entry, (ii)
with_execution_scope,- one shared session across the chain;
- same kernels, same warm cache state, same explicit CPU thread count with the effective thread count recorded (single-worker backend for the overhead baseline, per the umbrella’s 1T requirement);
- report session-entry count, executor-admission time, and body time separately;
- GPU: report enqueue vs synchronized completion distinctly.
Do not attribute an executor-admission saving to the API change. Benchmarks themselves live in tenferro-benchmark (#107) and strided-rs-benchmark-suite (#41); this document only fixes the protocol that this repository’s claims must follow.
Supersession of earlier session documents
session-oriented-concrete-apis.md (#1673) states “The one-shot API stays available and becomes a thin wrapper” and lists “the one-shot API is not broken for aesthetic consistency” as a non-goal. That was correct for its own scope (adding a session-explicit surface). This document supersedes those two clauses for operation-level session-opening APIs: #1926 removes them, because a thin wrapper that opens a session per call is still an implicit entry. What #1673 established and remains unchanged: the shape of the session surface (_in/_read methods, &mut dyn BackendSession, preserved Send bounds), the nested-entry prohibition, and cache parity.
exec-session.md remains the owner for “what a session is”. Its CpuBackend::with_backend_session code sample predates the permit-based run_backend_session_cached implementation and was corrected in the same change that added this document.
Non-goals
- No unrelated public API, backend, dependency, feature flag, or cache beyond the runtime-bound eager session surface required by #1926.
- No universal execution pipeline and no second operation registry.
- No claim that the migration improves performance without the measurement above.
- No frame that the deferred scope hook, the certification campaign or the benchmark repositories’ completion is required to close #1929; the 2026-09-27 revision moved that work to #1938 and left it with its own owners.
- No change to
Runtime::run_compiledregion formation, GPU device ordering, AD semantics, or numerical behavior. - No removal of the session surface itself, and no session handle inside
Tensor/TypedTensor/EagerTensorvalues.
Residual risks
The first two bullets below recorded risks before implementation; the current outcome and remaining capabilities are in the handoff.
- Scope. Slices 2–5 touch the hottest execution paths in three crates. Each slice must be independently revertible; a combined multi-crate rewrite would put numerical and AD behavior at risk with no smaller fallback.
- Eager AD boundary granularity. Whether one boundary per forward/backward group is expressible without changing AD value lifetime is unproven until slice 2 is attempted.
- Test churn. 740 source occurrences of
with_backend_sessionexist, most in tests/benches/examples. Removing one-shot spellings will rewrite many test call sites; the replacement must keep the tests meaningful rather than mechanically wrapping each call in its own boundary. - GPU release-mode enforcement stayed debug-only at the time; resolved by #1938 D6 (typed
Reenteredin every profile).
B2 work list: which operation methods invert, and what must not change
Derived by parsing the trait declarations in crates/tenferro-tensor/src/backend.rs and the impl Trait for Owner blocks for CpuBackend, CpuExecSession, CudaBackend, and WebGpuBackend. The parse expands the method-generating macros (delegate_with_pool_context!, delegate_with_pool!) used by CpuExecSession, so the counts below are not fn-line counts.
Invertible pairs: 31
A method pair inverts only when both halves exist and the read half is the provided one. Pairs: TensorElementwise 13 (add/sub/mul/neg/conj/ div/abs/sign/maximum/minimum/compare/select/clamp), TensorAnalytic 10 (exp/log/sin/cos/tanh/sqrt/rsqrt/pow/ expm1/log1p), TensorStructural 3 (transpose/reshape/ broadcast_in_dim), TensorReduction 4 (reduce_sum/reduce_prod/ reduce_max/reduce_min), TensorDot 1 (dot_general).
Not invertible: 13 required one-shots with no _read sibling
TensorStructural: cast, extract_diagonal, embed_diagonal, tril, triu. TensorIndexing: gather, scatter, slice, dynamic_slice, dynamic_update_slice, pad, concatenate, reverse.
These have a single spelling today. TensorIndexing’s methods are already session-safe — a session implementation does not open an entry — so they are not hidden entries, and adding _read siblings would be a new capability, not a migration. They stay as they are.
Per-backend _read coverage
| Backend object | of the 31 _read methods |
note |
|---|---|---|
CpuBackend |
31 present | none relies on the default |
CpuExecSession |
31 present | many generated by the delegation macros |
CudaBackend |
31 present | none relies on the default |
WebGpuBackend |
0 present | implements only the required one-shots, mostly unsupported!; dot_general/dot_general_with_conj are real |
So requiring _read is free for CPU and CUDA and needs 31 explicit implementations for WebGPU.
The default _read semantics are a tested contract, not an accident
crates/tenferro-tensor/src/tests/backend_default_read_tests.rs (1916 lines) pins the current defaults. Its central test, default_read_methods_delegate_owned_tensors_and_reject_views, asserts two behaviours for a backend that implements only the one-shot methods:
TensorRead::Tensorinputs delegate to the one-shot method (the test assertscallscontains"add"and"dot_general");TensorRead::Viewinputs are rejected with an error whose message contains"borrowed tensor views".
The rejection is uniform, not per family: the defaults for all 31 read halves return the owned tensor through read_tensor(op, input) and raise read_boundary_error(op) for a view. read_tensor is input.as_tensor().ok_or_else(|| read_boundary_error(op)), so the default is “delegate owned, reject views” everywhere. Materialization appears only where a method overrides the default with to_contiguous_read: dot_general_read, the _read_into outputs, and the cached dot paths.
Deleting a default therefore deletes a tested contract, and the view policy (accept-and-materialize versus reject-with-Unsupported) becomes the implementing backend’s explicit responsibility. “Do not change the acceptance surface” means: for every backend and every one of the 31 operations, views must be accepted or rejected exactly as the old default plus the backend’s override did. The affected test backends (DefaultReadBackend, DefaultOnlyBackend, DefaultOnlyExec, DefaultOnlyLinalgBackend, and the per-crate macro_rules! test backends) must migrate with it.
Staging
Step 1 — require _read, keep every caller working. Delete the provided defaults and make the read halves required, family by family. CPU and CUDA need no new code. Every other implementor, including the test backends, reproduces the old default in one line per operation, because the old default body is self.op(read_tensor(name, input)?, ..). That makes the two libraries below the read-half migration mechanical:
- expose
read_tensoras a#[doc(hidden)] pubbridge intenferro-tensor, alongside the bridges this crate already uses for cross-crate internals (with_cpu_exec_session,session_type_id,with_backend_session_cached), so an external implementor writesself.reduce_sum(read_tensor("reduce_sum", input)?, axes)instead of a six-line match; - for WebGPU, reproduce both branches of the old chain:
unsupported!with the same op literal where the owned branch ended in the one-shot’sunsupported!, andread_boundary_errorfor the view branch. Where the operation is real (dot_general), the old default body moves intodot_general_readunchanged.
Measured churn. Making only the four TensorReduction read halves required in a throwaway probe broke nine impl TensorReduction sites in four files (tenferro-ad/src/eager_backend.rs twice, tenferro-cpu/src/tests/cpu_tests/backend_misc.rs twice, tenferro-runtime/tests/integration/session_ops.rs, and tenferro-tensor/src/tests/backend_default_read_tests.rs) while the --workspace --all-targets check ran with default features. That is approximately linear, so the full 31 read halves imply on the order of seventy sites, each needing one line per operation. WebGPU’s 31 implementations are additionally required but are hidden from a default-feature check.
Per-family staging is viable and preferred for revertibility: the contract test file asserts each family separately, so a family slice updates only its own assertions. The ordering constraint is unchanged — a read half must become required before its one-shot sibling is deleted, or the still-provided default would recurse into itself.
Step 2 — delete the one-shot methods per family and migrate callers. Compile errors are the work list. Per family, the deletion also forces the generic helpers that called the one-shot on a &mut B to become session-taking, so each family slice carries the part of B3 it triggers. The pilot is TensorDot (one pair): dot_general has 33 non-test call sites in 19 files, 113 test/bench/example call sites, and 2 doctest call sites, and the provided dot_general_read, dot_general_cached, dot_general_read_into, dot_general_read_into_accum, and dot_general_with_conj[_read] bodies all reference it.
Ordering constraint. Deletion cannot be staged per family before Step 1 lands: while a read half is still provided, deleting its one-shot sibling makes the default recurse into itself.
Local gate limitation for compile_fail fixtures
trybuild compile_fail contracts compare rendered diagnostics against a committed .stderr, and two local conditions break that comparison independently of any source change:
- kache path remapping. The local build wrapper passes
--remap-path-prefix, so diagnostics render absolute/kache/...paths while the committed.stderrfiles carry crate-relative paths. Everycompile_failfixture in the affected crate then reports a mismatch. Running withRUSTC_WRAPPER=""restores the crate-relative form. - rustc version rendering. At least
tenferro-ad’stests/ui/eager_backend_owner_private.rsstill mismatches with the wrapper disabled: the committed.stderrexpects a two-part span (^^^^^^^^^^^^^------------with the label on its own line) and rustc 1.97.1 renders a single span. This reproduces with the working tree reverted, so it is pre-existing, not a consequence of any change here.
Consequences for the B1/B5 fixture sets:
passfixtures are unaffected —trybuilddoes not compare their output, so the session-surface contract intenferro-runtimeruns normally.- The
failfixtures that must prove the one-shot spellings are gone arecompile_fail, so their.stderrfiles are version-sensitive. They must be blessed in the same rustc that CI uses, and a local mismatch is not evidence that the expected error changed. Prefer fixtures whose diagnostic is structurally stable across versions (a missing-method or missing-trait-item error) over ones relying on span geometry. - A local
compile_failfailure therefore has to be checked against the pristine tree before it is attributed to a change. This was done once here:git stash+ rerun reproduced the same mismatch.
Step-1 mechanical aids and their limits
For the wider families the reproduction sites are numerous — 86 read halves across nine impls for the ten TensorAnalytic operations — so a throwaway generator was driven from cargo check E0046 output plus the help: implement the missing item lines. Four limits were hit; knowing them matters before the TensorElementwise family is attempted:
- Method-generating macros hide methods from a
^ fnscan. The implementation inventory must expanddelegate_with_pool_context!,delegate_with_pool!,panic_backend_methods!andunreachable_backend_methods!, or the estimate is wrong. - Some E0046 sites are reported inside a
macro_rules!definition: the GPUdelegate!shim and the per-cratepanic_*!factories. Generating a method body inside a macro definition corrupts the macro. Those sites must be edited by hand — add entries to the delegation macro invocation, or bodies inside the factory macro. - The trait declaration’s receiver must be dropped when deriving parameters, or the generated signature contains
&mut self:. - Locating an impl’s end by the first line equal to the impl’s indentation plus
}is not reliable in files with nested modules; in the tenferro-cpu test module it placed bodies outside the impl. Anchoring on the tail of the family’spanic_*!or delegation block is reliable, and is how the earlier families were patched.
Every generated site uses the delegating form (self.op(read_owned_tensor("<op>", input)?)), so Step 2 must inline the one-shot bodies into these read halves. That second visit is expected and mechanical: the one-shot bodies at these sites are panics, markers, or unsupported errors. WebGPU is the deliberate exception and was written directly in its final shape (unsupported! after evaluating the read input) across all four families it implements, so it never needs a second visit.
Step 1 complete
All 31 read halves of the invertible pairs are now required, verified by scanning the trait declarations: TensorElementwise 13, TensorAnalytic 10, TensorStructural 3, TensorReduction 4, TensorDot 1. The delegation direction is inverted: a session implementation is now the primary implementation, and no read half is derived from its one-shot sibling.
Four read-shaped methods remain provided and are deliberately outside the pair set, because each lacks a required one-shot sibling:
| Method | Provided default | Hidden session entry? |
|---|---|---|
TensorElementwise::rem_read |
delegates to rem |
no — rem is also provided and terminates in Unsupported |
TensorStructural::to_contiguous_read |
materializes compact host tensors, rejects views and device placement | no |
TensorReduction::reduce_sum_squares_read |
Unsupported, with a contract test asserting an explicit override is required |
no |
TensorDot::dot_general_with_conj_read |
materializes, then calls dot_general_with_conj |
no, but both reference the one-shot dot_general |
Step 2 must therefore also redirect the trait-internal defaults that reference a deleted one-shot: dot_general_with_conj (which calls dot_general), dot_general_with_conj_read, the dot_general_read_into* family, SessionCachedDot::dot_general_cached and its cached siblings, and the _read_into elementwise defaults. Those are trait-internal edits, not new implementor sites.
Evidence for Step 1: cargo fmt clean; cargo check --workspace --all-targets clean and warning-free; 5111 workspace tests pass with the single known pre-existing environmental trybuild mismatch in tenferro-ad.
Step 2 sizing (measured)
Call sites of the one-shot spellings, counted by \.(op)\( per family. The counts are an upper bound: a match may be a Tensor-side call on the session extension surface or a _read-adjacent helper, and the real work list is whatever cargo check reports after the methods are deleted.
| Family | non-test lib | tests/benches/examples | doctests | total |
|---|---|---|---|---|
TensorDot (dot_general) |
35 | 133 | 2 | 170 |
TensorStructural (transpose, reshape, broadcast_in_dim) |
78 | 157 | 12 | 247 |
TensorReduction (reduce_sum, reduce_prod, reduce_max, reduce_min) |
87 | 300 | 23 | 410 |
TensorAnalytic (10 operations) |
131 | 484 | 48 | 663 |
TensorElementwise (13 operations) |
337 | 1088 | 120 | 1545 |
Step 2 is therefore per-family in the order above: smallest first, so the migration pattern is proven before the families that dominate the diff. TensorDot is the pilot. Each family slice deletes the one-shot trait items, inlines the one-shot bodies into the read halves that currently delegate to them, redirects the trait-internal defaults listed above, and migrates callers. The trait-internal redirect belongs to the same slice as the deletion, because a default that still calls a deleted method cannot compile.
Step 2 pilot: attempted, measured, reverted
The TensorDot deletion was attempted and reverted. The tree at 5e9c7519a (Step 1 complete) is the last verified state. Recording what the attempt established, because it changes how Step 2 should be run.
Identification is the first hard part, not the edit. Of the 35 non-test \.dot_general\( matches, only five are the deleted trait method: two read halves that had to absorb the removed body, and three session read calls (tenferro-runtime/src/tensor.rs, tenferro-einsum/src/concrete.rs, tenferro-ad/src/eager_exec.rs). Every other match is a different API that must not be touched — EagerTensor::dot_general, TracedTensor::dot_general, capabilities.dot_general(), CpuGeneralContractionProvider::dot_general (&self, provider request), and gemm::dot_general free functions. A raw \.dot_general\( grep is therefore mostly false positives, and the same will hold for add, reshape, and the rest.
The test-side sites are three shapes. Session-receiver calls already inside a with_backend_session closure (rewrite to _read plus TensorRead::from_tensor wrapping); owner-receiver calls (backend.dot_general(..), including a multi-line backend\n .dot_general( form) that need a boundary; and read halves that still delegate to the removed method and must absorb its body.
Six concrete codemod failure modes, all hit:
- Diagnostics that point inside a
macro_rules!definition — the GPUdelegate!shims and the per-cratepanic_backend_methods!/unreachable_backend_methods!factories. Editing there corrupts the macro. - Factory invocation lines (
dot_general(...) -> Ret;entries) that must be deleted individually. - Receivers that are not simple identifiers:
&mut Bgeneric session parameters, multi-linereceiver\n .method(forms,&mut dyn BackendSession. - Argument lists with nested calls, macros and struct literals (
black_box(..),DotGeneralConfig { .. }), which need real paren matching — the indent/brace heuristics failed here repeatedly. E0407stays hidden while a crate’s earlier caller errors exist, so the work list arrives in stages rather than once.- String, raw-string and char literals break naive sanitizers.
Recommendation for Step 2. Migrate per operation rather than per family, so each slice is one trait item and its call sites; migrate the library sites first, then tests crate by crate; and keep the deletion in the same slice as its migrations so the tree compiles after every slice. If a codemod is used again, it should take an explicit file allowlist, refuse any site whose diagnostic points into a macro_rules! definition, and require cargo check --workspace --all-targets to pass after each file rather than once per family.
Step 2 complete: 31 of 31 pairs deleted
Every paired one-shot in the Step-2 inventory is gone, one operation per commit, in the smallest-first order the sizing table implies: the 10 analytic operations, the 13 elementwise ones, the three structural ones, the four reductions, and finally TensorDot::dot_general. dot_general_with_conj remains, because it is one of the 13 required one-shots with no _read sibling, and its body now calls dot_general_read.
The deletion exposed two failures that a compile-and-test pass alone would not have caught, and both are now guarded:
- A real implementation can hide behind the one-shot.
WebGpuExecSession’stranspose_readhad been written asunsupported!while the working device transpose lived in the one-shot, so deletingtransposesilently downgraded a supported operation. The read halves of every deleted operation were therefore re-checked against the removed body, and WebGPU transpose now keeps itsstructural::transposeroute. - A migrated read half can call itself. When a read half delegated to the one-shot and the one-shot was removed first, rewriting the delegate produced
fn op_read(..) { .. self.op_read(..) }in test fixtures. A repo-wide scan for self-recursive*_readbodies is now clean; the fixtures that had one carry the removed behaviour directly (delegate to the wrapped backend, return the fixture result, or keep the same rejection after materializing a view).
The migration tooling that made 31 slices tractable:
- the work list came from
cargo checkdiagnostics, never from a name grep, so the look-alike APIs (EagerTensor::dot_general,TracedTensor::transpose, provider capability queries,gemm::*free functions) were never touched; - call sites were rewritten from the diagnostic’s own span, with a string/comment-aware delimiter matcher over the original text, and only arguments whose
_readparameter is aTensorReadwere wrapped; - the structural half of each slice (trait item, CUDA body absorption into the read half, CPU session delegation, runtime extension, eager dispatcher) was driven by a per-operation table, and every count was asserted before writing;
cargo fmt --all,cargo check --workspace --all-targets, the workspace doctests and the focused suites ran per operation, and the full workspace suite once at the end.
Verification for the completed Step 2: cargo check --workspace --all-targets is clean and warning-free; cargo test --workspace passes 5207 tests with the single pre-existing environmental tenferro-ad trybuild span mismatch (eager_backend_capability_contract); the workspace doctests pass; and no *_read method recurses into itself.
The remaining Phase-B work is unchanged: B3 (drop BackendSession, TensorBackendOps and BackendCachedDot from the TensorBackend supertraits and keep only the session implementations), B4 (extension owner routes), B5 (audit gate, REPOSITORY_RULES.md owner, batched repository gate), then the Phase-C paired benchmark against the recorded baseline.
B3 reconnaissance: what removing the owner supertraits surfaces
Dropping BackendSession, TensorBackendOps and BackendCachedDot from TensorBackend was measured once, then reverted, so the next session starts from a known work list instead of rediscovering it.
Inside tenferro-tensor the change is small and self-contained:
impl<T> SessionCachedDot for T where T: TensorBackendhas to nameTensorDotitself once the blanket access throughTensorBackendOpsis gone;default_backend_sessionand the default bodies ofBackendSessionHost::with_backend_session[_cached]are what make an owner usable as a session. Removing them makeswith_backend_sessiona required method, which in turn forces the test fixtures that writeimpl BackendSessionHost for Fixture {}to open a real session (with_session_entry_guard(|| f(self))) instead of borrowing the owner’s own operation implementations.
Everything else is in tenferro-runtime, which is the crate that still treats the owner as an operation host:
exec/dispatch.rscallsdot_general_read/dot_general_with_conj_readon a&mut B: TensorBackend, so the FFI dispatch functions must take a session (or open one) rather than the owner;HostExecution(TensorBackendOps + TensorDeviceTransfer) and the host dispatch table lose their blanket source, soexecute_host_dispatchand itsB: TensorBackendre-exports inexec.rshave to become session-taking;BackendCachedDot::dot_general_read_cached/..._with_conj_read_cachedneed the cached trait bound back explicitly, or a session route;reclaim_exec_slot_with_backendusesTensorBuffer::reclaim_bufferon the owner, which was previously reachable throughTensorBackendOps.
The public boundary this reaches is TensorBackend-bounded generic API: 103 sites in 21 files, of which the runtime dispatch layer, tenferro-ad’s eager execution, and the ext/* extension modules dominate. B3 therefore splits into (i) the supertrait trim plus the tenferro-tensor fixture changes above, (ii) the runtime dispatch/execution layer moved to session-taking, and (iii) the owner-side operation implementations and BackendCachedDot impls deleted once nothing calls them. (i) alone does not compile the workspace, so it must ship in the same commit as (ii); (iii) is what makes the removal observable.
B3 depends on B4: the session-region path still falls back to the owner
A second B3 attempt converted the runtime dispatch layer and then stopped, because it ran into the extension route. The finding is an ordering constraint that the slice list does not show.
What converts cleanly: exec/dispatch.rs’s host table and the execute_*_host functions become &mut dyn BackendSession (the session path already calls them), the four owner-based evaluators in segment.rs/exec.rs open one session and use execute_segment_in_session / execute_value_segment_in_session / execute_ffi_instruction_exec, and the owner wrappers (execute_host_instruction, execute_ffi_instruction[_cached], reclaim_exec_slot_with_backend, reclaim_last_use_inputs_backend) disappear.
What does not convert yet is the extension operation route. Two facts pin it:
execute_prepared_extension_instructionbuildsErasedExecutionContext, which requires aSized + 'statictype, so it needs the owner (B: TensorBackend + 'static), not&mut dyn BackendSession. The session equivalent (execute_prepared_extension_instruction_in_session) exists and is used byexecute_ffi_instruction_exec, but the owner route is still reachable.segment_is_session_compatibleexcludes FFI ops that are not session compatible, and the session-region evaluator falls back toexecute_ffi_instruction_cached(backend, ..)for them. With the owner-based FFI dispatch deleted, that fallback has no callee.
So B3’s owner-side operation impls cannot be deleted while the extension fallback still runs operations on the owner, and the FFI dispatch table cannot become session-only while that fallback exists. B4 (extension routes primarily session-based, owner route limited to the runtime-formed region fallback) is therefore a prerequisite for B3(iii), not a follow-up.
Why the owner is still needed: capability identity, not execution. A second look at the extension routes narrows the reason. execute_linalg_extension_reads_owner (and the fft equivalent) already opens a session internally (backend.with_backend_session(..)) and runs the same session executor, so the execution is session-based today. What needs the owner is ErasedExecutionContext<'_, B> with B: TensorBackend + 'static, which the extension runtime uses to identify the concrete backend for capability dispatch. The session-based equivalent already exists (BackendSession::session_type_id, since #1938 D7 replaced by the opaque BackendSession::native_session token; with_cpu_exec_session, with_cuda_exec_session), and execute_prepared_extension_instruction_in_session uses it, but the owner path is still reachable.
It is reachable rather than dead because session support is per operation: linalg_session_supported returns true for the whole family on CPU, and false on CUDA for FullPivLu, FullPivLuSolve and general eig (extension.rs:1060, issue #1665). For those operations the scheduler keeps the owner path, and that path’s ErasedExecutionContext cannot be built from &mut dyn BackendSession.
So B4’s deliverable is precise: move extension capability dispatch from the owner’s type identity to the session’s, and give the remaining per-op exceptions a session route (or an explicit documented refusal), after which no extension operation needs the owner and the runtime’s owner extension path can be deleted. That is the prerequisite for B3.
The workable order is:
- B4: make the extension runtime registers session primary, move capability dispatch to
session_type_id, and keep the owner route only where the runtime forms a region explicitly. - B3(i)+(ii): trim the supertraits, delete
default_backend_session, thread a session through the dispatch layer — after B4 the only remaining owner bound is gone and the FFI table can be session-only. - B3(iii): delete the owner-side operation impls (
TensorElementwiseand the other families forCpuBackend,CudaBackend,WebGpuBackend), theBackendCachedDotimpls and the GPUdelegate!shims, then shrinkscripts/session-entry-allowlist.json.
B4 landed: the owner entry now forms the session
8468debd5 made execute_reads optional in define_extension_runtime!. A family that registers the session route (execute_in_session + session_supported) now gets an owner entry that calls with_backend_session itself and hands the extension nothing but &mut dyn BackendSession; supplying both routes, or neither, is a compile error. The einsum, linalg and fft owner routes (execute_*_reads_owner, execute_einsum_extension_reads) are deleted, the linalg session-context variant that only the owner route used is deleted, and the three tests that drove the owner routes now drive the session route. No operation runs on the owner through the extension path any more.
That removes the reason B3 could not start. What remains for B3(ii) is a mechanical, now-safe conversion of the runtime dispatch layer, with these exact call sites:
| Site | Today | After |
|---|---|---|
segment.rs:287, :390, :495, :579 |
execute_ffi_instruction_cached(backend, ..) |
session run, or the extension fallback |
segment.rs:296, :301, :400, :411, :504, :589 |
reclaim_last_use_inputs_backend(slots, inst, backend) |
reclaim_last_use_inputs_exec inside the run |
segment.rs:300, :409 |
execute_host_instruction(backend, ..) |
execute_host_instruction_exec inside the run |
exec.rs:649–:663, :694–:710 |
the same pair in the unsegmented evaluators | the same run structure |
exec.rs:873 |
execute_ffi_instruction(backend, ..) fallback |
deleted with the owner wrapper |
runtime/execution.rs:1270–:1287 |
owner host/ffi/reclaim | the run structure |
Two things to keep in mind while doing it:
- the fallback instruction is always an extension op:
is_session_compatible_instructionreturnstruefor non-FFI ops and forDotGeneral/DotGeneralWithConj(is_exec_session_ffi_op), so afalseresult can only come from an extension whosesupports_session()is false. The fallback therefore callsexecute_extension_instructionand can report an internal error for anything else, and it needs noBackendCachedDotbound. - the unsegmented evaluators execute instruction-by-instruction, so the conversion must group each maximal run of session-compatible instructions into one session (
with_backend_session_cached) and run the fallback outside it. Opening a session per instruction would be a needless change in session-entry count for the very path the Phase-C baseline measures.
B3(iii) work list: deleting the owner implementations
bed79ade0 removed the supertraits, so nothing requires the owner-side operation implementations any more; deleting them is the last step, and the measured shape of that step is recorded here because it is a call-site migration, not a deletion.
Deleting the eight CPU owner impls (BackendSession, TensorElementwise, TensorAnalytic, TensorStructural, TensorReduction, TensorDot, BackendCachedDot, TensorIndexing) leaves 261 errors in the CPU crate alone, across about a dozen test/bench files. They split into four shapes, and only the first two are mechanical:
| Shape | Count (CPU crate) | Rewrite |
|---|---|---|
paired one-shot with a _read sibling (slice, pad, gather, reverse, concatenate, cast, copy_read_into, to_contiguous_read, reduce_*_read) |
~120 | recv.op(args) → recv.with_backend_session(\|__s\| __s.op_read(args)) |
session method with the same name (dot_general_read_into_accum, elementwise_read_into, grouped_gemm_cached) |
~30 | recv.op(args) → recv.with_backend_session(\|__s\| __s.op(args)) |
cached dot family (dot_general[_with_conj]_cached, _read_cached, grouped_gemm_cached) |
~20 | the owner form takes (&mut cache, cache_slot, ..) and the session form takes (cache_slot, ..), so the cache becomes the receiver of with_backend_session_cached |
receivers that are not a plain identifier (CpuBackend::new().op(..), trait-qualified BackendCachedDot::op(&mut backend, ..)) and &mut dyn sites |
~15 | hand edits |
Two codemods were used and are worth reusing, outside the repository:
- a diagnostic-driven wrapper that reacts to
E0599 no method namedop`` and rewrites only the reported call span, one edit per file per round so later line numbers stay valid; it cleared ~80 of the first two shapes in a few rounds and left a short manual list; - a cached-shape rewriter for the third shape. It must match “first argument is the cache” only when the call is not already cache-bound, or it re-wraps its own output — the version that ran here oscillated and was reverted.
Remaining work, in order: finish the CPU crate (about 20 hand sites), repeat for the CUDA and WebGPU owner impls (which also removes the delegate! shim and moves their bodies to the module functions the sessions already call), migrate the downstream call sites in tenferro-ad, tenferro-linalg, tenferro-einsum and the runtime tests/benches, delete install_with_pool_context* once the CPU owner impls are gone, and re-bless scripts/session-entry-allowlist.json (the CPU entries it lists are exactly the ones that disappear).
B1 fail fixtures: the deleted spellings are pinned by compile_fail doctests
The B1 contract file covers the surviving surface with trybuild pass fixtures. The fail side is pinned with rustdoc compile_fail examples instead of trybuild .stderr files, because compile_fail only requires compilation to fail and therefore does not depend on the compiler’s span rendering or on any local build wrapper rewriting paths:
tenferro_tensor::BackendSessiondocuments thatexec.add(a, b)insidewith_backend_sessionno longer compiles;tenferro_cpu’s crate docs pin the deleted owner spellings for one operation per family (add,mul,exp,reduce_sum,transpose,dot_general).
The pass side is a trybuild contract that originally skipped itself when the NEXTEST environment variable was set. That skip was wrong for this repository: CI’s workspace profile is cargo nextest run --workspace plus cargo test --doc --workspace, so a nextest-only skip made the surviving-surface fixtures a CI target that never ran. The skip is removed after verifying the contract under nextest here (1 passed, about 60 s cold, about 3 s warm), so the fixtures and the compile_fail examples both execute in CI.
The fixtures whose targets B3 removes have landed with that slice: tenferro_tensor::BackendSessionHost pins the deleted default_backend_session factory, and tenferro_cpu’s crate docs pin the owner-level BackendCachedDot bound in addition to the one-shot spellings. The positive counterpart is the pass fixture for the cache-aware session route (with_backend_session_cached plus a session _cached contraction), which still compiles, so the two sides together pin that the cached route moved from the owner to the session rather than disappearing.