Session-Oriented Concrete APIs (issue #1673)

Status: design (pre-implementation review target).

Overview

Add a session-explicit execution surface for concrete Tensor/TypedTensor operations (including extension-crate operations) so a caller can reuse one borrowed BackendSession across a sequence of operations instead of paying a session entry per operation. The tensor value types remain context-free and are never owned by or lifetime-bound to a session.

The one-shot API (a.add(&b, &mut backend)) stays available and becomes a thin wrapper around the session implementation where practical.

Superseded in part by #1926. This document’s scope is adding the session-explicit surface, and it deliberately keeps the one-shot spelling. #1926 goes further and removes session-opening operation-level APIs, because a thin wrapper that opens a session per call is still an implicit entry. The session surface defined here, the nested-entry prohibition, and cache parity remain in force; only “the one-shot API stays available” changes.

Motivation (measured)

Session entry for managed CPU execution is ~3 µs/op (release, pinned, 1 worker; issue #1673 data comment). For trivial unary/binary ops (~3 µs total) it is ~100% of the cost; for a 10-op trivial chain the current API pays 10 session entries (~30 µs) where one entry (~3 µs + 10 op bodies) suffices. For compute-heavy ops (solve, ~28 µs after #1675) the session share is ~10%, so the surface targets short chains of cheap ops; it is not a GPU small-op-launch fix.

Gate (from the issue): with T_one_shot = 10-op trivial chain via the current one-shot API (10 session entries) and T_one_session = the same chain via one session entry, require T_one_shot / T_one_session >= 2 (predicted ~6). The baseline is explicit: current one-shot API, 10 session entries.

Verified execution model (corrects the 2026-08-14 design review)

The design review assumed with_backend_session runs the closure synchronously on the calling thread and therefore proposed dropping the Send bounds. That is false for CPU:

  • CpuBackend::with_backend_session → run_backend_session_cached → enter_managed_session/enter → install_scoped (crates/tenferro-cpu/src/provider.rs) → CpuContext::install (docs/design/exec-session.md, “Backend Mapping / CPU”) → rayon::ThreadPool::install — a real thread handoff to a Rayon worker, even for 1-thread pools (verified empirically: caller thread id differs from the closure thread id).
  • The default (non-overriding) backend path (default_backend_session, crates/tenferro-tensor/src/backend.rs) runs the closure on the calling thread.

Decision: keep the Send bounds (R: Send, f: ... + Send) on with_backend_session / with_backend_session_cached. They are a soundness requirement for the CPU managed session (a non-Send closure would execute on another thread). This is an inherent property of the current CPU managed session, not a migration accident: downstream users of the session-explicit surface must capture only Send state, which the one-shot API already requires of every closure it takes.

Dropping Send would require redesigning CPU session execution to run closures on the calling thread with per-op pool entry — a performance/ semantics change to the hottest segmented-execution path. That redesign is out of scope and tracked separately if ever desired; it must NOT be assumed as a prerequisite.

Evaluation-wide scope

A2 (one execution scope around a whole evaluation) remains deferred; see exec-session.md.

Nested-entry prohibition: mechanism, not convention

Operations receiving a BackendSession must never call with_backend_session internally. Enforcement by backend family:

Every rejection below is a typed SessionEntryError returned before the closure runs (#1938 D6); none of them panics.

  • CPU: fresh_execution_owner() (crates/tenferro-cpu/src/arbiter.rs) returns None when EXECUTION_OWNER is set on the thread (or an owned Rayon scope has an active owner), and admission reports SessionEntryError::Reentered. Covers the one-shot-inside-session case, including the pinned one-worker managed pool handoff.
  • Portable guard: with_session_entry_guard sets a thread-local in-session flag (Drop-guard restored, so panics in f still restore) and returns Reentered if it was already set, in every build profile.
  • CUDA / WebGPU: both overrides wrap their closure in the portable guard.

Generic spelling

Already settled by the current signatures: entries hand out &mut dyn BackendSession (no S: BackendSession + ?Sized generic on the trait methods). The _in-surface methods take session: &mut dyn BackendSession directly.

Cache access parity

with_backend_session_cached is the runtime-cache-aware entry (used by the scheduler/eager paths); with_backend_session is the canonical user entry and the one-shot concrete helpers delegate to it. The session-explicit surface uses the plain entry; extension plans (EinsumPlan etc.) own their long-lived caches. The canonical user entry must not silently become the slow entry: with_backend_session may forward to the cached path when a cache is available, but the public contract is the plain entry.

Surface design (prototype plan)

Layering (same vocabulary, no duplicate validation):

one-shot API: &mut TensorBackend ──enter one session──▶ session implementation
session API:  &mut dyn BackendSession ─────────────────▶ session implementation
  1. TensorSessionOpsExt / TypedTensorSessionOpsExt with *_in methods (naming: _in preferred; exact names decided in the prototype). Implemented by calling the session’s own op traits directly; broadcasting, dtype promotion, validation, and typed errors identical to the one-shot path (shared helpers — no second vocabulary).
  2. One-shot methods delegate to the _in implementations where practical.
  3. Concrete-only extensions (einsum plans first: EinsumPlan::prepare + execute_in_session) compose standard session ops; no ExtensionOp/SemanticProgram/runtime registration required at that tier.
  4. Runtime-integrated extensions keep the semantic ExtensionOp + AD rules; runtime preparation produces an executor that runs on the borrowed BackendSession like the concrete path.

Non-goals

  • No session/runtime handle inside Tensor/TypedTensor.
  • No change to the EagerTensor ownership/AD model.
  • No unbounded or globally owned sessions; closure-scoped borrowing stays the default.
  • No second operation vocabulary or duplicated validation.
  • GPU small-op launch overhead is a separate concern (fusion/static execution), not addressed by session reuse.
  • The one-shot API is not broken for aesthetic consistency. (#1926 supersedes this specific non-goal for session-opening operation-level APIs; the rest stand.)

Acceptance criteria

  • One session entry covers a mixed sequence of standard and extension concrete operations.
  • Standard broadcasting, dtype promotion, validation, placement, and typed errors identical to the one-shot path.
  • One-shot methods delegate to the same implementation where practical.
  • Tensor/TypedTensor remain context-free.
  • Concrete-only extension authors need no runtime/graph/AD machinery.
  • Nested-entry prohibition enforced per backend family (originally a CPU release panic plus a debug assert; since #1938 D6 a typed SessionEntryError::Reentered before the closure runs on CPU, CUDA and WebGPU).
  • Send bounds preserved and documented as a soundness requirement.
  • The 10-op trivial-chain gate passes (≥2x, predicted ~6x).
  • No material regression for large operations.

Prototype specification (PR B: the _in surface)

Scope decision for the first implementation PR. The design-review gate rules apply: this section is the pre-implementation design document; the implementation must not start until it has a reviewer-gpt verdict.

Measured state of the one-shot path before #1926

A binary op already pays more than one session entry today (crates/tenferro-runtime/src/tensor.rs):

  • add → broadcast_binary → broadcast_to per operand (each broadcast_to can enter 0-2 sessions: reshape + broadcast_in_dim, or zero when shapes already match via duplicate) → then with_backend_session(|exec| exec.add_read(...)) for the op itself. Worst case: 5 sessions per binary op (both operands need reshape+broadcast = 2 each + 1 final add, e.g. [1,2] + [2,1] → [2,2]); the common one-sided reshape+broadcast case is 3; equal-shape case is 1.
  • unary_fn (exp etc.) and reduce_sum: 1 session each.

This is the before-side of the measurement; the owner-level one-shot spelling it describes was deleted by #1926, so the numbers are now the reference that docs/testing/session-route-baseline.json compares against.

So the session-explicit surface must own the broadcast step too, or a binary op still pays 2 entries inside one session.

Surface (prototype op set)

crates/tenferro-runtime/src/lib.rs, next to TensorOpsExt:

pub trait TensorSessionOpsExt {
    fn add_in(&self, rhs: &Tensor, session: &mut dyn BackendSession) -> Result<Tensor>;
    fn mul_in(&self, rhs: &Tensor, session: &mut dyn BackendSession) -> Result<Tensor>;
    fn exp_in(&self, session: &mut dyn BackendSession) -> Result<Tensor>;
    fn reduce_sum_in(&self, axes: &[usize], session: &mut dyn BackendSession) -> Result<Tensor>;
}

and TypedTensorSessionOpsExt<T: TensorScalar> with the same four methods returning TypedTensor<T>. Exact op set is deliberately small: add_in/mul_in (binary + broadcast, the multi-session flagship), exp_in (unary), reduce_sum_in (reduction).

Naming is transitional: the _in suffix exists only because the session-explicit form coexists with the one-shot TensorOpsExt during the migration — same receiver type (Tensor) + same method name on two in-scope traits is an ambiguous-method-resolution error in Rust (and a prelude glob would break every call site). The final canonical names drop the suffix: once the one-shot API is migrated away (a release-boundary breaking change), the session-explicit methods become the plain add/exp/reduce_sum. The _in names must not be treated as the permanent public spelling.

Validation and broadcast sharing

  • Broadcast plan computation (broadcast_shapes, broadcast_input_plan) is pure and shared. The error mapping (broadcast_error) is not shared: the dynamic surface uses broadcast_error_to_validation (crates/tenferro-runtime/src/tensor.rs) and the typed surface has its own manual mapping (typed_tensor.rs). Each _in surface reuses its existing local mapping; no third mapping is added.
  • Dynamic Tensor helper (owned): session-level broadcast_to_in(input, target_shape, session) using session.reshape + session.broadcast_in_dim with the same plan logic as the existing broadcast_to, preserving the duplicate() path for equal shapes and broadcast sources (owned copy semantics unchanged). add_in/mul_in = broadcast_to_in both operands + session.add/ session.mul, all inside the one borrowed session.
  • Typed TypedTensor<T> helper (borrowed, read-based): the typed one-shot path keeps inputs borrowed and dispatches to the *_read methods (typed_tensor.rs); the typed _in surface mirrors that with reshape_read/broadcast_in_dim_read/add_read/mul_read on the session and into_typed_result for the output (dtype fixed by T, no promotion logic). Do NOT funnel the typed path through the owned dynamic helper — that would introduce copies or change allocation semantics.
  • No second validation vocabulary; no new error kinds. One-shot/session parity tests cover values and structured errors for both surfaces.

One-shot delegation

One-shot add/mul/exp/reduce_sum become backend.with_backend_session(|s| ..._in(...)) where practical. This drops the worst-case binary-op session count from 5 to 1 — a side benefit that must be measured as a no-regression check, not assumed. dot_general/matmul are excluded from this PR (cache-ownership decision pending) and keep their current path.

Explicitly out of scope for PR B

  • dot_general/matmul (deferred until cache ownership/parity with with_backend_session_cached / SessionCachedDot is decided; the current one-shot matmul already uses the plain entry, so this is purely a cache-ownership decision).
  • convert/cast and the rest of the op vocabulary (follow-up PRs).
  • EinsumPlan::execute_in_session and all extension-crate tiers (PR C).
  • GPU session surfaces; CUDA/WebGPU nested-entry enforcement.
  • Removing Send bounds (soundness requirement, see above).

Public API and tests (explicit before merge)

Both public traits (TensorSessionOpsExt, TypedTensorSessionOpsExt) get runnable doc examples (AGENTS.md requirement — compile and run, no ignore). Integration tests cover, per surface: equal-shape op, real broadcast, invalid broadcast (structured error parity with the one-shot path), dtype/error parity, typed output dtype validation (into_typed_result), and a test that an _in chain executes inside exactly one session entry.

Performance gate (measured, not assumed)

New criterion bench in crates/tenferro-runtime/benches/ (criterion is already a workspace dependency), with an explicit [[bench]] target and harness = false in Cargo.toml (mirroring elementwise_fusion.rs):

  • Exact chain (10 ops): three repetitions of add → exp → mul (9 ops) followed by a final reduce_sum([0]) (10th). reduce_sum is final so no intermediate op changes shape.
  • No-broadcast arm: 1×8 f64 constant-filled operands (broadcast is the duplicate path).
  • Broadcast arm: 1×1 f64 operands against 1×8 (real reshape+broadcast).
  • One-shot arm: the chain via TensorOpsExt (each op enters its own session — 10 entries). Session arm: the same chain via TensorSessionOpsExt inside one with_backend_session (1 entry).
  • Validate one result outside the timed region (finite, correct shape); black_box inputs and outputs inside the iter.
  • Protocol: record base-commit (pre-delegation) and post-change medians; gate T_one_shot / T_one_session >= 2 (predicted ~6); also a representative large-op one-shot before/after pair for the no-regression criterion. Report pinned (idle core), matching the session floor methodology; the gate is evaluated on the pinned numbers.

Prototype specification (PR C: ConcreteEinsumPlan::execute_in_session)

The first extension-crate example of the session-explicit concrete surface (design tier 1: concrete-only extensions composed from standard primitives — no ExtensionOp, SemanticProgram, runtime registration, or AD rules).

Key finding: the one-shot einsum already enters a session

ConcreteEinsumPlan::execute etc. (crates/tenferro-einsum/src/concrete.rs, public plan type is ConcreteEinsumPlan) already do backend.with_backend_session(|exec| eager_einsum_exec(exec, inputs, &self.tree)); the eager core functions (crates/tenferro-einsum/src/eager.rs::eager_einsum_exec*) all take &mut dyn BackendSession already. execute_*_in_session is a thin addition: plan validation (pure — validate_input_metadata/input_specs need no backend) + the same core call on the caller’s borrowed session — no new session entry. One-shot methods delegate to the _in_session variants.

Surface (complete signatures)

ConcreteEinsumPlan gains a _in_session mirror of every one-shot execute method (all mechanical: validate + call the eager core on the borrowed session). Signatures mirror the one-shot forms with backend: &mut B → session: &mut dyn BackendSession:

pub fn execute_in_session<'a, I>(
    &self, inputs: I, session: &mut dyn BackendSession,
) -> Result<Tensor>
where I: AsRef<[&'a Tensor]>;

pub fn execute_typed_in_session<'a, T: TensorScalar, I>(
    &self, inputs: I, session: &mut dyn BackendSession,
) -> Result<TypedTensor<T>>
where I: AsRef<[&'a TypedTensor<T>]>;

pub fn execute_read_in_session<'a, I>(
    &self, inputs: I, session: &mut dyn BackendSession,
) -> Result<Tensor>
where I: AsRef<[TensorRead<'a>]>;

pub fn execute_into_in_session<'a, I>(
    &self, inputs: I, session: &mut dyn BackendSession, out: TensorWrite<'_>,
) -> Result<()>
where I: AsRef<[&'a Tensor]>;

pub fn execute_typed_into_in_session<'a, 'out, T: TensorScalar, I, O>(
    &self, inputs: I, session: &mut dyn BackendSession, out: O,
) -> Result<()>
where I: AsRef<[&'a TypedTensor<T>]>, O: Into<TypedTensorWrite<'out, T>>;

pub fn execute_read_into_in_session<'a, I>(
    &self, inputs: I, session: &mut dyn BackendSession, out: TensorWrite<'_>,
) -> Result<()>
where I: AsRef<[TensorRead<'a>]>;

pub fn execute_read_into_accum_in_session<'a, I>(
    &self, inputs: I, session: &mut dyn BackendSession,
    accumulation: DotGeneralAccumulation, out: TensorWrite<'_>,
) -> Result<()>
where I: AsRef<[TensorRead<'a>]>;

One-shot execute* become backend.with_backend_session(|s| self.execute*_in_session(inputs, s)).

Validation ordering change (accepted): the one-shot path currently validates before entering the session, so invalid calls are entry-free. After delegation, validation happens inside the session entry — an invalid call pays one entry (~3 µs) before returning the identical error. Accepted (error paths are exceptional); covered by an error-parity test asserting the same structured error both ways.

Naming: _in_session matches the existing prepared-operation vocabulary (crates/tenferro-runtime/src/runtime/capability.rs: prepared_operation.execute_in_session); the runtime concrete noun ops use _in (PR B). Both are transitional; final canonical names drop the suffix.

Mixed chain (acceptance item)

backend.with_backend_session(|session| -> tenferro_einsum::Result<_> {
    let x = plan.execute_in_session(&[&a, &b], session)?;   // einsum Result
    let x = x.exp_in(session)?;                              // converts via From
    Ok(x.reduce_sum_in(&[0], session)?)
})

One entry covers standard ops + a prepared extension plan.

Bench (crates/tenferro-einsum/benches/, [[bench]] harness=false)

Exact workload: a prepared plan for "ij,jk->ik" on 8×8 f64 column-major inputs, executed 10 times per iter (10 calls). Backend: one-worker CpuBackend::with_threads(1) for all arms. Correctness: validate one result outside the timed region (shape [8,8], finite). black_box inputs and outputs inside the iter.

  • Arm A (one-shot): 10 × plan.execute([a, b], backend) — 10 session entries.
  • Arm B (in-session): the same 10 calls inside one with_backend_session(|s| ... execute_in_session(...)) — 1 entry.
  • Arm C (mixed): execute_in_session + exp_in + reduce_sum_in ×10 inside one session vs the one-shot equivalents.
  • Reported: pinned (idle core) medians for A/B/C; the A/B ratio and the mixed-chain ratio; a baseline/post-delegation pair for the one-shot arms. The PR-B trivial-chain ≥2 gate is a separate acceptance; for einsum the recorded evidence is the measured ratios and the delegation baseline, not a predicted number.

Acceptance (explicit before merge)

  • Runnable doctests (no ignore/no_run) for all seven _in_session methods.
  • Parity tests: _in_session results == one-shot results (values) for execute/execute_read/execute_typed; structured-error parity for invalid inputs (both orderings); typed dtype conversion; execute_into/accum output validation; and a session-counting test proving _in_session adds no nested entry (one entry for a mixed einsum + _in chain).

Out of scope for PR C

  • Native-kernel extension capabilities (FFT, decompositions, sparse — design tier 2) and runtime-integrated extension execution on the borrowed session (tier 3).
  • Top-level einsum()-trait _in_session variants (the prepared plan is the recommended repeated-execution API; trait variants are follow-up if the plan-level surface proves out).
  • Any change to the semantic/graph/AD einsum paths.

Migration specification (issue #1680, Phase 1: complete the _in surface)

Extends the PR-B prototype to the FULL concrete op vocabulary, keeping the transitional _in/_in_session names (final plain-name rename + one-shot removal is Phase 2, a release-boundary breaking change). Additive and non-breaking.

Op inventory (PR-B already covers add/mul/exp/reduce_sum)

TensorOpsExt (dynamic Tensor) — 26 new ops: - convert, cast (dtype) - sub, div, rem, pow, maximum, minimum (binary + broadcast) - neg, abs, sign, conj, log, expm1, log1p, sin, cos, tanh, sqrt, rsqrt (unary) - compare (binary + broadcast, CompareDir) - where_select, clamp (ternary + broadcast) - matmul (dot_general) - reshape, transpose (structural)

TypedTensorOpsExt<T> (typed) — 24 new ops: the dynamic set minus convert, cast, and where_select (typed where_select lives on the separate mask trait), plus broadcast_in_dim.

Per-family implementation (mirror the PR-B patterns exactly)

  • Dynamic (owned): broadcast_to_in + broadcast_binary_in exist from PR B; add broadcast_ternary_in for where_select/clamp. Each _in = broadcast via the _in helpers + the session op, all on the borrowed session; error mapping reuses broadcast_error_to_validation.
  • Typed ternary: add broadcast_ternary_in_read built from broadcast_shapes + three broadcast_to_in_read calls (the existing typed broadcast_ternary_read re-enters a session via broadcast_to_read and must NOT be reused); typed clamp_in uses it before session.clamp_read.
  • Unary: session.<op> directly (neg_in → session.neg, log_in → session.log, etc.).
  • Structural: reshape_in → session.reshape, transpose_in → session.transpose.
  • Dtype: convert_in → session.convert, cast_in → session.cast (the checked lattice / explicit projection live in the session ops).
  • matmul: matmul_in → session.dot_general with the standard matmul config. Cache-ownership decision (resolves the PR-B deferral): the concrete surface uses the plain with_backend_session/borrowed-session entry; with_backend_session_cached stays internal to the runtime scheduler/eager execution paths. No SessionCachedDot in the concrete surface.
  • Typed (borrowed, read-based): mirror PR-B’s typed path — reshape_read/broadcast_in_dim_read/<op>_read + into_typed_result, reusing the existing ReadInput and the typed manual broadcast-error mapping.
  • One-shot delegation: every one-shot method becomes backend.with_backend_session(|s| self.<op>_in(...)). Same accepted validation-ordering change as PR B/C (invalid calls pay one entry before the identical error; error-parity tested).

Tests

  • Per-family parity: _in == one-shot values (equal-shape + real broadcast arms), structured-error parity, typed dtype conversion, session-counting no-nested-entry (one entry for a chain using multiple new _in ops).
  • Runnable doctests (no ignore/no_run) on all new public methods.
  • One-shot delegation no-regression: the PR-B session_chain bench still passes and the one-shot arms do not regress (re-measure; the delegation can only reduce session entries).

Bench

Extend crates/tenferro-runtime/benches/session_chain.rs with a second chain using the NEW ops — one-shot vs single-session arms, pinned reporting. Keep the PR-B chain (add/exp/mul/reduce_sum) as the gate bench. Exact chain (10 ops, shapes valid throughout): [2,2] f64 through sub→log→pow→maximum→neg (elementwise, operands chosen so values stay positive through log), reshape to [4,1], transpose to [1,4], scalar-broadcast clamp, matmul by a [4,1] rhs to [1,1], then cast F64→F32. Validate a known result outside the timed region; black_box inputs/outputs inside.

Out of scope for Phase 1

  • One-shot trait removal and the _in→plain rename (Phase 2).
  • Top-level einsum-trait _in_session variants (plan-level surface suffices; revisit only if Phase-1 callers need them).
  • GPU (CUDA/WebGPU) nested-entry enforcement (tracked gap from PR A). ## Migration specification (issue #1680, Phase 2: single canonical core API)

The release-boundary breaking change for the tenferro-runtime core traits and the einsum plan surface. Removes the one-shot core API, renames the transitional _in/_in_session methods to the final plain names, and migrates every workspace call site of those APIs. End state for THIS phase: one session-explicit concrete API for the core op vocabulary and ConcreteEinsumPlan.

Scope boundary (explicit): other backend-taking concrete surfaces are NOT migrated in this phase and are enumerated as follow-ups — the einsum top-level einsum*-trait methods (crates/tenferro-einsum/src/concrete.rs traits), TensorFftExt/TensorReadFftExt (tenferro-fft), the linalg concrete traits (tenferro-linalg/src/tensor_ext.rs), and Tensor::index_select/Tensor::stack (tenferro-tensor shape_packing). This phase covers the runtime core + plan-level einsum; the others migrate in follow-up PRs toward the same end state.

Removals

  • TensorOpsExt (crates/tenferro-runtime/src/lib.rs) and its impl.
  • TypedTensorOpsExt<T> and its impl.
  • TypedTensorMaskOpsExt (typed where_select, implemented only for TypedTensor<bool> with a generic branch scalar) — replaced by a session-only mask trait TypedTensorMaskSessionOpsExt, implemented for TypedTensor<bool>, with where_select(&self, on_true: &TypedTensor<U>, on_false: &TypedTensor<U>, session) using broadcast_ternary_in_read + session.select_read + into_typed_result (preserves the compile-time bool-condition contract; NOT folded into the generic TypedTensorSessionOpsExt<T>).
  • ConcreteEinsumPlan’s seven backend-taking execute* methods (execute, execute_typed, execute_read, execute_into, execute_typed_into, execute_read_into, execute_read_into_accum).
  • The prelude re-exports only the session-explicit traits.

Renames (final plain names)

  • TensorSessionOpsExt: every *_in → the plain op name (add_in→add, convert_in→convert, matmul_in→matmul, …).
  • TypedTensorSessionOpsExt<T>: same rename.
  • ConcreteEinsumPlan: each execute*_in_session becomes the sole unsuffixed execute* (the backend-taking forms are removed, so no collision).
  • No receiver collisions: tensor-extension methods take &Tensor, backend ops take &mut self.

Call-site migration (~58 sites + tests/examples/docs/skills)

Mechanical rule, behavior-preserving: 1. Single-op call x.add(&y, &mut backend) → backend.with_backend_session(|s| x.add(&y, s))? (since #1938 D6 the entry is itself fallible, so the current spelling is ??); the closure’s result type is the op’s Result<Tensor> (annotate when the closure contains ?-chains so error types are unambiguous). 2. Consecutive session-capable ops on the same backend in one function are grouped in ONE with_backend_session; grouping stops at helpers or extension calls that still need &mut backend (avoid nested entry / borrow conflicts). 3. Reusable helpers that take &mut TensorBackend migrate to take &mut dyn BackendSession where they are pure op sequences. 4. Sites already inside a session call the ops directly. Diagnostics, README, docs/spec, tutorials, and shipped skills that reference the one-shot API migrate in the same PR. The Phase-1 one-shot-vs-_in parity tests become direct session tests against the independent expected values they already assert; the einsum mixed-session test gains independent expected values (drops the one-shot comparison arm).

Verification and gates (exact)

  • Workspace builds; full suites green (runtime, einsum, ad, cpu, dependents).
  • Scoped grep evidence (exclude intentionally retained internal spellings): no \.(add|sub|...)_in\( or \.execute\w*_in_session\( calls on the core/plan surfaces; PreparedOperation::execute_in_session (runtime capability.rs), EagerTensor::from_tensor_in (tenferro-ad), and the einsum top-level/FFT/linalg/shape-packing surfaces are explicit exclusions. Command: grep for _in(/_in_session( in tenferro-runtime, tenferro-einsum (plan scope), and the migrated call sites, review each hit against the exclusion list.
  • Benches: the one-shot arms are removed with the API; re-run the remaining single-session arms (PR-B gate chain, Phase-1 chain, einsum chain) and compare against the recorded Phase-1 medians (no regression; the one-shot numbers are historical). Report pinned.
  • Public API change is the intended release boundary; documented in the PR body/changelog.

Out of scope for Phase 2 (follow-up issues)

  • Einsum top-level traits, FFT, linalg, shape-packing one-shot surfaces.
  • GPU (CUDA/WebGPU) nested-entry enforcement.
  • Behavioral changes to any op (rename/removal only).

Migration specification (issue #1680, Phase 3: extension surfaces + GPU guard)

Completes the single-API end state for the remaining backend-taking concrete surfaces and wires the GPU nested-entry guard. All five areas in ONE designed-and-gated PR.

Capability decision (option b — built-in session dispatch): the concrete op methods take &mut dyn BackendSession and dispatch internally to the built-in exec sessions (the existing execute_linalg_extension_reads_on_session / execute_fft_extension_reads_session downcast patterns). This removes the caller’s manual with_cpu_exec_session/with_cuda_exec_session wrappers (the current linalg/fft doctests require them). A session that cannot run the op returns a typed capability error; callers never downcast themselves. Documented breaking restriction: the concrete op traits no longer accept arbitrary third-party LinalgBackend/FftBackend implementations — the public SPI traits stay (backend implementers still implement them), but the concrete op path is built-in-session only. The test-only custom backends (tenferro-linalg backend_errors.rs, tenferro-fft backend_capability.rs) migrate to the built-in sessions or are removed, and the extensibility claims in their docs are updated to the SPI-only scope.

1. einsum top-level traits (10 traits) + tensordot

Methods take &mut dyn BackendSession. Einsum-family internals already route through ConcreteEinsumPlan::execute* (Phase-2 finding-3 fix); tensordot is not a plan path — TensorTensordotExt/TypedTensorTensordotExt call session.dot_general/dot_general_read directly (concrete.rs:47-57, 79-91), the correct session form. Remove the backend-taking forms. Call sites (6 files) migrate per the Phase-2 rule.

2. FFT (TensorFftExt / TensorReadFftExt)

Methods take &mut dyn BackendSession; internal dispatch to the built-in FFT exec sessions (CPU/CUDA/WebGPU — all implement FftBackend); typed capability error otherwise. Inventory: 4 executor methods (fft/ifft/rfft/ irfft, lib.rs:236-333) + the 8 trait methods. The executor calls backend.execute_fft directly (no internal session entry); the manual external downcasts in callers/tests/benches (backend_capability.rs:335-345, fft_plan_cache.rs:51-70) are removed. Plan cache stays executor-owned. Call sites: 5 files + the test/bench downcast sites.

3. linalg concrete traits (TensorLinalgExt / TensorReadLinalgExt / TypedTensorLinalgExt — 64 methods)

Methods take &mut dyn BackendSession; internal dispatch downcasts once to the concrete exec session and runs the EXISTING generic composite bodies on it (they already operate on the borrowed LinalgBackend — tensor_ext.rs 1702-1817; base session ops via reshape_read/transpose_read/dot_general_read at 2514-2589). Preserve the direct solve_read_into write path (tensor_ext.rs:1914-1920 → cpu/backend.rs:718-741; do NOT replace with allocate+copy). Typed contract via typed_output (tensor_ext.rs:2197-2205), not into_typed_result. The caller’s with_cpu_exec_session wrappers are removed (internal dispatch replaces them). Validation, dtype promotion, error payloads unchanged. Call sites: 5 files.

4. shape_packing (Tensor::index_select / Tensor::stack)

Methods take &mut dyn BackendSession. Both already run inside a session closure (shape_packing.rs:138-145, 167-193) — replace the host parameter and remove the closure. Migration gate: scoped grep for .index_select(/ .stack( on the concrete surface; traced/eager APIs and historical docs/plans are explicitly excluded.

5. GPU nested-entry guard (CUDA / WebGPU)

(Historical specification; superseded by #1938 D6, where the guard returns a typed SessionEntryError::Reentered in every build profile and CPU reentry is typed instead of a panic.)

Extract the portable guard (thread-local in-session flag + panic-safe restore + debug assert) into a #[doc(hidden)] pub shared helper in tenferro-tensor (e.g. with_session_entry_guard). CUDA (cubecl/exec_session.rs:603) and WebGPU (webgpu/exec_session.rs:263) with_backend_session overrides wrap their closure in it — nested entry is detected in debug builds on every backend family. CPU keeps its release EXECUTION_OWNER panic; Send bounds untouched. Verification: package-targeted checks (cargo check -p tenferro-gpu --features cuda, --features webgpu) + new cfg-gated nested-entry tests for BOTH overrides; the shared helper’s own tests cover the flag semantics.

Baselines (recorded BEFORE implementation — required for the gate)

Pre-change pinned baselines recorded in the worklog: - FFT: crates/tenferro-fft/benches/fft_plan_cache.rs arms (direct_one_shot / executor_warm_cache, 1024 f64, 1 thread, pinned core 40) — measured: ~15.0/11.3 µs (small) and ~88.3/73.1 µs (large) pre-change. - Linalg: representative ops via the current caller pattern (with_backend_session + with_cpu_exec_session), 8×8 f64 diagonal (diag 2..=9), 1 thread, pinned core 40 — measured pre-change: solve 14.4 µs (incl. the direct solve_read_into path) and svd 18.6 µs (2000 iters, release). matmul shares the same session path as solve and is not a separate baseline. - Post-migration: same arms re-run; no-regression threshold = within the machine’s interleaved noise band (±1% on interleaved runs); recorded in the Phase-3 worklog.

Migration rules (reuse Phase-2), tests, doctests

Phase-2 rule (single-op wrap / group / session-typed helpers / already-in- session direct / closure result types; docs/spec/tutorials/skills/snippet generators migrate). Runnable doctests on every changed public method. Parity/error tests per area (values + structured errors; linalg typed conversion via typed_output; solve_read_into direct-path test). The runtime/ einsum chain benches are untouched (core API unchanged).

Out of scope

  • Low-level backend SPI traits (TensorBackend capability traits, the LinalgBackend/FftBackend trait definitions) — backend contracts, not concrete op surfaces; third-party backend implementers still implement them, but the concrete op path is built-in-session only (documented above).
  • GPU runtime execution (needs a CUDA host; CI gate item).