Session-Oriented Concrete APIs (issue #1673)
Status: design (pre-implementation review target).
Overview
Add a session-explicit execution surface for concrete Tensor/TypedTensor operations (including extension-crate operations) so a caller can reuse one borrowed BackendSession across a sequence of operations instead of paying a session entry per operation. The tensor value types remain context-free and are never owned by or lifetime-bound to a session.
The one-shot API (a.add(&b, &mut backend)) stays available and becomes a thin wrapper around the session implementation where practical.
Superseded in part by #1926. This document’s scope is adding the session-explicit surface, and it deliberately keeps the one-shot spelling. #1926 goes further and removes session-opening operation-level APIs, because a thin wrapper that opens a session per call is still an implicit entry. The session surface defined here, the nested-entry prohibition, and cache parity remain in force; only “the one-shot API stays available” changes.
Motivation (measured)
Session entry for managed CPU execution is ~3 µs/op (release, pinned, 1 worker; issue #1673 data comment). For trivial unary/binary ops (~3 µs total) it is ~100% of the cost; for a 10-op trivial chain the current API pays 10 session entries (~30 µs) where one entry (~3 µs + 10 op bodies) suffices. For compute-heavy ops (solve, ~28 µs after #1675) the session share is ~10%, so the surface targets short chains of cheap ops; it is not a GPU small-op-launch fix.
Gate (from the issue): with T_one_shot = 10-op trivial chain via the current one-shot API (10 session entries) and T_one_session = the same chain via one session entry, require T_one_shot / T_one_session >= 2 (predicted ~6). The baseline is explicit: current one-shot API, 10 session entries.
Verified execution model (corrects the 2026-08-14 design review)
The design review assumed with_backend_session runs the closure synchronously on the calling thread and therefore proposed dropping the Send bounds. That is false for CPU:
CpuBackend::with_backend_session→run_backend_session_cached→enter_managed_session/enter→install_scoped(crates/tenferro-cpu/src/provider.rs) →CpuContext::install(docs/design/exec-session.md, “Backend Mapping / CPU”) →rayon::ThreadPool::install— a real thread handoff to a Rayon worker, even for 1-thread pools (verified empirically: caller thread id differs from the closure thread id).- The default (non-overriding) backend path (
default_backend_session,crates/tenferro-tensor/src/backend.rs) runs the closure on the calling thread.
Decision: keep the Send bounds (R: Send, f: ... + Send) on with_backend_session / with_backend_session_cached. They are a soundness requirement for the CPU managed session (a non-Send closure would execute on another thread). This is an inherent property of the current CPU managed session, not a migration accident: downstream users of the session-explicit surface must capture only Send state, which the one-shot API already requires of every closure it takes.
Dropping Send would require redesigning CPU session execution to run closures on the calling thread with per-op pool entry — a performance/ semantics change to the hottest segmented-execution path. That redesign is out of scope and tracked separately if ever desired; it must NOT be assumed as a prerequisite.
Evaluation-wide scope
A2 (one execution scope around a whole evaluation) remains deferred; see exec-session.md.
Nested-entry prohibition: mechanism, not convention
Operations receiving a BackendSession must never call with_backend_session internally. Enforcement by backend family:
Every rejection below is a typed SessionEntryError returned before the closure runs (#1938 D6); none of them panics.
- CPU:
fresh_execution_owner()(crates/tenferro-cpu/src/arbiter.rs) returnsNonewhenEXECUTION_OWNERis set on the thread (or an owned Rayon scope has an active owner), and admission reportsSessionEntryError::Reentered. Covers the one-shot-inside-session case, including the pinned one-worker managed pool handoff. - Portable guard:
with_session_entry_guardsets a thread-local in-session flag (Drop-guard restored, so panics infstill restore) and returnsReenteredif it was already set, in every build profile. - CUDA / WebGPU: both overrides wrap their closure in the portable guard.
Generic spelling
Already settled by the current signatures: entries hand out &mut dyn BackendSession (no S: BackendSession + ?Sized generic on the trait methods). The _in-surface methods take session: &mut dyn BackendSession directly.
Cache access parity
with_backend_session_cached is the runtime-cache-aware entry (used by the scheduler/eager paths); with_backend_session is the canonical user entry and the one-shot concrete helpers delegate to it. The session-explicit surface uses the plain entry; extension plans (EinsumPlan etc.) own their long-lived caches. The canonical user entry must not silently become the slow entry: with_backend_session may forward to the cached path when a cache is available, but the public contract is the plain entry.
Surface design (prototype plan)
Layering (same vocabulary, no duplicate validation):
one-shot API: &mut TensorBackend ──enter one session──▶ session implementation
session API: &mut dyn BackendSession ─────────────────▶ session implementation
TensorSessionOpsExt/TypedTensorSessionOpsExtwith*_inmethods (naming:_inpreferred; exact names decided in the prototype). Implemented by calling the session’s own op traits directly; broadcasting, dtype promotion, validation, and typed errors identical to the one-shot path (shared helpers — no second vocabulary).- One-shot methods delegate to the
_inimplementations where practical. - Concrete-only extensions (einsum plans first:
EinsumPlan::prepare+execute_in_session) compose standard session ops; noExtensionOp/SemanticProgram/runtime registration required at that tier. - Runtime-integrated extensions keep the semantic
ExtensionOp+ AD rules; runtime preparation produces an executor that runs on the borrowedBackendSessionlike the concrete path.
Non-goals
- No session/runtime handle inside
Tensor/TypedTensor. - No change to the
EagerTensorownership/AD model. - No unbounded or globally owned sessions; closure-scoped borrowing stays the default.
- No second operation vocabulary or duplicated validation.
- GPU small-op launch overhead is a separate concern (fusion/static execution), not addressed by session reuse.
- The one-shot API is not broken for aesthetic consistency. (#1926 supersedes this specific non-goal for session-opening operation-level APIs; the rest stand.)
Acceptance criteria
- One session entry covers a mixed sequence of standard and extension concrete operations.
- Standard broadcasting, dtype promotion, validation, placement, and typed errors identical to the one-shot path.
- One-shot methods delegate to the same implementation where practical.
Tensor/TypedTensorremain context-free.- Concrete-only extension authors need no runtime/graph/AD machinery.
- Nested-entry prohibition enforced per backend family (originally a CPU release panic plus a debug assert; since #1938 D6 a typed
SessionEntryError::Reenteredbefore the closure runs on CPU, CUDA and WebGPU). Sendbounds preserved and documented as a soundness requirement.- The 10-op trivial-chain gate passes (≥2x, predicted ~6x).
- No material regression for large operations.
Prototype specification (PR B: the _in surface)
Scope decision for the first implementation PR. The design-review gate rules apply: this section is the pre-implementation design document; the implementation must not start until it has a reviewer-gpt verdict.
Measured state of the one-shot path before #1926
A binary op already pays more than one session entry today (crates/tenferro-runtime/src/tensor.rs):
add→broadcast_binary→broadcast_toper operand (eachbroadcast_tocan enter 0-2 sessions:reshape+broadcast_in_dim, or zero when shapes already match viaduplicate) → thenwith_backend_session(|exec| exec.add_read(...))for the op itself. Worst case: 5 sessions per binary op (both operands need reshape+broadcast = 2 each + 1 final add, e.g.[1,2] + [2,1]→[2,2]); the common one-sided reshape+broadcast case is 3; equal-shape case is 1.unary_fn(exp etc.) andreduce_sum: 1 session each.
This is the before-side of the measurement; the owner-level one-shot spelling it describes was deleted by #1926, so the numbers are now the reference that docs/testing/session-route-baseline.json compares against.
So the session-explicit surface must own the broadcast step too, or a binary op still pays 2 entries inside one session.
Surface (prototype op set)
crates/tenferro-runtime/src/lib.rs, next to TensorOpsExt:
pub trait TensorSessionOpsExt {
fn add_in(&self, rhs: &Tensor, session: &mut dyn BackendSession) -> Result<Tensor>;
fn mul_in(&self, rhs: &Tensor, session: &mut dyn BackendSession) -> Result<Tensor>;
fn exp_in(&self, session: &mut dyn BackendSession) -> Result<Tensor>;
fn reduce_sum_in(&self, axes: &[usize], session: &mut dyn BackendSession) -> Result<Tensor>;
}and TypedTensorSessionOpsExt<T: TensorScalar> with the same four methods returning TypedTensor<T>. Exact op set is deliberately small: add_in/mul_in (binary + broadcast, the multi-session flagship), exp_in (unary), reduce_sum_in (reduction).
Naming is transitional: the _in suffix exists only because the session-explicit form coexists with the one-shot TensorOpsExt during the migration — same receiver type (Tensor) + same method name on two in-scope traits is an ambiguous-method-resolution error in Rust (and a prelude glob would break every call site). The final canonical names drop the suffix: once the one-shot API is migrated away (a release-boundary breaking change), the session-explicit methods become the plain add/exp/reduce_sum. The _in names must not be treated as the permanent public spelling.
Validation and broadcast sharing
- Broadcast plan computation (
broadcast_shapes,broadcast_input_plan) is pure and shared. The error mapping (broadcast_error) is not shared: the dynamic surface usesbroadcast_error_to_validation(crates/tenferro-runtime/src/tensor.rs) and the typed surface has its own manual mapping (typed_tensor.rs). Each_insurface reuses its existing local mapping; no third mapping is added. - Dynamic
Tensorhelper (owned): session-levelbroadcast_to_in(input, target_shape, session)usingsession.reshape+session.broadcast_in_dimwith the same plan logic as the existingbroadcast_to, preserving theduplicate()path for equal shapes and broadcast sources (owned copy semantics unchanged).add_in/mul_in=broadcast_to_inboth operands +session.add/session.mul, all inside the one borrowed session. - Typed
TypedTensor<T>helper (borrowed, read-based): the typed one-shot path keeps inputs borrowed and dispatches to the*_readmethods (typed_tensor.rs); the typed_insurface mirrors that withreshape_read/broadcast_in_dim_read/add_read/mul_readon the session andinto_typed_resultfor the output (dtype fixed byT, no promotion logic). Do NOT funnel the typed path through the owned dynamic helper — that would introduce copies or change allocation semantics. - No second validation vocabulary; no new error kinds. One-shot/session parity tests cover values and structured errors for both surfaces.
One-shot delegation
One-shot add/mul/exp/reduce_sum become backend.with_backend_session(|s| ..._in(...)) where practical. This drops the worst-case binary-op session count from 5 to 1 — a side benefit that must be measured as a no-regression check, not assumed. dot_general/matmul are excluded from this PR (cache-ownership decision pending) and keep their current path.
Explicitly out of scope for PR B
dot_general/matmul(deferred until cache ownership/parity withwith_backend_session_cached/SessionCachedDotis decided; the current one-shot matmul already uses the plain entry, so this is purely a cache-ownership decision).convert/castand the rest of the op vocabulary (follow-up PRs).EinsumPlan::execute_in_sessionand all extension-crate tiers (PR C).- GPU session surfaces; CUDA/WebGPU nested-entry enforcement.
- Removing
Sendbounds (soundness requirement, see above).
Public API and tests (explicit before merge)
Both public traits (TensorSessionOpsExt, TypedTensorSessionOpsExt) get runnable doc examples (AGENTS.md requirement — compile and run, no ignore). Integration tests cover, per surface: equal-shape op, real broadcast, invalid broadcast (structured error parity with the one-shot path), dtype/error parity, typed output dtype validation (into_typed_result), and a test that an _in chain executes inside exactly one session entry.
Performance gate (measured, not assumed)
New criterion bench in crates/tenferro-runtime/benches/ (criterion is already a workspace dependency), with an explicit [[bench]] target and harness = false in Cargo.toml (mirroring elementwise_fusion.rs):
- Exact chain (10 ops): three repetitions of
add → exp → mul(9 ops) followed by a finalreduce_sum([0])(10th).reduce_sumis final so no intermediate op changes shape. - No-broadcast arm: 1×8 f64 constant-filled operands (broadcast is the
duplicatepath). - Broadcast arm: 1×1 f64 operands against 1×8 (real reshape+broadcast).
- One-shot arm: the chain via
TensorOpsExt(each op enters its own session — 10 entries). Session arm: the same chain viaTensorSessionOpsExtinside onewith_backend_session(1 entry). - Validate one result outside the timed region (finite, correct shape);
black_boxinputs and outputs inside the iter. - Protocol: record base-commit (pre-delegation) and post-change medians; gate
T_one_shot / T_one_session >= 2(predicted ~6); also a representative large-op one-shot before/after pair for the no-regression criterion. Report pinned (idle core), matching the session floor methodology; the gate is evaluated on the pinned numbers.
Prototype specification (PR C: ConcreteEinsumPlan::execute_in_session)
The first extension-crate example of the session-explicit concrete surface (design tier 1: concrete-only extensions composed from standard primitives — no ExtensionOp, SemanticProgram, runtime registration, or AD rules).
Key finding: the one-shot einsum already enters a session
ConcreteEinsumPlan::execute etc. (crates/tenferro-einsum/src/concrete.rs, public plan type is ConcreteEinsumPlan) already do backend.with_backend_session(|exec| eager_einsum_exec(exec, inputs, &self.tree)); the eager core functions (crates/tenferro-einsum/src/eager.rs::eager_einsum_exec*) all take &mut dyn BackendSession already. execute_*_in_session is a thin addition: plan validation (pure — validate_input_metadata/input_specs need no backend) + the same core call on the caller’s borrowed session — no new session entry. One-shot methods delegate to the _in_session variants.
Surface (complete signatures)
ConcreteEinsumPlan gains a _in_session mirror of every one-shot execute method (all mechanical: validate + call the eager core on the borrowed session). Signatures mirror the one-shot forms with backend: &mut B → session: &mut dyn BackendSession:
pub fn execute_in_session<'a, I>(
&self, inputs: I, session: &mut dyn BackendSession,
) -> Result<Tensor>
where I: AsRef<[&'a Tensor]>;
pub fn execute_typed_in_session<'a, T: TensorScalar, I>(
&self, inputs: I, session: &mut dyn BackendSession,
) -> Result<TypedTensor<T>>
where I: AsRef<[&'a TypedTensor<T>]>;
pub fn execute_read_in_session<'a, I>(
&self, inputs: I, session: &mut dyn BackendSession,
) -> Result<Tensor>
where I: AsRef<[TensorRead<'a>]>;
pub fn execute_into_in_session<'a, I>(
&self, inputs: I, session: &mut dyn BackendSession, out: TensorWrite<'_>,
) -> Result<()>
where I: AsRef<[&'a Tensor]>;
pub fn execute_typed_into_in_session<'a, 'out, T: TensorScalar, I, O>(
&self, inputs: I, session: &mut dyn BackendSession, out: O,
) -> Result<()>
where I: AsRef<[&'a TypedTensor<T>]>, O: Into<TypedTensorWrite<'out, T>>;
pub fn execute_read_into_in_session<'a, I>(
&self, inputs: I, session: &mut dyn BackendSession, out: TensorWrite<'_>,
) -> Result<()>
where I: AsRef<[TensorRead<'a>]>;
pub fn execute_read_into_accum_in_session<'a, I>(
&self, inputs: I, session: &mut dyn BackendSession,
accumulation: DotGeneralAccumulation, out: TensorWrite<'_>,
) -> Result<()>
where I: AsRef<[TensorRead<'a>]>;One-shot execute* become backend.with_backend_session(|s| self.execute*_in_session(inputs, s)).
Validation ordering change (accepted): the one-shot path currently validates before entering the session, so invalid calls are entry-free. After delegation, validation happens inside the session entry — an invalid call pays one entry (~3 µs) before returning the identical error. Accepted (error paths are exceptional); covered by an error-parity test asserting the same structured error both ways.
Naming: _in_session matches the existing prepared-operation vocabulary (crates/tenferro-runtime/src/runtime/capability.rs: prepared_operation.execute_in_session); the runtime concrete noun ops use _in (PR B). Both are transitional; final canonical names drop the suffix.
Mixed chain (acceptance item)
backend.with_backend_session(|session| -> tenferro_einsum::Result<_> {
let x = plan.execute_in_session(&[&a, &b], session)?; // einsum Result
let x = x.exp_in(session)?; // converts via From
Ok(x.reduce_sum_in(&[0], session)?)
})One entry covers standard ops + a prepared extension plan.
Bench (crates/tenferro-einsum/benches/, [[bench]] harness=false)
Exact workload: a prepared plan for "ij,jk->ik" on 8×8 f64 column-major inputs, executed 10 times per iter (10 calls). Backend: one-worker CpuBackend::with_threads(1) for all arms. Correctness: validate one result outside the timed region (shape [8,8], finite). black_box inputs and outputs inside the iter.
- Arm A (one-shot): 10 ×
plan.execute([a, b], backend)— 10 session entries. - Arm B (in-session): the same 10 calls inside one
with_backend_session(|s| ... execute_in_session(...))— 1 entry. - Arm C (mixed):
execute_in_session+exp_in+reduce_sum_in×10 inside one session vs the one-shot equivalents. - Reported: pinned (idle core) medians for A/B/C; the A/B ratio and the mixed-chain ratio; a baseline/post-delegation pair for the one-shot arms. The PR-B trivial-chain ≥2 gate is a separate acceptance; for einsum the recorded evidence is the measured ratios and the delegation baseline, not a predicted number.
Acceptance (explicit before merge)
- Runnable doctests (no
ignore/no_run) for all seven_in_sessionmethods. - Parity tests:
_in_sessionresults == one-shot results (values) for execute/execute_read/execute_typed; structured-error parity for invalid inputs (both orderings); typed dtype conversion;execute_into/accumoutput validation; and a session-counting test proving_in_sessionadds no nested entry (one entry for a mixed einsum +_inchain).
Out of scope for PR C
- Native-kernel extension capabilities (FFT, decompositions, sparse — design tier 2) and runtime-integrated extension execution on the borrowed session (tier 3).
- Top-level
einsum()-trait_in_sessionvariants (the prepared plan is the recommended repeated-execution API; trait variants are follow-up if the plan-level surface proves out). - Any change to the semantic/graph/AD einsum paths.
Migration specification (issue #1680, Phase 1: complete the _in surface)
Extends the PR-B prototype to the FULL concrete op vocabulary, keeping the transitional _in/_in_session names (final plain-name rename + one-shot removal is Phase 2, a release-boundary breaking change). Additive and non-breaking.
Op inventory (PR-B already covers add/mul/exp/reduce_sum)
TensorOpsExt (dynamic Tensor) — 26 new ops: - convert, cast (dtype) - sub, div, rem, pow, maximum, minimum (binary + broadcast) - neg, abs, sign, conj, log, expm1, log1p, sin, cos, tanh, sqrt, rsqrt (unary) - compare (binary + broadcast, CompareDir) - where_select, clamp (ternary + broadcast) - matmul (dot_general) - reshape, transpose (structural)
TypedTensorOpsExt<T> (typed) — 24 new ops: the dynamic set minus convert, cast, and where_select (typed where_select lives on the separate mask trait), plus broadcast_in_dim.
Per-family implementation (mirror the PR-B patterns exactly)
- Dynamic (owned):
broadcast_to_in+broadcast_binary_inexist from PR B; addbroadcast_ternary_infor where_select/clamp. Each_in= broadcast via the_inhelpers + the session op, all on the borrowed session; error mapping reusesbroadcast_error_to_validation. - Typed ternary: add
broadcast_ternary_in_readbuilt frombroadcast_shapes+ threebroadcast_to_in_readcalls (the existing typedbroadcast_ternary_readre-enters a session viabroadcast_to_readand must NOT be reused); typedclamp_inuses it beforesession.clamp_read. - Unary:
session.<op>directly (neg_in→session.neg,log_in→session.log, etc.). - Structural:
reshape_in→session.reshape,transpose_in→session.transpose. - Dtype:
convert_in→session.convert,cast_in→session.cast(the checked lattice / explicit projection live in the session ops). - matmul:
matmul_in→session.dot_generalwith the standard matmul config. Cache-ownership decision (resolves the PR-B deferral): the concrete surface uses the plainwith_backend_session/borrowed-session entry;with_backend_session_cachedstays internal to the runtime scheduler/eager execution paths. NoSessionCachedDotin the concrete surface. - Typed (borrowed, read-based): mirror PR-B’s typed path —
reshape_read/broadcast_in_dim_read/<op>_read+into_typed_result, reusing the existingReadInputand the typed manual broadcast-error mapping. - One-shot delegation: every one-shot method becomes
backend.with_backend_session(|s| self.<op>_in(...)). Same accepted validation-ordering change as PR B/C (invalid calls pay one entry before the identical error; error-parity tested).
Tests
- Per-family parity:
_in== one-shot values (equal-shape + real broadcast arms), structured-error parity, typed dtype conversion, session-counting no-nested-entry (one entry for a chain using multiple new_inops). - Runnable doctests (no ignore/no_run) on all new public methods.
- One-shot delegation no-regression: the PR-B
session_chainbench still passes and the one-shot arms do not regress (re-measure; the delegation can only reduce session entries).
Bench
Extend crates/tenferro-runtime/benches/session_chain.rs with a second chain using the NEW ops — one-shot vs single-session arms, pinned reporting. Keep the PR-B chain (add/exp/mul/reduce_sum) as the gate bench. Exact chain (10 ops, shapes valid throughout): [2,2] f64 through sub→log→pow→maximum→neg (elementwise, operands chosen so values stay positive through log), reshape to [4,1], transpose to [1,4], scalar-broadcast clamp, matmul by a [4,1] rhs to [1,1], then cast F64→F32. Validate a known result outside the timed region; black_box inputs/outputs inside.
Out of scope for Phase 1
- One-shot trait removal and the
_in→plain rename (Phase 2). - Top-level einsum-trait
_in_sessionvariants (plan-level surface suffices; revisit only if Phase-1 callers need them). - GPU (CUDA/WebGPU) nested-entry enforcement (tracked gap from PR A). ## Migration specification (issue #1680, Phase 2: single canonical core API)
The release-boundary breaking change for the tenferro-runtime core traits and the einsum plan surface. Removes the one-shot core API, renames the transitional _in/_in_session methods to the final plain names, and migrates every workspace call site of those APIs. End state for THIS phase: one session-explicit concrete API for the core op vocabulary and ConcreteEinsumPlan.
Scope boundary (explicit): other backend-taking concrete surfaces are NOT migrated in this phase and are enumerated as follow-ups — the einsum top-level einsum*-trait methods (crates/tenferro-einsum/src/concrete.rs traits), TensorFftExt/TensorReadFftExt (tenferro-fft), the linalg concrete traits (tenferro-linalg/src/tensor_ext.rs), and Tensor::index_select/Tensor::stack (tenferro-tensor shape_packing). This phase covers the runtime core + plan-level einsum; the others migrate in follow-up PRs toward the same end state.
Removals
TensorOpsExt(crates/tenferro-runtime/src/lib.rs) and its impl.TypedTensorOpsExt<T>and its impl.TypedTensorMaskOpsExt(typedwhere_select, implemented only forTypedTensor<bool>with a generic branch scalar) — replaced by a session-only mask traitTypedTensorMaskSessionOpsExt, implemented forTypedTensor<bool>, withwhere_select(&self, on_true: &TypedTensor<U>, on_false: &TypedTensor<U>, session)usingbroadcast_ternary_in_read+session.select_read+into_typed_result(preserves the compile-time bool-condition contract; NOT folded into the genericTypedTensorSessionOpsExt<T>).ConcreteEinsumPlan’s seven backend-takingexecute*methods (execute, execute_typed, execute_read, execute_into, execute_typed_into, execute_read_into, execute_read_into_accum).- The prelude re-exports only the session-explicit traits.
Renames (final plain names)
TensorSessionOpsExt: every*_in→ the plain op name (add_in→add,convert_in→convert,matmul_in→matmul, …).TypedTensorSessionOpsExt<T>: same rename.ConcreteEinsumPlan: eachexecute*_in_sessionbecomes the sole unsuffixedexecute*(the backend-taking forms are removed, so no collision).- No receiver collisions: tensor-extension methods take
&Tensor, backend ops take&mut self.
Call-site migration (~58 sites + tests/examples/docs/skills)
Mechanical rule, behavior-preserving: 1. Single-op call x.add(&y, &mut backend) → backend.with_backend_session(|s| x.add(&y, s))? (since #1938 D6 the entry is itself fallible, so the current spelling is ??); the closure’s result type is the op’s Result<Tensor> (annotate when the closure contains ?-chains so error types are unambiguous). 2. Consecutive session-capable ops on the same backend in one function are grouped in ONE with_backend_session; grouping stops at helpers or extension calls that still need &mut backend (avoid nested entry / borrow conflicts). 3. Reusable helpers that take &mut TensorBackend migrate to take &mut dyn BackendSession where they are pure op sequences. 4. Sites already inside a session call the ops directly. Diagnostics, README, docs/spec, tutorials, and shipped skills that reference the one-shot API migrate in the same PR. The Phase-1 one-shot-vs-_in parity tests become direct session tests against the independent expected values they already assert; the einsum mixed-session test gains independent expected values (drops the one-shot comparison arm).
Verification and gates (exact)
- Workspace builds; full suites green (runtime, einsum, ad, cpu, dependents).
- Scoped grep evidence (exclude intentionally retained internal spellings): no
\.(add|sub|...)_in\(or\.execute\w*_in_session\(calls on the core/plan surfaces;PreparedOperation::execute_in_session(runtime capability.rs),EagerTensor::from_tensor_in(tenferro-ad), and the einsum top-level/FFT/linalg/shape-packing surfaces are explicit exclusions. Command: grep for_in(/_in_session(in tenferro-runtime, tenferro-einsum (plan scope), and the migrated call sites, review each hit against the exclusion list. - Benches: the one-shot arms are removed with the API; re-run the remaining single-session arms (PR-B gate chain, Phase-1 chain, einsum chain) and compare against the recorded Phase-1 medians (no regression; the one-shot numbers are historical). Report pinned.
- Public API change is the intended release boundary; documented in the PR body/changelog.
Out of scope for Phase 2 (follow-up issues)
- Einsum top-level traits, FFT, linalg, shape-packing one-shot surfaces.
- GPU (CUDA/WebGPU) nested-entry enforcement.
- Behavioral changes to any op (rename/removal only).
Migration specification (issue #1680, Phase 3: extension surfaces + GPU guard)
Completes the single-API end state for the remaining backend-taking concrete surfaces and wires the GPU nested-entry guard. All five areas in ONE designed-and-gated PR.
Capability decision (option b — built-in session dispatch): the concrete op methods take &mut dyn BackendSession and dispatch internally to the built-in exec sessions (the existing execute_linalg_extension_reads_on_session / execute_fft_extension_reads_session downcast patterns). This removes the caller’s manual with_cpu_exec_session/with_cuda_exec_session wrappers (the current linalg/fft doctests require them). A session that cannot run the op returns a typed capability error; callers never downcast themselves. Documented breaking restriction: the concrete op traits no longer accept arbitrary third-party LinalgBackend/FftBackend implementations — the public SPI traits stay (backend implementers still implement them), but the concrete op path is built-in-session only. The test-only custom backends (tenferro-linalg backend_errors.rs, tenferro-fft backend_capability.rs) migrate to the built-in sessions or are removed, and the extensibility claims in their docs are updated to the SPI-only scope.
1. einsum top-level traits (10 traits) + tensordot
Methods take &mut dyn BackendSession. Einsum-family internals already route through ConcreteEinsumPlan::execute* (Phase-2 finding-3 fix); tensordot is not a plan path — TensorTensordotExt/TypedTensorTensordotExt call session.dot_general/dot_general_read directly (concrete.rs:47-57, 79-91), the correct session form. Remove the backend-taking forms. Call sites (6 files) migrate per the Phase-2 rule.
2. FFT (TensorFftExt / TensorReadFftExt)
Methods take &mut dyn BackendSession; internal dispatch to the built-in FFT exec sessions (CPU/CUDA/WebGPU — all implement FftBackend); typed capability error otherwise. Inventory: 4 executor methods (fft/ifft/rfft/ irfft, lib.rs:236-333) + the 8 trait methods. The executor calls backend.execute_fft directly (no internal session entry); the manual external downcasts in callers/tests/benches (backend_capability.rs:335-345, fft_plan_cache.rs:51-70) are removed. Plan cache stays executor-owned. Call sites: 5 files + the test/bench downcast sites.
3. linalg concrete traits (TensorLinalgExt / TensorReadLinalgExt / TypedTensorLinalgExt — 64 methods)
Methods take &mut dyn BackendSession; internal dispatch downcasts once to the concrete exec session and runs the EXISTING generic composite bodies on it (they already operate on the borrowed LinalgBackend — tensor_ext.rs 1702-1817; base session ops via reshape_read/transpose_read/dot_general_read at 2514-2589). Preserve the direct solve_read_into write path (tensor_ext.rs:1914-1920 → cpu/backend.rs:718-741; do NOT replace with allocate+copy). Typed contract via typed_output (tensor_ext.rs:2197-2205), not into_typed_result. The caller’s with_cpu_exec_session wrappers are removed (internal dispatch replaces them). Validation, dtype promotion, error payloads unchanged. Call sites: 5 files.
4. shape_packing (Tensor::index_select / Tensor::stack)
Methods take &mut dyn BackendSession. Both already run inside a session closure (shape_packing.rs:138-145, 167-193) — replace the host parameter and remove the closure. Migration gate: scoped grep for .index_select(/ .stack( on the concrete surface; traced/eager APIs and historical docs/plans are explicitly excluded.
5. GPU nested-entry guard (CUDA / WebGPU)
(Historical specification; superseded by #1938 D6, where the guard returns a typed SessionEntryError::Reentered in every build profile and CPU reentry is typed instead of a panic.)
Extract the portable guard (thread-local in-session flag + panic-safe restore + debug assert) into a #[doc(hidden)] pub shared helper in tenferro-tensor (e.g. with_session_entry_guard). CUDA (cubecl/exec_session.rs:603) and WebGPU (webgpu/exec_session.rs:263) with_backend_session overrides wrap their closure in it — nested entry is detected in debug builds on every backend family. CPU keeps its release EXECUTION_OWNER panic; Send bounds untouched. Verification: package-targeted checks (cargo check -p tenferro-gpu --features cuda, --features webgpu) + new cfg-gated nested-entry tests for BOTH overrides; the shared helper’s own tests cover the flag semantics.
Baselines (recorded BEFORE implementation — required for the gate)
Pre-change pinned baselines recorded in the worklog: - FFT: crates/tenferro-fft/benches/fft_plan_cache.rs arms (direct_one_shot / executor_warm_cache, 1024 f64, 1 thread, pinned core 40) — measured: ~15.0/11.3 µs (small) and ~88.3/73.1 µs (large) pre-change. - Linalg: representative ops via the current caller pattern (with_backend_session + with_cpu_exec_session), 8×8 f64 diagonal (diag 2..=9), 1 thread, pinned core 40 — measured pre-change: solve 14.4 µs (incl. the direct solve_read_into path) and svd 18.6 µs (2000 iters, release). matmul shares the same session path as solve and is not a separate baseline. - Post-migration: same arms re-run; no-regression threshold = within the machine’s interleaved noise band (±1% on interleaved runs); recorded in the Phase-3 worklog.
Migration rules (reuse Phase-2), tests, doctests
Phase-2 rule (single-op wrap / group / session-typed helpers / already-in- session direct / closure result types; docs/spec/tutorials/skills/snippet generators migrate). Runnable doctests on every changed public method. Parity/error tests per area (values + structured errors; linalg typed conversion via typed_output; solve_read_into direct-path test). The runtime/ einsum chain benches are untouched (core API unchanged).
Out of scope
- Low-level backend SPI traits (TensorBackend capability traits, the LinalgBackend/FftBackend trait definitions) — backend contracts, not concrete op surfaces; third-party backend implementers still implement them, but the concrete op path is built-in-session only (documented above).
- GPU runtime execution (needs a CUDA host; CI gate item).