Linear Algebra
tenferro-linalg is the public tensor linalg extension of the workspace. Its role is to validate contracts, build traced and eager linalg operations, register linalg execution with the runtime, and provide optional linalg AD rules behind the autodiff feature. It is not the backend execution contract itself.
Position in the Workspace
tenferro-tensor
TensorBackend, CPU kernels, linalg execution
|
v
tenferro-runtime
traced graph runtime and extension dispatch
|
v
tenferro-linalg
linalg extension API, runtime registration, and optional AD rules
Responsibilities
tenferro-linalg owns:
- shape and option validation
TensorLinalgExt,TensorReadLinalgExt, andTypedTensorLinalgExtfor backend-explicit primal execution- fixed-arity result tuples and typed real/complex output mappings
- composite lowering
- traced extension op payloads and runtime registration
- optional eager helpers and linalg AD rules behind
autodiff
tenferro-linalg does not own:
- backend-specific kernel interfaces
- CPU/GPU direct dispatch branches
- standalone structural view operations
- mandatory AD dependencies for primal-only users
Kernel Basis vs Composite API
The public API is larger than the backend kernel basis.
Kernel-oriented operations
These lower directly to TensorBackend linalg methods:
solvesolve_triangularqrsvdlu_factorcholeskyeigeneig
Composite operations
These remain in tenferro-linalg and lower through structural ops, core tensor ops, scalar/analytic ops, and the kernel basis above:
matrix_powercondtensorinvtensorsolvemulti_dotvecdotvanderinvand structured*_exwrappers built from solve/factorization kernels
The key rule is that public API breadth does not imply backend kernel breadth.
Injectable CPU kernels
On the CPU backend the kernel basis can be replaced op by op without replacing the backend: tenferro_linalg::cpu_kernels::CpuLinalgKernels (installed with install_linalg_kernels, stored as a typed extension of the CpuProviderBundle) has one method per primitive the CPU session dispatches (cholesky, triangular_solve, lu, full_piv_lu, solve, svd, svd_full, svd_values, qr, rank_revealing_qr, eigh, eigh_values, eig, eig_values). Each owned and _read hook of CpuExecSession asks the installed kernels first, inside its single operation entry, with the entered CpuExecutionContext; a kernel returns CpuLinalgOutcome::Unsupported before producing anything to fall through to the built-in faer/LAPACK kernel. Kernels must return exactly the built-in outputs (order, shapes, dtypes, pivot and ordering conventions, trailing batch axes), so composites and AD rules are unchanged. The Householder family, lu_factor, prepared LU solves and _into outputs keep the built-in kernels.
Built-in CPU providers
The built-in kernels are the extracted, tensor-free tlinalg (faer) and tlinalg-blas (LAPACK/BLAS) crates of the tlinalg-rs workspace. Neither depends on the other, and each owns its own error, scalar and scratch vocabulary; the interface they are adapted to is this crate’s, in cpu/tlinalg.rs, cpu/tlinalg_blas.rs, cpu/tlinalg_error.rs and cpu/tlinalg_workspace.rs. Every family is called once per batch with borrowed strided descriptors of dims [rows, cols, batch...]; the provider owns the kernels, their scratch and the batch loop, and returns compact column-major, batch-contiguous outputs. The host keeps shape validation with tenferro’s error shapes, empty and zero-batch results, pooled output allocation, tensor construction, negative-stride rejection (a borrowed operand with a negative stride is gathered once into a compact copy), the QR gauges, the rank decision of rank-revealing QR and its non-finite input screen (both routes screen before the provider runs), and the compact Householder QR state compositions (append, from-factors, R, Q columns) built on the providers’ reflector kernels.
Batch policy. On the faer route every batched family follows the effective CpuBatchPolicy through one host-resolved lane plan (cpu/tlinalg.rs lane_plan): Auto fans out over the context’s budget once the item-count thresholds allow, forced OuterParallel needs a context that can fan out, Sequential and ProviderItems stay on one lane, and WholeBatchVendor is a typed error. tlinalg runs the lanes as tasks on the pool the context hands it, each lane sequential; it never calls back into tenferro from a lane, so it needs no outer-lane execution contexts. Before the extraction only packed LU/solve followed the policy and every other faer family looped over its batch serially; those families now follow it too, and a forced strategy the route cannot serve is now a typed error for them as well. Under Auto they fan out only when one item is small, max(rows, cols) <= 64 (the crate-private AUTO_FAN_OUT_MAX_ITEM_DIM, wrapped around the plan by lane_plan_for_item): lanes pay off for many small matrices that cannot use the budget one at a time, while a large matrix is better served by faer’s own parallelism within the item, and the item-count thresholds cannot tell the two apart. Above the guard Auto keeps the pre-extraction behaviour, one lane with the context’s parallelism per item. The guard is provisional, pending a work-model thread policy for CPU linalg (#2000); forced strategies and the packed-LU family ignore it. On the LAPACK route the provider loops over the batch serially with one workspace query per call and vendor-owned threading; packed LU keeps its strategy admission and the other families ignore the policy, as before.
Concrete, Read, And Typed Boundary
Owned dynamic tensors call the matching LinalgBackend owned hook. Borrowed tensor reads use the _read surface and preserve strided layouts until the selected provider performs documented same-placement canonicalization. Typed tensors erase only the borrowed input through TensorScalar::tensor_read, then validate each dynamic backend output while downcasting it to the fixed typed tuple.
LinalgScalar is sealed to f32, f64, Complex32, and Complex64. It uses TensorScalar::Real for singular values, Hermitian eigenvalues, determinant log-magnitudes, and norms, and supplies a Complex associated type for general eigendecomposition. These adapters never perform an implicit device transfer; unsupported layout or placement combinations remain typed backend errors.
Shape Convention
All matrix APIs use column-major tensor layout with:
- first two dimensions = matrix dimensions
- remaining dimensions = batch dimensions
This is the column-major counterpart to PyTorch’s trailing matrix convention.
AD Boundary
AD formulas live in tenferro-linalg behind the autodiff feature.
That feature boundary is deliberate:
- primal-only users can depend on
tenferro-linalgwithout AD dependencies - AD users enable
tenferro-linalg/autodiffand pass the owned rule set into an explicittenferro_ad::AdContext - the process-global registration API is retained only as a compatibility bridge
Some public APIs are naturally primal-only, especially structured status/result surfaces such as factorization contracts with pivots or info metadata.
Solve Lowering
Traced solve and tracked eager solve emit one LuFactorSolve extension op with outputs (x, packed_lu, pivots). The CPU backend implements it as one fused kernel: each matrix is factored in a pooled scratch buffer and solved in place, and that scratch buffer becomes the packed LU output, so a primal only program does the same work as the plain Solve kernel. The backend default, used by CUDA, composes lu_factor with lu_solve_prepared.
Linearization emits LuSolvePrepared on the saved factors, and the transpose rule solves the adjoint system with the same factors. Reverse mode therefore factors A exactly once.
LuFactorSolve is never pruned to Solve. Traced AD prunes unused extension outputs of the source program before differentiating it; a prune to Solve would drop the saved factors and make the adjoint refactor A. The untracked eager surface, which never differentiates, still constructs Solve directly.
Batched CPU LAPACK kernels and the faer packed LU kernels follow one loop discipline: one workspace query per call, scratch reused across the batch, results written directly into the batched output, and no per matrix tensor or Vec allocation. lu_solve_prepared applies the pivots inside the provider solve (one getrs per matrix for every transpose and conjugate flag).
Current Implementation Status
The architectural boundary is now active rather than transitional:
tenferro-tensorowns backend-facing structured linalg kernels throughTensorBackendtenferro-linalgowns extension APIs, runtime registration, eager helpers, and feature-gated linalg differentiation rules
Current debt is mainly about capability breadth and composite coverage:
- some composite families still bottom out in CPU-only kernels because GPU capability is not implemented yet
- public primal parity is broader than VJP/JVP/HVP parity for several newer families
- some structured results are intentionally primal-only
Non-Goals
tenferro-linalg is not trying to be a literal mirror of torch.linalg at the backend boundary. PyTorch-style API families may exist publicly, but backend contracts are intentionally smaller and more Rust-structured.
For the broader family-level parity and backlog view, see reference/pytorch-dense-cpu-parity.md.