CPU Execution and NUMA Placement

tenferro distinguishes a CPU thread count from CPU affinity. A thread count limits parallelism; affinity determines which logical CPUs may execute those threads. On NUMA machines, setting only the count does not keep work in one memory-locality domain.

For session versus worker-dispatch overhead and ways to amortize it, see CPU Session Entry and Rayon Dispatch Cost.

What a NUMA Node Means

For this API, a NUMA node is an operating-system topology domain identified by its sparse OS node ID and its set of logical CPUs. A node is usable only through the intersection of that OS CPU set with the process affinity mask. Therefore:

  • NumaNodeId(2) means OS node 2, not the third entry in a dense array;
  • CPUs excluded by a container, scheduler, cpuset, or taskset are never added back by tenferro;
  • an OS node with an empty process-visible intersection is unavailable; and
  • AllAllowed means the complete process affinity mask, not every CPU installed in the host.

Inspect the process-visible topology before selecting a node:

use tenferro_cpu::{CpuBackend, CpuPlacement};

let backend = CpuBackend::new();
for node in backend.topology().nodes() {
    println!("OS node {}: {:?}", node.id(), node.cpus().as_usize_vec());
}

if backend.supports_placement(CpuPlacement::AllAllowed) {
    let all = backend.for_placement(CpuPlacement::AllAllowed)?;
    println!("{:?}", all.execution_info());
} else {
    println!("{:?}", backend.execution_info());
}

Managed Placement Is a faer Contract

The initial capability matrix is deliberately conservative:

Runtime CPU provider Auto NumaNode(id) AllAllowed
faer and tenferro-native kernels managed all-allowed engine managed pinned node engine managed pinned all-allowed engine
OpenBLAS, MKL, Accelerate, or another external BLAS provider-default, process-wide exclusive unsupported unsupported

On platforms where tenferro cannot set and verify worker affinity, faer’s Auto mode uses an unpinned compatibility context. Explicit managed placement still returns an error instead of silently weakening the request.

CpuPlacement::NumaNode and CpuPlacement::AllAllowed are supported by CpuBackendKind::Faer. tenferro creates a fixed Rayon engine for the resolved CPU set and confines every worker to that whole set when the engine is constructed; workers share the domain mask instead of owning one CPU each. The verified workers report CpuExecutorAffinity::TenferroDomainVerified. They share the set because a thread created by a provider (BLAS/LAPACK, or your own library call) inherits the creating thread’s mask: a one-CPU worker mask would confine the provider’s entire thread team to one CPU. CpuBackend clones are cheap handles: they share topology, engines, arbitration, and engine-owned caches.

use tenferro_cpu::{CpuBackend, CpuBackendKind, CpuPlacement};

let coordinator = CpuBackend::with_threads_and_kind(4, CpuBackendKind::Faer)?;
if let Some(node) = coordinator.topology().nodes().first() {
    let local = coordinator.for_placement(CpuPlacement::NumaNode(node.id()))?;
    let another_handle = local.clone();
    assert_eq!(local.resolved_placement(), another_handle.resolved_placement());
}

Auto resolves to AllAllowed for faer. An all-allowed engine can use all CPUs granted to the process, so splitting work by NUMA node does not prevent a separate all-node computation. Overlapping placements are serialized by the process-wide arbiter, including placements created by independently constructed backends; disjoint node placements may execute concurrently. Other top-level executions remain subject to these overlap rules.

CPU backend execution is not reentrant. Do not call a backend clone or another CPU backend directly from CpuBackend::install or a backend session, and do not make backend calls from Rayon tasks spawned inside one. A managed scope rejects same-thread, spawned, stolen, and shared-context re-entry with a typed SessionEntryError::Reentered (reported through Error::SessionEntry by direct operations such as install) before a second permit is acquired and before the nested callback runs. Work moved to an unrelated executor cannot always inherit that diagnostic marker and may instead wait for the outer permit, so waiting for it from the outer execution can deadlock. Finish the outer backend execution before launching new top-level backend calls. Ordinary Rayon work that does not re-enter a CPU backend remains supported.

Caller-managed executor domains

A host scheduler that already owns independent Rayon pools can register each pool as a caller-managed external domain. Use RayonCpuDomainExecutor::new(Arc<rayon::ThreadPool>) with ExternalCpuDomain::new_caller_managed; the generic CpuDomainExecutor trait remains available for non-Rayon executors. No CPU set is declared or inferred. Select a registered domain with CpuBackend::for_domain(CpuDomainId).

Caller-managed domains use faer and tenferro-native kernels only. tenferro retains the supplied executor, enters only that executor, applies the declared thread budget, and never creates or shuts down another pool. Entry from an exclusive coarse task already running on the same Rayon pool is supported. Distinct caller-managed domains do not enter the process-wide CPU-set arbiter, so their pools may overlap even when the process affinity mask is identical. The host owns cross-domain scheduling and oversubscription policy.

A second public backend entry into the same active domain is rejected, including entry from another worker of that pool. Provider bundles with external workers or uncontrolled thread counts, and BLAS/LAPACK configurations that cannot stay inside the supplied executor, fail during backend construction. Diagnostics report CpuAdmissionMode::CallerManaged; resolved placement and domain CPUs are None because tenferro has no verified placement claim.

Scoped direct faer calls

Downstream code that needs a faer routine not exposed by tenferro can use tenferro_cpu::FaerParallelismExt on the active BackendSession:

backend.with_backend_session(|session| {
    session.with_faer_parallelism(|parallel| {
        faer_operation(..., parallel)
    })
})??;

The callback receives the same policy as an internal faer operation: bounded Par::rayon(n) for managed multi-thread inner execution, and Par::Seq for one-thread or already-inner/outer-worker execution. The capability is lexical; it does not expose the Rayon pool, executor handle, mutable CPU context, or a value that can outlive the callback. Calling tenferro backend/session methods from the callback remains unsupported and retains the existing reentrancy diagnostics. This guarantee applies to direct faer/Rayon-compatible calls, not to workers created internally by OpenBLAS, MKL, Accelerate, or OpenMP.

External BLAS Providers

OpenBLAS, Intel MKL, Apple Accelerate, and OpenMP-backed BLAS implementations own their worker creation. Provider workers inherit the calling worker’s mask, which for a managed domain is the domain CPU set, but tenferro cannot bound the provider’s fan-out and a provider that installs its own affinity policy (for example KMP_AFFINITY with an explicit list) overrides the inherited mask. Consequently external BLAS backends accept only CpuPlacement::Auto; explicit NumaNode and AllAllowed requests return CpuPlacementError.

Auto for CpuBackendKind::Blas uses provider-default execution under a process-wide exclusive permit. This prevents tenferro-managed CPU work from overlapping a provider call whose worker CPU set is unknown. Provider variables such as OPENBLAS_NUM_THREADS, MKL_NUM_THREADS, and OMP_NUM_THREADS still control counts where supported, but they do not upgrade the provider to a tenferro-managed affinity contract.

If strict NUMA placement is required, select CpuBackendKind::Faer. tenferro makes no placement promise for threads a provider creates; if an application configures and pins a BLAS provider independently, that remains an application/provider responsibility.

Intel OpenMP worker affinity on Linux

When MKL first creates its workers from a tenferro worker, the new threads inherit that worker’s mask — the whole domain CPU set — so a 4-thread MKL operation starts with the intended CPUs available. A reported MKL_Get_Max_Threads() == 4 still does not by itself prove where the provider placed its workers, and the provider’s own settings decide how many run.

Providers that install their own affinity policy override the inherited mask. If your deployment sets one, keep it inside the CPUs the process may use and check /proc/<pid>/task/<tid>/status (Cpus_allowed_list) during execution:

taskset -c 1-4 env RAYON_NUM_THREADS=4 MKL_NUM_THREADS=4 OMP_NUM_THREADS=4 \
  MKL_DYNAMIC=FALSE \
  KMP_AFFINITY='granularity=fine,proclist=[1,2,3,4],explicit,norespect' \
  ./your-program

Replace both CPU lists with CPUs available to your application, and never list CPUs outside the application’s intended allocation. This is an Intel OpenMP setting, not a portable OpenBLAS/Accelerate prescription or a new tenferro guarantee. The norespect form was originally a workaround for an inherited one-CPU worker mask; managed domains no longer produce that mask, so it is only needed when you deliberately pin OpenMP yourself.

Fallible backend constructors return CpuBackendError. Configuration failures appear as CpuBackendError::Tensor, while topology discovery and engine placement failures remain inspectable through CpuBackendError::placement_error.

Batched Operation Strategy

Batched work (strided-batched and grouped GEMM, packed LU factor/solve) runs its independent items one of five ways, chosen by a CpuBatchPolicy:

CpuBatchStrategy Items run as
Auto (default) a route chosen from CpuBatchThresholds and the provider
Sequential one after another, with no parallelism
OuterParallel tenferro lanes on the backend’s own Rayon pool, each sequential
ProviderItems one after another, each using the provider’s own parallelism
WholeBatchVendor one vendor cblas_?gemm_batch call, with no tenferro fan-out

Set a default with CpuBackend::with_batch_policy, and override it for part of a session with tenferro_cpu::with_batch_policy(session, policy, |s| ...); wrapping a single call is a per-operation choice. The innermost scope wins and the previous policy is restored on return, error or unwind. Nothing is process-global.

Auto keeps the pre-existing behavior: a grouped batch of small GEMMs (every dimension at most vendor_batch_max_item_dim, default 16) may use the vendor batch call when the build links one (blas-openblas, blas-mkl), while a strided-batched contraction runs one provider GEMM per item and reaches the vendor batch call only through WholeBatchVendor (with OpenBLAS at one thread the per-item route was measured faster for items of 8 and larger); outer lanes are used when the batch has at least outer_min_items items and outer_min_items_per_lane per lane. A forced strategy never overrides a safety rule or invents a missing route: Sequential with a provider that declares its own threading (the built-in BLAS), OuterParallel on a one-thread backend or for strided-batched contractions, and WholeBatchVendor with faer or without a linked vendor routine all return an Unsupported error before any output is written. A contraction whose axes are all batch axes is an elementwise product and never lowers to per-element GEMMs.

CPU Affinity Is Not NUMA Memory Placement

Pinned workers restrict where computation may run. They do not make tensor allocation NUMA-local, choose a first-touch policy, migrate existing pages, or configure page interleaving. Input and output pages may therefore remain remote from the selected node. Applications that require memory locality must arrange allocation/first-touch or OS memory policy separately and measure the result on their deployment topology.

Fresh CPU results record the selected execution domain in Placement::cpu_affinity. This is routing and locality metadata: it identifies the domain that produced the result and can guide later scheduling. It is not proof that the storage was allocated by that domain, that its pages are pinned or resident there, or that an allocation-domain owner changed. Metadata-only views and reshapes retain their storage metadata, while caller-owned _into destinations are never retagged.

The worker budget is not mapped onto individual CPUs: every worker is confined to the same domain CPU set and the operating system schedules the pool inside it. That is not a promise to prefer physical cores over SMT siblings; tenferro does not infer core/sibling topology for a domain.

Where Elementwise Rayon Runs

For a faer backend, elementwise, analytic, reduction, structural, indexing, and faer GEMM work execute inside the selected fixed Rayon engine. With the default Auto placement this is the all-allowed process CPU set; with NumaNode(id) it is that node’s process-visible CPU set.

For a BLAS backend, a graph keeps one exclusive coordinator permit. Native tenferro segments and BLAS/LAPACK provider calls both cross the selected domain executor exactly once. The BLAS call runs inside that admitted operation region, but the provider runtime owns its worker fan-out; it does not use the executor’s Rayon team as BLAS workers. Thus an elementwise operation adjacent to BLAS does not run on an unconstrained global Rayon pool, and no provider path bypasses domain admission.

Supported Host instructions, native instructions, and session-capable GEMM FFI instructions share one backend session. Extension runtimes that cannot execute through BackendSession remain explicit session boundaries.

Diagnostics

Use CpuBackend::execution_info() for logs. CpuBackendKind::{Faer, Blas} is the stable public provider identity. provider_diagnostic() may mention a compiled provider such as OpenBLAS or a runtime-injected provider, but that string is diagnostic only and may change.

Runtime registration uses the opaque CpuRuntimeIdentity witness token for exact backend identity. Clones of one backend share the token; a newly constructed backend or a backend whose provider bundle, placement, or shared allocation domain changes receives a distinct token. The token carries no execution or storage authority and is not a provider/device identifier.

let backend = tenferro_cpu::CpuBackend::new();
let info = backend.execution_info();
println!("kind={:?} provider={}", info.backend_kind(), info.provider_diagnostic());
println!("mode={:?} workers={}", info.execution_mode(), info.worker_count());
println!("topology={:?} requested={:?} resolved={:?}", info.topology(),
    info.requested_placement(), info.resolved_placement());

See Parallelism and Caching for thread budgets, cache limits, and oversubscription guidance.