pub fn sum_with_per_step_rounding(values: &[Bf16]) -> Bf16Expand description
Sum a slice, rounding to bfloat16 after every step.
This is the weaker contract, and it is the one to avoid when the accumulation is promised in
f32: on a slice of three hundred ones it stalls at 256.0, because from there on bfloat16
spacing above one is 2.0 and adding one rounds back down.
ยงExamples
use tenferro_bf16_proof::{reduction::sum_with_per_step_rounding, Bf16};
let ones = vec![Bf16::from_f32(1.0); 300];
assert_eq!(sum_with_per_step_rounding(&ones).to_f32(), 256.0);