discounted_returns
std.seq.discounted_returns · Level L2Discounted returns of a reward sequence, as in reinforcement learning: a reverse Scan of discount_step.
Gₜ = rₜ + γ·Gₜ₊₁, Gₙ = 0
Signature
discounted_returns(r: f64[n], gamma: f64[]) → f64[n]
Structure
The function as NOVA stores it: one box per input, operation and output, and arrows that carry values. A double border marks another library function this one runs — called once, or by Scan once per element; select it to open that function.
- input
- operation
- constant
- call
- output
Verification
- Signature proven by NOVA’s shape solver, for every size.
- Equal to the reference
G = 0; for t from the end: G = r[t] + gamma*Gin exact rational arithmetic, on all 40 test cases. - All 163 float64 results inside the running error bound; the closest uses 86% of it.
- Interpreter and NumPy backend return bit-identical results.
Accuracy in detail
- correctly rounded (the float64 nearest the exact value)
- 87%
- bit-equal to the NumPy formula in float64
- 100%
- largest error, in units in the last place
- 264
Large ulp counts appear only where cancellation drives a result toward zero; the absolute error is still inside the bound.
Identity
sha256:df57cc292ea2776d9a95ede990fc63101bd35f1ef71d3c235e0752488e3eaad5The semantic hash of the graph. It changes when the program changes, and never when only its documentation does.