discounted_returns

std.seq.discounted_returns · Level L2

Discounted returns of a reward sequence, as in reinforcement learning: a reverse Scan of discount_step.

Gₜ = rₜ + γ·Gₜ₊₁, Gₙ = 0

Signature

discounted_returns(r: f64[n], gamma: f64[]) → f64[n]

Structure

The function as NOVA stores it: one box per input, operation and output, and arrows that carry values. A double border marks another library function this one runs — called once, or by Scan once per element; select it to open that function.

rf64[n]gammaf64[]0.0Scan←discount_stepGGf64[n]
  • input
  • operation
  • constant
  • call
  • output

Verification

  • Signature proven by NOVA’s shape solver, for every size.
  • Equal to the reference G = 0; for t from the end: G = r[t] + gamma*G in exact rational arithmetic, on all 40 test cases.
  • All 163 float64 results inside the running error bound; the closest uses 86% of it.
  • Interpreter and NumPy backend return bit-identical results.
Accuracy in detail
correctly rounded (the float64 nearest the exact value)
87%
bit-equal to the NumPy formula in float64
100%
largest error, in units in the last place
264

Large ulp counts appear only where cancellation drives a result toward zero; the absolute error is still inside the bound.

Identity

Calls
Called by
—
sha256:df57cc292ea2776d9a95ede990fc63101bd35f1ef71d3c235e0752488e3eaad5

The semantic hash of the graph. It changes when the program changes, and never when only its documentation does.