首页 > AI前沿 > Extending Pathwise Gradients to Discrete Random Variables via Finite-Order Relaxation

Extending Pathwise Gradients to Discrete Random Variables via Finite-Order Relaxation

arXiv机器学习 2026-10-06 13:30 5 阅读 查看原文

Pathwise gradients are preferred for continuous random variables because they are unbiased, low variance, and work with a single sample.

For discrete variables, however, the pathwise identity cannot generally be exact for every differentiable function.

We propose a general framework

We propose a general framework to construct finite-order exact pathwise gradient estimators for a range of common discrete variables such as Poisson.

The estimator is the least-norm solution among all solutions that are unbiased for polynomials of degree at most.

The resulting estimators preserve the hard forward sample, require no temperature tuning, and can be implemented in a few lines of codes.

Unique and minimizes weight variance

Against other admissible solutions, our estimator is unique and minimizes weight variance; in contrast, prior works use categorical variables or augmented representations to approximate non-categorical variables that induces excess variance and computations.

Non-asymptotic bias bound

To understand approximation bias for functions beyond the prescribed class, we also derive a non-asymptotic bias bound.

Experiments

In experiments our low order methods match or improve tuned baselines across linear, nonlinear and hierarchical latent-variable models, while out-speeding competitors in every runtime benchmark.