Back to archive
#ai#papers#aigen#robotics#training

Diffusion policy

A robot needs to avoid a person. Both the left and right routes may be good, but their average would lead straight into the obstacle. You need a model capable of proposing different coherent behaviors, rather than just one averaged movement.

A diffusion policy determines actions by gradually transforming a noise sample into a plan conditioned on observations. A policy is a way of choosing actions. Successive denoising steps can produce one of several possible sequences.

In the paper on motion planning in crowds, §3, the result is a short list of future motion commands. The robot can receive a coherent turn, rather than a mixture of conflicting intentions.

This describes a modeling capability, rather than guaranteeing a safe decision. Several denoising steps cost more computation, and quality depends on the data and training objective. An action chunk describes the generated action sequence independently of the method that creates it.