Unaccelerated

A Free Monitor for Accelerated Diffusion Policies
AirLab · The Robotics Institute · Carnegie Mellon University
Controlled simulation from our film. Same initial state, disturbance, reference schedule, and playback speed. How to read this demo
A plan can outlive the call that made it.
Let its remaining actions check what comes next.
01 / THE PROBLEM
An accelerator can pass its own similarity gate and still break the task. The unexpired reference plan gives us a different question: does the action still agree with the plan?



Original frames and measurements from Figure 1. Both accelerated conditions pass their own gate; one shared plan-deviation threshold, θ = 0.15, separates them.
Action-chunking diffusion policies predict more actions than they immediately execute. After a full reference call, those unexecuted actions remain available. We use them to monitor accelerated calls at matching future action indices, with no additional reference forward pass for the check.
The monitor combines normalized action deviation with a near-repetition check. A veto executes the corresponding stored reference slice and permanently demotes the accelerator for the episode. Candidates are selected using the tail of their on-policy deviation distribution, without task-outcome labels.
The paper studies 114 accelerator configurations across four policy families, then tests a frozen design on a fifth, held-out insertion setting. The results quantify effective compute savings, paired success-rate differences, and the limits of monitoring with an aging reference.
02 / THE METHOD
Read the unused part of a reference plan. Check two failure modes. Make a veto change the rest of the episode.
The next executed actions must be compared with the same future indices of the saved plan.
Read actions 2 and 3 from the saved reference. No new reference forward pass is needed for this check.
Small deviation δ and no near repetition: let the accelerated action through.
Conceptual curves, following Figure 2b. Near repetition compares full emitted chunks over the previous eight control calls; its threshold is calibrated from reference replay.
On a veto, execute the saved reference slice and step down to a more conservative candidate. The demotion is never undone within the episode.
Visit its own states. Measure the 95th-percentile deviation. Select the cheapest passing candidate.
114 configurations studied · θ = 0.15 · no task-outcome labels used for selection03 / THE EVIDENCE
We report compute savings together with success-rate changes. Uncertainty stays in the picture.
Diffusion Policy on ManiSkill3 insertion.
6,000 paired episodes · 71.3% reference success.
−0.2 ± 1.2 pp success-rate change
Change relative to the unaccelerated reference, in percentage points (pp). Shaded band: ±2 pp equivalence margin. The monitored arm's 95% paired interval lies inside the band.
Table II · 1,500 paired episodes per family · 95% paired intervals
| Policy / benchmark | Reference success | Success-rate change | Compute speedup | Within ±2 pp? |
|---|---|---|---|---|
| Cosmos LIBERO | 98.2% | +0.1 ± 1.1 pp | 1.55× | Yes |
| π0.5 LIBERO-10 | 97.0% | −0.2 ± 0.9 pp | 1.67× | Yes |
| Diffusion Policy PushT | 58.0% | +1.2 ± 2.6 pp | 1.75× | Undecided |
| 3D Diffusion Policy MetaWorld | 94.2% | +0.1 ± 1.4 pp | 3.09× | Yes |
What the comparison establishes. On three of these four families, the same reference schedule without monitoring is already within the margin. The held-out insertion task above provides the stronger comparison.
What remains uncertain. PushT's interval is too wide to establish equivalence. Compute speedup includes reference calls; it does not measure task duration, peak latency, or energy.
All closed-loop evaluations reported in the paper are in simulation. A reference plan uses an older observation; a small persistent bias may remain below the per-call threshold. Guarding every call does not guarantee zero errors or distribution-free safety.
The first four policy families informed the design. Only the fifth family was held out after the design was frozen. Reference calls are refreshed at m = ⌊H / Hexec⌋ to preserve coverage; delaying refresh creates uncovered calls and changes the compute tradeoff.
04 / THE FILM · 1080p
From the physical motivation
to the monitor and its evidence.
About the moving comparison. This controlled MuJoCo illustration uses synthetic action streams and choreographed object attachment, with object contact disabled. The repository monitor drives the monitored arm; release positions follow the simulated end effector. This illustration exercises deviation-triggered fallback and persistent demotion; the near-repetition veto is disabled here. These two trials illustrate the mechanism and are not benchmark success-rate measurements.
About the real-robot footage. Shared hardware excerpts provide context and do not show deployment of our method. Original footage, authors, licenses, and source playback speeds are documented in sources and acknowledgments .
THE UNEXPIRED PLAN

@misc{zhao2026unexpiredplan,
title = {The Unexpired Plan: A Free Monitor for Accelerated Diffusion Policies},
author = {Yi Zhao and Sebastian Scherer},
year = {2026},
eprint = {2610.05747},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
doi = {10.48550/arXiv.2610.05747},
url = {https://arxiv.org/abs/2610.05747}
}arXiv preprint · arXiv:2610.05747 [cs.RO] · Read the PDF ↗