AirLab

The Unexpired Plan.

A Free Monitor for Accelerated Diffusion Policies

Yi ZhaoSebastian Scherer

AirLab · The Robotics Institute · Carnegie Mellon University

ONE SCENE. ONE DISTURBANCE.A DIFFERENT RESPONSE.
Monitor offSame reference schedule
Target missed
Unexpired PlanReference plan stored
Target reached
00:00 / 00:13

Controlled simulation from our film. Same initial state, disturbance, reference schedule, and playback speed. How to read this demo

A plan can outlive the call that made it.
Let its remaining actions check what comes next.

↓

01 / THE PROBLEM

Agreement can
hide failure.

An accelerator can pass its own similarity gate and still break the task. The unexpired reference plan gives us a different question: does the action still agree with the plan?

One task, three conditions

REFERENCE

Unaccelerated

Unaccelerated reference: successful final task state from Figure 1
150/150task successes
The reference behavior97.0% across all ten tasks
ACCELERATED

Benign shortcut

Benign acceleration: successful final task state from Figure 1
150/150task successes
Own gate accepts 0.995 ≥ 0.97Plan deviation 0.014 → passes
ACCELERATED

Harmful shortcut

Harmful acceleration: failed final task state from Figure 1
0/150task successes
Own gate accepts 0.843 ≥ 0.80Plan deviation 0.502 → vetoes

Original frames and measurements from Figure 1. Both accelerated conditions pass their own gate; one shared plan-deviation threshold, θ = 0.15, separates them.

Read the research overview

Action-chunking diffusion policies predict more actions than they immediately execute. After a full reference call, those unexecuted actions remain available. We use them to monitor accelerated calls at matching future action indices, with no additional reference forward pass for the check.

The monitor combines normalized action deviation with a near-repetition check. A veto executes the corresponding stored reference slice and permanently demotes the accelerator for the episode. Candidates are selected using the tail of their on-policy deviation distribution, without task-outcome labels.

The paper studies 114 accelerator configurations across four policy families, then tests a frozen design on a fifth, held-out insertion setting. The results quantify effective compute savings, paired success-rate differences, and the limits of monitoring with an aging reference.

02 / THE METHOD

Already computed.
Still useful.

Read the unused part of a reference plan. Check two failure modes. Make a veto change the rest of the episode.

01

Align the actions in time.

The next executed actions must be compared with the same future indices of the saved plan.

INTERACTIVE SCHEMATIC
Saved referenceH = 8 · Hexec = 2 · m = 4
Accelerated call 1Compare indices 2–3
ExecutedUnexpiredCompared now
1

Read actions 2 and 3 from the saved reference. No new reference forward pass is needed for this check.

02

Check drift. Check repetition.

Schematic action curves: reference and accelerated proposal agree
Stored referenceAccelerator

Small deviation δ and no near repetition: let the accelerated action through.

Conceptual curves, following Figure 2b. Near repetition compares full emitted chunks over the previous eight control calls; its threshold is calibrated from reference replay.

03

Remember the veto.

Fast accelerator
More conservative
Reference

On a veto, execute the saved reference slice and step down to a more conservative candidate. The demotion is never undone within the episode.

Veto fallback0 extra reference calls
CHOOSING AN ACCELERATOR

Visit its own states. Measure the 95th-percentile deviation. Select the cheapest passing candidate.

114 configurations studied · θ = 0.15 · no task-outcome labels used for selection

03 / THE EVIDENCE

Less computation.
Measured against the task.

We report compute savings together with success-rate changes. Uncertainty stays in the picture.

THE HELD-OUT TEST · TABLE IV

A fifth family.
A frozen design.

Diffusion Policy on ManiSkill3 insertion.
6,000 paired episodes · 71.3% reference success.

1.96×effective compute speedup

−0.2 ± 1.2 pp success-rate change

Change relative to the unaccelerated reference, in percentage points (pp). Shaded band: ±2 pp equivalence margin. The monitored arm's 95% paired interval lies inside the band.

Across the original four families

Table II · 1,500 paired episodes per family · 95% paired intervals

Policy / benchmarkReference successSuccess-rate changeCompute speedupWithin ±2 pp?
Cosmos LIBERO98.2%+0.1 ± 1.1 pp1.55×Yes
π0.5 LIBERO-1097.0%−0.2 ± 0.9 pp1.67×Yes
Diffusion Policy PushT58.0%+1.2 ± 2.6 pp1.75×Undecided
3D Diffusion Policy MetaWorld94.2%+0.1 ± 1.4 pp3.09×Yes

What the comparison establishes. On three of these four families, the same reference schedule without monitoring is already within the margin. The held-out insertion task above provides the stronger comparison.

What remains uncertain. PushT's interval is too wide to establish equivalence. Compute speedup includes reference calls; it does not measure task duration, peak latency, or energy.

Scope and limitations

All closed-loop evaluations reported in the paper are in simulation. A reference plan uses an older observation; a small persistent bias may remain below the per-call threshold. Guarding every call does not guarantee zero errors or distribution-free safety.

The first four policy families informed the design. Only the fifth family was held out after the design was frozen. Reference calls are refreshed at m = ⌊H / Hexec⌋ to preserve coverage; delaying refresh creates uncovered calls and changes the compute tradeoff.

04 / THE FILM · 1080p

See the full story.

From the physical motivation
to the monitor and its evidence.

About the moving comparison. This controlled MuJoCo illustration uses synthetic action streams and choreographed object attachment, with object contact disabled. The repository monitor drives the monitored arm; release positions follow the simulated end effector. This illustration exercises deviation-triggered fallback and persistent demotion; the near-repetition veto is disabled here. These two trials illustrate the mechanism and are not benchmark success-rate measurements.

About the real-robot footage. Shared hardware excerpts provide context and do not show deployment of our method. Original footage, authors, licenses, and source playback speeds are documented in sources and acknowledgments .

THE UNEXPIRED PLAN

Read the plan.
Remember the veto.

AirLab, Carnegie Mellon University

Reference this work

@misc{zhao2026unexpiredplan,
  title         = {The Unexpired Plan: A Free Monitor for Accelerated Diffusion Policies},
  author        = {Yi Zhao and Sebastian Scherer},
  year          = {2026},
  eprint        = {2610.05747},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO},
  doi           = {10.48550/arXiv.2610.05747},
  url           = {https://arxiv.org/abs/2610.05747}
}

arXiv preprint · arXiv:2610.05747 [cs.RO] · Read the PDF ↗

Original paper figure

Original author-provided figure. Pinch or scroll to inspect details. Open vector figure ↗