Preprint · 2026

Belief-Informed Hybrid Controlwith Almost-Sure Target-Set Convergence

Clinton Enwerem, Saleh Kemal, John S. Baras, and Calin Belta

Institute for Systems Research, University of Maryland, College Park

Under parameter uncertainty, an informative control action may temporarily drive a system away from its target. We combine belief-space receding-horizon action selection with a bound on cumulative slack in an expected-decrease constraint, allowing these deviations while preserving target-set convergence under the stated feasibility and prediction conditions.

Bimanual Assembly Demonstration

From Acquisition to Release. A Drake simulation with overview and insertion views compares belief-informed one-step selection with reference tracking. Motion plays at 2.25× speed. The closing card shows the manuscript’s endpoint comparison, including dashed outlines of the desired part pose.

Belief-Space Receding-Horizon Action Selection

We consider a hybrid system with continuous flows, discrete mode transitions, and an unknown parameter that remains fixed during an execution. Each candidate feedback action specifies a feedback law and a stopping rule. We maintain a probability distribution over the unknown parameter and update this belief from the observed action outcome.

  1. Predict Progress

    For each candidate action, we predict the expected endpoint value of a nonnegative target-set progress function using the current parameter belief.

  2. Enforce Admissibility

    An action must satisfy an expected-decrease constraint with nonnegative slack. A scalar controller state tracks the remaining slack so its cumulative sum stays bounded.

  3. Anticipate Information

    In two-step selection, we update the predicted belief for each possible observation before evaluating the next feedback action. We execute the first action, observe its outcome, and replan.

Almost-Sure Target-Set Convergence

Under distance-comparison bounds, correct conditional prediction, and recursive feasibility, we prove almost-sure convergence to the target set at decision times. We also bound the sum of expected progress-function values and the expected time to enter a target neighborhood. Upper confidence bounds extend the analysis to bounded model samples with summable error probabilities.

The dual-control connection is explicit: the predicted observation changes the parameter belief, which changes both the cost and admissibility of the subsequent feedback action.

Learning an Unknown Control Direction

Planar trajectories from the initial state (1, 0) toward the origin: a test input distinguishes the two unknown control directions before contracting feedback follows

In a two-dimensional regulation problem, the unknown input sign determines which feedback gain contracts the state toward the origin. Two-step selection first applies a test input to identify this sign, then applies the corresponding contracting feedback. We verify recursive feasibility along both resulting trajectories; the state norm falls below 0.01 within 11 decisions for either sign. A one-step selector instead applies zero input and leaves no admissible action at the next decision.

Planar Control with an Unknown Input Sign. The axes are the state coordinates x1 and x2; the star marks the target at the origin and the black dot marks the initial state. Blue solid and green dashed curves show two-step selection for control-direction signs θ = +1 and θ = −1, respectively. The first test interval reveals the sign through a change in radius, after which the selected feedback contracts both trajectories toward the target. The dotted circle shows the constant-radius motion produced by the rotational drift with zero input.

One-Step Selection

First action selected over prior probability and normalized cumulative slack bound using one-step selection

Two-Step Selection

First action selected over the same grid using posterior-conditioned two-step selection
Legend for hold, test, negative and positive feedback gains, empty admissible set, and no feasible second decision
Information Changes the First Action. The maps use the same plant, feedback gains, initial state, and expected-decrease constraint. Across 399 initial prior–slack pairs, the selectors differ at 55 points. At 40, two-step selection replaces hold with the test input. At the remaining 15, it identifies that no admissible first action has a feasible second decision. Stars mark the nominal case: prior probability 0.5 and normalized cumulative slack bound 0.12.

Feedback Selection in Bimanual Assembly

We instantiate one-step selection on a RealHand A7 Black bimanual robot with two L6 hands in Drake. Physical finger contacts support the acquisition and transport of two free-body parts. Following insertion, selected clearance feedback drives hand withdrawal while preserving the assembled part pose.

Belief-Informed

OverviewReverse oblique view of the assembled part with the hands withdrawn under belief-informed control
Insertion detailInsertion detail under belief-informed control, closely aligned with the dashed desired-pose outline

0.188 mm position error

Reference Tracking

OverviewReverse oblique view of the reference-tracking endpoint with residual hand–housing contact
Insertion detailReference-tracking insertion detail showing displacement from the dashed desired-pose outline

5.20 mm position error

Manuscript Endpoint Comparison. At nominal friction 0.7, both methods start from the same state. We evaluate them at our method’s completion time of 30.96 s. Its infinity-norm relative-position error is 96.4% lower than the reference-tracking baseline’s error. The insertion-detail views use a common spatial scale; dashed outlines mark the desired part pose.
Manuscript outcomes at 30.96 s. Mint denotes belief-informed selection; salmon denotes reference tracking.
QuantityBelief-informedReference tracking
Infinity-norm relative-position error0.188 mm5.20 mm
Minimum hand–assembly clearance7.14 mm0 mm
Summed hand–housing contact-force magnitude0 N0.704 N

Effect of Belief Refinement

At lower friction, 0.35, conditioning on acquisition observations produces an empty admissible set. With the prior held fixed, the alternative selector admits an action that violates the conditional expected-decrease constraint. This comparison shows how posterior conditioning changes action admissibility; the planar example establishes recursive feasibility and demonstrates informative two-step lookahead.

BibTeX

@misc{enwerem2026beliefinformedhybrid,
  title  = {Belief-Informed Hybrid Control with Almost-Sure Target-Set Convergence},
  author = {Enwerem, Clinton and Kemal, Saleh and Baras, John S. and Belta, Calin},
  year   = {2026},
  url    = {https://clintonenwerem.com/belief-hybrid-control/}
}