SafeVantageVantage-Aware Memory for Reliable Embodied Decisions

Sean Hardesty Lewis1, Zuyi Guo2, Benwang Chen2, Zirui Li3, Hongyi Lin4, and Heye Huang2,†
1Massachusetts Institute of Technology 2Korea Advanced Institute of Science and Technology 3Nanyang Technological University 4Tsinghua University †Corresponding author

Abstract

Embodied agents must decide both where to look and when the evidence warrants a decision. SafeVantage is a semantic memory and active-perception policy that records the views supporting each claim. The policy uses this evidence state to select corroborating views and to output Yes, No, or Abstain.

On a held-out ProcTHOR test spanning 232 unseen houses and 7,424 paired episodes per method and action budget, SafeVantage raises macro-F1 from 0.4845 to 0.6044 and reduces travel from 7.23 to 4.94 m at eight actions, while risk falls from 0.3964 to 0.3842. At twelve actions, macro-F1 rises from 0.5855 to 0.6559, risk falls from 0.3685 to 0.3649, and travel falls from 8.42 to 7.44 m.

On a separate 96-house cohort, removing supporting-view identity, target-ray alignment, and yaw diversity lowers macro-F1 by 4.35 points at eight actions and 1.15 at twelve. Equal-input HM3D comparisons show lower selective risk, and ScanNet interventions show that answer quality decreases when supporting views are removed and recovers when they are restored.

SafeVantage links semantic claims to supporting views and selective decisions
Figure 1. Overview of SafeVantage: a vantage-aware memory retains the viewpoints supporting each claim and uses the evidence to select the next view and make the final Yes, No, or Abstain decision.

A memory for deciding what an embodied agent has seen, what remains uncertain, and which view should come next.

Method

For each scene, SafeVantage captures posed observations containing RGB, optional depth, camera pose, and time. A semantic claim is represented by its supporting views, support count, and optional 3D target estimate. This record preserves which observations justify a claim and distinguishes one weak observation from a repeated spatial pattern.

The memory separates target visibility, property or relation sufficiency, search coverage, and answer correctness. Before a provisional detection, the policy balances spatial novelty, category-room evidence, and travel. After a detection, it seeks a corroborating view from a distinct position and uses the resulting evidence state for the final decision. The selective head emits Yes, No, or Abstain under calibrated thresholds.

Viewpoint support and search coverage provide complementary evidence
Figure 2. Viewpoint support and search coverage provide complementary evidence.

Results

Across 232 unseen ProcTHOR houses, SafeVantage improves macro-F1 from 0.4845 to 0.6044 at eight actions and from 0.5855 to 0.6559 at twelve, compared with validation-selected equal-budget baselines. It achieves lower risk and higher answer rates at both budgets, with 31.7% less travel at eight actions.

One held-out episode comparing category-room prior with SafeVantage
Figure 3. One held-out episode for β€œIs there a paper towel in the room?”
Ablation of vantage-aware memory components
Figure 4. Removing vantage-aware state lowers macro-F1.

The trajectory and ablations show how retained viewpoints help turn partial detections into corroborated decisions. HM3D comparisons and ScanNet view-restoration studies further support the value of this evidence. On 233 additional unseen houses, SafeVantage improves macro-F1 by 7.3 and 4.3 percentage points at eight and twelve actions, with lower risk and travel at both budgets.

How to cite

For manuscripts, talks, and other work citing the current version, use:

Lewis, S. H., Guo, Z., Chen, B., Li, Z., Lin, H., & Huang, H. (2026).
SafeVantage: Vantage-Aware Memory for Reliable Embodied Decisions.
https://safevantage.github.io/