Abstract
Embodied agents must decide both where to look and when the evidence warrants a decision. SafeVantage is a semantic memory and active-perception policy that records the views supporting each claim. The policy uses this evidence state to select corroborating views and to output Yes, No, or Abstain.
On a held-out ProcTHOR test spanning 232 unseen houses and 7,424 paired episodes per method and action budget, SafeVantage raises macro-F1 from 0.4845 to 0.6044 and reduces travel from 7.23 to 4.94 m at eight actions, while risk falls from 0.3964 to 0.3842. At twelve actions, macro-F1 rises from 0.5855 to 0.6559, risk falls from 0.3685 to 0.3649, and travel falls from 8.42 to 7.44 m.
On a separate 96-house cohort, removing supporting-view identity, target-ray alignment, and yaw diversity lowers macro-F1 by 4.35 points at eight actions and 1.15 at twelve. Equal-input HM3D comparisons show lower selective risk, and ScanNet interventions show that answer quality decreases when supporting views are removed and recovers when they are restored.
Method
For each scene, SafeVantage captures posed observations containing RGB, optional depth, camera pose, and time. A semantic claim is represented by its supporting views, support count, and optional 3D target estimate. This record preserves which observations justify a claim and distinguishes one weak observation from a repeated spatial pattern.
The memory separates target visibility, property or relation sufficiency, search coverage, and answer correctness. Before a provisional detection, the policy balances spatial novelty, category-room evidence, and travel. After a detection, it seeks a corroborating view from a distinct position and uses the resulting evidence state for the final decision. The selective head emits Yes, No, or Abstain under calibrated thresholds.
Results
Across 232 unseen ProcTHOR houses, SafeVantage improves macro-F1 from 0.4845 to 0.6044 at eight actions and from 0.5855 to 0.6559 at twelve, compared with validation-selected equal-budget baselines. It achieves lower risk and higher answer rates at both budgets, with 31.7% less travel at eight actions.


The trajectory and ablations show how retained viewpoints help turn partial detections into corroborated decisions. HM3D comparisons and ScanNet view-restoration studies further support the value of this evidence. On 233 additional unseen houses, SafeVantage improves macro-F1 by 7.3 and 4.3 percentage points at eight and twelve actions, with lower risk and travel at both budgets.
How to cite
For manuscripts, talks, and other work citing the current version, use:
Lewis, S. H., Guo, Z., Chen, B., Li, Z., Lin, H., & Huang, H. (2026).
SafeVantage: Vantage-Aware Memory for Reliable Embodied Decisions.
https://safevantage.github.io/