Paper

Foresight: Failure Detection for Long-HorizonRobotic Manipulation with Action-ConditionedWorld Model Latents

D_research 2026. 8. 13. 23:10

Paper Info

  • Title: Foresight: Failure Detection for Long-Horizon Robotic Manipulation with Action-Conditioned World Model Latents
  • Venue: arXiv 2026
  • Keywords: Long-Horizon Robotic Manipulation, Failure Detection, Action-Conditioned World Model, World Model Latent, Runtime Monitoring, Conformal Prediction

Summary

실제 로봇이 수행하는 manipulation task는 하나의 단순한 행동으로 끝나는 경우보다 여러 개의 subgoal이 연속적으로 연결되는 long-horizon 형태가 많다. 이러한 작업에서는 작은 action error나 state deviation이 즉시 task failure로 나타나지 않고 여러 단계 동안 누적될 수 있으며, 최종적으로 failure가 명확하게 드러났을 때에는 이미 recovery가 어려운 상태일 수 있다. 따라서 long-horizon robotic manipulation에서는 최종 task 결과가 실패했는지를 확인하는 것보다, 현재 진행 중인 trajectory가 실패 방향으로 이동하고 있는지를 runtime에서 판단하는 것이 중요하다.

 

하지만 long-horizon failure detection은 일반적인 image-based anomaly detection처럼 단순하게 다루기 어렵다. 동일한 visual state라도 robot이 이전에 어떤 action을 수행했고 현재 task의 어느 단계에 있는지에 따라 의미가 달라지기 때문이다.

 

예를 들어 물체가 테이블 위에 놓여 있는 동일한 observation이라도,

1) 아직 grasp를 수행하기 전이라면 정상적인 상태일 수 있지만,

2) 이미 lift action을 수행한 뒤라면 missed grasp를 의미할 수 있으며,

3) placement action을 완료한 뒤라면 다시 정상적인 상태일 수 있다.

 

따라서 현재 image가 일반적인 성공 trajectory와 얼마나 비슷한지만으로는 long-horizon failure를 충분히 판단하기 어렵다. 중요한 것은 현재까지의 observation이 robot이 수행한 action과 task progression에 비추어 정상적인 상태인지를 판단하는 것이다.

 

기존 runtime failure detection은 다양한 signal을 활용해왔다.

 

1) Policy uncertainty나 predicted action을 이용하는 방법,

2) Policy 내부 latent representation을 이용하여 failure score를 학습하는 방법,

3) VLM을 통해 robot behavior나 visible mistake를 판단하는 방법,

4) Observation embedding이나 video world-model latent를 이용하여 정상 trajectory에서 벗어난 상태를 탐지하는 방법

 

하지만 이러한 접근은 long-horizon manipulation에서 각각 한계를 가진다.

Policy uncertainty나 policy-internal representation을 사용하는 방법은 robot policy의 logits, hidden state, token probability와 같은 내부 정보에 의존하기 때문에 특정 policy architecture에 종속될 수 있다. 서로 다른 VLA나 visuomotor policy를 사용할 경우 detector가 사용하는 representation 역시 달라질 수 있어 범용적인 runtime monitor로 적용하기 어렵다.

반면 VLM이나 observation embedding 기반 방법은 현재 robot이 어떤 장면에 놓여 있는지는 판단할 수 있지만, 해당 상태가 어떤 action의 결과로 나타난 것인지를 충분히 고려하지 못한다. 특히 long-horizon task에서는 현재 observation만 보면 정상처럼 보이더라도 이전 action을 고려하면 이미 잘못된 progression일 수 있다.

최근 world model을 failure detection에 활용하는 연구도 등장하고 있지만, 일반적인 video world model latent는 주로 observation sequence 자체를 표현한다. 따라서 현재 scene이 어떤 상태인지에 대한 정보는 포함할 수 있지만, 현재 상태에서 robot이 수행하려는 action에 따라 앞으로 어떤 변화가 발생해야 하는지를 직접 반영하는 데에는 제한적이다.

 

논문은 여기서 long-horizon failure detection에서 중요한 것은 단순한 visual abnormality가 아니라 action과 observation progression 사이의 관계라고 본다. 작은 deviation이 발생하더라도 현재 action에 맞는 정상적인 trajectory를 유지하고 있다면 failure가 아닐 수 있고, 반대로 시각적으로는 자연스러운 장면이라도 robot action이 의도한 progression과 맞지 않는다면 failure의 초기 신호일 수 있다.

 

또 다른 문제는 정확한 failure onset을 annotation하기 어렵다는 것이다. 수백에서 수천 timestep으로 구성된 trajectory에서는 작은 deviation이 어느 시점부터 실제 failure로 이어졌는지를 명확하게 정의하기 어렵다. 따라서 각 timestep에 failure 여부를 직접 annotation하지 않고, 최종 task success/failure와 같은 trajectory-level label만으로도 runtime detector를 학습할 수 있는 방법이 필요하다.

 

즉 기존 방법의 핵심 한계는 다음과 같이 정리할 수 있다.

Observation 자체만으로는 동일한 상태가 action과 task stage에 따라 달라지는 의미를 구분하기 어렵고, policy 내부 signal에 의존하는 방법은 특정 policy에 종속되며, long-horizon task에서는 정확한 failure 발생 시점을 annotation하기도 어렵다.

 

따라서 논문이 제기하는 핵심 문제는 다음과 같다.

현재 observation과 policy가 수행하려는 action을 함께 고려하여 long-horizon trajectory의 정상적인 progression을 판단하고, 정확한 failure timestamp 없이도 runtime failure를 탐지할 수 있는가?

 

논문은 이를 해결하기 위해 Foresight를 제안한다.

Figure 1

 

Foresight는 현재 observation과 policy가 생성한 action을 action-conditioned world model에 입력하고, 이를 통해 생성된 latent representation을 이용하여 long-horizon trajectory의 failure score를 추정하는 runtime failure detection framework이다.

 

Foresight의 핵심 가정은 단순한 visual representation보다 현재 상태에서 예정된 action을 수행했을 때 예상되는 future state를 표현하는 latent가 failure detection에 더 적합하다는 것이다. 이를 위해 Foresight는 V-JEPA 2-AC를 action-conditioned world model backbone으로 사용한다. 먼저 pretrained visual encoder가 현재 observation context를 latent representation으로 변환하고, action-conditioned predictor가 해당 latent와 policy가 생성한 future action chunk를 이용하여 앞으로 나타날 latent state를 예측한다. 이 과정에서 두 종류의 representation이 만들어진다.

 

1) Hidden latent: 현재 observation 자체를 표현하는 latent,

2) Predicted latent: 현재 observation에서 예정된 action을 수행했을 때 world model이 예상하는 future state를 표현하는 latent

 

Foresight는 주로 action-conditioned predicted latent를 failure detector의 입력으로 사용한다. 단순히 현재 image에 어떤 물체가 존재하는지를 표현하는 것이 아니라, 현재 상태와 robot action을 함께 고려한 execution-aware representation을 얻기 위한 것이다.

그러나 long-horizon failure는 하나의 timestep에서 발생하는 isolated anomaly라기보다 여러 timestep에 걸쳐 누적되는 trajectory-level pattern일 수 있다. 따라서 Foresight는 각 predicted latent를 독립적으로 판단하지 않고 지금까지 생성된 latent sequence를 causal sequence model에 입력한다.

 

논문은 MLP, LSTM, Causal Transformer의 세 detector를 비교한다. Causal Transformer는 현재 timestep까지의 latent sequence만 사용하여 failure score를 계산하기 때문에 실제 runtime에서도 미래 observation을 사용하지 않는다. 이를 통해 state transition, task progression, accumulated deviation과 같이 여러 timestep에 걸쳐 나타나는 failure pattern을 학습한다.

Detector 학습에는 정확한 failure 발생 timestep이 아니라 rollout 전체의 최종 success/failure label만 사용한다. 즉 사람이 “몇 번째 timestep에서 failure가 시작되었는가”를 annotation하지 않아도, successful trajectory와 failed trajectory의 latent sequence 차이로부터 timestep별 failure score를 학습한다.

 

Runtime에서는 detector가 출력한 failure score를 하나의 고정 threshold와 비교하지 않는다. Long-horizon task에서는 task stage에 따라 정상적인 score 수준이 달라질 수 있기 때문이다. 이를 위해 Foresight는 Functional Conformal Prediction(FCP)을 사용하여 successful calibration rollout에서 시간에 따른 정상 score 범위를 추정하고, timestep별 time-varying threshold를 생성한다. Failure score가 해당 threshold를 처음 초과하면 failure alarm을 발생시킨다. 따라서 task 초기와 후반의 정상적인 uncertainty 수준이 서로 다르더라도 동일한 threshold를 강제로 적용하지 않고, successful trajectory에서 관측되는 temporal score distribution을 기준으로 failure를 판단한다.

 

실험은 LIBERO-Long, ManiSkill-Long, BEHAVIOR-1K의 세 simulation benchmark와 ReactorX 및 Franka를 이용한 real-world task에서 수행한다. LIBERO-Long은 평균 약 253 simulation step의 tabletop manipulation을 포함하고, ManiSkill-Long은 평균 약 1,484 step이 필요한 multi-stage manipulation을 포함한다. BEHAVIOR-1K는 navigation과 manipulation이 함께 필요한 household task로 successful rollout이 평균 8,557 simulation step에 이르며, 가장 긴 task는 평균 13,657 step을 요구한다.

또한 OpenVLA, π0-FAST, π0.5, ACT, SmolVLA, GR00T N1.5 등 서로 다른 policy와 ReactorX, Franka, mobile manipulator 등 다양한 embodiment에서 평가하여 특정 policy나 robot에만 적용되는 detector인지 확인한다.

 

Foresight-Transformer는 simulation에서 LIBERO-Long 0.94, ManiSkill-Long 0.80, BEHAVIOR-1K 0.78의 balanced accuracy를 기록하며 기존 failure detection 방법보다 높은 성능을 보인다. 특히 가장 긴 BEHAVIOR-1K에서는 기존 최고 baseline의 balanced accuracy 0.64보다 높은 0.78을 기록하여, action-conditioned world-model latent가 긴 trajectory에서 효과적인 failure representation이 될 수 있음을 보여준다. Real-world에서도 ReactorX/ACT에서 ROC-AUC 0.93, ReactorX/π0.5에서 0.87, Franka/GR00T N1.5에서 0.89를 기록하며 네 가지 setting 중 세 setting에서 가장 높은 성능을 보인다. 또한 MLP detector는 일부 real-world task에서 ROC-AUC 0.50~0.59 수준에 머무는 반면 LSTM과 Transformer는 더 높은 성능을 보이며, long-horizon failure detection에서 단일 timestep보다 trajectory sequence를 함께 모델링하는 것이 중요함을 보여준다.

 

결과적으로 Foresight의 목적은 현재 image가 정상적인지 여부만 판단하는 것이 아니라, 현재 observation과 robot action을 이용해 예상되는 future dynamics를 latent space에서 표현하고, 해당 representation의 temporal progression을 이용하여 long-horizon task failure를 탐지하는 것이다.

논문의 기여는 action-conditioned world-model latent를 이용한 policy-interface-agnostic failure detection framework, 정확한 failure timestamp 없이 trajectory-level success/failure label만으로 학습하는 causal sequence detector, 그리고 functional conformal prediction을 이용한 time-varying runtime threshold로 정리할 수 있다.

Critic

1. Expected Response와 Actual Response의 직접적인 비교 부재

Foresight는 action-conditioned world model을 사용하지만, world model이 예측한 future latent와 이후 실제로 관측된 latent 사이의 residual을 직접 failure signal로 사용하지 않는다. Action이 representation에 반영되기는 하지만 “현재 action에 대해 실제 robot response가 예상과 얼마나 다른가”라는 expected-actual consistency 자체를 명시적으로 모델링하는 구조는 아니다.

 

2. Physical Sensor Response 미활용

Foresight는 visual observation과 policy action을 중심으로 failure를 판단하며 joint velocity, acceleration, force/torque, tactile signal과 같은 physical sensor response는 직접 사용하지 않는다. 따라서 mass, friction, COM, actuator condition과 같이 image보다 robot의 물리적 반응에서 먼저 나타나는 변화는 조기에 포착하기 어려울 수 있다.

 

3. Trajectory-Level Label에 따른 Label Ambiguity

Foresight는 정확한 failure timestamp 없이 최종 success/failure label만 이용할 수 있다는 장점이 있지만, failed rollout의 초반 정상 구간까지 동일한 trajectory-level failure supervision의 영향을 받는다. 실제 failure precursor가 후반에 발생하더라도 초기 정상 trajectory와 명확하게 분리되지 않기 때문에 weak supervision에 따른 label noise가 존재한다.

 

4. Failure Detection과 Early Prediction의 차이

논문은 failure가 task 종료 전에 탐지될 수 있음을 qualitative example로 보여주지만, 주요 정량 평가는 rollout-level ROC-AUC와 balanced accuracy를 중심으로 이루어진다. 이러한 지표는 successful rollout과 failed rollout을 얼마나 잘 구분하는지는 평가하지만, 실제 failure보다 얼마나 이른 시점에 warning을 발생시키는지는 직접 측정하지 않는다. 따라서 early failure prediction 자체에 대한 정량적 검증은 제한적이다.

 

5. World Model Error와 Failure Signal의 혼재

Foresight의 detector는 action-conditioned world model이 생성한 latent representation에 의존한다. 따라서 새로운 object, visual condition 또는 dynamics에서 world model 자체의 prediction이 부정확해질 경우 robot execution이 정상적이어도 latent가 달라질 수 있다. 실제 failure와 world-model prediction error를 별도로 구분하는 mechanism은 제공하지 않는다.

 

6. 완전한 Cross-Policy Generalization의 한계

Foresight는 policy logits나 hidden state를 사용하지 않기 때문에 interface 측면에서는 policy-independent하지만, 학습된 detector가 새로운 policy에도 동일하게 generalize하는 것은 아니다. 실제 실험에서도 π0.5에서 ACT로의 transfer는 높은 성능을 보이는 반면 ACT에서 π0.5로의 transfer는 크게 감소한다. Training policy가 target policy의 recovery pattern이나 다양한 behavior를 포함하지 못하면 추가적인 학습이 필요할 수 있다.

 

7. Action-Conditioned Predictor 재학습 필요

Pretrained V-JEPA 2 visual encoder는 frozen 상태로 사용하지만 action-conditioned predictor는 benchmark의 robot rollout을 이용하여 별도로 학습한다. 따라서 새로운 embodiment나 action space에 적용할 때 pretrained world model을 그대로 사용하는 것은 아니며, 해당 robot의 action-conditioned dynamics를 학습하기 위한 추가 데이터와 training cost가 필요하다.

 

8. 높은 World Model 계산 비용

Foresight의 failure detector 자체는 매우 가볍지만 world-model feature extraction이 전체 연산량의 대부분을 차지한다. V-JEPA 2-AC 기반 feature extractor는 약 1.3B parameter 규모이며 H200 GPU에서도 한 번의 inference에 약 183 ms가 필요하다. Action chunk 단위 monitoring에서는 사용할 수 있지만 빠른 closed-loop control이나 contact-sensitive manipulation에서는 runtime overhead가 문제가 될 수 있다.

 

9. Calibration Distribution 의존성

Functional conformal prediction은 successful calibration rollout에서 정상 score distribution을 추정하여 threshold를 설정한다. 따라서 deployment 환경의 object, visual condition, robot dynamics 또는 task distribution이 calibration data와 크게 달라지면 기존 threshold의 false-positive control이 유지되지 않을 수 있다. 논문 역시 conformal guarantee가 calibration distribution과 deployment condition의 일치에 의존한다고 명시한다.

 

10. Failure 원인에 대한 설명 부족

Foresight는 현재 trajectory가 failure로 진행되고 있는지를 scalar failure score로 제공하지만, 해당 failure가 grasp error, object slip, incorrect action, navigation error 또는 environment dynamics 변화 중 무엇 때문인지는 직접 구분하지 않는다. 따라서 실제 robot recovery까지 연결하려면 failure detection 이후 원인을 판단하는 localization이나 diagnosis 과정이 추가적으로 필요하다.