『The Optimizer Becomes a Policy: Inside RP-1's Learned Planner for Robot World Models』のカバーアート

The Optimizer Becomes a Policy: Inside RP-1's Learned Planner for Robot World Models

The Optimizer Becomes a Policy: Inside RP-1's Learned Planner for Robot World Models

無料で聴く

ポッドキャストの詳細を見る
Pantheon Industries introduces Reinforced Planning (RP-1), the first fully learned planner that iteratively improves robot action plans from scratch using Reinforcement Learning over a frozen World Model. RP-1 reimagines planning as a learned policy: starting with a candidate action sequence, it uses a frozen World Model to imagine outcomes, a learned critic to evaluate closeness to goal (value-based scoring instead of latent distance), and a neural planner to revise the sequence over multiple iterations. This replaces hand-designed search heuristics (CEM, MPPI, Adam) with learned, reusable plan-improvement rules. Key results across ThreeRoom, OGBench, and Reacher benchmarks: RP-1 beats SOTA latent methods on 48 of 48 head-to-head comparisons, achieves 2.2x higher success rate on hard and long-horizon tasks, is 67x faster at scale (50 robot arms on H200 GPUs), and requires 1000x fewer World Model queries than MPPI/CEM.
adbl_web_anon_alc_button_suppression_t1
まだレビューはありません