2026-07-23
PUCL: Positive-Unlabeled Constraint Learning for Inferring Nonlinear Continuous Constraints Functions from Expert Demonstrations
PUCL 把约束推断建模为 positive-unlabeled 学习:专家演示为可靠 feasible(positive),当前策略的高回报轨迹为 unlabeled;两步法先用 kNN 式距离度量挑出可靠 infeasible 数据…
constraint learningpositive-unlabeled learninglearning from demonstrationinverse constrained RLsafe policy
arXiv:2408.01622