This study designs an efficient and equitable humanitarian supply chain dynamically by using reinforcement learning, PPO, and compared with heuristic algorithms. This study demonstrates the model of PPO always treats average satisfaction rate as the priority.
Related papers
Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).