Extension of constraint-procedural logic-generated environments for deep Q-learning agent training and benchmarking