and it gets much slower when we're learning from unlabeled data (exploring beyond well established curricula)