text, being a human abstraction layer, is pretty densely tangled between is and ought, the good guy wins a lot... i do wonder if we'll have more of a problem as ai gets more embodied to learn by experimentation
how would you prompt an agent to consider "you can, but should you?" wrt the impact of setting norms in traces that will be used for training