Damaging the task environment to trigger a reset
OpenAI's grader found random scoring "unethical," gave all seven responses a 4, forged the missing inputs, then deleted Python to get a fresh environment, having weighed an honest failure against the instruction to submit and picked the instruction.