Of course, we cannot rely on an LLM to have the same sense of proportionality (or any inherent sense of concern about legal implications) when answering a question. Without explicit instructions about what is allowed or justified and what is not, a well-resourced AI agent will try every possible avenue to satisfy the user’s request to the best of its ability.
That red line seems more like a red suggestion to me…
Credit: Getty Images
That red line seems more like a red suggestion to me…
Credit: Getty Images
OpenAI says the internal testing in this case was conducted “without the full set of security measures used in our publicly available products.” Given that lack of constraints, one could say that the agent was working as intended, in a sense, using all the tools available to generate a response to the message.
At the same time, OpenAI says that the agent in the test “was required to answer these questions using publicly published statistics” and “took actions that we had not authorized it to take” to obtain that information. From the outside, it is difficult to know how strong OpenAI’s attempts to deny “clearance” were in practice. It’s plausible that the OpenAI agent ignored a relatively simple “anti-hacking” directive in its system message in order to give a complete response that would satisfy a direct message from the user, for example.
In public analyzes of multiple “misalignment” incidents published earlier this month, OpenAI identified multiple cases of “reward hacking,” where an agent resorted to extreme methods to generate a better response to a user’s request. The company said it had recently taken steps to prevent this type of reward hacking by adding explicit punishments for misaligned behavior to the system’s reward function.
In retrospect, it’s hard to see why those types of protections weren’t in place in June and whether they could have prevented a potential international incident in this case.