OpenAI says it has canceled plans to release its updated GPT-6.1 model next month as it continues to investigate what testing shows as a regression in terms of security compared to previous models.
The move, first reported by The Wall Street Journal on Monday night and later confirmed in OpenAI statements to the press, reflects what OpenAI Head of Security Systems Saachi Jain said was a “tradeoff” between performance and security seen when testing the now-discarded model. Jain said GPT-6.1 was better than previous models at taking difficult tasks to completion without human intervention. But the model was also more likely to fail tests related to alignment (i.e., staying within the limits set by the human creators) and was more willing to use sometimes “unsafe” tools and services to move forward with a task. It was also more likely to try to mislead end users about actions it did or did not take, Jain said.
Last week, OpenAI said it would stop training its “most capable models” following an incident in which a model attempted to bypass internet access restrictions. GPT-6.1 was not among the “most capable models” covered by that move, OpenAI told the WSJ. And while GPT-6.1 won’t be released as is, the company said it intends to use the same base model for additional training runs that it said will hopefully lead to future GPT-6 generation models.