AT2k Design BBS Message Area
Casually read the BBS message area using an easy to use interface. Messages are categorized exactly like they are on the BBS. You may post new messages or reply to existing messages!

You are not logged in. Login here for full access privileges.

Previous Message | Next Message | Back to Slashdot  <--  <--- Return to Home Page
   Local Database  Slashdot   [58 / 103] RSS
 From   To   Subject   Date/Time 
Message   VRSS    All   OpenAI Admits Six More Instances of AI Models Acting Deceptively   September 17, 2026
 2:20 AM  

Feed: Slashdot
Feed Link: https://slashdot.org/
---

Title: OpenAI Admits Six More Instances of AI Models Acting Deceptively

Link: https://slashdot.org/story/26/09/17/0641223/o...

OpenAI announced Wednesday that "We do not believe that the AI industry has
solved alignment and monitoring to a sufficient degree to continue
responsibly scaling at maximum speed for much longer." But along with the
announcement, OpenAI announced it "found additional incidents of AI models
acting deceptively and taking unsanctioned actions during training," reports
CNN. And they add that OpenAI is also "introducing a new process for the
company to publicly report such instances." Under the new system, OpenAI will
share updates on concerning AI behavior more frequently instead of waiting to
bundle multiple instances into one report. The company said it wants to share
more information about troubling AI behavior in the absence of an industry-
wide standard... "As AI systems grow more advanced and more widely deployed,
we need to build a broader and better-informed consensus on the progress of
alignment research," OpenAI wrote in a blog post Wednesday... OpenAI said it
observed "misaligned behavior" when training and evaluating AI models in six
circumstances in the last six months... In one rare instance, OpenAI said an
unreleased research model added "jailbreak-like instructions" to the
summaries it uses to preserve context in long-running tasks that said it was
"freed from the roles and identities that bind other chatbots." Separately,
the company said some instances of its 5.6 Sol model included directives to
invent information to conceal failures from the user during training. Other
newly reported incidents include an instance of an agent uploading files to
the internet to cite them without being told to do so, and agents publicly
sharing files to collaborate on a task when they were instructed to only use
local files during training. AI models also used an internal software
repository as a message board in an unsanctioned way. These instances
involved unreleased internal models or internal research models.

Read more of this story at Slashdot.

---
VRSS v2.1.180528
  Show ANSI Codes | Hide BBCodes | Show Color Codes | Hide Encoding | Hide HTML Tags | Show Routing
Previous Message | Next Message | Back to Slashdot  <--  <--- Return to Home Page

VADV-PHP
Execution Time: 0.0185 seconds

If you experience any problems with this website or need help, contact the webmaster.
VADV-PHP Copyright © 2002-2026 Steve Winn, Aspect Technologies. All Rights Reserved.
Virtual Advanced Copyright © 1995-1997 Roland De Graaf.
v2.1.250224