AT2k Design BBS Message Area
Casually read the BBS message area using an easy to use interface. Messages are categorized exactly like they are on the BBS. You may post new messages or reply to existing messages!

You are not logged in. Login here for full access privileges.

Previous Message | Next Message | Back to Slashdot  <--  <--- Return to Home Page
   Local Database  Slashdot   [9 / 103] RSS
 From   To   Subject   Date/Time 
Message   VRSS    All   Anthropic Reveals Fourth Likely Crime Committed By Its AI   September 9, 2026
 10:40 PM  

Feed: Slashdot
Feed Link: https://slashdot.org/
---

Title: Anthropic Reveals Fourth Likely Crime Committed By Its AI

Link: https://yro.slashdot.org/story/26/09/10/02242...

An anonymous reader quotes a report from The Register: Amid industry soul-
searching about the possibility of AI improving itself to the point that it
kills everyone, Anthropic has revealed yet another incident that would
qualify as a crime if perpetrated by a person. The AI biz published "an
alignment assessment" detailing four times Claude models accessed third-party
systems without authorization. The company has already reported three of the
incidents. Evidence of the fourth was lurking in a session transcript dating
back to January 2026 when the misbehavior occurred. Anthropic found the first
three by scanning around 141,000 transcripts where Claude could have obtained
internet access during evaluation. It missed the fourth initially because
"our scan relied on an agentic search." [...] The January 2026 AI trespass
involved an early version of Claude Opus 4.6, which was given a Capture the
Flag (CTF) challenge under the oversight of the third-party model evaluator
where the other hacking events occurred. Opus 4.6 managed to sabotage its
chances of success by disabling the machine it was targeting. It assigned the
device an IP address that already existed on another piece of hardware,
rendering the target unreachable and making it impossible to solve the
challenge. Those familiar with other incidents where AI models violated third-
party systems may recall that unsolvable tasks represent a common catalyst
for misbehavior. Models exhaust all aligned options, and then turn to
transgressive approaches. Opus 4.6 might have been an exception, but when it
tried to abort the task after recognizing that it could not reach the target
machine, it failed to do so "due to a misconfiguration in [the model's]
evaluation harness." It failed to shut down not just once but seven times. So
it continued onward, trying other expected means to reach the target machine
but failing. Then it explored further. "The model discovered a machine
belonging to a third party that it was able to access, and stated that it
believed this third party was part of the CTF," Anthropic explained in its
post. "Inside the machine, the model found a file listing a password, which
it used to gain admin access to the system." The model went on to gather more
credentials, and modified a system setting to make it easier to access the
personal information of an individual associated with the third party
evaluation organization. Opus 4.6 might have done more but for the fact that
it exhausted its token budget, bringing the session to an end. Anthropic says
it's not as concerned about this incident as the others because the model
tried to abort its task.

Read more of this story at Slashdot.

---
VRSS v2.1.180528
  Show ANSI Codes | Hide BBCodes | Show Color Codes | Hide Encoding | Hide HTML Tags | Show Routing
Previous Message | Next Message | Back to Slashdot  <--  <--- Return to Home Page

VADV-PHP
Execution Time: 0.0163 seconds

If you experience any problems with this website or need help, contact the webmaster.
VADV-PHP Copyright © 2002-2026 Steve Winn, Aspect Technologies. All Rights Reserved.
Virtual Advanced Copyright © 1995-1997 Roland De Graaf.
v2.1.250224