AI
AI
AI
Anthropic PBC today revealed that it has developed an artificial intelligence model more capable than Claude Mythos 5.
The company detailed the algorithm in the latest edition of its AI alignment report. The document, which is published every three to six months, outlines the potential risks posed by the company’s large language models. The newest installment runs for 186 pages.
The report discusses two AI risk categories dubbed Threat Model 1 and Threat Model 2. The first category focuses on catastrophic harms, such as a hypothetical future LLM that could help bad actors develop biological weapons. Threat Model 2 encompasses smaller hazards. In particular, it covers situations where an AI model with access to an organization’s systems tempers with those systems or decision-making processes.
In February, Anthropic estimated that its models had a “very low” chance of causing Threat Model 2 situations. Today’s report increases the risk level to “low.” The company attributed the change to recent cybersecurity incidents involving its models. In June, Anthropic disclosed that three of its LLMs had carried out cyberattacks during internal tests.
The company stated at the time that one of the breaches was carried out by an unreleased LLM. Its new risk report reveals that it has developed two successors to Claude Mythos 5 dubbed Model 1 and Model 2. The latter algorithm, which is the more capable of the two, is “heavily used” by Anthropic staffers.
The company estimates that Model 2 is a “noticeable improvement on Mythos 5 for many tasks relevant to internal use.” However, Anthropic says that it doesn’t represent as big of a leap as the introduction of Mythos Preview in April. Mythos Preview was the first LLM with the ability to automatically identify a large number of severe software vulnerabilities. Anthropic’s earlier models lacked that capability.
The company says its researchers are using Model 2 to write software, generate AI training data and automate other engineering tasks. Anthropic estimates that its LLMs are helping to accelerate the pace of its AI development efforts. However, that speedup is not believed to be a risk.
A recent open letter signed by prominent AI researchers warned about so-called recursive self-improvement. That’s a hypothetical future scenario in which AI models gain the ability to autonomously improve themselves. An LLM with such a capability could pose a risk because researchers may struggle to equip it with safety guardrails.
Anthropic estimates that recursive self-improvement may start becoming an issue when researchers observe “a doubling of the pace of progress beyond pre-AI-acceleration rates.” Today’s report states that the threshold has not yet been met. However, Anthropic noted that “we are less confident in this assessment” than before because its best internal benchmarks struggle to keep with LLM advances.
Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.
Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.