Clockwork.io bags $31M in funding to keep AI inference and training workloads running like … clockwork
Clockwork Systems Inc., the data center infrastructure startup that helps to maximize the efficiency of artificial intelligence chip clusters, has raised $31 million in fresh funding and announced the launch of a new feature called TorchSnap that helps to minimize wasted compute.
Today’s round was co-led by Seligman Ventures, Wing Ventures and Premji Invest and also saw the return of existing backers New Enterprise Associates and e& Capital. It brings the startup’s total amount raised to date to $73 million.
At a time when the costs associated with AI compute are rapidly escalating, Clockwork says many organizations are now much less focused on securing raw graphics processing unit resources and more concerned with how to maximize their existing clusters. When running massive distributed AI workloads in clusters of thousands of GPUs, it’s common to see hardware failures occur on a daily basis. A good example is the experience of Meta Platforms Inc., which reported hardware issues every three hours on average during its Llama 3 model’s 54-day training run across a cluster of 16,384 GPUs.
In response to such failures, teams normally reload their saved progress from a snapshot, but this recovery process can take up to 90 minutes to complete, the startup said. During this time, all of the thousands of healthy GPUs are left sitting idle, and once they’re up and running again, the cluster often has to repeat work it has already completed.
This is why fault tolerance has become a major issue for data center operators and AI teams, and it’s something that Clockwork enables them to address. The company has developed a programmable software layer that sits between the GPUs and running AI workloads where it can synchronize GPU clusters and deliver nanosecond-accurate telemetry that identifies failures before they lead to a full cluster restart.
Clockwork Chief Executive Suresh Vasudevan said that though there’s no getting away from GPU failures when running such enormous clusters, teams shouldn’t have to endure losing hours of work those chips have performed. “Fault tolerance is a goodput multiplier: it keeps GPUs doing useful work instead of waiting for recovery or repeating work already done,” he said. “We built our software alongside enterprises and cloud providers operating some of the largest GPU fleets, so it handles the failures they actually see.”
With the launch of its new TorchSnap feature, Vasudevan says Clockwork is enhancing its cluster resilience capabilities further, adding a third protective layer alongside its existing LinkPass network failover tool and TorchPass GPU migration software. It works by capturing multinode snapshots of distributed AI inference workloads across each node within a cluster, without any need for developer code modifications, so running jobs can quickly be restarted wherever they left off, the CEO explained. Teams can also add “checkpointing logic” at the application level, so that less progress is lost and less computation repeated after a failure.
SemiAnalysis analyst Dylan Patel said greater fault tolerance is urgently needed for AI inference workloads. “Cluster fault tolerance used to be a training problem, but it is now an inference problem too,” he explained. “Clockwork.io keeps replicas serving through link flaps and network failures, and its extremely fast checkpoints accelerate weight transfer back into the rollout fleet, so neither direction stalls the run.”
Since its last funding round just over a year ago, Clockwork has seen rapid adoption of its Ai cluster efficiency-boosting software across public cloud infrastructure providers, neoclouds and enterprise fleets. For instance, Microsoft Corp.’s LinkedIn has deployed Clockwork’s LinkPass functionality across its entire fleet of GPUs to eliminate thousands of GPU-hours of downtime each month. LinkPass enables it to reroute AI traffic around any optical link and switch failures to keep jobs running in the event of hardware issues.
Another customer is Together AI Inc., which offers ClockWork’s TorchPass capability as a service on its GPU clusters. Meanwhile, the neocloud provider WhiteFiber Corp., which rents GPU access to enterprises, is leveraging Clockwork’s technology to audit and validate cluster reliability before new AI workloads enter production.
NEA Venture Partner Greg Papadopoulos said that the most reliable thing about GPUs is their unreliability. “Put enough GPUs into one machine and something is always failing, but the industry’s answer is still to just stop the whole job and reload a checkpoint,” he said. “That made sense for supercomputers thirty years ago, but it makes no sense at this scale. We backed the team early and invested again because this layer is becoming part of what an AI cluster is.”
Image: SiliconANGLE/Meta AI
A message from John Furrier, co-founder of SiliconANGLE:
Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.
- 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more
- 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network
Are you an AWS customer? Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: https://siliconangle.com/aws-marketplace/
About SiliconANGLE Media
Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.