CompaniesInvestorsPeople
Home
Loading

aVenture is in Beta: research coverage is expanding as we build, so please independently verify key details before making investment decisions.

aVenture is in Beta: research coverage is expanding as we build, so please independently verify key details before making investment decisions.

Get in Touch

  • Contact

  • Request a Demo

  • Request Data Updates

  • Add a Company

Research

  • Companies

  • Investors

  • People

aVenture

  • Download App

  • Pricing

Download the aVenture Research beta for iOS and iPadOSDownload aVenture Research on the Mac App Store

Resources

  • Documentation

  • Use Cases

  • CLI

  • MCP

  • Feature Requests

  • Sitemap

Member

Backed by

© aVenture Investment Company, 2026. All rights reserved.

San Francisco, CA, USA

Privacy Policy · Terms of Service · Privacy FAQ

aVenture Investment Company ("aVenture") is an independent research platform providing detailed analysis and data on startups, venture capital investments, and key industry individuals. It is not a registered investment adviser, broker-dealer, or investment advisor and does not provide investment advice or recommendations. The data provided by aVenture does not constitute recommendations or advice, whether by methodology, analysis, AI-generated content, or a statement written by a staff member of aVenture.

aVenture is not affiliated with any of the people, companies, organizations, government agencies, regulatory bodies, or investment funds we provide coverage for on this site unless explicitly stated otherwise. Users assume full responsibility for decisions made based on information obtained from this platform. Links to external websites do not imply endorsement or affiliation with aVenture. Any links that provide the ability to invest in a primary or secondary transaction in a company are for convenience only and do not constitute solicitations or offers to buy or sell an investment. Investors should exercise heightened precaution and due diligence when investing in private companies, especially those not independently audited.

While we strive to provide valuable insights with objectivity and professional diligence, we cannot guarantee the accuracy of the information provided on our platform. Before making any investment decisions, you should verify the accuracy of all pertinent details for your decision. To the fullest extent permitted by law, aVenture shall not be liable for any direct, indirect, incidental, consequential, or financial damages arising from use of this site, whether by consumers of its contents directly or by persons or organizations covered by our research, even if we are advised of the possibility. Our best-efforts processes and correction request forms do not create a warranty or duty of care.

Profiles on this platform may include content generated in part by large language models (LLMs, artificial intelligence) that aggregate publicly available sources (e.g., SEC EDGAR, public filings, press releases). Source attribution is provided where known; always verify statements and claims here against original sources before relying on any data. Content on our site may contain inaccuracies, omissions, or what are commonly called 'hallucinations' if generated in part or in full by AI / LLMs. The risk can also exist even when content is written by a human, as internal and third-party sources may also have inaccuracies for the same or different reasons. While we randomly audit a proportion of content, this is not exhaustive.

We recommend that an independent auditor be hired to verify the accuracy of the information before relying on it for any sensitive decisions. By accessing this platform, you agree not to rely solely on any information generated by AI, aggregated, or sourced or written otherwise on this site, for investment, financial, or other decisions. aVenture assumes no responsibility for inaccuracies, omissions, or hallucinations. You must independently verify all data from primary sources. Use of this platform constitutes your waiver of claims for reliance-based damages, including negligent misrepresentation. To report an error, request a correction, or dispute information about a company or individual, contact us via our request data updates form.

Loading
Loading
Home
News
Clockwork.io bags $31M in funding to keep AI inference and training workloads running like … clockwork

From SiliconAngle

By Mike Wheatley

October 5, 2026

Clockwork.io bags $31M in funding to keep AI inference and training workloads running like … clockwork

Clockwork.io bags $31M in funding to keep AI inference and training workloads running like … clockwork

Clockwork Systems Inc., the data center infrastructure startup that helps to maximize the efficiency of artificial intelligence chip clusters, has raised $31 million in fresh funding and announced the launch of a new feature called TorchSnap that helps to minimize wasted compute.

Today’s round was co-led by Seligman Ventures, Wing Ventures and Premji Invest and also saw the return of existing backers New Enterprise Associates and e& Capital. It brings the startup’s total amount raised to date to $73 million.

At a time when the costs associated with AI compute are rapidly escalating, Clockwork says many organizations are now much less focused on securing raw graphics processing unit resources and more concerned with how to maximize their existing clusters. When running massive distributed AI workloads in clusters of thousands of GPUs, it’s common to see hardware failures occur on a daily basis. A good example is the experience of Meta Platforms Inc., which reported hardware issues every three hours on average during its Llama 3 model’s 54-day training run across a cluster of 16,384 GPUs.

In response to such failures, teams normally reload their saved progress from a snapshot, but this recovery process can take up to 90 minutes to complete, the startup said. During this time, all of the thousands of healthy GPUs are left sitting idle, and once they’re up and running again, the cluster often has to repeat work it has already completed.

This is why fault tolerance has become a major issue for data center operators and AI teams, and it’s something that Clockwork enables them to address. The company has developed a programmable software layer that sits between the GPUs and running AI workloads where it can synchronize GPU clusters and deliver nanosecond-accurate telemetry that identifies failures before they lead to a full cluster restart.

Clockwork Chief Executive Suresh Vasudevan said that though there’s no getting away from GPU failures when running such enormous clusters, teams shouldn’t have to endure losing hours of work those chips have performed. “Fault tolerance is a goodput multiplier: it keeps GPUs doing useful work instead of waiting for recovery or repeating work already done,” he said. “We built our software alongside enterprises and cloud providers operating some of the largest GPU fleets, so it handles the failures they actually see.”

With the launch of its new TorchSnap feature, Vasudevan says Clockwork is enhancing its cluster resilience capabilities further, adding a third protective layer alongside its existing LinkPass network failover tool and TorchPass GPU migration software. It works by capturing multinode snapshots of distributed AI inference workloads across each node within a cluster, without any need for developer code modifications, so running jobs can quickly be restarted wherever they left off, the CEO explained. Teams can also add “checkpointing logic” at the application level, so that less progress is lost and less computation repeated after a failure.

SemiAnalysis analyst Dylan Patel said greater fault tolerance is urgently needed for AI inference workloads. “Cluster fault tolerance used to be a training problem, but it is now an inference problem too,” he explained. “Clockwork.io keeps replicas serving through link flaps and network failures, and its extremely fast checkpoints accelerate weight transfer back into the rollout fleet, so neither direction stalls the run.”

Since its last funding round just over a year ago, Clockwork has seen rapid adoption of its Ai cluster efficiency-boosting software across public cloud infrastructure providers, neoclouds and enterprise fleets. For instance, Microsoft Corp.’s LinkedIn has deployed Clockwork’s LinkPass functionality across its entire fleet of GPUs to eliminate thousands of GPU-hours of downtime each month. LinkPass enables it to reroute AI traffic around any optical link and switch failures to keep jobs running in the event of hardware issues.

Another customer is Together AI Inc., which offers ClockWork’s TorchPass capability as a service on its GPU clusters. Meanwhile, the neocloud provider WhiteFiber Corp., which rents GPU access to enterprises, is leveraging Clockwork’s technology to audit and validate cluster reliability before new AI workloads enter production.

NEA Venture Partner Greg Papadopoulos said that the most reliable thing about GPUs is their unreliability. “Put enough GPUs into one machine and something is always failing, but the industry’s answer is still to just stop the whole job and reload a checkpoint,” he said. “That made sense for supercomputers thirty years ago, but it makes no sense at this scale. We backed the team early and invested again because this layer is becoming part of what an AI cluster is.”

View original article on siliconangle.com

Most Recent

Amazon hires a veteran Microsoft AI leader to help build its tools for coding and work

A longtime Microsoft engineering leader who was previously a Copilot chief technology officer is joining Amazon Web Services to work on its competing AI tools for coding and workplace productivity. Pedram Rezaei, a nearly 20-year Microsoft veteran who was most recently a corporate vice president on

Oct 5, 2026

Namespace raises $42M to build out developer-focused compute cloud

Zurich-founded developer infrastructure startup Namespace Labs Inc. today announced it raised $42 million in Series B, led by Scale Venture Partners, to give developers powerful compute infrastructure to speed up software engineering work. NEA, 20VC, Essence, Burst Capital and Susa Ventures also joi

Oct 5, 2026

Nvidia and CoreWeave tackle the CPU bottleneck in agentic AI infrastructure

Nvidia and CoreWeave address CPU demands in agentic AI infrastructure, linking Vera processors to sandbox performance, execution capacity and security controls.

Oct 5, 2026

Orbital Robotics gets set to send up a pair of arms for International Space Station’s robots

A free-flying NASA robot aboard the International Space Station will soon demonstrate a system for capturing and servicing satellites in orbit, using a pair of autonomous robotic arms built by Seattle startup Orbital Robotics. “This demo will prove our ability to capture, manipulate and upgrade sate

Oct 5, 2026

Similar Posts

Aranya raises $11M to turn bare-metal servers into AI clusters in less than 48 hours

Aranya Inc., a startup that says it can transform raw bare-metal servers into custom, production-ready graphics processing unit clusters for artificial intelligence inference in less than two days, launched today with $11 million in funding. The company says it’s addressing a critical bottleneck in

Sep 1, 2026

AI factories enter the execution era as Cisco and Nvidia push rack-scale systems into production

AI factories enter the execution era as Cisco and Nvidia push rack-scale systems into production - SiliconANGLE Cisco and NVIDIA advance rack-scale AI infrastructure designed to accelerate deployment, improve operations and reduce time to first token.

Aug 25, 2026

Quintessent bags $40M to develop lasers for AI clusters

Quintessent bags $40M to develop lasers for AI clusters - SiliconANGLE

Aug 24, 2026

On-demand GPU infrastructure startup GMI Cloud raises $263M to fuel global expansion

Taiwan-based neocloud GMI Cloud Inc. wants to capitalize on what it says is an unprecedented demand for artificial intelligence computing power after raising a massive $663 million in funding from two sources.

Sep 30, 2026