Cloudflare moves into analytics workloads with a serverless alternative to dedicated data infrastructure
Cloudflare Inc. is moving beyond its bread-and-butter content delivery networks and cybersecurity offerings with the launch of a new serverless data platform called Cloudflare Basin, which aims to make advanced data analytics cheaper and easier for small businesses and developer teams.
Announced today, Cloudflare Basin is designed to eliminate the need for companies to set up and maintain the complex infrastructure that underpins most analytics workloads. It gives them a much simpler way to ingest, store, catalog and query analytical data, without having to worry about managing the underlying servers and systems that make this possible.
Data analytics has never been easy. It requires a dedicated team of experienced data engineers who know how to set up, manage and maintain the expensive server systems that store and organize all of that data. In addition, companies must have the skills to stitch together various different systems and create the data pipelines to get the information they want to analyze into the analytics engine.
These operations are prohibitively expensive, Cloudflare says. Companies have to pay extortionate egress fees just to be able to move their data from the public cloud to another system, and then there’s the salaries and system overhead to worry about. As a result, it’s usually only the largest organizations that can actually run any kind of sophisticated data analytics operation, which gives them a significant advantage over smaller competitors.
Cloudflare Basin is aimed at leveling the playing field. The company says it’s a more open and accessible alternative because it doesn’t require customers to invest in dedicated servers and pay to move their data around different clouds and systems. It’s built atop the open-source Apache Iceberg table format, which is designed for massive, petabyte-scale analytics datasets, and Cloudflare R2, the company’s proprietary, egress-free distributed object storage service.
The choice of Apache Iceberg is a smart move, because it enables analytics engines such as Apache Spark and DuckDB to read data directly at its source. Because of this, customers don’t need to copy information, move files from one location to another or change its format. They simply leave it where it is, with the only thing extracted being the actual insights generated when they query the data in place. All of the data collection, preparation and analysis work runs on Cloudflare’s global network.
The company is taking advantage of the economics of the Apache Iceberg data stack, which makes storage and compute more interchangeable, explained Michael Ni of Constellation Research. “Cloudflare already has the global infrastructure, serverless compute and an egress-free storage model,” he said. “And it already sits in the path of a lot of application, log and event data, so that reduces its data movement costs significantly.”
Cloudflare says this this is why Basin can be much more affordable and flexible than alternative models, with simplified pricing where customers are charged based on the amount of data they analyze.
Chief Technology Officer Dane Knecht said software developers shouldn’t need to know how to operate data infrastructure just to be able to query their own information. “With Cloudflare Basin, we are bringing the same serverless model that developers expect from Cloudflare to analytics,” he said. “No clusters to manage, no unnecessary data movement and open standards that keep customers in control of their data.”
Cloudflare’s move into data analytics is significant because it brings the company into competition with established data warehouse giants such as Snowflake Inc. and Databricks Inc. The company is best known for its content delivery networks and cybersecurity offerings, but these days it has much broader edge computing ambitions. With Cloudflare basin, it’s trying to leverage its massive and highly-interconnected global network to help companies analyze their data closer to where it’s created, bypassing centralized cloud data warehouse platforms to reduce latency and costs.
Although Cloudflare insists that Basin is not a blanket replacement for every data warehouse workload, it believes it still has a big opportunity in the analytics market. It points out that many analytics workloads are too small to justify the expense of setting up dedicated server clusters and moving data around, but that could change with Basin. “I don’t expect larger companies to rip out Databricks or Snowflake any time soon, but Cloudflare has the opportunity to capture workloads at the edge of the traditional data warehouse market and then move up the stack,” Ni said.
Image: SiliconANGLE/Dreamina
A message from John Furrier, co-founder of SiliconANGLE:
Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.
- 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more
- 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network
Are you an AWS customer? Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: https://siliconangle.com/aws-marketplace/
About SiliconANGLE Media
Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.