mikeaorlando.com Michael Orlando
  • Home
  • Start Here
  • About
  • Work & Ventures
  • Ideas & Writing
  • Collaborate
  • Search

▶Great Data Lake LLC — Data Brokerage and Value-Add Business

Great Data Lake LLC — Data Brokerage and Value-Add Business

Founded and ran a data brokerage from 2017 to 2022 that turned hundreds of terabytes of raw location and advertising data into clean, compliant, sellable datasets for defense, government, and finance customers.

Routing Notes

  • Parent Projects for Work
  • Signal Working Systems

Overview

Founded and operated a data company from 2017 to 2022 that processed hundreds of terabytes of raw, real-time data streams into clean, compliant, valuable datasets for a variety of industries. GDL sat at the intersection of data engineering and data commerce.

Context

Great Data Lake ran from 2017 to 2022. It started from a gap I kept seeing from the supply side: businesses increasingly needed large, real-time datasets, but what the market sold was raw, duplicated, poorly documented, and legally uncomfortable to touch. The value was never in having data — it was in data that arrived clean, described, reduced to what mattered, and compliant enough that a buyer’s own lawyers could say yes.

GDL began by building on Lumate’s data infrastructure and hard-won adtech knowledge, converting what had been a media-trading capability into a data brokerage. It was a B2B business built from scratch: fewer than ten carefully negotiated proprietary sources on the supply side, and customers in defense, government, advertising, automotive, community planning, and finance on the demand side.

The System

The business was a set of interlocking systems, each solving a real friction in how data changes hands:

  • Reduction as the product. Hundreds of terabytes of inbound data per month became tens of gigabytes per day of delivered product — deduplicated, fraud-filtered (or fraud-marked, so customers could decide), and cut to the fields customers actually used. Built on Spark and Databricks, running on AWS (S3 for storage, Redshift for warehousing).
  • Location accuracy as a solved problem, not a caveat. Raw GPS coordinates from mobile devices carry real noise — even two devices sitting next to each other report slightly different coordinates. I built a system that scored every incoming location observation for accuracy using a combination of k-nearest-neighbors and HyperLogLog probability models, comparing each point against a database of prior observations to tell precise GPS from lookup-based approximation. The scores were served from a hash-partitioned Redis cluster with sub-10-millisecond lookups — fast enough to append to live real-time-bidding requests — appending accuracy to roughly 25,000 records per second, sustained at 99.9% uptime for four years straight.
  • Identity linkage without cookies. Real-time-bidding ad requests carry IP addresses, device identifiers, timestamps, and user agents. I built collection and deduplication pipelines that turned that raw exhaust into usable device-and-IP linkage data — server-side, which made it more durable than cookie-based tracking, and statistically screened for the anomalies that usually mean ad fraud rather than real user behavior.
  • Places as a first-class data layer. For point-of-interest products, device location observations were matched against a licensed, NAICS-classified commercial database of roughly 120 million U.S. business locations, built as precise polygons rather than simple points — enough to say not just “near a business” but “inside this specific one.”
  • Derivative products. Each source became multiple targeted products, so one negotiated supply relationship served many customer needs at different price points and freshness levels — per-record streams down to daily batches.
  • Compliance as infrastructure. Processes for CCPA, COPPA, GDPR, and Privacy Shield ran on both sides of every transaction — internal handling plus obligations flowing through to buyers and sellers. For location data especially, cleaning required genuine GIS expertise, not just deduplication.
  • A sampling mechanism that protected the asset. Prospective buyers evaluated real data through the normal production delivery pipeline — without receiving the whole dataset. Sales friction went down without giving the product away.
  • A reseller and broker network. Several white-label partnerships put GDL’s sourcing and cleaning behind other companies’ sales relationships. One multi-year partnership alone ran roughly ten different datasets on a straightforward revenue-share model and produced millions in revenue, without GDL needing its own sales force for every deal.

Outcomes

  • A functioning data brokerage serving defense, government, advertising, and finance customers on subscription, with products including location datasets and the Rent Demand Pressure Quantification dataset for real estate developers.
  • Not every product found a buyer. A predictive demographic model built on graph databases and machine learning — trained on app-usage and location patterns, validated against Facebook’s own ad-targeting data — hit a genuinely high accuracy rate and worked technically. It never found commercial traction: the best-fit buyer was slow to evaluate it, and the volume of records clearing the accuracy threshold was too thin to build a business around. Technical success and market success are different problems, and this project only solved the first one.
  • Durable capability in data commerce — sourcing negotiations, compliance engineering, product derivation, and channel building — that continues to shape how I approach data-oriented ventures.
  • The company wound down as an operating entity in 2022; its patterns, relationships, and lessons carried forward into later work, the same long-tail arc told honestly in The Lumate Arc.

Artifacts

This is the accuracy-scoring system from the field, not a slide about it — a Tableau view from April 2020 plotting real location observations by their HyperLogLog accuracy score (“log HLL”), the same scoring pipeline described above:

Tableau visualization of mobile device location observations colored by HyperLogLog accuracy score, April 2020
Location-accuracy scoring in production, April 2020 — observations plotted by log HLL score.

There is no client-facing photography from this chapter — GDL sold data pipelines, not a storefront, and nobody thought to photograph a Redis cluster. What you can check independently: greatdatalake.com is archived on the Wayback Machine from March 2025, showing where the underlying geolocation-accuracy work went next — CentricIP™, an IP geolocation product built on the same accuracy lineage described above.

Patterns Worth Reusing

  • Sell the reduction, not the volume. Customers paid for what was removed — duplicates, fraud, irrelevance — as much as for what remained. Refinement is a product.
  • Accuracy is a metric, not an assumption. GPS coordinates look precise; they aren’t automatically. Building a real scoring system for data quality — instead of treating incoming data as ground truth — was what made the downstream products trustworthy enough to sell into regulated, real-money decisions.
  • One source, many products. Deriving multiple targeted products from each supply relationship multiplied revenue per negotiation, the highest-cost part of the business.
  • Compliance is a moat when it’s real. Making privacy law operational — not just contractual — was what let regulated buyers say yes quickly.
  • Let prospects touch production. The sampling system built trust precisely because it used the real delivery pipeline; demos that resemble production close faster than decks.
  • Technical success is not commercial success. The demographic-prediction pipeline worked and beat the incumbent benchmark. It still didn’t ship as a business, because the buyer’s evaluation cycle and the addressable data volume mattered more than the model’s accuracy. Both are real constraints, not excuses.

Explore

  • ✎Posts
  • ○Categories
  • ◉Projects for Fun
  • ▶Projects for Work
  • ◎Systems Across Domains
  • ▪Technical Tools
  • ◆Transferable Skills
  • ✕Value Multipliers
  • ◇Ventures

Michael Orlando

Systems, ventures, writing, and public proof arranged so the right people can find the right next step.

© 2026 Michael Orlando

Privacy & Data Handling

about.me Substack Credly Polywork