In 2016, SafeGraph’s founder Auren Hoffman wrote a blog post called “Where Should Machines Go To Learn?” It wasn’t a flashy post. No bold predictions about AGI. No hype. Just a clear-eyed argument about AI and machine learning as they existed then — image recognition, fraud detection, recommendation engines, self-driving cars. The questions weren’t about artificial general intelligence. They were about whether you could build a model that actually worked reliably on real-world data. Auren’s argument was simple: you couldn’t, not without better access to high-quality, organized data. And if we wanted machines to actually get smarter, someone needed to do the unglamorous work of organizing the world’s information so everyone else didn’t have to.
I joined SafeGraph in early 2018, two years after that post was published. Like a lot of people in data, I’d heard the “data is the new oil” framing for years, and I believed it. I knew machine learning lived or died on data quality, garbage in garbage out. But what drew me specifically to SafeGraph was the problem space: location data. The physical world. So many real applications, so many industries that needed to understand what was happening on the ground. When I read Auren’s post early on, the argument made complete sense. But the full weight of it hadn’t landed yet. The world hadn’t caught up to the thesis.
It has now.
He Was Right. More Right Than Even He Probably Expected.
Auren’s argument was essentially this: AI and ML progress was being held back not by a lack of smart people or clever algorithms, but by a lack of accessible, high-quality data. The big players like Google and Facebook had a massive advantage simply because their businesses generated data at scale. Everyone else was starting from scratch.
Fast forward to today. ChatGPT. Gemini. Claude. The research breakthroughs that produced these models were genuinely extraordinary. The researchers deserve enormous credit. And yet the models that pulled ahead did so because they also solved the data problem. Better algorithms raised the ceiling. Better data determined who actually reached it.
The proof isn’t just in the models themselves. Entirely new categories of companies were built to feed them. Scale AI, Mercor, and others created industries around data labeling and annotation: teaching frontier models to reason, recognize, and respond. These aren’t software companies. They’re data companies. And they became some of the most valuable players in the AI ecosystem almost overnight. Auren’s thesis was right, and the market confirmed it by building billion dollar businesses around it. Compute matters too, enormously. But data remained the other half of the equation that nobody could shortcut.
Auren’s other argument was about comparative advantage. Paraphrasing Ricardo: if some people focus on organizing the past and others focus on predicting the future, everyone wins. You shouldn’t have to master data infrastructure to build great AI. The same way you rent AWS instead of buying servers, you should be able to rent data instead of acquiring and cleaning it yourself.
That model is now the dominant architecture of the AI industry. Foundation model providers. Fine-tuning layers. Data infrastructure companies. Application builders. The stack separated exactly the way Auren described it would.
What Even He Couldn’t Have Predicted
Here’s what I don’t think anyone fully saw coming in 2016: the problem wouldn’t just be training data. It would be grounding data.
Here’s how I think about it: these models have PhD-level reasoning ability. But stale or inaccurate data makes them effectively blind. For example, a language model might understand words, but it doesn’t inherently understand physical dimensions or spatial relationships. In a logistics use case, that’s not an abstract problem. The agent fails its task. A delivery goes to the wrong place. A route doesn’t work. The intelligence was there. The grounding wasn’t.
As AI systems move from answering questions to actually doing things — booking appointments, routing deliveries, analyzing markets; they need to be grounded in accurate, current, structured information about physical places.
Some of our software OEM customers are now building native agentic workflows directly into their products; not prototypes but shipped features. And the conversations we’re having about data quality have changed too. The bar is higher. Customers are more demanding. They’re not just asking whether we have coverage. They’re asking how we know it’s right.
That’s the problem SafeGraph has been quietly solving for almost a decade.
The Bottleneck Shifted. The Opportunity Didn’t.
In 2016, the bottleneck was: “I want to build an AI system. Where do I get training data?”
In 2026, the bottleneck is: “I have a capable AI system. How do I make it useful in the real world?”
The answer, in many cases, involves data about physical places: accurate business listings, addresses, building footprints, geospatial context. The kind of data that seems mundane until your AI agent recommends a business that closed two years ago, or routes a delivery to a building that doesn’t exist at that address, or gives a site selection recommendation based on a competitive landscape that’s six months stale.
We used to talk about garbage in, garbage out as a performance problem. When AI agents are making real-world decisions, it becomes a trust problem. Bad data doesn’t just produce worse outputs; it produces wrong actions.
The companies building serious AI products on top of physical world data have figured this out. They’re not trying to build their own POI database from scratch. Data companies like SafeGraph exist to handle the curation, validation, and transformation work so that innovators can focus on what they’re actually building: the analytics, the applications, the intelligence layer.
The Library Is Open. Are You Using It?
Auren ended his 2016 post with a call to democratize data access, so innovators could spend less time organizing the past and more time predicting the future.
That mission hasn’t changed, but the stakes have. As AI agents step into the real world, ‘good enough’ data isn’t enough anymore. The underlying data has to actually be right — fresh, validated, and precise.
That’s what we’ve spent the last decade building at SafeGraph. If physical-world grounding is the missing piece in the AI system you’re deploying, I’d love to trade notes. Drop me a line at [email protected]
This post is a follow-up to “Where Should Machines Go To Learn?”, written by SafeGraph founder Auren Hoffman in December 2016.
Originally published as a LinkedIn post by Jason Richman. Read the original post on LinkedIn