The dream of enterprise semantics: Why this time is different

InfoWorld ·

The dream of enterprise semantics: Why this time is different

In most companies, the struggle to answer a new question fast has nothing to do with a lack of data. The problem is a lack of semantics. Or to put it another way, a lack of agreement about what the data means . Suppose a CFO wants to generate a new report on revenue by customer segment. It sounds simple, but in practice, for the employees tasked with executing the CFO’s directive, it means a multi-day argument about what a “customer segment” even is. Before they can make the report, they need to define the terms of the report. The dream of enterprise semantics, as I define it, is that someone who understands the business should be able to talk to an intelligent data agent and get the answer they need without knowing how the data is structured or stored. No more tickets. No more definitional arguments. For decades, we’ve tried to do better. We even created solutions for enterprise semantics, but the work required was too much. AI changes the equation. As the data agents at OpenAI and Anthropic have shown, the dream of enterprise semantics has come to life. We’ve tried this before The dream of machine-readable business meaning is not new. After inventing the World Wide Web in 1989, Tim Berners-Lee proposed an idea in 1999 called the Semantic Web . He argued that if everyone tagged their web pages according to machine-readable standards, internet search would become far more powerful. Standards such as RDF, OWL, and SKOS emerged to let people tag content according to ontologies that captured meaning and enabled reasoning. But the effort never took off. Building an ontology that actually worked for a business required rare hybrid expertise, months of workshops, and constant upkeep as the business changed. Most projects produced a carefully built ontology for one narrow domain, then collapsed once the team moved on or the business outgrew the map. The standards themselves survived: RDF, OWL, and many standard vocabularies form the foundation of serious data management today. The vision was right; the economics made execution impossible at scale. The labor required to build and maintain an ontology never paid for itself. Several startups built solutions using semantic approaches and solved real problems in narrow niches, but the category never found broad traction. Another approach was to define concepts in a data catalog. This form of BI semantics created precise definitions of key metrics such as revenue, income, and customers. People building solutions could then use the standard glossary definitions to keep dashboards and reports consistent and accurate. The problem is that these definitions make sense to people, not machines. A human reading a text definition of “customer” can fill in the rest from experience: which systems customers live in, how they relate to orders, regions, and revenue. An AI can’t fill in what it was never given. It needs entities, properties, and relationships: a customer is a person, belongs to a domain, and places orders. Give an agent that graph and it can build its own model of the business and infer new information from it. Ask how many customers you serve in Europe, and it can reason its way to an answer. Give it a definition with no structure behind it, and it will guess confidently instead. BI semantics help within a narrow scope, but they’re rarely built on the open standards that would let AI systems reason over them. The dream of enterprise semantics in 2026 The power of frontier large language models (LLMs) tempts people into thinking that simply loading enterprise data into these models gets you close to the dream of enterprise semantics. It doesn’t. LLMs still need to be told what the structured data means, how business concepts are defined, and which data is authoritative. They need what people now call context. Context has become a buzzword used loosely, but three specific types of context, taken together, form a layer that lets AI deliver on the dream of enterprise semantics: Data context: A single, unified metadata graph that connects structural schemas, data quality signals, end-to-end lineage, and usage profiling. Data context lets the AI know exactly what data exists. Semantic context: Moving beyond flat tables to formal ontologies and business relationships. Semantic context teaches the AI the meaning of data and explicit business rules, for instance how “customer” relates to “net revenue”, so it can reason accurately without hallucinating. Memory context: A persistent, shared record across the organization of human corrections, feedback, and tribal knowledge. Memory context allows agents to inherit the corrections experts have already made, instead of starting from scratch in every chat. These three layers let AI actually guide a business user to the answers and data they need. Without an open context layer, frontier models struggle. With an open context layer, our internal tests show that AI produces answers that are seven times more accurate while reducing query workloads by 86%. Gartner predicts that organizations that prioritize semantics in their AI-ready data could see 60% lower AI costs by 2027. How AI changes the economics of enterprise semantics Applying AI to data management changes the economics of building the open context layer in several ways. Building a fully populated, unified repository of context that maps all of an enterprise’s data used to be a huge amount of work. Trust me—I struggled for years to make this happen at Uber when I was chief data officer. Now, working through connectors to virtually every data source, AI agents automate the creation of technical metadata describing data structure, along with metadata for security, governance, quality, and lineage. Building an ontology used to require a human expert starting from scratch. Now AI can propose the ontology, using public standards where they fit, while humans review, refine, and approve what it creates. A domain expert who couldn’t build a thousand-concept ontology alone can validate what AI drafts. That’s a real shift in what’s achievable. Previous semantic projects decayed because businesses changed faster than humans could update ontologies by hand. Now AI can monitor the data environment continuously, catch drift, and propose updates for a human to review, so the context layer stays current without heroic effort. What finally becomes possible Now return to our CFO example. With the context layers in place, asking a new question leads to a different outcome. The AI agent finds the relevant tables, consults the ontology to give “customer segment” its precise organizational definition, and retrieves a note from memory that explains how post-acquisition figures require a specific adjustment the finance team already validated. Then it delivers the answer with a clear explanation of how it was calculated. That’s what data democratization actually means once context is in place. It means giving people access to answers, not just the data warehouse. A marketing analyst or a product manager doesn’t need to understand the technical stack. The context layer absorbs that complexity so they don’t have to. The Semantic Web had the right vision and the wrong tools. Thirty years later, the tools have caught up: AI can build ontologies faster than humans could manage alone, keep the context layer current as the business changes, and let agents build on frameworks that other agents helped construct. The bottleneck that kept this dream out of reach for three decades is gone, and organizational knowledge can now build on itself instead of decaying between projects. The CFO’s question has always deserved a correct answer. The infrastructure to deliver one is finally within reach. — New Tech Forum provides a venue for technology leaders—including vendors and other outside contributors—to explore and discuss emerging enterprise technology in unprecedented depth and breadth. The selection is subjective, based on our pick of the technologies we believe to be important and of greatest interest to InfoWorld readers. InfoWorld does not accept marketing collateral for publication and reserves the right to edit all contributed content. Send all inquiries to  doug_dineley@foundryco.com .

In most companies, the struggle to answer a new question fast has nothing to do with a lack of data. The problem is a lack of semantics. Or to put it another way, a lack of agreement about what the data means . Suppose a CFO wants to generate a new report on revenue by customer segment. It sounds simple, but in practice, for the employees tasked with executing the CFO’s directive, it means a multi-day argument about what a “customer segment” even is. Before they can make the report, they need to define the terms of the report. The dream of enterprise semantics, as I define it, is that someone who understands the business should be able to talk to an intelligent data agent and get the answer they need without knowing how the data is structured or stored. No more tickets. No more definitional arguments. For decades, we’ve tried to do better. We even created solutions for enterprise semantics, but the work required was too much. AI changes the equation. As the data agents at OpenAI and Anthropic have shown, the dream of enterprise semantics has come to life. We’ve tried this before The dream of machine-readable business meaning is not new. After inventing the World Wide Web in 1989, Tim Berners-Lee proposed an idea in 1999 called the Semantic Web . He argued that if everyone tagged their web pages according to machine-readable standards, internet search would become far more powerful. Standards such as RDF, OWL, and SKOS emerged to let people tag content according to ontologies that captured meaning and enabled reasoning. But the effort never took off. Building an ontology that actually worked for a business required rare hybrid expertise, months of workshops, and constant upkeep as the business changed. Most projects produced a carefully built ontology for one narrow domain, then collapsed once the team moved on or the business outgrew the map. The standards themselves survived: RDF, OWL, and many standard vocabularies form the foundation of serious data management today. The vision was right; the economics made execution impossible at scale. The labor required to build and maintain an ontology never paid for itself. Several startups built solutions using semantic approaches and solved real problems in narrow niches, but the category never found broad traction. Another approach was to define concepts in a data catalog. This form of BI semantics created precise definitions of key metrics such as revenue, income, and customers. People building solutions could then use the standard glossary definitions to keep dashboards and reports consistent and accurate. The problem is that these definitions make sense to people, not machines. A human reading a text definition of “customer” can fill in the rest from experience: which systems customers live in, how they relate to orders, regions, and revenue. An AI can’t fill in what it was never given. It needs entities, properties, and relationships: a customer is a person, belongs to a domain, and places orders. Give an agent that graph and it can build its own model of the business and infer new information from it. Ask how many customers you serve in Europe, and it can reason its way to an answer. Give it a definition with no structure behind it, and it will guess confidently instead. BI semantics help within a narrow scope, but they’re rarely built on the open standards that would let AI systems reason over them. The dream of enterprise semantics in 2026 The power of frontier large language models (LLMs) tempts people into thinking that simply loading enterprise data into these models gets you close to the dream of enterprise semantics. It doesn’t. LLMs still need to be told what the structured data means, how business concepts are defined, and which data is authoritative. They need what people now call context. Context has become a buzzword used loosely, but three specific types of context, taken together, form a layer that lets AI deliver on the dream of enterprise semantics: Data context: A single, unified metadata graph that connects structural schemas, data quality signals, end-to-end lineage, and usage profiling. Data context lets the AI know exactly what data exists. Semantic context: Moving beyond flat tables to formal ontologies and business relationships. Semantic context teaches the AI the meaning of data and explicit business rules, for instance how “customer” relates to “net revenue”, so it can reason accurately without hallucinating. Memory context: A persistent, shared record across the organization of human corrections, feedback, and tribal knowledge. Memory context allows agents to inherit the corrections experts have already made, instead of starting from scratch in every chat. These three layers let AI actually guide a business user to the answers and data they need. Without an open context layer, frontier models struggle. With an open context layer, our internal tests show that AI produces answers that are seven times more accurate while reducing query workloads by 86%. Gartner predicts that organizations that prioritize semantics in their AI-ready data could see 60% lower AI costs by 2027. How AI changes the economics of enterprise semantics Applying AI to data management changes the economics of building the open context layer in several ways. Building a fully populated, unified repository of context that maps all of an enterprise’s data used to be a huge amount of work. Trust me—I struggled for years to make this happen at Uber when I was chief data officer. Now, working through connectors to virtually every data source, AI agents automate the creation of technical metadata describing data structure, along with metadata for security, governance, quality, and lineage. Building an ontology used to require a human expert starting from scratch. Now AI can propose the ontology, using public standards where they fit, while humans review, refine, and approve what it creates. A domain expert who couldn’t build a thousand-concept ontology alone can validate what AI drafts. That’s a real shift in what’s achievable. Previous semantic projects decayed because businesses changed faster than humans could update ontologies by hand. Now AI can monitor the data environment continuously, catch drift, and propose updates for a human to review, so the context layer stays current without heroic effort. What finally becomes possible Now return to our CFO example. With the context layers in place, asking a new question leads to a different outcome. The AI agent finds the relevant tables, consults the ontology to give “customer segment” its precise organizational definition, and retrieves a note from memory that explains how post-acquisition figures require a specific adjustment the finance team already validated. Then it delivers the answer with a clear explanation of how it was calculated. That’s what data democratization actually means once context is in place. It means giving people access to answers, not just the data warehouse. A marketing analyst or a product manager doesn’t need to understand the technical stack. The context layer absorbs that complexity so they don’t have to. The Semantic Web had the right vision and the wrong tools. Thirty years later, the tools have caught up: AI can build ontologies faster than humans could manage alone, keep the context layer current as the business changes, and let agents build on frameworks that other agents helped construct. The bottleneck that kept this dream out of reach for three decades is gone, and organizational knowledge can now build on itself instead of decaying between projects. The CFO’s question has always deserved a correct answer. The infrastructure to deliver one is finally within reach. — New Tech Forum provides a venue for technology leaders—including vendors and other outside contributors—to explore and discuss emerging enterprise technology in unprecedented depth and breadth. The selection is subjective, based on our pick of the technologies we believe to be important and of greatest interest to InfoWorld readers. InfoWorld does not accept marketing collateral for publication and reserves the right to edit all contributed content. Send all inquiries to  doug_dineley@foundryco.com .

Источник: InfoWorld