Taxonomies: is it better to Build or Buy?

Driving consistency in how information is classified was once just a headache
for your Information Manager. Now it’s everyone’s problem, your AI initiatives included!

Taxonomies and ontologies create a common “language” for finding, classifying and sharing information. They provide the descriptive metadata that makes information assets genuinely useful: connecting people to the right data at the right time.

That value proposition hasn’t changed since we first wrote on this topic, well before the AI revolution. What has changed since is the stakes. A well-designed taxonomy was once a significant advantage in information retrieval. Today, it is a prerequisite for effective artificial intelligence. Without structured, consistently tagged information, AI systems, no matter how powerful, are working blind.

Judi Vernau first called this out this 25 years ago: “Our ability to create information has substantially outpaced our ability to retrieve relevant information”. Today, the gap between those two curves has never been wider. AI is both the reason why and, with the right foundations, the solution.

What is a Taxonomy?


A simple taxonomy is a hierarchical set of reference terms used for classification and information management. A faceted taxonomy goes further, representing multiple dimensions (discipline, document type, topic, geography, etc.) that together describe information from different angles.

In practice, mature working taxonomies are richer still. The relationships between terms create semantic networks that enable inference, context-aware search and knowledge discovery. At this level, a taxonomy becomes an ontology: a formal representation of the concepts and relationships in a domain.

Taxonomies can be contextual (describing the characteristics of information) or entity-based (capturing the names and codes of physical or organisational things, such as wells, equipment, basins, licenses, and companies. Most enterprise taxonomies need both.

Why Taxonomies Matter More Than Ever


The business case for taxonomies has always been strong. Classification, search, integration, knowledge retention, and data quality all improve substantially when a consistent taxonomy is in place. But the AI era has added a new and compelling dimension.

AI systems need structured data to work reliably. Large language models (LLMs) are impressive general reasoners, but when deployed in an enterprise setting they must draw on an organisation’s own information. That process (known as Retrieval Augmented Generation, or RAG) works far better when the underlying content is consistently classified and described. Without good metadata, AI search returns irrelevant results, misses critical documents, and produces unreliable answers. The problem of AI “hallucination” is significantly worse when the retrieval layer is weak.

Agentic AI amplifies the requirement. The newest generation of AI systems doesn’t just answer questions, it takes actions: drafting reports, triggering workflows, routing tasks, and much, much more. These agents rely on structured knowledge to navigate the enterprise. A taxonomy is, in effect, the map that tells an AI agent where things are and what they mean.

Knowledge graphs are the next step. The market has moved beyond simple hierarchical taxonomies toward knowledge graphs: rich semantic networks that connect entities, concepts, relationships and properties. These graphs are becoming the backbone of enterprise AI strategies, providing the structured context that allows AI to reason, not just retrieve. A well-designed taxonomy is the foundation from which a knowledge graph is built.

Regulatory and audit pressure is growing. Explainability requirements for AI decisions are increasing, driven by internal governance, regulatory frameworks and board-level scrutiny. Organisations need to demonstrate not just that AI produced an answer, but why. A structured taxonomy with defined relationships makes AI reasoning more traceable and auditable.

Data without metadata was always of limited value. In the AI era, it is a liability.

Key Benefits


The core benefits of a well-implemented taxonomy remain those we have long advocated, but each is now amplified by its AI implications:

Classification. Information is logically structured using metadata, making it usable regardless of which application or technology platform holds it. Metadata from a well-designed taxonomy is persistent; technology platforms and organisation structures are not. In an AI context, consistent classification is what allows an LLM to retrieve the right content rather than a plausible-sounding approximation.

Search and retrieval. Taxonomical relationships (hierarchical and otherwise) enable far greater search flexibility, improving both precision (accuracy) and recall (number of relevant results). Combined with AI-powered semantic search, a good taxonomy produces a step-change improvement in findability.

AI readiness. Well-tagged content is the raw material for RAG-based AI tools, knowledge graphs and agentic workflows. Organisations with mature taxonomies can deploy AI faster, with less risk and better outcomes than those starting from an unstructured baseline.

Integration. A common taxonomy allows information from multiple repositories to be brought together without requiring users to know where everything is stored. This is especially valuable as organisations work across cloud platforms, legacy systems and partner environments.

Knowledge retention. Taxonomies embed domain knowledge into the structure of information itself. When experienced staff leave, the knowledge encoded in the taxonomy and its classifications remains. This is increasingly important as the industry manages an ongoing skills transition alongside the great shift change.

Coherence and cost savings. Consistent reference terms eliminate the proliferation of niche naming conventions, pick-list values and local standards. Precise search and retrieval reduces wasted time. Better-structured data reduces the risk of missed decisions and operational errors.

One shared
understanding

How to Assess a Taxonomy


A taxonomy covering a significant part of a business is large and complex. The ultimate test remains whether it delivers business value when used as intended, but the criteria have expanded.

Information retrieval. Run a pilot using the taxonomy to classify a representative set of content and compare search results against the same content without the taxonomy. This was historically the most direct test of functional value.

AI performance. Test the taxonomy specifically as a foundation for AI-driven retrieval. Does classifying content with these terms improve the relevance and accuracy of AI-generated responses, and MCP tool relevance and precision? Does it reduce hallucinations? These are now becoming the core evaluation criteria.

Industry currency. Speak to companies in similar businesses who have practical, current experience using a domain- or enterprise-scale taxonomy. Domain-specific taxonomies developed with input from multiple organisations will reflect current terminology, emerging technologies and regulatory requirements more accurately than an in-house effort.

Extensibility and AI compatibility. Can the taxonomy be linked to or extended into a knowledge graph? Is it structured in a way that is compatible with modern semantic web standards? These questions are increasingly important as organisations move from basic classification toward AI-native information architectures.

Build or Buy?


The fundamental question is unchanged: given the recognised value of a taxonomy, what is the most effective way to acquire one? Our view hasn’t changed over the last 26 years, but the number and weighting of contributing factors have shifted.

The table below summarises the key variables:

Purchased Taxonomies


Acquisition cost:
Low


Maintenance cost:
Low


Time to value:
Weeks to Months (industry-proven)


Domain coverage:
High (structured, tested)


AI readiness:
High


Risk:
Low

Custom Built Taxonomies


Build cost:
High


Maintenance cost:
Medium – High


Time to value:
Years


Domain coverage:
Variable


AI readiness:
Uncertain


Risk:
High

The case for building


To build a taxonomy, a company must form a dedicated team of domain experts and information specialists. This is not (yet another) part-time task for a data manager. It requires sustained resource, external expertise, and significant budget. A realistic minimum for a first draft covering a meaningful scope is £250,000 to £500,000, with substantial ongoing costs thereafter.

Companies frequently underestimate this. The result is common: projects stall, coverage remains thin, and the taxonomy delivers less value than hoped while costing more than planned. Without ongoing maintenance it becomes out of date and untrusted in short order, and the entire investment is wasted.

Does AI make building easier?


Not as much as we might hope. LLMs can assist with aspects of taxonomy development: suggesting terms, identifying gaps, drafting definitions, dealing with translations, and such like. LLMs can accelerate ontology development when used alongside human experts.

However, LLM assistance does not change the fundamental calculus. Building a high-quality, domain-specific taxonomy still requires deep subject matter expertise, rigorous design principles and extensive validation. AI tools can speed up some tasks but cannot substitute for the accumulated domain knowledge embedded in a mature, industry-proven taxonomy. In fact, using an LLM to build a taxonomy from scratch introduces its own risks: models trained on generally available text may reflect outdated, imprecise or inconsistent industry terminology.

The better use of AI in this context is augmentation of an existing, high-quality taxonomy: extending coverage, suggesting additions, flagging inconsistencies, and adding new languages, rather than trying to replace it.

The case for buying


A commercially available, domain-specific taxonomy has typically been developed over many years, refined with input from multiple organisations in the sector, and tested in real-world information management environments. It reflects industry best practice rather than a single company’s view of the world.

The perception that an in-house taxonomy is inherently better aligned to the business rarely survives scrutiny. As the OSDU Forum has shown multiple times, organisations in the same sector share the vast majority of their information structures, and their working processes. Company-specific additions, where genuinely needed, typically represent less than 5% of the taxonomy and can be accommodated as extensions maintained in collaboration with the supplier.

A purchased taxonomy is also better positioned to support AI deployment. It arrives structured, validated and ready to use, immediately improving the quality of content classification and, by extension, the performance of AI retrieval and generation systems. The time-to-value advantage over a custom build is measured in years.

From an industry perspective, a shared taxonomy enables interoperability and data exchange between companies, partners and contractors. This matters for mergers, acquisitions and disposals, and increasingly for the data-sharing arrangements that underpin cross-supply chain AI initiatives.

Co-development


Where no suitable commercial offering covers the required scope, engaging a specialist taxonomy provider as a co-development partner remains a practical alternative. The supplier brings design experience, tools for taxonomy build and maintenance, and domain knowledge, while the client brings sector-specific context.

Intellectual property and funding arrangements can be structured to work for both parties, and the resulting taxonomy is far more likely to be fit for purpose than a fully in-house effort.

Completeness
Cost

Conclusion


The case for buying rather than building a taxonomy has always been compelling. In 2021, we argued it on the basis of cost, coverage, risk and time-to-value. All of those arguments remain valid.

What the AI era adds is urgency. Organisations that lack well-structured, consistently classified information are not simply missing a useful capability, they are unable to get reliable value from AI tools that their competitors are already deploying. The quality of your taxonomy is now directly linked to the quality of your AI.

For most organisations, the right path is clear: acquire a proven, domain-specific taxonomy, deploy it across your information environment, and use it as the foundation for AI-driven search, retrieval and knowledge management. The investment is modest relative to the value at stake, and the risk is far lower than the alternative.

The question isn’t really whether you can afford to buy a taxonomy. It is whether you can afford not to have one.

Flare Solutions has been building and deploying information management solutions for the energy industry since 1998. Our Sirus platform includes the Flare Taxonomies – the leading taxonomies for the energy industry – providing the structured information foundation that modern AI applications require.

Want to know more? Contact us today.