Ontology Pipeline: A Framework for Knowledge Engineering, by Jessica Talisman
Build semantic knowledge systems with trusted and AI-ready taxonomies, ontologies, and knowledge graphs.
Please order the PDF version if you live in one of these countries, as although we do our best to ship worldwide, our printer currently will not allow us to ship print copies to a number of countries including Mexico, India, Africa, Brazil, China, Chile, Costa Rica, Indonesia, Jordan, Libya, Malaysia, Norway, Switzerland, Saudi Arabia, Romania, Serbia, Peru, Venezuela, Israel, Egypt, Brunei, Turkey, and North Korea.Enter the librarians
The Ontology Pipeline
Controlled vocabulary
Taxonomy
Thesaurus
Ontology
Knowledge graph
The Ontology Pipeline: A semantic knowledge management framework
Consistency pays off
Simple Precision and Recall for Metrics
What a controlled vocabulary is not
How to build a controlled vocabulary
Corpus analysis: Learning from actual usage
Existing standards
Application schemas
Provenance matters
Resolving duplicates
Handling homonyms
Partial overlaps
ANSI Z39.19
Definitions: Giving terms a precise meaning
Writing effective definitions
Context clues
Back to definitions
Creating consistency through normalization
Infrastructure for thinking
Beyond the probabilistic fog
Statistical learning and domain precision
Disambiguation at scale
Next-level data hygiene
SKOS: Meaning for AI systems
From RAG to riches: Production AI
The context engineering revolution
The guidance: ANSI/NISO Z39.19
The governance imperative
Vocabularies for autonomous AI
The semantic foundation for trustworthy AI
Workflow and governance
Discovery
Conclusion
A systems thinking perspective
The enterprise metadata conversation
The library science approach
Semantics, metadata, and library science
The metadata difference
Where we diverge
What the enterprise can learn from library sciences
A systemic problem
The separation of concerns is not helping
Data models as semantic infrastructure
Why semantic layers fall short
Semantic-first governance
MDM and the semantic layer
The semantic crisis
The metadata application profile
The CatChums example
MAPs solve the semantic layer disconnect
The three-layer solution
Data quality and semantic precision
Threading semantics
Governing semantics with structured metadata
Practical implementations
Real-world implementation patterns
Conclusion
A different approach
Our goal
What a taxonomy is (and is not)
Classification, navigation, and reasoning
Before we get started
Ontology pipeline
Use cases and requirements
Coverage model
Gathering from diverse formats and file types
Sourcing across organizational boundaries
Leveraging external and industry taxonomies
The capture, collection, and reconciliation process
Mapping to industry standard vocabularies
Why domain coverage is important
Coverage model as a framework
Transforming flat vocabulary into a structured hierarchy
What we’re building
From coverage model to hierarchy
Starting with top-level categories
The Is-a test
Taxonomy levels
Writing definitions
What makes a good definition
Managing alternative labels
The acronym rule
Structuring your taxonomy in a spreadsheet
Reading the hierarchy
How AI systems utilize taxonomies
Building a branch example
Common pitfalls and how to avoid them
Quality checklist
Why SKOS?
What SKOS adds to your spreadsheet taxonomy
Anatomy of a SKOS taxonomy
From spreadsheet to SKOS: A column-by-column translation
Documentation properties in practice
Connecting through mapping properties
The complete financial literacy branch
SKOS taxonomies as AI infrastructure
Governance and lifecycle
Lifecycle management
Tooling
Quality checklist for SKOS taxonomies
The thesaurus horizon
Conclusion
Explicit concept hierarchies and relations
Lightweight inference support
Synonym and mapping handling for lexical reasoning
Interoperability within neuro-symbolic architectures
Not all ontologies are created equal
Foundation for shared understanding
A shared vocabulary
Support iterative development
Why SKOS is enough
The mighty thesaurus
Next steps
The journey from controlled vocabulary to thesaurus
Conclusion
Why the confusion exists
Ontology Is philosophizing
The field lacks institutional standardization
The semantic web stack is not a relic of the past
Ontology as a marketing concept
It’s a model, it’s a language, it’s an expression
We don’t see the problem
A brief history
Ontologies are symbolic AI
The semantic web and the problem it solves
The building blocks: RDF and URIs
RDF triple
RDFS is the first step toward semantics
RDFa: The embedding syntax
OWL: Logical machinery
Are these ontologies? The standards debate
Semantic standards and ontological commitment
The spectrum and the pipeline
What is an ontology?
What’s not an ontology?
The scenario
Ontology design heuristics
The three-box architecture as design process
Building our workflow ontology framework
Integrating the CBox into the ontology
The SKOS-XL route
The complete ontology design framework
The metadata layer
Metadata as integration
Building the classes
Building the properties
The metadata layer at work: A business use case
Validating the model
NTWF ontology coverage
Governance
Competency questions as a governance instrument
Ontologies and AI
AI contributes to ontology development
The NTWF graph as an AI system registry
Ontology is the ground truth
RAG over structured graphs
AI-assisted ontology maintenance
Ontology as organizational memory
Conclusion
A knowledge graph is an architecture
Knowledge graph architecture components
Taxonomies and thesauri as a conceptual model
Bee taxonomy
The ontology: Formal logic and reasoning
NTWF workflow ontology
Vocabulary versus ontology
Ontologies, neural networks, and AI
The knowledge base
Metadata schemas
Requirements from the Gene Ontology (GO)
The validation layer
The query layer
Natural Language Processing (NLP)
Vector databases
Knowledge graph embeddings
Graph algorithms
Knowledge graph component summary
The Ontology Pipeline stages at a glance
The landscape of architectures
Three common RDF architectures
The enterprise knowledge graph
R2RML overview
Reasoning
NLP pipelines
Query layer
Deployments
The enterprise pattern’s tradeoff
The domain knowledge graph
The linked data knowledge graph
Deployments
Infrastructure and availability
The linked data pattern’s tradeoff
Architecture as commitment
Scoping the graph
Competency questions as the scoping device
Scoping by pattern
Enterprise knowledge graph
Domain knowledge graph
Linked data knowledge graph
Staffing the graph
Core roles
Staffing by pattern
Phases of implementation
Production infrastructure
Production governance
Maintaining semantic integrity
Ontology evolution
Vocabulary maintenance
Validation as continuous assurance
The AI feedback loop
The cost of semantic debt
Architecture is a practice
Conclusion
Ontology Pipeline gives data leaders, semantic engineers, ontologists, taxonomists, knowledge managers, librarians, AI architects, and enterprise data teams a practical framework for turning messy organizational language into structured, machine-readable knowledge. Instead of treating ontology development or knowledge graph construction as a black box, this book provides a clear sequence: controlled vocabularies, metadata schemas, taxonomies, thesauri, ontologies, and knowledge graphs.
Analyze how language, definitions, labels, synonyms, metadata, and relationships shape the performance of artificial intelligence systems. The book explains why large language models, retrieval-augmented generation, semantic search, entity resolution, and information retrieval all depend on clean, governed, semantically enriched data. It also shows how library and information science methods can help organizations build scalable semantic knowledge management systems that support both human understanding and machine reasoning.
Design each stage of the Ontology Pipeline with practical guidance, examples, standards, and implementation patterns. Readers will explore controlled vocabulary development, SKOS taxonomies, thesaurus relationships, RDF, OWL, SHACL, SPARQL, competency questions, ontology governance, semantic validation, and knowledge graph architecture. The book connects these concepts to real enterprise needs, including data quality, semantic layers, AI governance, domain modeling, knowledge management, and production AI infrastructure.
Evaluate what it really takes to build and maintain a knowledge graph as architecture, not just as a product or database. Learn why we build, govern, and maintain a knowledge graph with attention to semantic debt, staffing, operational funding, ontology change management, validation workflows, and long-term trust. For teams investing in AI, data governance, data catalogs, metadata management, enterprise architecture, or semantic technology, this book provides the missing roadmap.
Apply the Ontology Pipeline to create a formal, explicit, shared model of what your organization knows. Whether you are building a semantic layer, improving RAG accuracy, designing a domain ontology, creating an enterprise knowledge graph, or trying to make AI outputs more reliable, Ontology Pipeline provides the vocabulary, structure, and processes to move from disconnected data to governed organizational knowledge.
Jessica Talisman is a Semantic Engineer, Information Architect, and knowledge infrastructure strategist with more than 25 years of experience spanning enterprise architecture, e-commerce content systems, digital libraries, and knowledge management.
She is the creator of the Ontology Pipeline™, a structured framework for building semantic knowledge infrastructure from first principles, moving progressively from controlled vocabularies to taxonomies, thesauri, ontologies, and fully realized knowledge graphs. She has led semantic architecture initiatives at Adobe, where she architected an RDF-based knowledge graph supporting the Digital Experience ecosystem, and at Amazon, where she worked in information architecture and taxonomy.
She is the founder of Contextually LLC, a consulting and coaching practice specializing in ontology modeling, NLP integration, knowledge graphs, and knowledge infrastructure design, and of The Knowledge Graph Academy, a cohort-based program that trains future semantic engineers and ontologists through a balance of theory and practice. She writes regularly on her Substack newsletter, Intentional Arrangement, where her work explores the relationship between semantic systems and AI.
Please complete all fields.