# Enter the Knowledge Graph: Workshop Resources
---
**Featured Image:**
[A person standing next to several large computer screens with charts] (https://www.archeredu.com/wp-content/uploads/2026/05/HEMJ-Why-Speed-Is-the-New-Personalization-in-Higher-Ed-Marketing-600x245-1.jpg)
---
**Author:** Ray Martinez
**Published:** September 1, 2026
**Updated:** October 1, 2026
---
This post contains resources from Ray Martinez's Content Marketing World 2026 presentation.
I ran this workshop against archeredu.com before I ran it in front of a room. Our custom lexicon has 18 terms that define how we talk about enrollment marketing. Google extracted 14 of them cleanly. Not one came back with a Knowledge Graph ID or a Wikipedia link.
That gap between extraction and understanding is what the ninety minutes were about. Extraction means a machine pulled your term off the page. Grounding means it knows what the term refers to. Four of our terms came back perfectly and still failed, because clean extraction with empty metadata leaves a machine holding a string with no identity and nothing to cite.
The templates and code are below, along with the spec. The method behind them is in two earlier pieces: [how to identify and prioritize entity gaps] (https://searchengineland.com/schema-ai-search-identify-prioritize-entity-gaps-482728) , and how to turn [entity gaps into organic growth] (https://searchengineland.com/schema-ai-search-entity-gaps-organic-growth-489462) .
## Build the roster before the calendar.
An entity map is a roster, not a keyword list. It is the set of things your institution is, sells, competes on, and wants to be known for, written down with consistent names and sorted into layers.
Most teams already have a roster and don’t know it. It is in your top nav, your sales deck, your glossary and FAQ pages, and the acronyms your team uses in Slack that never appear in writing. Nobody invents ten entities from a blank page. You are taking inventory of what you already say.
Keywords and entities overlap here, and that trips people up. Most of your roster is both. A keyword is how something gets asked. An entity is what gets asked about.
### The roster, in a format a machine will read
One row per entity. An identifier, the name, the layer it belongs to, and a plain-English definition. That last field does the work, and you almost certainly wrote it already on a glossary page. It is not a place for keywords.
**ARCHER-ENTITY-CORPUS.JSONL**
{"id":"AEL-003","name":"Online Program Management","layer":"custom_lexicon","embedding_text":"A service model in which an external partner provides marketing, recruitment, and student support for a university's online degree programs, usually under a revenue-share agreement."} {"id":"AEL-004","name":"OPM Alternative","layer":"custom_lexicon","embedding_text":"A fee-for-service engagement model that delivers enrollment marketing without the long-term revenue-share contract of a traditional OPM."} {"id":"AEL-006","name":"Generative Engine Optimization","layer":"custom_lexicon","embedding_text":"The practice of structuring content and entity data so that generative answer engines retrieve, understand, and cite a source."} {"id":"AEL-010","name":"Stealth Applicant","layer":"custom_lexicon","embedding_text":"A prospective student who researches a program and submits an application without ever completing a request-for-information form, leaving no attributable inquiry in the funnel."} {"id":"AEL-011","name":"Cost per Enrollment","layer":"measurement","embedding_text":"Total marketing spend divided by the number of enrolled students, used to compare channel efficiency across an enrollment cycle."}
**One record per line.** The wrapping above is the page, not the format. JSONL is newline-delimited JSON, so a pretty-printed file will not parse. If you are unsure which serialization to use, [Knowledge Graph Navigator] (https://knowledgegraphnavigator.com) (a reference site I maintain) covers JSONL, JSON-LD, RDFa, and N-Quads side by side.
## Close what the machine misread, and name what you cannot close
Every term lands in one of three buckets. It comes back credited in full, credited wrong, or not on the sleeve at all.
Credited wrong is the one worth your attention. Our term “stealth applicant” describes a prospective student who researches and applies without ever filling out an RFI form. Google returned “applicant,” typed as a person. The behavior that is the entire point of the term did not survive the trip.
**WHAT CAME BACK FROM ANALYZEENTITIES**
{ "name": "applicant", "type": "PERSON", "salience": 0.031, "metadata": {}, "mentions": [ { "text": { "content": "Stealth Applicant", "beginOffset": 412 }, "type": "COMMON" } ] }
The empty metadata object is the finding. When a term resolves to something Google already knows, that object carries a Knowledge Graph identifier and a Wikipedia URL. Ours carried nothing. Sixteen of our eighteen terms came back the same way.
### Fix one: define the term so it can be reassembled.
Publish the definition as a DefinedTermSet on a glossary hub. This is what a machine reads instead of guessing from body copy.
**ON /GLOSSARY/: THE DEFINITION ITSELF**
{ "@context": "https://schema.org", "@type": "DefinedTermSet", "@id": "https://www.archeredu.com/glossary/#archer-enrollment-lexicon", "name": "Archer Enrollment Lexicon", "hasDefinedTerm": [ { "@type": "DefinedTerm", "@id": "https://www.archeredu.com/glossary/stealth-applicant/#term", "name": "Stealth Applicant", "termCode": "AEL-010", "description": "A prospective student who researches a program and submits an application without ever completing a request-for-information form, leaving no attributable inquiry in the funnel.", "inDefinedTermSet": "https://www.archeredu.com/glossary/#archer-enrollment-lexicon" } ] }
### Fix two: point every page that uses the term to that definition
A definition nobody links to is a page nobody reads. Use about for what the page is actually about, and mentions for the supporting cast.
**ON AN ARTICLE THAT USES THE TERM**
{ "@context": "https://schema.org", "@type": "Article", "@id": "https://www.archeredu.com/hemj/stealth-applicants-enrollment-funnel/#article", "headline": "Stealth applicants are already in your funnel", "about": { "@id": "https://www.archeredu.com/glossary/stealth-applicant/#term" }, "mentions": [ { "@id": "https://www.archeredu.com/glossary/enrollment-funnel/#term" }, { "@id": "https://www.archeredu.com/glossary/request-for-information/#term" } ] }
### Fix three: ground the term to something outside your own site
Extraction can succeed. Grounding can still fail. sameAs tells a machine that your term is the same thing as an entry it already trusts. Only two of our eighteen terms had anywhere to point.
**GROUNDING A TERM THAT DOES HAVE AN EXTERNAL ANCHOR**
{ "@type": "DefinedTerm", "@id": "https://www.archeredu.com/glossary/online-program-management/#term", "name": "Online Program Management", "alternateName": ["OPM"], "termCode": "AEL-003", "sameAs": [ "https://en.wikipedia.org/wiki/Online_program_management", "https://www.wikidata.org/wiki/Q<ID>" ] }
**Verify before you ship.** Confirm the exact Wikipedia title and look up the real Wikidata Q-number rather than copying the placeholder. A sameAs pointing at the wrong entity is worse than no sameAs, because you have now asserted something false in a format built for assertions. For terms with no external anchor yet, publish the definition first and add sameAs when an entry appears.
### The rule that governs all three
Marked-up claims have to match the visible copy on the page. If they do not, you are lying to a machine that keeps receipts.
## Where the program-level properties come from
Schema.org covers a program page thinly. It gives you a name, a provider, and a credential, then stops short of the things a prospective student actually decides on. We published an extension to close that gap, and it is open at [schema.archeredu.com] (https://schema.archeredu.com) .
It is 85 properties: 23 from schema.org and 62 custom ones, prioritized by what answer engines appear to use. The namespace is declared in the context and then used like any other prefix.
**A PROGRAM PAGE USING BOTH VOCABULARIES**
{ "@context": { "schema": "https://schema.org/", "archeredu": "https://schema.archeredu.com/" }, "@type": "schema:EducationalOccupationalProgram", "@id": "https://example.edu/nursing/pmc-pmhnp/#program", "schema:name": "Post-Master's Certificate: Psychiatric Mental Health Nurse Practitioner", "schema:provider": { "@id": "https://example.edu/#organization" }, "schema:timeToComplete": "P15M", "archeredu:programmaticAccreditation": { "@type": "schema:Organization", "name": "Commission on Collegiate Nursing Education", "sameAs": "https://www.aacnnursing.org/CCNE" } }
Two details matter more than they look. Program duration goes in as an ISO 8601 value, so “15 months” becomes P15M and stops being a string a machine has to parse out of a sentence. And accreditation resolves to an organization with its own sameAs rather than sitting on the page as text.
The full vocabulary file is at [schema.archeredu.com/education-schema.jsonld] (https://schema.archeredu.com/education-schema.jsonld) , and the property guidance by degree level is on the site.
## Measure coverage, not just rankings
We scored 26 entities across 2,770 crawled pages: our own 290, and two competitors at 1,909 and 571. Counting mentions would have rewarded whoever published the most, so we measured semantic depth against the pooled distribution instead. Parity with the peer pool is 1.0.
On generative engine optimization, our depth index came back at 9.04 against roughly 0.06 for both competitors. On the full roster, we scored 58. They scored 67 and 62.
| Entity AEL-006 | Depth index | Keyword pages | Tier |
|---|
| Archer | 9.04 | 12 | Strong |
| Competitor A | 0.06 | 0 | Partial |
| Competitor B | 0.07 | 1 | Partial |
Uncomfortable publishing, which is why it is here. The audit also found a competitor with 25 pages built around a concept we treat as a flagship service, where we had zero. Not thin pages. Zero. Our DefinedTerm markup for five other clusters was already written and had never shipped as pages.
### The four moves, and the spec behind them
**Seed:** find the pages genuinely about the term, pooled across every crawl, with a minimum of two. **Center:** average those pages into a single point that represents what the term means in this market. **Measure:** how many of each site’s pages sit near that point, against how many should. **Tier:** strong, partial, or missing, scored on whichever is stronger, the semantic signal or the keyword evidence.
You do not have to write that yourself. Hand the block below to Antigravity, Claude Code, or Codex. It comes out at roughly 300 lines of Python either way, and the tool you pick matters less than the spec you hand it.
**THE SESSION SHEET**
Build a quarterly entity coverage audit. Python, no framework. CORPUS Input: JSONL, one row per entity. Fields: id, name, layer, embedding_text. Validate on load. Fail loudly on a missing field. API CONTRACT Google Cloud Natural Language, analyzeEntities. Service-account auth, credentials from env, never inline. Retry with exponential backoff on 429 and 5xx. Cap at 5 attempts. Capture the full metadata object per entity. Empty is a result, not an error. SCORING RULES, in priority order 1. Exact match on the full term. 2. Multi-word substring match. 3. Known acronym match from an explicit alias list. 4. Stoplist of generic single words, applied before 2 and 3. Without it "Agency of Record" scores on any page containing the word "record". OUTPUTS One CSV for humans: entity, site, tier, depth index, keyword pages. One N-Quads file for machines, under a named graph URI that encodes the run date. GUARDRAILS Rate-limit pause between calls. --dry-run validates the corpus and spends no API calls. Hard cap on total calls. Exit non-zero rather than truncate. DEFINITION OF DONE Re-runnable quarterly with a new graph URI per run. A diff report between any two runs, by entity and by tier.
**The diff is the point.** Without a run-over-run comparison, you have a script. With a run-over-run comparison, you have an instrument.
## The resources
### If you are building the strategy
- [Roster template] (https://docs.google.com/spreadsheets/d/1zTNGSVDU1AE3bXjEqfCXbPCCw51isDwMxjsexFfnoDw/edit?usp=sharing) , the three-layer worksheet from Exercise One: core, services, custom lexicon
- [Queue template] (https://docs.google.com/spreadsheets/d/1IgFLsmiBOPIjRfzGnXSJCmEb8VutncYii8xGwB0RRIQ/edit?usp=sharing) , covering the entity gap a piece closes, the pillar it belongs to, the competitor position, the priority tier, and the internal links it ships with
- [Workshop slides] (https://docs.google.com/presentation/u/0/d/19na2B-r2SEal6_mJRJc8L8Wix19crJcY-OrnwpJ_V2c/edit)
### If you are building the pipeline
- [The session sheet] (https://docs.google.com/document/d/1SCxWs9Ry7hq8xmGcrzE3gucjkm5f02QmSeNotLoSrxY/edit?usp=sharing) , the block above as a file you can hand to an agent
- [Example corpus rows] (https://drive.google.com/file/d/1bYuqCrs1syg6TSyx10rwQCIN0CFFMEQZ/view?usp=sharing) , the five JSONL rows above plus five more
- [Google Cloud Natural Language API documentation] (https://cloud.google.com/natural-language/docs) , the analyzeEntities endpoint and the metadata object where a Knowledge Graph ID either appears or does not
### Reference
- [schema.archeredu.com] (https://schema.archeredu.com) , our published extension for higher education program discovery: 85 properties with guidance by degree level
- [Knowledge Graph Navigator] (https://knowledgegraphnavigator.com) , a reference site I maintain. Start with Languages and Formats if you are deciding how to serialize your roster, or Core Concepts if terms like named graph and N-Quads are new
- [schema.org/DefinedTermSet] (https://schema.org/DefinedTermSet) and [schema.org/sameAs] (https://schema.org/sameAs)
[Wikidata] (https://www.wikidata.org) , to check whether your term has anything to point at yet
## Treat the graph as infrastructure
None of this is an overnight play for AI citations. Coverage is a balance sheet. It is not a campaign metric. The work you ship this quarter keeps working next quarter.
Your first week is five days. Write the roster on Monday. Read your own pages on Tuesday and find out which terms your site actually defines. Check two competitors on Wednesday. On Thursday, pick the three entities nobody owns. Write one headline that connects two of them on Friday.
We had already written the hard part and never built the pages. Most teams have.
If you want this run against your own institution, that is [what we do] (https://www.archeredu.com/ai-ready-organic-strategy/) .