Extracts the terms a document set actually uses and builds a canonical glossary of one term per concept, flagging synonyms that must collapse and homonyms that must split, with each definition traced to the file:line where it was read or marked as unverified rather than invented. It treats the glossary as an index, not as a store - the definition is repeated at first use in every document, because a term whose meaning lives only in a remote glossary recreates the working-memory tax the corpus exists to remove. Use when a document set needs a canonical vocabulary, or when the same concept is being called several things.