Skip to content

Term Counter (Glossary QA)

The term counter is a quality-assurance tool for glossaries. It finds every occurrence of the glossary’s terms in a source text and, when you also send a translation, checks whether the target-language entries of those terms appear in it. Use it to verify that a translation followed the glossary, whether it was produced by the platform, another tool, or a human translator.

The counter doesn’t translate anything and doesn’t call a language model: results are deterministic and don’t consume translation quota.

Endpoint: POST /translation/glossaries/{id}/terms/count

  • id: The glossary ID

Sent as JSON.

Field Type Required Description
sourceLanguage string yes Language code of sourceText. Must be one of the glossary’s languages
sourceText string yes Text to search, up to 100,000 characters
targetLanguage string with targetText Language code of targetText. Must be one of the glossary’s languages
targetText string with targetLanguage Translation to check, up to 100,000 characters

targetLanguage and targetText are sent together or not at all.

  • Letter case follows each entry’s caseSensitivity setting. By default, uppercase letters of an entry must be uppercase in the text, and lowercase letters may be either case: IT matches “IT” but not “it”, and term matches “Term” and “TERM”.
  • Grammatical forms count as the term. Plurals, cases and conjugations are recognized through each word’s dictionary form (lemma): “policies” counts as policy, “mice” as mouse, “Rechenzentren” as Rechenzentrum, “машины” as машина, “luces” as luz, “والكتاب” as كتاب. In multi-word entries every word may change: “машинного обучения” counts as машинное обучение. Each match reports whether it is the exact entry or an inflected form. See Word Forms.
  • Words the dictionary doesn’t know (product names, abbreviations, jargon) fall back to a list of common endings per language: “APIs” counts as API. Abbreviations are never reduced to a dictionary word, so IT doesn’t match the pronoun “it”.
  • Whole words only. A term inside a longer word doesn’t count: API isn’t found in “RAPID”. Chinese, Japanese and Thai have no spaces between words, and Korean attaches particles to words, so for these languages a word segmenter decides where words start and end. See Languages Without Spaces.
  • Full-width and half-width characters match their normal forms: “API” counts as API.
  • Separators: spaces, hyphens and line breaks between the words of a multi-word entry are interchangeable, so machine learning matches “Machine-learning”.
  • Overlaps: where entries overlap, the longest one wins. “API keys” counts once, as API key, not also as API.
  • Accents must match, except in text written entirely in capitals, where they are often dropped (“ETAT” matches état). A form that changes a letter’s accent is still a form of the word: “Äpfel” counts as Apfel.
  • Target side: only the target-language entries of terms found in the source are checked.

Success Response (200)

The example checks a German translation against a glossary with the terms IT, API key (German: API-Schlüssel) and iPhone.

Request:

{
"sourceLanguage": "en",
"sourceText": "Our IT team rotates API keys for every iPhone app. WE FIXED IT.",
"targetLanguage": "de",
"targetText": "Unser IT-Team rotiert API-Schlüssel für jede Iphone-App."
}

Response:

{
"status": "ok",
"timestamp": "2025-01-12T22:31:48.856Z",
"data": {
"source": {
"language": "en",
"count": 3,
"uniqueTerms": 3,
"matches": [
{ "termId": "cmterm_it", "glossaryTerm": "IT", "caseSensitivity": "auto", "form": "exact", "text": "IT", "start": 4, "end": 6 },
{ "termId": "cmterm_apikey", "glossaryTerm": "API key", "caseSensitivity": "auto", "form": "inflected", "text": "API keys", "start": 20, "end": 28 },
{ "termId": "cmterm_iphone", "glossaryTerm": "iPhone", "caseSensitivity": "auto", "form": "exact", "text": "iPhone", "start": 39, "end": 45 }
],
"ambiguousCount": 1,
"ambiguousMatches": [
{ "termId": "cmterm_it", "glossaryTerm": "IT", "caseSensitivity": "auto", "form": "exact", "text": "IT", "start": 60, "end": 62, "reason": "all_caps_context" }
]
},
"target": {
"language": "de",
"count": 2,
"uniqueTerms": 2,
"matches": [
{ "termId": "cmterm_it", "glossaryTerm": "IT", "caseSensitivity": "auto", "form": "exact", "text": "IT", "start": 6, "end": 8 },
{ "termId": "cmterm_apikey", "glossaryTerm": "API-Schlüssel", "caseSensitivity": "auto", "form": "exact", "text": "API-Schlüssel", "start": 22, "end": 35 }
],
"ambiguousCount": 0,
"ambiguousMatches": []
},
"terms": [
{
"termId": "cmterm_it",
"sourceTerm": "IT",
"targetTerm": "IT",
"sourceCount": 1,
"sourceAmbiguousCount": 1,
"targetCount": 1,
"targetAmbiguousCount": 0,
"targetCaseMismatches": [],
"targetSpellingVariants": []
},
{
"termId": "cmterm_apikey",
"sourceTerm": "API key",
"targetTerm": "API-Schlüssel",
"sourceCount": 1,
"sourceAmbiguousCount": 0,
"targetCount": 1,
"targetAmbiguousCount": 0,
"targetCaseMismatches": [],
"targetSpellingVariants": []
},
{
"termId": "cmterm_iphone",
"sourceTerm": "iPhone",
"targetTerm": "iPhone",
"sourceCount": 1,
"sourceAmbiguousCount": 0,
"targetCount": 0,
"targetAmbiguousCount": 0,
"targetCaseMismatches": [
{ "text": "Iphone", "start": 45, "end": 51 }
],
"targetSpellingVariants": []
}
]
}
}

Reading the example:

  • “IT” in “WE FIXED IT” is ambiguous: in text written in capitals, the abbreviation can’t be told apart from the pronoun “it”.
  • “API keys” is one match for API key (an inflected form), not also a match for API.
  • The translation wrote “Iphone”: it isn’t counted (targetCount: 0) and is reported in targetCaseMismatches, so the glossary wasn’t followed for iPhone.

source and target have the same shape. target and terms are null when no target is sent.

Field Type Description
language string Language code of the text
count integer Number of confident matches (total occurrences)
uniqueTerms integer Number of distinct glossary terms among the confident matches
matches array Confident matches, in text order
ambiguousCount integer Number of ambiguous matches
ambiguousMatches array Matches that can’t be decided from the text alone, each with a reason

Each match:

Field Type Description
termId string ID of the glossary term
glossaryTerm string The entry as stored in the glossary
caseSensitivity string The entry’s letter case setting applied to this match
form string exact: the entry as written (up to letter case and separators). inflected: another grammatical form of it (“keys”, “Rechenzentren”, 食べた)
text string The text exactly as it appears in the input
start, end integer Position in the input text, in Unicode code points (end is exclusive)
reason string Only in ambiguousMatches, see below

Ambiguity reasons:

reason Meaning Example
all_caps_context An entry with capital letters, found in text written in capitals IT in “WE FIXED IT”
initial_capital_only An entry whose only capital is its first letter, found in lowercase. Only with the auto setting Apple found as “apple”
lemma_only A word that shares only its dictionary form with the entry, which can be a different word “see” for the entry saw (the tool)
segmentation_boundary The term was found inside a longer word (Chinese, Japanese, Thai, Korean) 学习 (learning) inside “学习者” (learner)

Each item of terms compares one term found in the source with its translation:

Field Type Description
termId string ID of the glossary term
sourceTerm string Source-language entry
targetTerm string or null Target-language entry; null when the term has no entry in the target language
sourceCount, sourceAmbiguousCount integer Confident and ambiguous occurrences in the source
targetCount, targetAmbiguousCount integer Confident and ambiguous occurrences in the target
targetCaseMismatches array Occurrences of the target entry written with the wrong letter case (text, start, end)
targetSpellingVariants array Occurrences of the target entry in a different spelling of the same word (text, start, end). Japanese only, for example サーバ for サーバー

A term was followed when targetCount covers sourceCount. Compare the counts the way your content needs: a translation can legitimately use a term fewer times, for example when two sentences are merged.

Status When
400 Invalid body; targetLanguage without targetText or the other way around; a language code that isn’t supported; or a language that isn’t one of the glossary’s languages, for example Target language 'ru' is not in glossary cmglossary123. Available languages: en, es, de
401 Missing or invalid API key
404 The glossary doesn’t exist or belongs to another organization
Terminal window
curl -X POST "https://platform.algebras.ai/api/v1/translation/glossaries/cmglossary123/terms/count" \
-H "X-Api-Key: your_api_key_here" \
-H "Content-Type: application/json" \
-d '{
"sourceLanguage": "en",
"sourceText": "Our IT team rotates API keys for every iPhone app.",
"targetLanguage": "de",
"targetText": "Unser IT-Team rotiert API-Schlüssel für jede Iphone-App."
}'

A glossary keeps one form of each term per language, so that translations stay consistent. The counter still has to recognize the term when grammar changes its form. It does this in two passes:

  1. The entry as written, with letter case, separators and a list of common endings per language taken into account (“APIs” for API, “Straßen” for Straße).
  2. Dictionary forms (lemmas) for the rest of the text. Each word of the text and of the entry is reduced to its dictionary form, and the sequences are compared word by word.
Language Text Entry Result
English “policies”, “mice”, “children”, “went” policy, mouse, child, go inflected
German “Rechenzentren”, “Äpfel” Rechenzentrum, Apfel inflected
Russian “машины”, “машинного обучения” машина, машинное обучение inflected
Spanish “luces” luz inflected
Arabic, Hebrew “والكتاب”, “והספר” كتاب, ספר inflected (attached prefixes)
Japanese 食べた 食べる inflected
English “see”, “seeing” saw (the tool) ambiguous, lemma_only

The last row shows the limit of dictionary forms: different words can share one. When the entry word isn’t a dictionary form itself and the two forms share little of their beginning, the match is reported as ambiguous with the reason lemma_only instead of being counted.

Dictionary forms are available for about 50 languages, including all major European languages, Russian, Ukrainian, Turkish, Arabic, Hebrew, Persian, Hindi, Indonesian and Malay. For other languages, only the first pass applies.

Japanese often writes the same word in more than one way: サーバ and サーバー (server), いく and 行く (to go). Because a glossary exists to keep terminology consistent, a different spelling of the term isn’t counted. On the target side it’s listed in targetSpellingVariants, so the translation can be corrected to the glossary’s spelling. Conjugated forms of the same spelling (食べた for 食べる) do count.

Chinese, Japanese and Thai are written without spaces between words, and Korean attaches particles directly to words. For these languages the counter finds the term in the text and then asks a word segmenter whether the term starts and ends on word boundaries:

Language Text Entry Result
Chinese 我们使用机器学习模型 机器学习 match
Chinese 学习者 (learner) 学习 (learning) ambiguous, segmentation_boundary
Japanese 機械学習モデルを 機械学習 match
Thai ผู้เรียนรู้ เรียนรู้ ambiguous, segmentation_boundary
Korean 머신러닝은 좋다 머신러닝 match, without the particle 은

A term split into several words by the segmenter still matches, as long as its first and last characters are on word boundaries: 机器学习 is two dictionary words, 机器 and 学习. Segmenters can be wrong about new vocabulary, so a term found inside a longer word is reported as ambiguous rather than dropped.

Full-width letters and digits, common in Chinese and Japanese text, match their normal forms: “APIキー” contains API.

For Lao, Khmer, Burmese and Tibetan, terms are found anywhere in the text, without a boundary check.

Capitalization can change what a word means: “IT” (information technology) and “it” (the pronoun), “mW” (milliwatt) and “MW” (megawatt). It can also be irrelevant: “Email” at the start of a sentence is still “email”. The caseSensitivity setting tells the term counter which applies to a glossary entry.

The setting belongs to each language entry of a term, so the English and German entries of the same term can differ. It’s used only by the term counter and doesn’t change how translations are generated.

caseSensitivity Rule Use for Example
auto (default) Uppercase letters must match; lowercase letters may be either case. If only the first letter is capitalized, a lowercase occurrence is reported as ambiguous for review Most terms IT ≠ “it” · term = “Term” = “TERM” · Apple → “apple”: ambiguous
capitals Uppercase letters must match; lowercase letters may be either case. No ambiguity for an initial capital: it is always meaningful Brand and product names Apple ≠ “apple” · iPhone = “IPHONE”
exact Every letter must match exactly Units, symbols, code identifiers mW ≠ “MW” · pH ≠ “PH”
ignore Any capitalization counts; matches are never ambiguous Ordinary words that can’t be confused with other words email = “Email” = “EMAIL”

In every mode, accents must match, except in text written entirely in capitals.

Entry Text auto capitals exact ignore
IT “it works” no match no match no match match
IT “Contact the IT team” match match match match
IT “WE FIXED IT” ambiguous ambiguous ambiguous match
Apple “an apple” ambiguous no match no match match
term “Term”, “TERM” match match no match match
iPhone “IPHONE” match match no match match
mW “5 MW” match match no match match

The mW row shows why exact exists: with any other mode, “MW” (megawatt) would be counted as mW (milliwatt).

On the target side, an occurrence rejected only because of its letter case is listed in targetCaseMismatches. With exact, a translation that writes “Iphone” or “IPHONE” for iPhone shows up there.

Send caseSensitivity with a definition when you create, bulk-create or update a term. Omitting it on create means auto; omitting it on update keeps the current value; sending auto resets an entry to the default. Every response that returns definitions includes the value.

Create a unit term that must match exactly in every language:

Terminal window
curl -X POST "https://platform.algebras.ai/api/v1/translation/glossaries/cmglossary123/terms" \
-H "X-Api-Key: your_api_key_here" \
-H "Content-Type: application/json" \
-d '{
"definitions": [
{ "language": "en", "term": "mW", "definition": "Milliwatt", "caseSensitivity": "exact" },
{ "language": "de", "term": "mW", "definition": "Milliwatt", "caseSensitivity": "exact" }
]
}'

Change only the mode of the English entry of an existing term. The term text and definition stay as they are:

Terminal window
curl -X PUT "https://platform.algebras.ai/api/v1/translation/glossaries/cmglossary123/terms/cmterm123" \
-H "X-Api-Key: your_api_key_here" \
-H "Content-Type: application/json" \
-d '{ "definitions": [{ "language": "en", "caseSensitivity": "capitals" }] }'

An unknown value returns 400 with caseSensitivity must be one of: auto, capitals, exact, ignore.

The setting is also available in the app, in the Letter case column of a glossary (see Glossaries), and through the MCP term tools (create_glossary_term, bulk_create_glossary_terms, update_glossary_term).