Term Counter (Glossary QA)
The term counter is a quality-assurance tool for glossaries. It finds every occurrence of the glossary’s terms in a source text and, when you also send a translation, checks whether the target-language entries of those terms appear in it. Use it to verify that a translation followed the glossary, whether it was produced by the platform, another tool, or a human translator.
The counter doesn’t translate anything and doesn’t call a language model: results are deterministic and don’t consume translation quota.
Count Terms
Section titled “Count Terms”Endpoint: POST /translation/glossaries/{id}/terms/count
Path Parameters
Section titled “Path Parameters”id: The glossary ID
Request Body
Section titled “Request Body”Sent as JSON.
| Field | Type | Required | Description |
|---|---|---|---|
sourceLanguage |
string | yes | Language code of sourceText. Must be one of the glossary’s languages |
sourceText |
string | yes | Text to search, up to 100,000 characters |
targetLanguage |
string | with targetText |
Language code of targetText. Must be one of the glossary’s languages |
targetText |
string | with targetLanguage |
Translation to check, up to 100,000 characters |
targetLanguage and targetText are sent together or not at all.
How Terms Are Matched
Section titled “How Terms Are Matched”- Letter case follows each entry’s
caseSensitivitysetting. By default, uppercase letters of an entry must be uppercase in the text, and lowercase letters may be either case:ITmatches “IT” but not “it”, andtermmatches “Term” and “TERM”. - Grammatical forms count as the term. Plurals, cases and conjugations are recognized through each word’s dictionary form (lemma): “policies” counts as
policy, “mice” asmouse, “Rechenzentren” asRechenzentrum, “машины” asмашина, “luces” asluz, “والكتاب” asكتاب. In multi-word entries every word may change: “машинного обучения” counts asмашинное обучение. Each match reports whether it is theexactentry or aninflectedform. See Word Forms. - Words the dictionary doesn’t know (product names, abbreviations, jargon) fall back to a list of common endings per language: “APIs” counts as
API. Abbreviations are never reduced to a dictionary word, soITdoesn’t match the pronoun “it”. - Whole words only. A term inside a longer word doesn’t count:
APIisn’t found in “RAPID”. Chinese, Japanese and Thai have no spaces between words, and Korean attaches particles to words, so for these languages a word segmenter decides where words start and end. See Languages Without Spaces. - Full-width and half-width characters match their normal forms: “API” counts as
API. - Separators: spaces, hyphens and line breaks between the words of a multi-word entry are interchangeable, so
machine learningmatches “Machine-learning”. - Overlaps: where entries overlap, the longest one wins. “API keys” counts once, as
API key, not also asAPI. - Accents must match, except in text written entirely in capitals, where they are often dropped (“ETAT” matches
état). A form that changes a letter’s accent is still a form of the word: “Äpfel” counts asApfel. - Target side: only the target-language entries of terms found in the source are checked.
Response
Section titled “Response”Success Response (200)
The example checks a German translation against a glossary with the terms IT, API key (German: API-Schlüssel) and iPhone.
Request:
{ "sourceLanguage": "en", "sourceText": "Our IT team rotates API keys for every iPhone app. WE FIXED IT.", "targetLanguage": "de", "targetText": "Unser IT-Team rotiert API-Schlüssel für jede Iphone-App."}Response:
{ "status": "ok", "timestamp": "2025-01-12T22:31:48.856Z", "data": { "source": { "language": "en", "count": 3, "uniqueTerms": 3, "matches": [ { "termId": "cmterm_it", "glossaryTerm": "IT", "caseSensitivity": "auto", "form": "exact", "text": "IT", "start": 4, "end": 6 }, { "termId": "cmterm_apikey", "glossaryTerm": "API key", "caseSensitivity": "auto", "form": "inflected", "text": "API keys", "start": 20, "end": 28 }, { "termId": "cmterm_iphone", "glossaryTerm": "iPhone", "caseSensitivity": "auto", "form": "exact", "text": "iPhone", "start": 39, "end": 45 } ], "ambiguousCount": 1, "ambiguousMatches": [ { "termId": "cmterm_it", "glossaryTerm": "IT", "caseSensitivity": "auto", "form": "exact", "text": "IT", "start": 60, "end": 62, "reason": "all_caps_context" } ] }, "target": { "language": "de", "count": 2, "uniqueTerms": 2, "matches": [ { "termId": "cmterm_it", "glossaryTerm": "IT", "caseSensitivity": "auto", "form": "exact", "text": "IT", "start": 6, "end": 8 }, { "termId": "cmterm_apikey", "glossaryTerm": "API-Schlüssel", "caseSensitivity": "auto", "form": "exact", "text": "API-Schlüssel", "start": 22, "end": 35 } ], "ambiguousCount": 0, "ambiguousMatches": [] }, "terms": [ { "termId": "cmterm_it", "sourceTerm": "IT", "targetTerm": "IT", "sourceCount": 1, "sourceAmbiguousCount": 1, "targetCount": 1, "targetAmbiguousCount": 0, "targetCaseMismatches": [], "targetSpellingVariants": [] }, { "termId": "cmterm_apikey", "sourceTerm": "API key", "targetTerm": "API-Schlüssel", "sourceCount": 1, "sourceAmbiguousCount": 0, "targetCount": 1, "targetAmbiguousCount": 0, "targetCaseMismatches": [], "targetSpellingVariants": [] }, { "termId": "cmterm_iphone", "sourceTerm": "iPhone", "targetTerm": "iPhone", "sourceCount": 1, "sourceAmbiguousCount": 0, "targetCount": 0, "targetAmbiguousCount": 0, "targetCaseMismatches": [ { "text": "Iphone", "start": 45, "end": 51 } ], "targetSpellingVariants": [] } ] }}Reading the example:
- “IT” in “WE FIXED IT” is ambiguous: in text written in capitals, the abbreviation can’t be told apart from the pronoun “it”.
- “API keys” is one match for
API key(aninflectedform), not also a match forAPI. - The translation wrote “Iphone”: it isn’t counted (
targetCount: 0) and is reported intargetCaseMismatches, so the glossary wasn’t followed foriPhone.
Response Fields
Section titled “Response Fields”source and target have the same shape. target and terms are null when no target is sent.
| Field | Type | Description |
|---|---|---|
language |
string | Language code of the text |
count |
integer | Number of confident matches (total occurrences) |
uniqueTerms |
integer | Number of distinct glossary terms among the confident matches |
matches |
array | Confident matches, in text order |
ambiguousCount |
integer | Number of ambiguous matches |
ambiguousMatches |
array | Matches that can’t be decided from the text alone, each with a reason |
Each match:
| Field | Type | Description |
|---|---|---|
termId |
string | ID of the glossary term |
glossaryTerm |
string | The entry as stored in the glossary |
caseSensitivity |
string | The entry’s letter case setting applied to this match |
form |
string | exact: the entry as written (up to letter case and separators). inflected: another grammatical form of it (“keys”, “Rechenzentren”, 食べた) |
text |
string | The text exactly as it appears in the input |
start, end |
integer | Position in the input text, in Unicode code points (end is exclusive) |
reason |
string | Only in ambiguousMatches, see below |
Ambiguity reasons:
reason |
Meaning | Example |
|---|---|---|
all_caps_context |
An entry with capital letters, found in text written in capitals | IT in “WE FIXED IT” |
initial_capital_only |
An entry whose only capital is its first letter, found in lowercase. Only with the auto setting |
Apple found as “apple” |
lemma_only |
A word that shares only its dictionary form with the entry, which can be a different word | “see” for the entry saw (the tool) |
segmentation_boundary |
The term was found inside a longer word (Chinese, Japanese, Thai, Korean) | 学习 (learning) inside “学习者” (learner) |
Each item of terms compares one term found in the source with its translation:
| Field | Type | Description |
|---|---|---|
termId |
string | ID of the glossary term |
sourceTerm |
string | Source-language entry |
targetTerm |
string or null | Target-language entry; null when the term has no entry in the target language |
sourceCount, sourceAmbiguousCount |
integer | Confident and ambiguous occurrences in the source |
targetCount, targetAmbiguousCount |
integer | Confident and ambiguous occurrences in the target |
targetCaseMismatches |
array | Occurrences of the target entry written with the wrong letter case (text, start, end) |
targetSpellingVariants |
array | Occurrences of the target entry in a different spelling of the same word (text, start, end). Japanese only, for example サーバ for サーバー |
A term was followed when targetCount covers sourceCount. Compare the counts the way your content needs: a translation can legitimately use a term fewer times, for example when two sentences are merged.
Errors
Section titled “Errors”| Status | When |
|---|---|
400 |
Invalid body; targetLanguage without targetText or the other way around; a language code that isn’t supported; or a language that isn’t one of the glossary’s languages, for example Target language 'ru' is not in glossary cmglossary123. Available languages: en, es, de |
401 |
Missing or invalid API key |
404 |
The glossary doesn’t exist or belongs to another organization |
Example Request
Section titled “Example Request”curl -X POST "https://platform.algebras.ai/api/v1/translation/glossaries/cmglossary123/terms/count" \ -H "X-Api-Key: your_api_key_here" \ -H "Content-Type: application/json" \ -d '{ "sourceLanguage": "en", "sourceText": "Our IT team rotates API keys for every iPhone app.", "targetLanguage": "de", "targetText": "Unser IT-Team rotiert API-Schlüssel für jede Iphone-App." }'import requests
glossary_id = "cmglossary123"url = f"https://platform.algebras.ai/api/v1/translation/glossaries/{glossary_id}/terms/count"headers = { "X-Api-Key": "your_api_key_here", "Content-Type": "application/json"}data = { "sourceLanguage": "en", "sourceText": "Our IT team rotates API keys for every iPhone app.", "targetLanguage": "de", "targetText": "Unser IT-Team rotiert API-Schlüssel für jede Iphone-App."}
response = requests.post(url, headers=headers, json=data)
if response.status_code == 200: for term in response.json()["data"]["terms"]: followed = term["targetCount"] >= term["sourceCount"] print(f'{term["sourceTerm"]} -> {term["targetTerm"]}: {"ok" if followed else "missing"}', [m["text"] for m in term["targetCaseMismatches"]])else: print(f"Error: {response.status_code}") print(response.json())const glossaryId = "cmglossary123";const url = `https://platform.algebras.ai/api/v1/translation/glossaries/${glossaryId}/terms/count`;
const response = await fetch(url, { method: "POST", headers: { "X-Api-Key": "your_api_key_here", "Content-Type": "application/json", }, body: JSON.stringify({ sourceLanguage: "en", sourceText: "Our IT team rotates API keys for every iPhone app.", targetLanguage: "de", targetText: "Unser IT-Team rotiert API-Schlüssel für jede Iphone-App.", }),});
if (response.ok) { const { data } = await response.json(); for (const term of data.terms) { const followed = term.targetCount >= term.sourceCount; console.log(`${term.sourceTerm} -> ${term.targetTerm}: ${followed ? "ok" : "missing"}`); }} else { const error = await response.json(); console.error("Error:", response.status, error);}package main
import ( "bytes" "encoding/json" "fmt" "io" "net/http")
func main() { glossaryID := "cmglossary123" url := fmt.Sprintf("https://platform.algebras.ai/api/v1/translation/glossaries/%s/terms/count", glossaryID)
payload := map[string]string{ "sourceLanguage": "en", "sourceText": "Our IT team rotates API keys for every iPhone app.", "targetLanguage": "de", "targetText": "Unser IT-Team rotiert API-Schlüssel für jede Iphone-App.", }
jsonData, _ := json.Marshal(payload)
req, _ := http.NewRequest("POST", url, bytes.NewBuffer(jsonData)) req.Header.Set("X-Api-Key", "your_api_key_here") req.Header.Set("Content-Type", "application/json")
client := &http.Client{} resp, err := client.Do(req) if err != nil { panic(err) } defer resp.Body.Close()
body, _ := io.ReadAll(resp.Body) fmt.Printf("%d\n%s\n", resp.StatusCode, string(body))}Word Forms
Section titled “Word Forms”A glossary keeps one form of each term per language, so that translations stay consistent. The counter still has to recognize the term when grammar changes its form. It does this in two passes:
- The entry as written, with letter case, separators and a list of common endings per language taken into account (“APIs” for
API, “Straßen” forStraße). - Dictionary forms (lemmas) for the rest of the text. Each word of the text and of the entry is reduced to its dictionary form, and the sequences are compared word by word.
| Language | Text | Entry | Result |
|---|---|---|---|
| English | “policies”, “mice”, “children”, “went” | policy, mouse, child, go |
inflected |
| German | “Rechenzentren”, “Äpfel” | Rechenzentrum, Apfel |
inflected |
| Russian | “машины”, “машинного обучения” | машина, машинное обучение |
inflected |
| Spanish | “luces” | luz |
inflected |
| Arabic, Hebrew | “والكتاب”, “והספר” | كتاب, ספר |
inflected (attached prefixes) |
| Japanese | 食べた | 食べる |
inflected |
| English | “see”, “seeing” | saw (the tool) |
ambiguous, lemma_only |
The last row shows the limit of dictionary forms: different words can share one. When the entry word isn’t a dictionary form itself and the two forms share little of their beginning, the match is reported as ambiguous with the reason lemma_only instead of being counted.
Dictionary forms are available for about 50 languages, including all major European languages, Russian, Ukrainian, Turkish, Arabic, Hebrew, Persian, Hindi, Indonesian and Malay. For other languages, only the first pass applies.
Spelling Variants (Japanese)
Section titled “Spelling Variants (Japanese)”Japanese often writes the same word in more than one way: サーバ and サーバー (server), いく and 行く (to go). Because a glossary exists to keep terminology consistent, a different spelling of the term isn’t counted. On the target side it’s listed in targetSpellingVariants, so the translation can be corrected to the glossary’s spelling. Conjugated forms of the same spelling (食べた for 食べる) do count.
Languages Without Spaces
Section titled “Languages Without Spaces”Chinese, Japanese and Thai are written without spaces between words, and Korean attaches particles directly to words. For these languages the counter finds the term in the text and then asks a word segmenter whether the term starts and ends on word boundaries:
| Language | Text | Entry | Result |
|---|---|---|---|
| Chinese | 我们使用机器学习模型 | 机器学习 |
match |
| Chinese | 学习者 (learner) | 学习 (learning) |
ambiguous, segmentation_boundary |
| Japanese | 機械学習モデルを | 機械学習 |
match |
| Thai | ผู้เรียนรู้ | เรียนรู้ |
ambiguous, segmentation_boundary |
| Korean | 머신러닝은 좋다 | 머신러닝 |
match, without the particle 은 |
A term split into several words by the segmenter still matches, as long as its first and last characters are on word boundaries: 机器学习 is two dictionary words, 机器 and 学习. Segmenters can be wrong about new vocabulary, so a term found inside a longer word is reported as ambiguous rather than dropped.
Full-width letters and digits, common in Chinese and Japanese text, match their normal forms: “APIキー” contains API.
For Lao, Khmer, Burmese and Tibetan, terms are found anywhere in the text, without a boundary check.
Letter Case Matching
Section titled “Letter Case Matching”Capitalization can change what a word means: “IT” (information technology) and “it” (the pronoun), “mW” (milliwatt) and “MW” (megawatt). It can also be irrelevant: “Email” at the start of a sentence is still “email”. The caseSensitivity setting tells the term counter which applies to a glossary entry.
The setting belongs to each language entry of a term, so the English and German entries of the same term can differ. It’s used only by the term counter and doesn’t change how translations are generated.
caseSensitivity |
Rule | Use for | Example |
|---|---|---|---|
auto (default) |
Uppercase letters must match; lowercase letters may be either case. If only the first letter is capitalized, a lowercase occurrence is reported as ambiguous for review | Most terms | IT ≠ “it” · term = “Term” = “TERM” · Apple → “apple”: ambiguous |
capitals |
Uppercase letters must match; lowercase letters may be either case. No ambiguity for an initial capital: it is always meaningful | Brand and product names | Apple ≠ “apple” · iPhone = “IPHONE” |
exact |
Every letter must match exactly | Units, symbols, code identifiers | mW ≠ “MW” · pH ≠ “PH” |
ignore |
Any capitalization counts; matches are never ambiguous | Ordinary words that can’t be confused with other words | email = “Email” = “EMAIL” |
In every mode, accents must match, except in text written entirely in capitals.
How Each Mode Treats the Same Text
Section titled “How Each Mode Treats the Same Text”| Entry | Text | auto |
capitals |
exact |
ignore |
|---|---|---|---|---|---|
IT |
“it works” | no match | no match | no match | match |
IT |
“Contact the IT team” | match | match | match | match |
IT |
“WE FIXED IT” | ambiguous | ambiguous | ambiguous | match |
Apple |
“an apple” | ambiguous | no match | no match | match |
term |
“Term”, “TERM” | match | match | no match | match |
iPhone |
“IPHONE” | match | match | no match | match |
mW |
“5 MW” | match | match | no match | match |
The mW row shows why exact exists: with any other mode, “MW” (megawatt) would be counted as mW (milliwatt).
On the target side, an occurrence rejected only because of its letter case is listed in targetCaseMismatches. With exact, a translation that writes “Iphone” or “IPHONE” for iPhone shows up there.
Setting the Mode
Section titled “Setting the Mode”Send caseSensitivity with a definition when you create, bulk-create or update a term. Omitting it on create means auto; omitting it on update keeps the current value; sending auto resets an entry to the default. Every response that returns definitions includes the value.
Create a unit term that must match exactly in every language:
curl -X POST "https://platform.algebras.ai/api/v1/translation/glossaries/cmglossary123/terms" \ -H "X-Api-Key: your_api_key_here" \ -H "Content-Type: application/json" \ -d '{ "definitions": [ { "language": "en", "term": "mW", "definition": "Milliwatt", "caseSensitivity": "exact" }, { "language": "de", "term": "mW", "definition": "Milliwatt", "caseSensitivity": "exact" } ] }'import requests
glossary_id = "cmglossary123"url = f"https://platform.algebras.ai/api/v1/translation/glossaries/{glossary_id}/terms"headers = { "X-Api-Key": "your_api_key_here", "Content-Type": "application/json"}data = { "definitions": [ {"language": "en", "term": "mW", "definition": "Milliwatt", "caseSensitivity": "exact"}, {"language": "de", "term": "mW", "definition": "Milliwatt", "caseSensitivity": "exact"} ]}
response = requests.post(url, headers=headers, json=data)print(response.status_code, response.json())const glossaryId = "cmglossary123";const url = `https://platform.algebras.ai/api/v1/translation/glossaries/${glossaryId}/terms`;
const response = await fetch(url, { method: "POST", headers: { "X-Api-Key": "your_api_key_here", "Content-Type": "application/json", }, body: JSON.stringify({ definitions: [ { language: "en", term: "mW", definition: "Milliwatt", caseSensitivity: "exact" }, { language: "de", term: "mW", definition: "Milliwatt", caseSensitivity: "exact" }, ], }),});
console.log(response.status, await response.json());package main
import ( "bytes" "encoding/json" "fmt" "io" "net/http")
func main() { glossaryID := "cmglossary123" url := fmt.Sprintf("https://platform.algebras.ai/api/v1/translation/glossaries/%s/terms", glossaryID)
payload := map[string]interface{}{ "definitions": []map[string]string{ {"language": "en", "term": "mW", "definition": "Milliwatt", "caseSensitivity": "exact"}, {"language": "de", "term": "mW", "definition": "Milliwatt", "caseSensitivity": "exact"}, }, }
jsonData, _ := json.Marshal(payload)
req, _ := http.NewRequest("POST", url, bytes.NewBuffer(jsonData)) req.Header.Set("X-Api-Key", "your_api_key_here") req.Header.Set("Content-Type", "application/json")
client := &http.Client{} resp, err := client.Do(req) if err != nil { panic(err) } defer resp.Body.Close()
body, _ := io.ReadAll(resp.Body) fmt.Printf("%d\n%s\n", resp.StatusCode, string(body))}Change only the mode of the English entry of an existing term. The term text and definition stay as they are:
curl -X PUT "https://platform.algebras.ai/api/v1/translation/glossaries/cmglossary123/terms/cmterm123" \ -H "X-Api-Key: your_api_key_here" \ -H "Content-Type: application/json" \ -d '{ "definitions": [{ "language": "en", "caseSensitivity": "capitals" }] }'An unknown value returns 400 with caseSensitivity must be one of: auto, capitals, exact, ignore.
The setting is also available in the app, in the Letter case column of a glossary (see Glossaries), and through the MCP term tools (create_glossary_term, bulk_create_glossary_terms, update_glossary_term).