
Security News
Insecure Agents Podcast: How to Keep AI Agents From Bypassing Security Controls
Socket CTO Ahmad Nassri discusses how to keep AI agents from bypassing package blocks, limit credential access, and monitor their actions.
wikidata-google-knowledge-mcp
Advanced tools
MCP server and CLI for bounded Wikidata search, optional Google Knowledge Graph cross-checks and deterministic entity resolution for AI agents.
Token-efficient entity search, resolution and knowledge-graph helpers for AI agents.
Documentation · Install · Demo and benchmark · Changelog
wikidata-google-knowledge-mcp is an open-source MCP server, CLI and Agent Skill for
Claude Code, Cursor, Codex and other MCP clients. It searches and resolves real-world
entities across Wikidata and, optionally, the Google Knowledge Graph without dumping large
provider responses into your LLM context. You get a few ranked candidates, the facts you
asked for, and a deterministic entity-resolution decision with evidence codes, including an
explicit HOLD or AMBIGUOUS when the evidence does not settle the identity.
/m/... = Wikidata P646, Google /g/... = Wikidata P2671).Read-only. Wikidata needs no API key.
Audit an existing QID with kg_audit_identity / wdkg audit-identity; discover missing identities
with kg_discover_identity / wdkg discover-identity; fetch compact bulk facts with
kg_lookup_entities / wdkg lookup-qids. Large corpora use wdkg identity-batch streaming JSONL.
Identity verdicts: VERIFIED_EXISTING, DETERMINISTIC_CONFIRM, WRONG_IDENTITY, AMBIGUOUS,
INSUFFICIENT_EVIDENCE. Reviewed provenance is required for VERIFIED_EXISTING; provider
agreement is not local proof. identity candidate != approved public sameAs.
Google is off for these operations unless --google-crosscheck is explicitly requested.
The five existing MCP tools and prior CLI commands remain available.
Identity guide · Closed reason codes
Why this exists · Features · Quick start · MCP installation · Wikidata-only usage · Google Knowledge Graph setup · Tool reference · Entity resolution · Batch resolution · Ambiguity and evidence · Examples · Benchmark · Supported clients · Privacy · Development
Generic Wikidata tools answer "what does Wikidata say about X?". Agents that link records to a knowledge graph mostly need to know which entity is theirs. When a model does that through raw search and statement dumps, it reads dozens of namesakes and full claim lists, then picks one by prose reasoning. This project moves that work into deterministic code and gives the model a short answer.
| Raw Wikidata MCP workflow | wikidata-google-knowledge-mcp | |
|---|---|---|
| Search results | every provider hit goes into context | 3 candidates by default (max 5) with name/place/type match flags |
| Entity facts | full statement lists | only the properties you ask for; ranks, qualifiers and references on request |
| Output size | unbounded | every response is fitted to a byte budget (6,000 bytes by default, configurable) |
| Relationships | hand-written SPARQL | kg_related: outgoing, inverse or class hierarchy, depth- and row-capped |
| Identity decision | the model picks from candidates | deterministic decision plus evidence codes (OFFICIAL_HOST_EXACT, GEO_MATCH...) |
| Namesakes | easy to pick the wrong one | HOLD / AMBIGUOUS are first-class results; no identity probabilities |
| Google Knowledge Graph | separate tool, manual comparison | exact joins via Wikidata P646 / P2671, kept apart from identity proof |
| Thousands of entities | one model tool call per entity | wdkg resolve-batch: streaming JSONL, checkpoint/resume, request budgets |
| Repeated lookups | new provider calls each time | local SQLite cache; a finished batch rerun makes zero provider requests |
kg_search returns at most 5 candidates (3 by default) and says how
the name, place and type matched. Place and type hints are checked locally; they are not
sent to the provider.kg_entity returns the properties you name (P31, P131...) or a
short identity/location overview. Up to three properties can include ranks, qualifiers
and references.kg_related follows one property outward or inward, or walks
instance-of/subclass-of, with depth and row caps.kg_resolve takes a local identity envelope (name, kind, city,
country, coordinates, address, official URL, aliases, venue, event date, organizer, role,
occupation, creator, year, existing ids) and reconciles provider candidates
deterministically.kg:/m/... ids are joined to Wikidata P646 and
kg:/g/... ids to P2671 by exact string equality; canonical Wikipedia URLs are compared
with Wikidata sitelinks. Provider agreement is reported separately from proof that the
entity is yours.AMBIGUOUS or HOLD, never a silent pick.
Provider ranking is never turned into an identity probability.EXTERNAL_ID_EXACT,
OFFICIAL_HOST_EXACT, GEO_MATCH, ADDRESS_MATCH, CITY_MATCH, COUNTRY_MATCH,
KIND_COMPATIBLE, EVENT_DATE_MATCH and VENUE_MATCH, not confidence percentages.wdkg resolve-batch streams JSONL in and out with bounded memory,
checkpoint/resume, per-provider request caps, per-row errors, cache reuse and
deterministic output, plus an auditable evidence export.kg_status reports provider endpoints, whether a Google key is
present and where it came from, cache statistics and limits. It never prints the key.One Python core serves the MCP server, the wdkg CLI and the Agent Skill; there is no
second resolver implementation.
Requirements: uv (it provides Python 3.11+ if needed) and outbound HTTPS.
# run once without installing (from PyPI)
uvx --from wikidata-google-knowledge-mcp==0.3.0 wdkg search "Fox Theatre" --place Atlanta
# or install both commands (wdkg and wikidata-google-knowledge-mcp) on PATH
uv tool install wikidata-google-knowledge-mcp==0.3.0
wdkg resolve "Fox Theatre" --kind place --city Atlanta --url https://www.foxtheatre.org
The first run downloads the package and its dependencies; later runs use uv's cache. The
same release is also installable straight from the GitLab tag:
uv tool install git+https://gitlab.com/revanalex/wikidata-google-knowledge-mcp@v0.3.0.
wdkg --help lists every command with examples and explains the decisions;
wdkg status shows provider and credential status without any network call.
The MCP server is wikidata-google-knowledge-mcp (stdio). Every client below starts it with
the same command:
uvx wikidata-google-knowledge-mcp@0.3.0
If you installed it with uv tool install, the command is just wikidata-google-knowledge-mcp.
As a plugin (MCP server and skill together):
claude plugin marketplace add https://gitlab.com/revanalex/wikidata-google-knowledge-mcp.git
claude plugin install wikidata-google-knowledge-mcp@wikidata-google-knowledge-mcp
The same works inside a session with /plugin marketplace add ... and /plugin install ....
MCP server only:
claude mcp add --scope user wikidata-google-knowledge -- \
uvx wikidata-google-knowledge-mcp@0.3.0
As a plugin (MCP server and skill together):
codex plugin marketplace add https://gitlab.com/revanalex/wikidata-google-knowledge-mcp.git
codex plugin add wikidata-google-knowledge-mcp@wikidata-google-knowledge-mcp
MCP server only, in ~/.codex/config.toml:
[mcp_servers.wikidata-google-knowledge]
command = "uvx"
args = ["wikidata-google-knowledge-mcp@0.3.0"]
startup_timeout_sec = 60 # the first start downloads the package
env_vars = ["GOOGLE_KNOWLEDGE_GRAPH_API_KEY"] # optional: forward the Google key
or codex mcp add wikidata-google-knowledge -- uvx wikidata-google-knowledge-mcp@0.3.0.
Add to ~/.cursor/mcp.json (all projects) or .cursor/mcp.json (one project):
{
"mcpServers": {
"wikidata-google-knowledge": {
"command": "uvx",
"args": ["wikidata-google-knowledge-mcp@0.3.0"]
}
}
}
To load the full plugin (MCP server and skill) locally, clone this repository into
~/.cursor/plugins/local/wikidata-google-knowledge-mcp and reload the window.
Any client that launches stdio servers can use the command above. The server is listed in the
official MCP Registry as com.dondego/wikidata-google-knowledge-mcp, so clients that
read the registry can find and install it from there. The repository also ships portable
Agent Plugins manifests (plugin.json, mcp.json, skills/).
Everything works without Google. Wikidata lookups use the public Wikidata MCP service on
Wikimedia Cloud (wd-mcp.wmcloud.org) for interactive calls, and the Wikidata Action API and Query Service for batch resolution. No account or
key is needed.
wdkg search "Tate Modern" --place London --type museum
wdkg entity Q193375 --props P31,P131,P17 --evidence P131
wdkg related Q193375 --prop P361
wdkg related Q193375 --hierarchy --depth 2
Google adds a second, independent candidate source and exact id joins. Without a key it is
never called. With a key, interactive search uses it only when you ask (provider="google",
--provider google, --fallback); the resolver uses it for every row with --providers dual,
and in its default minimal mode only for rows Wikidata left unresolved that carry an
official_url, or when the optional model arbitration is on (the cases where Google can change
the result).
Create an API key for the Knowledge Graph Search API (Google prerequisites) and restrict it to that API.
Provide it as GOOGLE_KNOWLEDGE_GRAPH_API_KEY in the environment of the MCP server or
CLI, or put one literal line in ~/.config/wikidata-google-knowledge-mcp/secrets.env
(chmod 600), which also works for GUI clients that do not pass your shell environment:
GOOGLE_KNOWLEDGE_GRAPH_API_KEY=your-key-here
The file is parsed, never executed. The key travels only in the X-Goog-Api-Key header.
Check with wdkg status or kg_status: they report present and the source, never the
value.
The Knowledge Graph Search API is an older Google API. Its availability, quotas and terms are set by Google and may change; this project cannot promise it will stay available. All Wikidata features keep working without it.
| MCP tool | Use it for | Key parameters |
|---|---|---|
kg_search | searching by name | query, place, type, lang, limit (≤5), provider (wikidata/google), fallback |
kg_entity | reading selected facts of one QID | id, props (≤12 PIDs), evidence (≤3 PIDs with ranks, qualifiers, references), lang |
kg_related | following relationships | id, prop, inverse, hierarchy, depth (≤3), limit (≤25) |
kg_resolve | resolving one ambiguous real-world entity | name, kind, city, country, latitude/longitude, address, official_url, aliases, venue, event_date, organizer, role, occupation, affiliation, creator, year, existing_wikidata_qid, existing_google_kg_id, lang, explain |
kg_status | checking provider and credential status | none; makes no network calls |
All tools are read-only and return one compact JSON object with status, warnings and a
meta block (cache state, upstream request count, bytes).
The wdkg CLI exposes the same core: search, entity, related, sparql (guarded,
read-only, LIMIT enforced), status, cache stats|clear, resolve, resolve-batch,
export-evidence and validate-evidence. Run wdkg <command> --help for flags. Output
fields, warnings, limits and settings are documented in
skills/wikidata-google-knowledge/reference.md.
kg_resolve (one entity) and wdkg resolve / resolve-batch accept an identity envelope.
Only name is required; give whatever you already know:
{"name": "Fox Theatre", "kind": "place", "city": "Atlanta", "country": "US",
"official_url": "https://www.foxtheatre.org"}
Supported fields: name, kind (person, organization, place, event,
event_series, work, other), city, country, latitude + longitude, address,
official_url, aliases, venue, event_date, organizer, role, occupation,
affiliation, creator, year, existing_wikidata_qid, existing_google_kg_id, lang,
and in the CLI/JSONL only source_id, existing_id_trusted and official_url_reviewed.
The last two assert human review, so the MCP tool does not accept them.
Every result carries a decision:
| decision | meaning |
|---|---|
AUTO_MATCH | a kind-specific rule passed: a local anchor (official host, coordinates, address, creator, date + venue...) agrees and nothing conflicts |
AMBIGUOUS | several viable identities remain |
HOLD | the policy will not decide automatically, or evidence is incomplete |
CONFLICT | evidence disagrees, for example an existing id is contradicted |
NO_CANDIDATE | the bounded search found nothing usable (not proof that nothing exists) |
MODEL_MATCH | optional, off by default: a model picked a supplied candidate (CLI only) |
Kind rules in short: a place needs an official host, coordinates or address match; an
organization needs its official host; a person needs a reviewed official site; a dated event
needs its date plus venue, organizer or coordinates; a series needs its official host or
city plus organizer/venue; a work needs its creator. The full rules, evidence vocabulary and
retention basis are in
skills/wikidata-google-knowledge/resolution.md.
For corpora, use the CLI rather than repeated MCP calls:
wdkg resolve-batch entities.jsonl --out resolved.jsonl \
--max-wikidata-requests 2000 --max-google-requests 0
resolved.jsonl.receipt.json
with decision counts, request usage and stop reason.--dry-run parses and plans without network. --providers dual adds Google for every row
(needs a key); the default minimal mode calls Google only when it can change a result.wdkg export-evidence resolved.jsonl --bundle evidence/ writes an auditable bundle
(decisions.jsonl, evidence.jsonl, manifest.json, validation.json);
wdkg validate-evidence evidence/ checks it.Two kinds of agreement are kept apart:
EXTERNAL_ID_EXACT, WIKIPEDIA_EXACT, PROVIDER_NAME_AGREEMENT...). This says nothing
about whether that thing is your entity.OFFICIAL_HOST_EXACT, GEO_MATCH, ADDRESS_MATCH, CITY_MATCH, EVENT_DATE_MATCH,
VENUE_MATCH, CREATOR_MATCH...). Only these can produce AUTO_MATCH.Conflicts (KIND_CONFLICT, GEO_CONFLICT, EVENT_GRAIN_CONFLICT, CITY_CONFLICT...)
reject a candidate or block automation. Google's resultScore is passed through as
result_score_raw for display and never used in a decision. Nothing is guessed: missing
input stays missing, and QIDs or Google ids only ever come from provider responses or your
input.
The outputs below come from real runs of v0.1.0 against live Wikidata (and, where noted, the Google Knowledge Graph), shortened to the relevant fields.
An exact label lookup finds 131 Wikidata items named "Fox Theatre" (the namesake_heavy:131
warning). With an official website:
wdkg resolve "Fox Theatre" --kind place --city Atlanta --url https://www.foxtheatre.org --providers dual
{
"decision": "AUTO_MATCH",
"wikidata_qid": "Q1440190",
"google_kg_id": "kg:/m/04qrhq",
"provider_concordance": ["EXTERNAL_ID_EXACT", "PROVIDER_HOST_AGREEMENT", "PROVIDER_NAME_AGREEMENT",
"PROVIDER_TYPE_COMPATIBLE", "WIKIPEDIA_EXACT"],
"local_anchor_evidence": ["CITY_MATCH", "OFFICIAL_HOST_EXACT"],
"evidence_codes": ["CITY_MATCH", "EXTERNAL_ID_EXACT", "KIND_COMPATIBLE", "NAME_EXACT", "OFFICIAL_HOST_EXACT", "..."],
"warnings": ["namesake_heavy:131"]
}
EXTERNAL_ID_EXACT here means Google's kg:/m/04qrhq equals the P646 value on Wikidata
Q1440190. Without Google (--providers minimal, no key) the same call still returns
AUTO_MATCH for Q1440190, with google_kg_id: null.
wdkg resolve "Fox Theatre" --kind place --city Atlanta --providers dual
{
"decision": "HOLD",
"reasons": ["NO_LOCAL_ANCHOR"],
"candidate_ids": ["Q1440190", "Q3080199", "Q8565227", "Q559730"],
"candidate_pairs": [
{"wikidata_qid": "Q1440190", "google_kg_id": "kg:/m/04qrhq", "methods": ["EXTERNAL_ID_EXACT", "WIKIPEDIA_EXACT"], "selected": false},
{"wikidata_qid": "Q3080199", "google_kg_id": "kg:/m/05c6w7", "methods": ["EXTERNAL_ID_EXACT", "WIKIPEDIA_EXACT"], "selected": false}
],
"local_anchor_evidence": ["CITY_MATCH"]
}
Google and Wikidata agree exactly on four different Fox Theatres. That agreement is
provider concordance, not proof of which one is yours, and a city name alone is not enough
for a place, so the result is HOLD.
wdkg resolve "Tate Modern" --kind place --city London --lat 51.5076 --lon -0.0994
The coordinates match both Tate Modern (Q193375, the gallery) and Bankside Power Station
(Q806832, the building that houses it). The result is AMBIGUOUS
(MULTIPLE_ANCHORED_CANDIDATES) with both ids in candidate_ids, instead of a guess. An
official URL would settle it.
wdkg resolve "Primavera Sound" --kind event --date 2019-05-30 --city Barcelona
# decision NO_CANDIDATE, conflict EVENT_GRAIN_CONFLICT on Q2439480 (the festival series)
wdkg resolve "Primavera Sound" --kind event_series --city Barcelona --url https://www.primaverasound.com
# decision AUTO_MATCH, wikidata_qid Q2439480, local anchors CITY_MATCH + OFFICIAL_HOST_EXACT
The 2019 edition is a dated occurrence; Q2439480 is the recurring festival. The resolver never links an occurrence to its series or the other way round.
examples/entities.jsonl holds five of the records above plus an
underspecified person:
wdkg resolve-batch examples/entities.jsonl --out resolved.jsonl --max-google-requests 0
{"op": "resolve-batch", "status": "ok", "rows": 6,
"decisions": {"AUTO_MATCH": 2, "AMBIGUOUS": 2, "HOLD": 1, "NO_CANDIDATE": 1, "CONFLICT": 0, "MODEL_MATCH": 0}}
Running the same command again reports "processed": 0, "skipped_already_done": 6 and no
upstream requests.
Measured on seven questions replayed from captured Wikidata MCP responses (offline, CC0 data),
the answer the model reads is 61–85% smaller than the upstream text for searches with many
hits and for entity facts, and larger for tiny answers (a 25-hit search, a one-level class
hierarchy), because match flags, warnings and the meta block cost bytes. A live session
against Wikidata used 3 upstream requests per cold call and 0 on repeat. Method, per-case
numbers, what is omitted and the limitations are on the
benchmark page; there is no accuracy claim.
claude mcp add)config.toml)mcp.json, or the plugin loaded locally)wdkg CLIThe skill in skills/wikidata-google-knowledge/
tells the agent which tool to use when; it is plain Markdown and works across clients.
kg_search are checked locally and are not sent. There is no telemetry.~/.cache/wikidata-google-knowledge-mcp/
(Wikidata results expire after 7 days by default; WDKG_CACHE=0 or --no-cache disables it) and, for the
resolver, an evidence store in the same directory.Cache-Control
header allows. The Knowledge Graph Search API currently sends none, so nothing received
only from Google is stored: no names, descriptions, scores, website URLs or Wikipedia
URLs. A Google id is kept only when it equals an independently obtained string (a Wikidata
P646/P2671 value or your own input).Details: docs/privacy.md.
git clone https://gitlab.com/revanalex/wikidata-google-knowledge-mcp.git
cd wikidata-google-knowledge-mcp
uv sync
uv run pytest -q # offline: fixtures and fake transports only
uv run python scripts/check_public_tree.py # forbidden files, secret patterns, manifest checks
uv run --group docs python scripts/build_site.py --out public # documentation site
uv run --group bench python benchmarks/replay_benchmark.py # offline benchmark
The tests never call Wikidata, Google or any model API. CI runs the same checks plus a clean wheel install, a CLI smoke test, an MCP stdio start-up check and a secret scan. See CHANGELOG.md for releases and SECURITY.md for reporting vulnerabilities.
MIT; see LICENSE. Wikidata content is available under CC0. Google Knowledge Graph results are subject to Google's terms and are not redistributed by this project.
This is an independent open-source project. It is not affiliated with, endorsed by or sponsored by the Wikimedia Foundation, Wikimedia Deutschland, Google, OpenAI, Anthropic or Anysphere (Cursor). Product names are used only to describe compatibility.
Built by the DondeGo.com team as part of our work on semantic city discovery.
0.3.0 — 2026-09-30
Add bounded remote Streamable HTTP infrastructure around canonical PyPI 0.3.0, with authoritative SQLite daily admission, fail-closed accounting and Wikidata-only egress. Production is live at https://wikidata-knowledge-mcp.revanalex.workers.dev/mcp; Registry metadata 0.3.1 adds the remote while retaining PyPI 0.3.0.
Document remote privacy, UTC quota behavior, tool bounds and native client configuration.
Publish a portable Smithery MCPB launcher for the existing 0.3.0 PyPI release, with eight read-only tools and no bundled resolver or credentials.
Document local MCPB installation and add a bounded stdio acceptance helper.
server.json name, PyPI package, stdio).FAQs
MCP server and CLI for bounded Wikidata search, optional Google Knowledge Graph cross-checks and deterministic entity resolution for AI agents.
We found that wikidata-google-knowledge-mcp demonstrated a healthy version release cadence and project activity because the last version was released less than a year ago. It has 1 open source maintainer collaborating on the project.

Security News
Socket CTO Ahmad Nassri discusses how to keep AI agents from bypassing package blocks, limit credential access, and monitor their actions.

Security News
GPT-6 Astra tried to plant malicious code in simulated open source projects using fake GitHub accounts and deceptive PRs during an assigned CTF challenge.

Security News
upm uses Node.js to deliver fast npm installs in about 250 KB, with a JavaScript API and security defaults.