@cyanheads/census-mcp-server
Query U.S. Census Bureau data, variables, and geography via MCP. STDIO or Streamable HTTP.
7 Tools
Tools
8 tools covering the full Census data workflow — from dataset discovery and variable search through geography resolution and ranked comparisons:
census_list_datasets | Browse available Census Bureau datasets (ACS5, ACS1, Population Estimates, Decennial, County Business Patterns, Economic Census, Nonemployer Statistics) with vintage years and dataset codes. |
census_list_geographies | List the geography levels supported by a dataset and year, with parent requirements and example FIPS values. |
census_search_variables | Keyword search across variable labels and concept groups. On ACS, returns estimate and margin-of-error codes together. |
census_get_variable | Fetch full metadata for one or more variable codes — label, concept, predicate type, universe, MOE sibling. |
census_list_predicate_values | List the codes a filter dimension accepts (EMPSZES, LFO, POPGROUP, NAICS2017…), from the dataset dictionary or a live wildcard enumeration. |
census_resolve_geography | Convert place names (e.g., "King County, WA") or street addresses to Census FIPS identifiers via TIGERweb and Census Geocoder. |
census_query_data | Query a Census dataset for variables at a specific geography. Returns estimates with MOE, suppression codes resolved to readable reasons, and predicate filtering for the business datasets. |
census_compare_geographies | Rank and compare variables across multiple geographies — all counties in a state, all states nationally, or a named set. Sorted table output, with the same predicate filtering. |
census_list_datasets
Browse available Census Bureau datasets.
- Returns dataset codes, names, descriptions, and available vintage years
- Covers ACS5, ACS5 Data Profiles, ACS5 Subject Tables, ACS1, ACS1 Data Profiles, Population Estimates, Decennial Redistricting (P.L. 94-171), Decennial DHC, County Business Patterns (
cbp), Economic Census (ecnbasic), and Nonemployer Statistics (nonemp)
- Each description names the filter predicates the dataset requires and the geography levels it publishes — both vary by dataset
- Accepts an optional keyword filter
- Dataset codes (e.g.,
acs/acs5) are the values to pass to other tools
available_years is exhaustive, not a sample: any other year fails with year_not_available before a request goes out, naming the years that do work. It is narrower than what the Census API hosts — pep/charv reaches its 2020-2022 estimates through the YEAR filter inside the 2023 vintage, and the cbp/nonemp vintages left out reject the NAME column every query here sends
census_search_variables
Search Census variables by keyword.
- Full-text search across label and concept fields with relevance scoring (exact concept match > label match > partial)
- On ACS datasets, returns estimate (E suffix) and margin-of-error (M suffix) codes together so both can be requested in one query — no other family publishes margins of error, and an E-final code there is an ordinary code
- Also surfaces the predicate codes a dataset filters on, such as
NAICS2017 in cbp
- Configurable limit (default 20, max 100);
total_matches indicates how many matched before the limit
- Cache-backed: variables.json is fetched once per dataset+year with a configurable TTL (default 24h)
census_list_predicate_values
List the codes a filter dimension accepts, so a predicates map can be written without guessing.
- Two routes, picked by where the answer lives: a dimension with a published value list is read from the dataset dictionary, one without is enumerated live by wildcarding it on the data endpoint.
NAICS* and POPGROUP always publish one (thousands of codes — narrow them with query); on the current vintages EMPSZES, LFO, RCPSZES, TAXSTAT, and TYPOP publish none, so the live route is the only place their codes appear
- A dictionary value list is a classification shared across Census products, not a record of what one dataset serves —
dec/ddhca declares 5,543 POPGROUP codes and publishes 2,996, cbp declares 6,694 NAICS2017 codes and publishes 2,003. The declared list is checked against the dataset's own published rows and the dead codes are dropped; source says whether that check ran and the notice says how many were withheld. A keyword that matched only withheld codes names them, so "total population" on dec/ddhca reports that 001 is declared and serves nothing rather than reading like a typo
- Keyword
query matches code and label; results are sorted by code and a truncated list is disclosed rather than passed off as complete
ecnbasic publishes TAXSTAT and TYPOP per industry, so within_naics scopes the enumeration — and the notice says the result is complete for that industry alone. A per-industry dimension is left unchecked for the same reason, since an unscoped check would withhold codes a scoped query does return
- Live enumerations are cached per dataset, year, dimension, industry scope, and probe measure
census_resolve_geography
Convert place names and addresses to Census FIPS identifiers.
- Named places (e.g., "King County, WA", "Seattle, WA", "California") resolved via TIGERweb MapServer
- Street addresses resolved to tract level via Census Geocoder
- Auto-detects the geography level — state for an abbreviation or spelled-out state name, county for "County"/"Borough"/"Parish", tract for "Tract", otherwise place falling back to county;
geography_type overrides it
- Also resolves metropolitan/micropolitan statistical areas, combined statistical areas, and consolidated cities — never auto-detected, since their names overlap city names, so each needs an explicit
geography_type. The value is the level's own Census API name, so it feeds geography_level unchanged
- Optional
county_fips pins a tract name to one county, since a tract name is unique only inside its county. Only county and tract sit within a county, so it restricts resolution to those two levels rather than being dropped on a layer that cannot apply it
- Prefers an exactly-named match, so "Kansas City, MO" does not resolve to North Kansas City
- Never picks between matches: anything still matching more than one geography comes back as
ambiguous_name, with every candidate carrying the code resolving it would have returned, plus the state that separates same-named places
- Returns
state_fips (→ parent_fips) and fips_summary (→ geography_fips) ready to pass to other tools; a statistical area omits state_fips, since it can span several states and takes no parent
census_query_data
Query a Census dataset for one or more variables at a specific geography.
- Requires FIPS codes — use
census_resolve_geography first for place names
- Use
geography_fips: "*" to return all geographies at the level within the parent
- The level and its parents are checked against the dataset's own geography metadata before the query runs: a missing
parent_fips returns parent_required naming what to add, and a parent the level does not sit within returns parent_not_accepted naming the input to drop — neither reaches the API as an opaque 400
parent_fips and county_fips are zero-padded to the widths the Census matches on, so "5" and "05" both find Arkansas; either also takes "*", which is what reaches every block group in a state. geography_fips takes its width from geography_level and is passed through as given
- Each row carries both
geography_fips (bare level code, round-trips back into this tool) and geography_geoid (level plus parents, nationally unique)
- A query that matches nothing returns
no_data with dataset-aware recovery, not a retried upstream error
- Optional
predicates map for the datasets that filter on one — {"NAICS2017": "5112"} narrows a cbp count to software publishers, and census_list_predicate_values supplies the codes. Keys are validated against the dataset's own variables before the query
- Dimensions left unset are named in a notice and their applied default is echoed per row in
applied_filters. That label is load-bearing: cbp defaults NAICS2017 to the all-industries total, but dec/ddhca defaults POPGROUP to one population group and ecnbasic defaults its NAICS dimension to a single sector, so an unfiltered value can read like a total without being one. A dimension that publishes no label attribute (pep/charv YEAR, the nonemp NAICS codes before 2012) has no default to echo, and the notice says so rather than leaving it looking undefaulted
- One geography can come back on more than one row:
pep/charv publishes an April 1 estimates base alongside its July 1 estimate, and MONTH is what separates them — not YEAR, which both rows carry. Each row names its record in a record field and on its rendered heading, and the notice gives the predicate that pins one ({"MONTH": "7"})
- Suppression codes (geography too small, data not collected, etc.) resolved to human-readable reasons
- A cell that holds text rather than a number keeps it, under
value, so a null estimate says which of three things it is: suppressed is a number the Census withheld, a value alongside it is text (GEO_ID returns "0500000US53033"), and neither is an empty cell
- Variable labels enriched from cache and surfaced alongside estimates
- Requires
CENSUS_API_KEY
census_compare_geographies
Rank and compare variables across multiple geographies.
- Fetches all geographies at a level (e.g., all WA counties) in one API call, then sorts and slices
- Optional
within parameter to constrain to a parent FIPS; omit for national comparison
- Optional
geographies list to filter to specific geographies — full GEOIDs ("53033", "06037") work across states; bare level codes ("033") need within to disambiguate. Entries matching no row, and bare codes that matched more than one state, are named in a notice
- Same pre-query level and parent validation as
census_query_data, reported against within / within_county
- Configurable sort variable, direction, and limit (default 50, max 500)
- Same
predicates map as census_query_data, applied to every geography — without it the ranking runs on whatever default the API picks, named in the notice and echoed per row in applied_filters
- A dataset that publishes several records per geography is refused rather than ranked twice: a rank is a statement about one geography, so
pep/charv without a pinned record fails with ambiguous_rows naming MONTH and the code to pass. With one pinned, each geography ranks once and the row says which record it is
- Suppressed values sorted to end of results and labeled rather than passed through as negative sentinels
- Same
value field as census_query_data for a text cell; text has no ordering, so sorting on a column of it leaves every row tied
- Requires
CENSUS_API_KEY
Features
Built on @cyanheads/mcp-ts-core:
- Declarative tool definitions — single file per tool, framework handles registration and validation
- Unified error handling — handlers throw, framework catches, classifies, and formats with recovery hints
- Structured logging with optional OpenTelemetry tracing
- STDIO and Streamable HTTP transports
Census-specific:
- In-process variable cache with configurable TTL — variables.json fetched once per dataset+year, searched client-side
- Three-API backend: Census Data API for data queries, TIGERweb for named-place resolution, Census Geocoder for address-to-tract
- Automatic retry with backoff on all external API calls
- FIPS formatting helpers — zero-padded state, county, and tract codes ready to pass between tools
Agent-friendly output:
- Workflow-oriented tool surface —
fips_summary and state_fips return values are ready to pass as geography_fips and parent_fips to the next tool
- Suppression codes decoded — Census negative sentinel values (e.g.,
-666666666) surfaced as human-readable reasons instead of raw numbers
- Recovery hints on errors — ambiguous geography names include candidate lists; missing API key errors include registration URL
Getting started
API key: Register a free key at api.census.gov/data/key_signup.html. Variable search and geography resolution work without a key; data queries (census_query_data, census_compare_geographies) require one.
Add the following to your MCP client configuration file:
{
"mcpServers": {
"census-mcp-server": {
"type": "stdio",
"command": "bunx",
"args": ["@cyanheads/census-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info",
"CENSUS_API_KEY": "your-census-api-key"
}
}
}
}
Or with npx (no Bun required):
{
"mcpServers": {
"census-mcp-server": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@cyanheads/census-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info",
"CENSUS_API_KEY": "your-census-api-key"
}
}
}
}
Or with Docker:
{
"mcpServers": {
"census-mcp-server": {
"type": "stdio",
"command": "docker",
"args": [
"run", "-i", "--rm",
"-e", "MCP_TRANSPORT_TYPE=stdio",
"-e", "CENSUS_API_KEY=your-census-api-key",
"ghcr.io/cyanheads/census-mcp-server:latest"
]
}
}
}
For Streamable HTTP, set the transport and start the server:
MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 CENSUS_API_KEY=... bun run start:http
Prerequisites
Installation
git clone https://github.com/cyanheads/census-mcp-server.git
- Navigate into the directory:
cd census-mcp-server
bun install
cp .env.example .env
Configuration
CENSUS_API_KEY | Required for data queries. Register free at api.census.gov/data/key_signup.html. | — |
CENSUS_DEFAULT_YEAR | Default vintage year when no year is specified. | 2024 |
CENSUS_VARIABLE_CACHE_TTL_HOURS | Hours to cache variables.json per dataset+year in memory. | 24 |
MCP_TRANSPORT_TYPE | Transport: stdio or http. | stdio |
MCP_HTTP_PORT | Port for HTTP server. | 3010 |
MCP_AUTH_MODE | Auth mode: none, jwt, or oauth. | none |
MCP_LOG_LEVEL | Log level (debug, info, notice, warning, error). | info |
OTEL_ENABLED | Enable OpenTelemetry instrumentation. | false |
See .env.example for the full list of optional overrides.
Running the server
Local development
bun run rebuild
bun run start:stdio
bun run start:http
Run checks and tests:
bun run devcheck
bun run test
bun run lint:mcp
Docker
docker build -t census-mcp-server .
docker run --rm -e CENSUS_API_KEY=your-key -p 3010:3010 census-mcp-server
The Dockerfile defaults to HTTP transport, stateless session mode, and logs to /var/log/census-mcp-server. OpenTelemetry peer dependencies are installed by default — build with --build-arg OTEL_ENABLED=false to omit them.
Project structure
src/index.ts | createApp() entry point — registers tools and initializes services. |
src/config/server-config.ts | Census-specific env var parsing and validation with Zod. |
src/mcp-server/tools/definitions/ | Tool definitions (*.tool.ts). |
src/services/census-api/ | Census Data API client — data queries, suppression code mapping, retry logic. |
src/services/geography/ | Geography resolution — TIGERweb named-place lookup and Census Geocoder address-to-tract. |
src/services/variable-cache/ | In-process variables.json cache with TTL and keyword search. |
tests/ | Vitest tests mirroring src/ structure. |
Development guide
See CLAUDE.md for development guidelines and architectural rules. The short version:
- Handlers throw, framework catches — no
try/catch in tool logic
- Use
ctx.log for request-scoped logging, ctx.state for tenant-scoped storage
- Register new tools via the barrel in
src/mcp-server/tools/definitions/index.ts
- Wrap external API calls: validate raw → normalize to domain type → return output schema; never fabricate missing fields
Contributing
Issues and pull requests are welcome. Run checks and tests before submitting:
bun run devcheck
bun run test
License
Apache-2.0 — see LICENSE for details.